[c-nsp] IOS XR deadlocking just wondering if this was already fixed?

Drew Weaver drew.weaver at thenap.com
Fri Jul 31 10:13:17 EDT 2026


Howdy,

We don't have a current TAC on the one remaining XR device we use and I can't find any information on whether this is a known bug or not or if it has been fixed.

I'm basically just trying to discern whether it's known and has been fixed prior to paying for a support contract to download the image.

Version 24.4.2

Jul 13 12:22:45  local7 1674: LC/0/0/CPU0:Jul 13 12:22:45.219 EDT: dumper[69551]: %OS-SYSLOG-3-LOG_ERR : Dumping core /misc/scratch/core/hvic_7_7851.by.6.20260713-122245.xr-vm_node0_0_CPU0.cd973.
core.gz
Jul 13 12:22:46  local7 1676: LC/0/0/CPU0:Jul 13 12:22:46.856 EDT: dumper[69577]: %OS-SYSLOG-3-LOG_ERR : hvic_7_7851 signature 4c74a4dde0b54dceb8232625ee6f7cb6
Jul 13 12:22:55  local7 1678: LC/0/0/CPU0:Jul 13 12:22:55.600 EDT: dumper[155]: %OS-COREHELPER-6-CORE_COPIED : Copied core hvic_7_7851.by.6.20260713-122245.xr-vm_node0_0_CPU0.cd973.core.gz to 0/R
P0/CPU0:/misc/disk1
Jul 13 12:22:56  local7 1680: LC/0/0/CPU0:Jul 13 12:22:56.006 EDT: dumper[155]: %OS-COREHELPER-6-DELETE_CORE : Deleted core file hvic_7_7851.by.6.20260713-122245.xr-vm_node0_0_CPU0.cd973.core.gz.
Jul 13 12:22:59  local7 1682: RP/0/RP0/CPU0:Jul 13 12:22:59.414 EDT: logger[68242]: %OS-SYSLOG-4-LOG_WARNING : PAM detected crash by process hvic_7 on 0_0_CPU0. All necessary files for debug have
been collected and saved at 0/RP0/CPU0  : harddisk:/cisco_support/PAM-crash-xr_0_0_CPU0-hvic_7-2026Jul13-122256.tgz (Please copy tgz file out of the router and send to Cisco support. This tgz file will be r
emoved after 14 days.)

It seems like this issue happened and then that started a memory leak or some other cascading issue that led to this larger issue.

Jul 24 05:38:53  local7 1700: RP/0/RP1/CPU0:Jul 24 05:38:53.716 EDT: gsp[339]: %OS-gsp-6-MSG_GROUP_UNRESPONSIVE : group 2021 im_attr_owners is unresponsive for at least 10 seconds
Jul 24 05:38:53  local7 1702: LC/0/0/CPU0:Jul 24 05:38:53.716 EDT: gsp[350]: %OS-gsp-6-MSG_GROUP_UNRESPONSIVE : group 2021 im_attr_owners is unresponsive for at least 10 seconds
Jul 24 05:38:53  local7 1704: RP/0/RP0/CPU0:Jul 24 05:38:53.721 EDT: gsp[397]: %OS-gsp-6-MSG_GROUP_UNRESPONSIVE : group 2021 im_attr_owners is unresponsive for at least 10 seconds
Jul 24 05:40:28  local7 1706: RP/0/RP0/CPU0:Jul 24 05:40:28.758 EDT: ospfv3[1025]: %ROUTING-OSPFv3-6-HA_INFO : Process 1: LWM close to standby, NSR connecting
Jul 24 05:40:31  local7 1708: RP/0/RP0/CPU0:Jul 24 05:40:31.258 EDT: rmf_svr[222]: %PKT_INFRA-FM-3-FAULT_MAJOR : ALARM_MAJOR :RP-RED-LOST-NSRNR :DECLARE :0/RP0/CPU0:
Jul 24 05:41:28  local7 1710: RP/0/RP1/CPU0:Jul 24 05:41:28.751 EDT: ospfv3[1025]: %ROUTING-OSPFv3-6-HA_INFO : Process 1: Transitioning to nsr available
Jul 24 05:41:38  local7 1712: RP/0/RP0/CPU0:Jul 24 05:41:38.751 EDT: rmf_svr[222]: %PKT_INFRA-FM-3-FAULT_MAJOR : ALARM_MAJOR :RP-RED-LOST-NSRNR :CLEAR :0/RP0/CPU0:
Jul 24 05:43:53  local7 1714: LC/0/0/CPU0:Jul 24 05:43:53.830 EDT: gsp[350]: %OS-gsp-6-MSG_GROUP_UNRESPONSIVE : group 2021 im_attr_owners is unresponsive for at least 10 seconds

... more of these ...

Jul 24 05:55:35  local7 1732: RP/0/RP0/CPU0:Jul 24 05:55:35.356 EDT: sysdb_shared_nc[323]: %SYSDB-SYSDB-6-TIMEOUT_EDM : EDM request for 'oper/intf_mgbl/gl/full/' from 'show_interface' (jid 68249,
node 0/RP0/CPU0). No response from 'intf_mgbl' (jid 1227, node 0/RP0/CPU0) within the timeout period (100 seconds)
Jul 24 06:10:36  local7 1774: RP/0/RP0/CPU0:Jul 24 06:10:36.558 EDT: sysdb_shared_nc[323]: %SYSDB-SYSDB-3-EVENT_TIMEOUT : client 'intf_mgbl' (jid 1227, 0/RP0/CPU0) failed to pickup sysdb EDM even
t within 100 seconds.  Please collect 'show tech sysdb' for detailed information

Jul 24 06:10:47  local7 1776: RP/0/RP0/CPU0:Jul 24 06:10:47.826 EDT: sysmgr_control[66660]: %OS-SYSMGR-4-PROC_RESTART_NAME : User drew (vty0) requested a restart of process lpts_fm at 0/RP0/CPU0
^^^ me trying to recover it
Jul 24 06:12:16  local7 1778: RP/0/RP0/CPU0:Jul 24 06:12:16.637 EDT: sysdb_shared_nc[323]: %SYSDB-SYSDB-6-TIMEOUT_EDM : EDM request for 'oper/intf_mgbl/gl/full/' from 'show_interface' (jid 66628,
node 0/RP0/CPU0). No response from 'intf_mgbl' (jid 1227, node 0/RP0/CPU0) within the timeout period (100 seconds)
Jul 24 06:13:54  local7 1780: LC/0/0/CPU0:Jul 24 06:13:54.501 EDT: gsp[350]: %OS-gsp-6-MSG_GROUP_UNRESPONSIVE : group 2021 im_attr_owners is unresponsive for at least 10 seconds
Jul 24 06:13:54  local7 1782: RP/0/RP1/CPU0:Jul 24 06:13:54.612 EDT: gsp[339]: %OS-gsp-6-MSG_GROUP_UNRESPONSIVE : group 2021 im_attr_owners is unresponsive for at least 10 seconds

Jul 24 06:14:13  local7 1788: RP/0/RP0/CPU0:Jul 24 06:14:13.794 EDT: sysmgr_control[67183]: %OS-SYSMGR-4-PROC_RESTART_NAME : User drew (vty0) requested a restart of process lpts_pa at 0/RP0/CPU0
^^^ me trying to recover it

Jul 24 06:15:38  local7 1790: RP/0/RP0/CPU0:Jul 24 06:15:38.104 EDT: sysdb_shared_nc[323]: %SYSDB-SYSDB-6-TIMEOUT_EDM : EDM request for 'oper/intf_mgbl/gl/full/' from 'show_interface' (jid 67144,
node 0/RP0/CPU0). No response from 'intf_mgbl' (jid 1227, node 0/RP0/CPU0) within the timeout period (100 seconds)
Jul 24 06:17:18  local7 1792: RP/0/RP0/CPU0:Jul 24 06:17:18.244 EDT: sysdb_shared_nc[323]: %SYSDB-SYSDB-6-TIMEOUT_EDM : EDM request for 'oper/intf_mgbl/gl/full/' from 'show_interface' (jid 67437,
node 0/RP0/CPU0). No response from 'intf_mgbl' (jid 1227, node 0/RP0/CPU0) within the timeout period (100 seconds)
Jul 24 06:18:37  local7 1794: RP/0/RP0/CPU0:Jul 24 06:18:37.745 EDT: reload[67832]: %MGBL-SCONBKUP-6-INTERNAL_INFO : Reload debug script successfully spawned
Jul 24 06:18:54  local7 1796: LC/0/0/CPU0:Jul 24 06:18:54.616 EDT: gsp[350]: %OS-gsp-6-MSG_GROUP_UNRESPONSIVE : group 2021 im_attr_owners is unresponsive for at least 10 seconds

Jul 24 06:19:32  local7 1804: RP/0/RP1/CPU0:Jul 24 06:19:32.672 EDT: rmf_svr[200]: %HA-REDCON-4-FAILOVER_REQUESTED : failover has been requested by operator, waiting to initiate
^^^ me trying to recover it

RSP Failover seemed to just copy the deadlocked sysdb database (or whatever) to the other RSP.

After a full reload of everything in the router it seems to have stopped (for now).

On or around July 6th we began using the ancient check_ifstatus nagios plugin to poll interface status on this ASR9902 once every 5 minutes.

Based upon my previous experience with SNMP on this platform I would assign a 40% probability that the rather light SNMP polling caused this issue.

Anyway please let me know if you've seen this.
Take care,
-Drew




More information about the cisco-nsp mailing list