↓ Skip to main content
  1. Network Articles/
  2. IS-IS Articles/

IS-IS Link Failure and Path Switchover

Table of Contents

IS-IS Link Failure and Path Switchover

A link-state routing protocol keeps traffic flowing when one path fails, as long as another one exists. In IS-IS the sequence is:

  1. Notice the failure
  2. The router that noticed rebuilds its own LSP and floods it
  3. Every router that receives the LSP runs SPF again
  4. The new result lands in the forwarding table (FIB)

Step 1, noticing, takes by far the most time. IS-IS has three ways of noticing, and which one fires decides whether traffic stops for 0.14 seconds or 28 seconds. This article changes only how the failure is caused, in one fixed topology, and measures the difference on IOS XR.

How it is noticedTriggerSpeed
Interface downThe router’s own port loses the linkFastest
An LSP arrivesThe LSP from the router that noticed comes round the other pathFast
The holding time expiresThe neighbour’s IIHs stop arriving and the timer runs outSlowest (30 s by default)

How the holding time itself is derived is covered in IS-IS Hello and Holding Time, and the SPF and LSP generation delays in IS-IS Convergence Timers.

Verification on real devices

Lab environment

Lab: two disjoint paths between R1 and R6

Six Cisco IOS XR routers (XRd 26.1.1) with two disjoint paths between R1 and R6. The upper one (R1-R2-R3-R6) costs 30 in total and the lower one (R1-R4-R5-R6) costs 60, so traffic normally takes the upper path.

When the upper path becomes unusable the intermediate hops swap wholesale from R2 and R3 to R4 and R5, which makes the switchover obvious in a traceroute. Every router runs level 2 only in area 49.0001 and every link is point-to-point.

Measurements are taken with a ping already running before the failure: ping 6.6.6.6 source 1.1.1.1 count 600 interval 100 timeout 1 from both R1 and R6. The number of lost packets gives the outage: a probe that gets no answer waits the full timeout 1 second before the next one, so the count of lost packets is effectively the number of seconds.

The STEPs

STEPChangeWhat it shows
0Initial stateThat the upper path is used
1Stop the link in CMLWhat happens when the link “breaks”
2Restore the linkWhat happens on the way back
3shutdown at both endsWhen both routers know about the fault themselves
4Remove the shutdownRecovery
5shutdown at one end onlyWhen only the shut side knows
6Remove the shutdownRecovery
7Stop the R2 nodeWhen a whole device dies
8Start R2 (final state)Recovery

The normal path (STEP 0)

R1 reaches R6 over the upper path at a metric of 30.

STEP 0: traceroute and route on R1
RP/0/RP0/CPU0:R1#traceroute 6.6.6.6 source 1.1.1.1 timeout 1 probe 2 maxttl 6

Type escape sequence to abort.
Tracing the route to 6.6.6.6

 1  10.0.12.2 10 msec  5 msec 
 2  10.0.23.3 11 msec  8 msec 
 3  10.0.36.6 23 msec  * 
RP/0/RP0/CPU0:R1#show route isis
i L2 6.6.6.6/32 [115/30] via 10.0.12.2, 00:01:50, GigabitEthernet0/0/0/0

Trigger ①: the interface goes down (STEP 3)

With shutdown on both R1 and R2, both routers learn about it immediately as a fault on their own port. The syslog reason is Interface state down.

STEP 3: syslog on R1
RP/0/RP0/CPU0:Sep 23 22:28:13.138 UTC: isis[1003]: %ROUTING-ISIS-5-ADJCHANGE : ISIS (1): Adjacency to R2 (GigabitEthernet0/0/0/0) (L2) Down, Interface state down 

Only one of the 600 pings was lost. Measured from the capture timestamps, the gap between the last packet on the upper path and the first one on the lower path is 0.14 seconds.

STEP 3: ping from R1 to R6 (600 packets)
Type escape sequence to abort.
Sending 600, 100-byte ICMP Echos to 6.6.6.6 timeout is 1 seconds:
Success rate is 99 percent (599/600), round-trip min/avg/max = 8/14/174 ms

The route moves to the lower path (metric 60).

STEP 3: the route on R1
i L2 6.6.6.6/32 [115/60] via 10.0.14.4, 00:01:24, GigabitEthernet0/0/0/1

Trigger ②: an LSP arrives (STEP 5)

With shutdown on the R2 side only, R2 is the only one that knows. R1’s interface stays up.

Even so, only one packet was lost (1.00 second rather than 0.14). R1 did not wait for the holding time.

The reason is that the LSP R2 rebuilt comes round the other path to R1. R2 sends it to R3, and it travels on through R6, R5 and R4 to R1. R1 runs SPF on that LSP and learns that the upper path is gone.

R1’s adjacency is still up at that point. It only goes down once the holding time expires, which the timestamps make clear.

Time (UTC)Event
22:35:56.437R2 drops the adjacency with Interface state down (the side that was shut)
22:35:57.45The first packet appears on the lower path (R1 has switched its forwarding)
22:36:23.728R1 drops the adjacency with Holdtime expired (26.3 seconds later)
STEP 5: syslog on R1 (the adjacency only drops 26 seconds later)
RP/0/RP0/CPU0:Sep 23 22:36:23.728 UTC: isis[1003]: %ROUTING-ISIS-5-ADJCHANGE : ISIS (1): Adjacency to R2 (GigabitEthernet0/0/0/0) (L2) Down, Holdtime expired 

Dropping an adjacency and switching a path are two different things. Forwarding switches as soon as the LSP arrives. This works because a detour exists; where two routers are joined by a single link there is no path for the LSP either, so the holding time has to expire.

Trigger ③: the holding time expires (STEP 1 and 7)

A link can stop carrying frames without going down — one direction of the optics is dark, or a switch in between silently discards. The router’s interface stays up, so nothing is noticed until the IIHs stop arriving and the holding time expires.

Stopping the link in CML produces exactly that state. The XRd interface stays up and the syslog reason is Holdtime expired.

STEP 1: syslog on R1 (the link stopped in CML)
RP/0/RP0/CPU0:Sep 23 22:20:29.445 UTC: isis[1003]: %ROUTING-ISIS-5-ADJCHANGE : ISIS (1): Adjacency to R2 (GigabitEthernet0/0/0/0) (L2) Down, Holdtime expired 

28 packets were lost, an outage of about 28 seconds: the default holding time of 30 seconds minus the time already elapsed since the last IIH.

STEP 1: ping from R1 to R6 (28 of 600 lost)
Type escape sequence to abort.
Sending 600, 100-byte ICMP Echos to 6.6.6.6 timeout is 1 seconds:
!!!!!!!!!!!!!!!!!!............................!!!!!!!!!!!!!!!!!!!!!!!!
Success rate is 95 percent (572/600), round-trip min/avg/max = 9/13/66 ms

Stopping the whole R2 node in STEP 7 lost the same 28 packets. A device disappearing is not signalled as a link failure either, so both neighbours, R1 and R3, fall back on the holding time.

The difference the failure mode makes

How each failure is noticed and how long the outage lasts
STEPHow the failure was causedLost packetsOutageTrigger
3shutdown at both ends1 / 6000.14 sInterface state down
5shutdown at one end1 / 6001.00 sThe peer’s LSP came round the detour
1Link stopped in CML28 / 600about 28 sHoldtime expired
7R2 node stopped28 / 600about 28 sHoldtime expired

The values were the same in both directions, R1 to R6 and R6 to R1. All three recovery operations (STEP 2, 4 and 6) delivered all 600 packets: nothing is lost when traffic moves back to the upper path.

Where an outage of tens of seconds is unacceptable, detection must not rely on the holding time. Shortening the hello interval is covered in IS-IS Hello and Holding Time, and sub-second detection with BFD in IS-IS and BFD.

Design notes

  • Some failures never bring the link down. While the interface stays up and only the frames are lost, IS-IS cannot notice until the holding time expires — close to 30 seconds by default
  • In a lab, “unplugging the cable” may not be what actually happens. Stopping a link in CML leaves the XRd interface up, so the test really measures the holding time path
  • An adjacency being up does not mean that path is in use. In STEP 5 the adjacency stayed up while forwarding had already moved to the detour

Verification configs and show output

Everything below was collected from all six routers at every STEP. The verification config is the ..._run.txt file (the final state is the one from the last STEP).

FileContents
..._show.txtshow version / show route isis / show isis interface / show isis neighbors detail / show isis topology / show isis database detail / show isis spf-log / show isis lsp-log / show isis adjacency-log and more
..._ping.txtThe 600-packet ping from R1 and R6 (already running before the failure)
..._clear.txtA record of the interface counters cleared for that STEP
..._log.txtshow logging limited to the range of that STEP
..._run.txtshow running-config at that STEP (the verification config for that STEP)
..._commit.cfgOnly the configuration actually committed in that STEP
..._trace.txtshow isis trace. The trace buffer accumulates from boot, so the STEP 8 file covers the whole test (only the six STEP 8 files are attached)

STEP 0: Initial state

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclear—logrun
R2show—clear—logrun
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 1: Stop the link in CML

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclear—logrun
R2show—clear—logrun
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 2: Restore the link

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclear—logrun
R2show—clear—logrun
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 3: shutdown at both ends

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclearcommitlogrun
R2show—clearcommitlogrun
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 4: Remove the shutdown

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclearcommitlogrun
R2show—clearcommitlogrun
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 5: shutdown at one end only

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclear—logrun
R2show—clearcommitlogrun
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 6: Remove the shutdown

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclear—logrun
R2show—clearcommitlogrun
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 7: Stop the R2 node

Routershow outputpingclearcommitted configsyslogrunning-config
R1showpingclear—logrun
R2——————
R3show—clear—logrun
R4show—clear—logrun
R5show—clear—logrun
R6showpingclear—logrun

STEP 8: Start R2 (final state)

Routershow outputpingclearcommitted configsyslogrunning-configtrace
R1showpingclear—logruntrace
R2show—clear—logruntrace
R3show—clear—logruntrace
R4show—clear—logruntrace
R5show—clear—logruntrace
R6showpingclear—logruntrace

Packet captures were taken per STEP (R1-R2 is the link that was stopped, so nothing could be retrieved for STEP 1 and 7).

STEPR1-R2R1-R4
0pcappcap
1—pcap
2—pcap
3pcappcap
4pcappcap
5pcappcap
6pcappcap
7—pcap
8—pcap

References

StandardTitleSummary
ISO/IEC 10589:2002 (2nd edition)Intermediate System to Intermediate System intra-domain routeing information exchange protocolDefines how an adjacency is held (holdingTimer) and how an LSP is rebuilt when one goes down. This article draws on section 8.2 (maintaining adjacencies) and section 7.3 (LSP generation and flooding).
RFC 5303Three-Way Handshake for IS-IS Point-to-Point AdjacenciesThe three-way handshake on point-to-point links, which detects the state where only one side has failed.

Related articles