IS-IS Link Failure and Path Switchover
A link-state routing protocol keeps traffic flowing when one path fails, as long as another one exists. In IS-IS the sequence is:
- Notice the failure
- The router that noticed rebuilds its own LSP and floods it
- Every router that receives the LSP runs SPF again
- The new result lands in the forwarding table (FIB)
Step 1, noticing, takes by far the most time. IS-IS has three ways of noticing, and which one fires decides whether traffic stops for 0.14 seconds or 28 seconds. This article changes only how the failure is caused, in one fixed topology, and measures the difference on IOS XR.
| How it is noticed | Trigger | Speed |
|---|---|---|
| Interface down | The router’s own port loses the link | Fastest |
| An LSP arrives | The LSP from the router that noticed comes round the other path | Fast |
| The holding time expires | The neighbour’s IIHs stop arriving and the timer runs out | Slowest (30 s by default) |
How the holding time itself is derived is covered in IS-IS Hello and Holding Time, and the SPF and LSP generation delays in IS-IS Convergence Timers.
Verification on real devices
Lab environment
Six Cisco IOS XR routers (XRd 26.1.1) with two disjoint paths between R1 and R6. The upper one (R1-R2-R3-R6) costs 30 in total and the lower one (R1-R4-R5-R6) costs 60, so traffic normally takes the upper path.
When the upper path becomes unusable the intermediate hops swap wholesale from R2 and R3 to R4 and R5,
which makes the switchover obvious in a traceroute. Every router runs level 2 only in area 49.0001
and every link is point-to-point.
Measurements are taken with a ping already running before the failure:
ping 6.6.6.6 source 1.1.1.1 count 600 interval 100 timeout 1 from both R1 and R6. The number of lost
packets gives the outage: a probe that gets no answer waits the full timeout 1 second before the next one,
so the count of lost packets is effectively the number of seconds.
The STEPs
| STEP | Change | What it shows |
|---|---|---|
| 0 | Initial state | That the upper path is used |
| 1 | Stop the link in CML | What happens when the link “breaks” |
| 2 | Restore the link | What happens on the way back |
| 3 | shutdown at both ends | When both routers know about the fault themselves |
| 4 | Remove the shutdown | Recovery |
| 5 | shutdown at one end only | When only the shut side knows |
| 6 | Remove the shutdown | Recovery |
| 7 | Stop the R2 node | When a whole device dies |
| 8 | Start R2 (final state) | Recovery |
The normal path (STEP 0)
R1 reaches R6 over the upper path at a metric of 30.
RP/0/RP0/CPU0:R1#traceroute 6.6.6.6 source 1.1.1.1 timeout 1 probe 2 maxttl 6
Type escape sequence to abort.
Tracing the route to 6.6.6.6
1 10.0.12.2 10 msec 5 msec
2 10.0.23.3 11 msec 8 msec
3 10.0.36.6 23 msec *
RP/0/RP0/CPU0:R1#show route isis
i L2 6.6.6.6/32 [115/30] via 10.0.12.2, 00:01:50, GigabitEthernet0/0/0/0Trigger ①: the interface goes down (STEP 3)
With shutdown on both R1 and R2, both routers learn about it immediately as a fault on their own port.
The syslog reason is Interface state down.
RP/0/RP0/CPU0:Sep 23 22:28:13.138 UTC: isis[1003]: %ROUTING-ISIS-5-ADJCHANGE : ISIS (1): Adjacency to R2 (GigabitEthernet0/0/0/0) (L2) Down, Interface state down Only one of the 600 pings was lost. Measured from the capture timestamps, the gap between the last packet on the upper path and the first one on the lower path is 0.14 seconds.
Type escape sequence to abort.
Sending 600, 100-byte ICMP Echos to 6.6.6.6 timeout is 1 seconds:
Success rate is 99 percent (599/600), round-trip min/avg/max = 8/14/174 msThe route moves to the lower path (metric 60).
i L2 6.6.6.6/32 [115/60] via 10.0.14.4, 00:01:24, GigabitEthernet0/0/0/1Trigger ②: an LSP arrives (STEP 5)
With shutdown on the R2 side only, R2 is the only one that knows. R1’s interface stays up.
Even so, only one packet was lost (1.00 second rather than 0.14). R1 did not wait for the holding time.
The reason is that the LSP R2 rebuilt comes round the other path to R1. R2 sends it to R3, and it travels on through R6, R5 and R4 to R1. R1 runs SPF on that LSP and learns that the upper path is gone.
R1’s adjacency is still up at that point. It only goes down once the holding time expires, which the timestamps make clear.
| Time (UTC) | Event |
|---|---|
| 22:35:56.437 | R2 drops the adjacency with Interface state down (the side that was shut) |
| 22:35:57.45 | The first packet appears on the lower path (R1 has switched its forwarding) |
| 22:36:23.728 | R1 drops the adjacency with Holdtime expired (26.3 seconds later) |
RP/0/RP0/CPU0:Sep 23 22:36:23.728 UTC: isis[1003]: %ROUTING-ISIS-5-ADJCHANGE : ISIS (1): Adjacency to R2 (GigabitEthernet0/0/0/0) (L2) Down, Holdtime expired Dropping an adjacency and switching a path are two different things. Forwarding switches as soon as the LSP arrives. This works because a detour exists; where two routers are joined by a single link there is no path for the LSP either, so the holding time has to expire.
Trigger ③: the holding time expires (STEP 1 and 7)
A link can stop carrying frames without going down — one direction of the optics is dark, or a switch in
between silently discards. The router’s interface stays up, so nothing is noticed until the IIHs stop
arriving and the holding time expires.
Stopping the link in CML produces exactly that state. The XRd interface stays up and the syslog reason
is Holdtime expired.
RP/0/RP0/CPU0:Sep 23 22:20:29.445 UTC: isis[1003]: %ROUTING-ISIS-5-ADJCHANGE : ISIS (1): Adjacency to R2 (GigabitEthernet0/0/0/0) (L2) Down, Holdtime expired 28 packets were lost, an outage of about 28 seconds: the default holding time of 30 seconds minus the time already elapsed since the last IIH.
Type escape sequence to abort.
Sending 600, 100-byte ICMP Echos to 6.6.6.6 timeout is 1 seconds:
!!!!!!!!!!!!!!!!!!............................!!!!!!!!!!!!!!!!!!!!!!!!
Success rate is 95 percent (572/600), round-trip min/avg/max = 9/13/66 msStopping the whole R2 node in STEP 7 lost the same 28 packets. A device disappearing is not signalled as a link failure either, so both neighbours, R1 and R3, fall back on the holding time.
The difference the failure mode makes
| STEP | How the failure was caused | Lost packets | Outage | Trigger |
|---|---|---|---|---|
| 3 | shutdown at both ends | 1 / 600 | 0.14 s | Interface state down |
| 5 | shutdown at one end | 1 / 600 | 1.00 s | The peer’s LSP came round the detour |
| 1 | Link stopped in CML | 28 / 600 | about 28 s | Holdtime expired |
| 7 | R2 node stopped | 28 / 600 | about 28 s | Holdtime expired |
The values were the same in both directions, R1 to R6 and R6 to R1. All three recovery operations (STEP 2, 4 and 6) delivered all 600 packets: nothing is lost when traffic moves back to the upper path.
Where an outage of tens of seconds is unacceptable, detection must not rely on the holding time. Shortening the hello interval is covered in IS-IS Hello and Holding Time, and sub-second detection with BFD in IS-IS and BFD.
Design notes
- Some failures never bring the link down. While the interface stays
upand only the frames are lost, IS-IS cannot notice until the holding time expires — close to 30 seconds by default - In a lab, “unplugging the cable” may not be what actually happens. Stopping a link in CML leaves the
XRd interface
up, so the test really measures the holding time path - An adjacency being up does not mean that path is in use. In STEP 5 the adjacency stayed up while forwarding had already moved to the detour
Verification configs and show output
Everything below was collected from all six routers at every STEP. The verification config is the
..._run.txt file (the final state is the one from the last STEP).
| File | Contents |
|---|---|
..._show.txt | show version / show route isis / show isis interface / show isis neighbors detail / show isis topology / show isis database detail / show isis spf-log / show isis lsp-log / show isis adjacency-log and more |
..._ping.txt | The 600-packet ping from R1 and R6 (already running before the failure) |
..._clear.txt | A record of the interface counters cleared for that STEP |
..._log.txt | show logging limited to the range of that STEP |
..._run.txt | show running-config at that STEP (the verification config for that STEP) |
..._commit.cfg | Only the configuration actually committed in that STEP |
..._trace.txt | show isis trace. The trace buffer accumulates from boot, so the STEP 8 file covers the whole test (only the six STEP 8 files are attached) |
STEP 0: Initial state
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | — | log | run |
| R2 | show | — | clear | — | log | run |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 1: Stop the link in CML
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | — | log | run |
| R2 | show | — | clear | — | log | run |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 2: Restore the link
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | — | log | run |
| R2 | show | — | clear | — | log | run |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 3: shutdown at both ends
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | commit | log | run |
| R2 | show | — | clear | commit | log | run |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 4: Remove the shutdown
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | commit | log | run |
| R2 | show | — | clear | commit | log | run |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 5: shutdown at one end only
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | — | log | run |
| R2 | show | — | clear | commit | log | run |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 6: Remove the shutdown
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | — | log | run |
| R2 | show | — | clear | commit | log | run |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 7: Stop the R2 node
| Router | show output | ping | clear | committed config | syslog | running-config |
|---|---|---|---|---|---|---|
| R1 | show | ping | clear | — | log | run |
| R2 | — | — | — | — | — | — |
| R3 | show | — | clear | — | log | run |
| R4 | show | — | clear | — | log | run |
| R5 | show | — | clear | — | log | run |
| R6 | show | ping | clear | — | log | run |
STEP 8: Start R2 (final state)
| Router | show output | ping | clear | committed config | syslog | running-config | trace |
|---|---|---|---|---|---|---|---|
| R1 | show | ping | clear | — | log | run | trace |
| R2 | show | — | clear | — | log | run | trace |
| R3 | show | — | clear | — | log | run | trace |
| R4 | show | — | clear | — | log | run | trace |
| R5 | show | — | clear | — | log | run | trace |
| R6 | show | ping | clear | — | log | run | trace |
Packet captures were taken per STEP (R1-R2 is the link that was stopped, so nothing could be retrieved for STEP 1 and 7).
| STEP | R1-R2 | R1-R4 |
|---|---|---|
| 0 | pcap | pcap |
| 1 | — | pcap |
| 2 | — | pcap |
| 3 | pcap | pcap |
| 4 | pcap | pcap |
| 5 | pcap | pcap |
| 6 | pcap | pcap |
| 7 | — | pcap |
| 8 | — | pcap |
References
| Standard | Title | Summary |
|---|---|---|
| ISO/IEC 10589:2002 (2nd edition) | Intermediate System to Intermediate System intra-domain routeing information exchange protocol | Defines how an adjacency is held (holdingTimer) and how an LSP is rebuilt when one goes down. This article draws on section 8.2 (maintaining adjacencies) and section 7.3 (LSP generation and flooding). |
| RFC 5303 | Three-Way Handshake for IS-IS Point-to-Point Adjacencies | The three-way handshake on point-to-point links, which detects the state where only one side has failed. |
Related articles
- What is IS-IS
- IS-IS NSAP Addresses and the NET (System ID)
- IS-IS Multiple Area Addresses (Multihoming), Merging and Splitting Areas
- IS-IS Level 1 and Level 2 (the hierarchy)
- IS-IS Packet Types and Header Format
- IS-IS Adjacency Formation and States
- IS-IS DIS and the Pseudonode
- IS-IS Network Types (broadcast / point-to-point)
- IS-IS Metrics (narrow and wide)
- IS-IS Authentication (hello-password and lsp-password)
- IS-IS LSPs and the Link-State Database
- The Main IS-IS TLVs
- IS-IS Flooding and LSDB Synchronisation
- IS-IS SPF Computation and Route Selection
- IS-IS ECMP (Equal-Cost Multipath)
- The IS-IS ATT Bit and the Level 1 Default Route
- IS-IS Route Leaking and the Up/Down Bit
- IS-IS Route Summarization
- The IS-IS Overload Bit
- IS-IS IPv6 Support and Multi-Topology (TLV 236 / MT ID 2)
- IS-IS Hello and Holding Time
- IS-IS Convergence Timers (SPF / LSP Generation)
- IS-IS Flooding Timers (LSP Interval, Retransmission, CSNP / PSNP)
- IS-IS Link Failure and Path Switchover
- IS-IS Redistribution (connected / static)
- IS-IS Redistribution of BGP Routes