↓ Skip to main content
  1. Network Articles/
  2. MPLS-VPN Articles/

MPLS VPN PE-CE Routing with eBGP

Table of Contents

PE-CE Routing with eBGP in MPLS VPN

The PE and the CE run eBGP with each other, and the CE advertises the prefixes of its site. Section 7 of RFC 4364 lists it as one of the ways a PE learns routes from a CE.

On the PE it is just a neighbor under the VRF.

On the PE (IOS XR)
router bgp 65001
 vrf CUST-A
  rd 65001:1
  address-family ipv4 unicast
   redistribute connected
  !
  neighbor 172.16.2.2
   remote-as 65102
   address-family ipv4 unicast
    route-policy PASS-CE2 in
    route-policy PASS-CE2 out
    soft-reconfiguration inbound always
   !
  !
 !
!

Routes learned from the CE need no redistribute. Once they are in the VRF’s BGP table they travel to the other PEs as VPNv4. What redistribute connected is for is the PE-CE subnet, which the CE does not advertise and which the provider puts into the VPN itself.

IOS XR requires a route policy on an eBGP neighbor. Without one nothing is sent or accepted, so a plain PASS-* policy goes on each session. The names differ per neighbor, because one name shared by several eBGP neighbors merges them into a single update group and makes advertised-routes unreadable.

The CE is an ordinary eBGP router that knows nothing about MPLS or VRFs.

On the CE (IOS XR)
router bgp 65102
 bgp router-id 12.12.12.12
 address-family ipv4 unicast
  network 10.1.2.0/24
  network 10.1.100.0/24
 !
 neighbor 172.16.2.1
  remote-as 65001
  address-family ipv4 unicast
   route-policy PASS-PE2 in
   route-policy PASS-PE2 out
  !
 !
!

What Static Routing Cannot Do: Redundancy and Attributes

AspectStaticeBGP
Liveness detectionNone; the PE keeps advertising after the CE diesYes; the route follows the session
The same prefix from several sitesNot possible (the PE cannot choose)Possible; BGP best path selection picks one
SteeringOnly by changing the PEThe customer can signal it with an attribute
A new prefix at the siteThe PE has to be reconfiguredThe CE just advertises it

The third row matters most: the customer decides where the traffic goes, without asking the provider to change anything.

Steering from the Customer Side with MED

MED tells a neighbouring AS which entry point to use for traffic coming into your own AS. A lower value wins.

Advertise the same prefix from several sites and give the preferred one the lower MED, and the traffic follows.

CE2 advertises with a MED (IOS XR)
route-policy MED-200
  set med 200
  pass
end-policy
!
router bgp 65102
 neighbor 172.16.2.1
  address-family ipv4 unicast
   route-policy MED-200 out
  !
 !
!

MED is only compared between routes from the same neighbouring AS. The sites advertising the same prefix therefore have to share an AS number for the comparison to happen at all, which is also what customers normally do.

Using one AS number at several sites brings a problem of its own: the routes of one site are dropped as an AS loop at the others. That is covered in when several MPLS VPN sites share one AS number.

The Order of Best Path Selection

Which site wins is decided by BGP best path selection. These are the steps that matter here (the numbering follows the table in BGP path attributes and best path selection; steps 1 and 3 are WEIGHT and locally originated routes, which are equal in this lab):

StepCriterionIts role here
2LOCAL_PREFThe same (100) in every STEP
4AS_PATH lengthThe same (one AS, 65102) in every STEP
5ORIGINThe same (IGP) in every STEP
6MULTI_EXIT_DISC (MED)The value the customer moves
7eBGP versus iBGPThe eBGP route wins. It comes after MED, which matters later
8IGP metric to the NEXT_HOPThe decider when the MEDs are equal

MED at step 6 comes before the eBGP-over-iBGP rule at step 7, and that is what produces the behaviour seen later, where the losing PE stops advertising its own route.

Lab Setup

Seven XRd routers make three sites of customer A, and sites 2 and 3 both advertise the same 10.1.100.0/24.

NodeASWhat it advertises
CE16510110.1.1.0/24 (the observer)
CE26510210.1.2.0/24 plus 10.1.100.0/24
CE36510210.1.3.0/24 plus 10.1.100.0/24 (the same address 10.1.100.1 as CE2)

Only the core link toward PE3 carries an OSPF cost of 10, so a tie breaks toward PE2. Everything is observed with a traceroute from CE1: a last hop of 172.16.2.2 means CE2 and 172.16.3.2 means CE3.

STEPs 1 to 3 touch the CEs only. The PEs are never reconfigured.

STEPChangeWhat it shows
0Initial state (no MED)Two paths; the tie breaks on the IGP metric to the next hop, toward PE2
1CE2 sets MED 200 and CE3 sets MED 100Traffic moves to PE3
2CE2 changes to MED 50Traffic moves back to PE2
3Both CEs drop the MEDBack to the state of STEP 0
4Shut down both ends of the PE2-CE2 linkThe session drops and traffic fails over to PE3
5Restore (final state)Back on PE2

STEP 0: A Tie Breaks on the IGP Metric to the Next Hop

PE1 has both paths.

PE1: two paths for 10.1.100.0/24 (STEP 0)
RP/0/RP0/CPU0:PE1#show bgp vpnv4 unicast rd 65001:1 10.1.100.0/24
Sun Oct  4 05:40:43.568 UTC
BGP routing table entry for 10.1.100.0/24, Route Distinguisher: 65001:1
Versions:
  Process           bRIB/RIB   SendTblVer
  Speaker                 17           17
Last Modified: Oct  4 05:38:43.023 for 00:02:00
Paths: (2 available, best #1)
  Not advertised to any peer
  Path #1: Received by speaker 0
  Not advertised to any peer
  65102, (received & used)
    2.2.2.2 (metric 3) from 2.2.2.2 (2.2.2.2)
      Received Label 24006 
      Origin IGP, metric 0, localpref 100, valid, internal, best, group-best, import-candidate, imported
      Received Path ID 0, Local Path ID 1, version 17
      Extended community: RT:65001:100 
      Source AFI: VPNv4 Unicast, Source VRF: CUST-A, Source Route Distinguisher: 65001:1
  Path #2: Received by speaker 0
  Not advertised to any peer
  65102, (received & used)
    4.4.4.4 (metric 12) from 4.4.4.4 (4.4.4.4)
      Received Label 24006 
      Origin IGP, metric 0, localpref 100, valid, internal, import-candidate, imported
      Received Path ID 0, Local Path ID 0, version 0
      Extended community: RT:65001:100 
      Source AFI: VPNv4 Unicast, Source VRF: CUST-A, Source Route Distinguisher: 65001:1

LOCAL_PREF (100), AS_PATH (65102) and MED (metric 0) are all equal, so what decides is the IGP metric to the next hop: 3 for 2.2.2.2 against 12 for 4.4.4.4, and the PE2 path is best.

The traffic follows.

traceroute from CE1 to 10.1.100.1 (STEP 0)
RP/0/RP0/CPU0:CE1#traceroute 10.1.100.1 source 10.1.1.1 timeout 1 probe 2 maxttl 6
Sun Oct  4 05:39:52.396 UTC

Type escape sequence to abort.
Tracing the route to 10.1.100.1

 1  172.16.1.1 7 msec  4 msec 
 2  10.0.13.3 [MPLS: Labels 24002/24006 Exp 0] 14 msec  12 msec 
 3  10.0.23.2 [MPLS: Label 24006 Exp 0] 15 msec  13 msec 
 4  172.16.2.2 24 msec  * 

STEP 1 and 2: The MED on the CE Alone Moves the Traffic

CE2 gets set med 200 and CE3 gets set med 100. Only the CEs are touched.

Configuration committed on CE2 (STEP 1)
route-policy MED-200
  set med 200
  pass
end-policy
!
router bgp 65102
 neighbor 172.16.2.1
  address-family ipv4 unicast
   route-policy MED-200 out
  !
 !
!
end

The lower MED, on CE3, wins.

CE1 to 10.1.100.1 (STEP 1)
RP/0/RP0/CPU0:CE1#traceroute 10.1.100.1 source 10.1.1.1 timeout 1 probe 2 maxttl 6
Sun Oct  4 05:45:17.767 UTC

Type escape sequence to abort.
Tracing the route to 10.1.100.1

 1  172.16.1.1 6 msec  4 msec 
 2  10.0.13.3 [MPLS: Labels 24001/24006 Exp 0] 19 msec  12 msec 
 3  10.0.34.4 [MPLS: Label 24006 Exp 0] 13 msec  11 msec 
 4  172.16.3.2 21 msec  * 
CE1 to 10.1.100.1 (STEP 2)
RP/0/RP0/CPU0:CE1#traceroute 10.1.100.1 source 10.1.1.1 timeout 1 probe 2 maxttl 6
Sun Oct  4 05:50:39.791 UTC

Type escape sequence to abort.
Tracing the route to 10.1.100.1

 1  172.16.1.1 6 msec  4 msec 
 2  10.0.13.3 [MPLS: Labels 24002/24006 Exp 0] 15 msec  14 msec 
 3  10.0.23.2 [MPLS: Label 24006 Exp 0] 16 msec  12 msec 
 4  172.16.2.2 15 msec  * 

In STEP 2 CE2 lowers its own MED to 50 and the traffic comes back to site 2. Not one line changed on the provider side. With static routing this is impossible without reconfiguring the PE.

The Losing Path Disappears from PE1

Look at PE1 again in STEP 1.

PE1: 10.1.100.0/24 (STEP 1)
RP/0/RP0/CPU0:PE1#show bgp vpnv4 unicast rd 65001:1 10.1.100.0/24
Sun Oct  4 05:46:11.165 UTC
BGP routing table entry for 10.1.100.0/24, Route Distinguisher: 65001:1
Versions:
  Process           bRIB/RIB   SendTblVer
  Speaker                 23           23
Last Modified: Oct  4 05:43:45.023 for 00:02:26
Paths: (1 available, best #1)
  Not advertised to any peer
  Path #1: Received by speaker 0
  Not advertised to any peer
  65102, (received & used)
    4.4.4.4 (metric 12) from 4.4.4.4 (4.4.4.4)
      Received Label 24006 
      Origin IGP, metric 100, localpref 100, valid, internal, best, group-best, import-candidate, imported
      Received Path ID 0, Local Path ID 1, version 23
      Extended community: RT:65001:100 
      Source AFI: VPNv4 Unicast, Source VRF: CUST-A, Source Route Distinguisher: 65001:1

There is only one path left. STEP 0 had two, and adding a MED difference removed one. The reason is on PE2, the losing side.

PE2: 10.1.100.0/24 (STEP 1)
RP/0/RP0/CPU0:PE2#show bgp vrf CUST-A 10.1.100.0/24
Sun Oct  4 05:46:46.335 UTC
BGP routing table entry for 10.1.100.0/24, Route Distinguisher: 65001:1
Versions:
  Process           bRIB/RIB   SendTblVer
  Speaker                 25           25
Last Modified: Oct  4 05:43:45.023 for 00:03:01
Paths: (3 available, best #1)
  Not advertised to any peer
  Path #1: Received by speaker 0
  Not advertised to any peer
  65102, (received & used)
    4.4.4.4 (metric 12) from 4.4.4.4 (4.4.4.4)
      Received Label 24006 
      Origin IGP, metric 100, localpref 100, valid, internal, best, group-best, import-candidate, imported
      Received Path ID 0, Local Path ID 1, version 25
      Extended community: RT:65001:100 
      Source AFI: VPNv4 Unicast, Source VRF: CUST-A, Source Route Distinguisher: 65001:1
  Path #2: Received by speaker 0
  Not advertised to any peer
  65102
    172.16.2.2 from 172.16.2.2 (12.12.12.12)
      Origin IGP, metric 200, localpref 100, valid, external
      Received Path ID 0, Local Path ID 0, version 0
      Extended community: RT:65001:100 
      Origin-AS validity: (disabled)

PE2 holds both the route from CE2 with MED 200 and the route from PE3 with MED 100, and the one it chose as best is the one from PE3. As the table earlier showed, MED (step 6) is compared before eBGP-over-iBGP (step 7).

Since its own eBGP route is no longer best, PE2 does not advertise it to the other PEs, which the output states as Not advertised to any peer.

The capture shows the same thing.

BGP UPDATEs between PE1 and P1 (STEP 1)
$ tshark -r mpls-vpn-pece-bgp-step1-pe1p1.pcap -Y 'bgp.type==2' -T fields \
    -e frame.number -e ip.src \
    -e bgp.update.path_attribute.multi_exit_disc \
    -e bgp.mp_reach_nlri_ipv4_prefix -e bgp.mp_unreach_nlri_ipv4_prefix
11	2.2.2.2	200	10.1.2.0	10.1.100.0
24	4.4.4.4	100	10.1.3.0,10.1.100.0	
BGP UPDATEs between PE1 and P1 (STEP 2)
$ tshark -r mpls-vpn-pece-bgp-step2-pe1p1.pcap -Y 'bgp.type==2' -T fields \
    -e frame.number -e ip.src \
    -e bgp.update.path_attribute.multi_exit_disc \
    -e bgp.mp_reach_nlri_ipv4_prefix -e bgp.mp_unreach_nlri_ipv4_prefix
7	2.2.2.2	50	10.1.2.0,10.1.100.0	
8	4.4.4.4			10.1.100.0

In STEP 1, PE2 (2.2.2.2) advertises its own 10.1.2.0/24 with MED 200 while withdrawing 10.1.100.0/24. In STEP 2 it is PE3 (4.4.4.4) that withdraws 10.1.100.0/24 instead. Whichever side loses withdraws.

The withdrawal is No.11 of the STEP 1 capture.

PE2 to PE1, an UPDATE (tshark -V, the MP_UNREACH_NLRI part)
Border Gateway Protocol - UPDATE Message
    Marker: ffffffffffffffffffffffffffffffff
    Length: 45
    Type: UPDATE Message (2)
    Withdrawn Routes Length: 0
    Total Path Attribute Length: 22
    Path attributes
        Path Attribute - MP_UNREACH_NLRI
            Flags: 0x90, Optional, Extended-Length, Non-transitive, Complete
                1... .... = Optional: Set
                .0.. .... = Transitive: Not set
                ..0. .... = Partial: Not set
                ...1 .... = Extended-Length: Set
                .... 0000 = Unused: 0x0
            Type Code: MP_UNREACH_NLRI (15)
            Length: 18
            Address family identifier (AFI): IPv4 (1)
            Subsequent address family identifier (SAFI): Labeled VPN Unicast (128)
            Withdrawn Routes
                BGP Prefix
                    Prefix Length: 112
                    Label Stack: 0 (withdrawn)
                    Route Distinguisher: 65001:1
                    MP Unreach NLRI IPv4 prefix: 10.1.100.0
Download the pcap of the packet in the tshark output above (No.11 UPDATE, PE2 withdrawing)

Best External and ADD-PATH Are What Keep the Backup

The behaviour follows the specification, but it is awkward in operation: PE1 knows no backup path at all, so if site 3 fails nothing can switch over until PE2 re-runs best path selection and advertises its own route again.

That is what BGP Best External, which advertises a non-best external route, and BGP ADD-PATH, which carries several paths for one prefix, are for. Their use with VPNv4 is covered in VPNv4 Best External.

STEP 4: The Session Drops and Traffic Fails Over

Both ends of the PE2-CE2 link go down (shutting one end leaves the interface up at the other, so a link failure is made at both ends).

PE1: 10.1.100.0/24 (STEP 4)
RP/0/RP0/CPU0:PE1#show bgp vpnv4 unicast rd 65001:1 10.1.100.0/24
Sun Oct  4 06:02:47.965 UTC
BGP routing table entry for 10.1.100.0/24, Route Distinguisher: 65001:1
Versions:
  Process           bRIB/RIB   SendTblVer
  Speaker                 34           34
Last Modified: Oct  4 05:59:53.023 for 00:02:55
Paths: (1 available, best #1)
  Not advertised to any peer
  Path #1: Received by speaker 0
  Not advertised to any peer
  65102, (received & used)
    4.4.4.4 (metric 12) from 4.4.4.4 (4.4.4.4)
      Received Label 24006 
      Origin IGP, metric 0, localpref 100, valid, internal, best, group-best, import-candidate, imported
      Received Path ID 0, Local Path ID 1, version 34
      Extended community: RT:65001:100 
      Source AFI: VPNv4 Unicast, Source VRF: CUST-A, Source Route Distinguisher: 65001:1

The path through PE2 is gone and the remaining one, through site 3, is best. A ping run after convergence gets every packet through, so traffic has moved to site 3 and keeps flowing. Note that this ping ran about 100 seconds after the link went down, so it does not measure whether packets are lost during the switchover.

ping from CE1 to 10.1.100.1 (STEP 4)
RP/0/RP0/CPU0:CE1#ping 10.1.100.1 source 10.1.1.1 count 10 timeout 1
Sun Oct  4 06:01:35.212 UTC
Type escape sequence to abort.
Sending 10, 100-byte ICMP Echos to 10.1.100.1 timeout is 1 seconds:
!!!!!!!!!!
Success rate is 100 percent (10/10), round-trip min/avg/max = 12/17/36 ms

With static routing the PE kept advertising after the CE died and the packets were simply dropped. With eBGP the route follows the session and the other site takes over. That is what having a redundant design means.

STEP 5: Restoring

Bringing the link back makes the PE2 path, closer in IGP metric, best again, and the state returns to that of STEP 0.

References

RFCTitleSections used
RFC 4364BGP/MPLS IP Virtual Private Networks (VPNs)7 (how PEs learn routes from CEs)
RFC 4271A Border Gateway Protocol 4 (BGP-4)9.1.2.2 (the order of best path selection), 5.1.4 (MULTI_EXIT_DISC)

Book: Luc De Ghein, MPLS Fundamentals (Cisco Press, 2006), Chapter 7

Verification Configs and show Output

Every STEP was captured on all seven routers, one file per router and kind. The verification config is the ..._run.txt file (the final state is the one from the last STEP).

FileContents
..._show.txtThe full state at that STEP: OSPF (show ospf neighbor, show ospf database), LDP (show mpls ldp neighbor brief, show mpls ldp bindings), MPLS forwarding (show mpls forwarding, show mpls label table), MP-BGP (show bgp vpnv4 unicast, show bgp vpnv4 unicast rd 65001:1 10.1.100.0/24, show bgp vpnv4 unicast labels), the VRF (show vrf all detail, show route vrf CUST-A, show cef vrf CUST-A) and the PE-CE eBGP session (the advertised-routes / routes / received routes set). 50 commands on a PE, 20 on a P, 16 on a CE
..._log.txtshow logging limited to that STEP
..._run.txtshow running-config at that STEP (the verification config)
..._ping.txtping and traceroute between the CEs
..._trace.txtshow bgp trace on the PEs
..._commit.cfgThe configuration actually committed in that STEP (only for the routers that changed)

STEP 0: Initial state (no MED)

Routershowsyslogrunning-configpingtracecommitted config
CE1showlogrunping--
PE1showlogrunpingtrace-
P1showlogrunping--
PE2showlogrunpingtrace-
CE2showlogrunping--
PE3showlogrunpingtrace-
CE3showlogrunping--

STEP 1: CE2 sets MED 200 and CE3 sets MED 100

Routershowsyslogrunning-configpingtracecommitted config
CE1showlogrunping--
PE1showlogrunpingtrace-
P1showlogrunping--
PE2showlogrunpingtrace-
CE2showlogrunping-commit
PE3showlogrunpingtrace-
CE3showlogrunping-commit

STEP 2: CE2 changes to MED 50

Routershowsyslogrunning-configpingtracecommitted config
CE1showlogrunping--
PE1showlogrunpingtrace-
P1showlogrunping--
PE2showlogrunpingtrace-
CE2showlogrunping-commit
PE3showlogrunpingtrace-
CE3showlogrunping--

STEP 3: Both CEs drop the MED

Routershowsyslogrunning-configpingtracecommitted config
CE1showlogrunping--
PE1showlogrunpingtrace-
P1showlogrunping--
PE2showlogrunpingtrace-
CE2showlogrunping-commit
PE3showlogrunpingtrace-
CE3showlogrunping-commit

STEP 4: Shut down both ends of the PE2-CE2 link

Routershowsyslogrunning-configpingtracecommitted config
CE1showlogrunping--
PE1showlogrunpingtrace-
P1showlogrunping--
PE2showlogrunpingtracecommit
CE2showlogrunping-commit
PE3showlogrunpingtrace-
CE3showlogrunping--

STEP 5: Restore (final state)

Routershowsyslogrunning-configpingtracecommitted config
CE1showlogrunping--
PE1showlogrunpingtrace-
P1showlogrunping--
PE2showlogrunpingtracecommit
CE2showlogrunping-commit
PE3showlogrunpingtrace-
CE3showlogrunping--

Packet captures were taken per STEP.

STEPPE1-P1
0pcap
1pcap
2pcap
3pcap
4pcap
5pcap

Related Articles