From Lit to Dark: The Step-By-Step Migration Guide (With Downtime Math)

From Lit to Dark: The Step-By-Step Migration Guide (With Downtime Math)

Migrating from lit waves or Ethernet services to your own dark fiber sounds simple, until you’re juggling optics, optical budgets, L2/L3 the-day-of changes, and customer-facing downtime. The trick is to treat it like a controlled surgical swap: pre-provision everything, prove the light path before you touch routing, and model downtime in minutes (and dollars) so business owners sign off with eyes wide open.

Optics & Budget L2/L3 Cutover Downtime Math Risk & Rollback Acceptance Tests

Why Move from Lit to Dark?

Control & Scale
Pick your optics, ring topology, protection scheme, and upgrade path (10G ➜ 25/100G) on your schedule.
Economics
Capex up front; lower long-run $/Gbps. IRU/lease options vs monthly lit rates with escalators.
Reliability
Engineer route diversity and restoration SLAs you can enforce—no mystery middle boxes.

Downtime Cost & Maintenance-Window Estimator

Model the real impact so stakeholders approve a realistic window. Use conservative inputs; compare with your change-freeze policy.

Tip: If planned minutes exceed your monthly SLO budget, split the cutover into two maintenance windows with pre-provisioning between them.

Step-By-Step Migration Playbook

1

Assess & Inventory

Understand the current lit service and what it really carries.

  • Document all circuits, C-tags/S-tags, VLANs, LAGs, QinQ, and any provider MAC limits.
  • Record routing: IGP areas, BGP peers, LACP hashing, VRRP/HSRP roles, BFD timers.
  • Capture measured throughput, jitter, and current MTTR/availability.
  • Confirm handoff types (e.g., 10G LR, 1G LX) and optics at each side.
2

Design the Dark Path

Optical budget, topology, and protection strategy.

  • Choose topology (point-to-point, ring, diverse routes) and failure domains.
  • Select optics (10G LR/ER/ZR, 25G LR, 100G LR4/ER4, DWDM w/ mux/demux if needed).
  • Calculate optical budget (below)—include fiber attenuation, splices, patch panels, and safety margin.
  • Decide L2/L3 approach: keep same subnets and swing, or renumber during cutover.
3

Pre-Provision & Test in Parallel

Bring up optics and links out-of-band before traffic moves.

  • Turn up dark pair(s); verify light with OTDR and power meter at both ends.
  • Patch into non-production ports; configure LACP, native VLANs, or sub-interfaces.
  • Stand up passive BFD + BGP/OSPF neighbors in shutdown or on maintenance VRFs.
  • Run Ixia/iperf loss/jitter tests; record baseline.
4

Maintenance Window & Cutover

Shortest possible data-plane interruption; deterministic rollback.

  • Freeze changes elsewhere. Announce user-facing maintenance plus buffer.
  • Disable lit LAG member(s) to drain; verify routing convergence on the dark path.
  • Enable dark LAG/BGP/OSPF; verify adjacencies, MAC move, ARP/ND aging.
  • Smoke tests: latency, jitter, path trace, application health checks.
  • If KPIs not met in X minutes, execute rollback (below).
5

Stabilize & Decommission

Close the loop and retire the lit service.

  • Run 24–72h enhanced monitoring; raise BFD aggressiveness only after stability.
  • Document final optics power, OTDR traces, and inventory (labels/patch plans).
  • Decommission lit handoffs; remove configs and update diagrams.

Optical Budget & Link Margin — Quick Check

Confirm your chosen optics will work with margin. Use spec-sheet values for Tx power and Rx sensitivity.

Rule of thumb: Aim for ≥ 3 dB end-to-end margin after accounting for aging and temperature.

L2/L3 Cutover Patterns That Reduce Pain

If You’re Layer-2 Heavy

  • Mirror port channels on dark side; match LACP rate/hash.
  • Keep VLAN IDs and trunk/native exactly the same; pre-provision STP priorities.
  • Use MAC move thresholds in security policies to avoid false positives during swing.

If You’re Routing at the Edge

  • Stand up parallel BGP or OSPF neighbors on dark interfaces, shut until window.
  • Use maintenance metrics: OSPF cost increase on lit, then enable dark and revert cost.
  • Conservative BFD timers for day-1 (e.g., 300/300/3 ms later, start 500/500/5 or disabled).

Risk & Rollback Matrix

Risk Early Warning Mitigation Rollback Trigger
Optical power too low/high Rx near sensitivity / ORL alarms Swap optics, add pads, clean connectors; re-test OTDR Rx < spec for 5 min or BER spikes
LACP/MTU mismatch One-way traffic, slow hash, drops Pre-check MTU end-to-end; match LACP rate/hash Packet loss > 0.1% beyond 3 min
Routing flap/storm Adjacency churn, CPU spike Stagger bring-up; dampening; increase hello/dead temporarily >2 flaps in 5 min
Unexpected latency/jitter App health fails, p95 latency jump Pin traffic via policy route; verify path diversity SLO breach for 2+ critical apps

Acceptance & Documentation Checklist

  1. Light Path: OTDR trace stored; Rx/Tx power recorded at both ends.
  2. Performance: Throughput, latency, jitter, and loss within targets under load.
  3. Resiliency: Link failure test proves convergence within target (e.g., < 200 ms with LACP + BFD).
  4. Monitoring: Optics DOM, link state, BFD, and routing neighbors graphed; alerts configured.
  5. Docs: Patch panels labeled; diagrams and IPAM updated; change record closed.

Cutover Runbook — Task Timer

Estimate total hands-on time; keep tasks atomic. The window should include this time plus validation and rollback buffers.

Every environment is different. Validate optics against vendor specs, confirm optical power and OTDR traces on the actual span, and rehearse your runbook on a non-production pair if possible. Align the maintenance window with your SLO budget, set a strict rollback timer, and keep detailed acceptance records so audits and future upgrades are straightforward.