On layer 2 networks, high availability can be achieved by:
- antique protocols, like the Spanning Tree Protocol,
- proprietary protocols, like Juniper’s virtual chassis, Cisco’s virtual port channel and other MC-LAG implementations,1
- standardized protocols, like Shortest Path Bridging Protocol or TRILL, or
- the underlying network (in the case of an overlay network, like VXLAN).
In data center topologies, right cabling is a time-consuming endeavor and is error prone. Prescriptive Topology Manager (PTM) is a dynamic cabling verification tool to help detect and eliminate such errors. It takes a graphviz-DOT specified network cabling plan (something many operators already generate), stored in a topology.dot file, and couples it with runtime information derived from LLDP to verify that the cabling matches the specification. The check is performed on every link transition on each node in the network. It also detects forwarding path failures using Bidirectional Forwarding Detection (BFD).
You can customize the topology.dot file to control ptmd at both the global/network level and the node/port level.
PTM runs as a daemon, named ptmd.
tl;dr - Many people love 2-node clusters because they seem conceptually simpler and 33% cheaper, but while it’s possible to construct good ones, most will have subtle failure modes
The first step towards creating any HA system is to look for and try to eliminate single points of failure, often abbreviated as SPoF.
Simple yet Powerful Turnkey Solution to Build Clouds and Manage Data Center Virtualization
Failure is inevitable. As engineers building and maintaining complex systems, we likely encounter failure in some form on a daily basis. Not every failure requires a postmortem, but if a failure impacts the bottom line of the business, it becomes important to follow a postmortem process. I say “follow a postmortem process” instead of “do a postmortem”, because a postmortem should have very specific goals designed to prevent future failures in your environment. Simply asking the five whys to try and determine the root cause is not enough.
Building an operating system for data centers