Concepts
Live migration and HA
How running servers move between hosts without stopping, how servers recover when a host fails, and the per-server policies that control both.
Servers keep their disks on network storage, so the host a server runs on is replaceable. Ankra Cloud uses that for two things: moving running servers off a host before maintenance (live migration) and restarting servers elsewhere when a host fails (high availability).
Live migration#
Before a host is serviced, Ankra drains it: each running server is migrated live to another host while it keeps
running. Memory and device state are copied over the storage network; the final switch pauses the guest briefly,
typically well under a second. The server keeps its addresses, disks and connections. While a migration runs, the
server's active_operation is a server.migrate, and start, stop and delete answer 409 until it finishes. A stopped
server moves cold (server.move), which you would not notice.
If a migration cannot finish, it is settled safely: the server keeps running on exactly one host, never two.
Migration policy#
PUT /v1/servers/{id}/migration-policy chooses how hard live migrations try to converge for a busy server:
| Policy | Behaviour |
|---|---|
default |
Up to 2 s of downtime over three attempts; may throttle the guest to converge. |
relaxed |
Prefers a finished move: up to 5 s over five attempts, with longer timeouts. |
strict |
Never disturbs the guest: at most 300 ms over two attempts, no throttling. If that is not enough the migration fails and the server stays where it is. |
High availability#
The control plane watches every compute host's heartbeat. When a host that runs servers stops reporting, it is first fenced: powered off and cut off from storage, so it can never write to your disks again. Only then are its servers moved to other hosts. Nothing is fenced while most of a zone looks silent at once, since that points at a control plane or network problem rather than failed hosts.
What happens to each server depends on its HA policy (PUT /v1/servers/{id}/ha-policy):
ha_policy |
After a host failure |
|---|---|
restart (default) |
A running server is started on another host (server.recover). |
none |
The server arrives on another host stopped (server.move); you start it when you are ready. |
A stopped server always arrives stopped. Recovery is a restart, not a live move: the guest boots again from its disks, so unsynced writes in the guest's memory are lost, as after a power cut.
Server groups (anti-affinity)#
To keep replicas of a service on different hosts, create a server group (POST /v1/server-groups) and create each
replica with its server_group_id. Members of a strict group never share a host; creating one more member than the
zone has hosts with room fails with 503. A soft group spreads members while it can and shares a host only when every
other host is full.
Managed services#
Load balancers run their two VMs on different hosts and fail over in about a second. Databases are single instances
restarted by the same HA mechanism. Network edges placed separate isolate each role on its own VM.