Databases
Ordering one, what each arrangement survives, and why replication is not a backup.
A managed database is ordered separately from any application, because it outlives them. You will redeploy an application a hundred times against the same database.
The platform creates it, keeps it running, patches it, and hands you credentials once. What is inside it is yours.
Choosing an arrangement
Three, and each adds exactly one capability over the one before.
Single. One copy. The cheapest thing that is still a database. No failover: if the machine underneath it goes, the database is unreachable until it comes back. Correct for development and staging, where an outage is an inconvenience.
Single with pooling. The same one copy, with connection pooling in front. It survives a flood of clients but not the loss of its machine. Choose it when your application opens many short-lived connections — which most web frameworks do — and you are not yet paying for availability.
Highly available. Several copies, one of which is the primary and accepts writes while the others follow it. If the primary fails, another takes over and your application reconnects to the same address. This is the only arrangement that survives losing a machine.
Why the number of copies is odd
Three or five, never two or four. Deciding which copy is in charge needs a majority to agree. An even number can split down the middle and agree on nobody — which makes two copies less available than one, because either failure leaves no majority and the database stops accepting writes to avoid corrupting itself.
Connecting
You are given one address. Use it and nothing else. It stays correct when the primary moves, which is the entire point: your application should never know or care which copy is in charge.
Credentials are shown once, when the database is created or when a user is added. They are not retrievable afterwards. If you lose them, create new ones.
Backups are not replication
This deserves its own heading because it is the most expensive misunderstanding in this document.
Replication copies every write to every copy, faithfully and immediately.
That includes the write you did not mean. A dropped table, a DELETE with a
mistaken condition, a migration that ran against the wrong database — all of it
is replicated perfectly to every copy within moments.
Replication protects you from hardware failing. Backups protect you from people and software being wrong. They are different problems and they need different machinery. You need both.
Backups are not on by default. Turn them on, choose the hour, and check once that a restore works before you need it to.
Edge cases worth knowing before you meet them
The database says it is running and your application cannot connect. An instance is marked running as soon as its parts are created, slightly before one of the copies has been elected primary. For the first half-minute or so this is optimistic. Retry.
Connections are refused under load. A database accepts a fixed number of connections. An application with many copies, each opening its own pool, will exhaust them long before the database is actually busy. Use the pooled arrangement, and lower the pool size in your application — a web application almost never needs the default its framework ships with.
A failover happened and in-flight transactions failed. This is correct. Anything uncommitted when the primary died is gone; a promotion cannot invent the result of a transaction nobody finished. Applications that matter should retry a failed transaction rather than assume it succeeded.
A failover happened and a very recent write is missing. Copies follow the primary with a small delay. If the primary dies in that window, the write it acknowledged may not have reached the copy that took over. The window is short but it is not zero. If your application cannot tolerate that at all, say so before you order — the trade is slower writes for no loss, and it has to be chosen deliberately.
A copy is far behind. Usually a long-running query on that copy, or a burst of writes larger than the link can carry. It catches up on its own. A copy that is too far behind is not promoted during a failover, because promoting it would discard more than the alternative.
The disk filled up. Everything stops, including the ability to delete rows — which needs to write. Watch the size; do not wait for it.
You deleted the database. It is gone, and so is everything in it. The confirmation is deliberately tedious.
What to read next
- Object storage for files rather than rows.
- When things go wrong.