We calculate the acceptable losses
For every system we record two numbers: the acceptable outage and the acceptable data loss. Everything else follows from them.
Six parts. The first is entirely about the calculations; the other five are assembled to fit the numbers, not the other way round.
For every system we record two numbers: the acceptable outage and the acceptable data loss. Everything else follows from them.
A schedule, version retention, an off-site copy and a monthly attempt to restore the archive on a separate environment.
We watch the machines, the links, free space and services. The alert reaches us before it reaches your staff.
A backup link with a different provider and automatic failover. For retail and warehousing this is mandatory.
UPS units that shut the servers down cleanly, a stock of critical spares and extended warranty on the key components.
The document for the day of an outage: the sequence of steps, the names of those responsible, phone numbers and the order in which systems come back up.
The five most common scenarios. Each has its own defence, and one cannot substitute for another.
Backups and a stock of spares save you. It happens more often than anything else and recovers more predictably.
An offline archive saves you. It finds ordinary backups and destroys them along with the original, which is why you need a copy with separate access.
Deep version retention saves you. The loss is noticed weeks later, so yesterday's backup is not enough.
A second provider saves you. Two links from the same provider usually follow the same route and die together.
A UPS with a clean shutdown saves you. A sudden loss of power harms the database more than the outage itself.
A backup only becomes a backup once you have managed to restore it. In our experience about one client in three has a «configured» scheme that does not come back up: the job has been failing for months, the wrong folder is being copied, or the archive is locked with a password nobody remembers. Checking takes one day.
The question is not whether you have backups but when you last restored one and how long it took. Start with a test: we take your copy, bring it up on a separate environment and time it. The result either confirms the scheme or shows the weak point, and both are useful.
It depends on the numbers you need. Coming back up within a working day is inexpensive; within an hour is noticeably dearer, because it needs a second site on hot standby. So the first thing we calculate is the acceptable downtime: usually it turns out half the systems are fine with the cheap scheme.
We do, under contract and within the stated deadlines. The procedure is written so that your own person could follow it if our engineer is unavailable. The document is kept by you.
Yes, that is exactly what the test restore is for. Every month we bring the archive up on a separate environment with a stopwatch running. Once a year it is worth staging a drill with your staff: it shows at once where the paperwork parts ways with reality.
Разберём вашу схему копирования, попробуем поднять архив и замерим настоящее время подъёма каждой системы.
Request received
It is already with a manager. You will get an answer within the working day, and urgent requests go to the duty engineer immediately.
There is no such city in the list. Check the spelling.