Storage economics

Backup storage refresh: Expand or replace?

A renewal quote forces an expand or replace decision. The mechanical questions that actually settle it.

7 min read
Older and newer server racks side by side in a dark data center aisle joined by glowing teal light paths

A backup storage refresh rarely announces itself as an architecture decision. It arrives as a support renewal quote with a date on it, and the choice gets made against that date: extend the existing platform, or replace it. The quote is a real input, but it prices one year of one line item. What it does not describe is data layout, migration effort, or how long the old system has to stay powered because of locks already written to it.

The comparison usually made is renewal cost against new purchase cost over the same number of years. That framing is clean and it hides most of the decision. Expansion is only cheap if the platform can absorb capacity without changing how it lays data out, and replacement is only expensive if the data has to move on a deadline set by something other than the finance calendar.

What the renewal quote does not price

A renewal covers hardware support and software subscription for the installed base. It does not cover power and rack space, the hours the team spends on the platform, or the price of capacity added late in the life of a system, which is frequently quoted differently from capacity bought at the start. It says nothing about the renewal after this one, and that is often the number that settles the case.

Two figures are worth extracting before any comparison begins. The first is the published end of sale and end of support date for the exact model installed, since a platform that can be renewed for one more year but not for three has already answered the question. The second is the price of an expansion bought at this point in the life of the system, quoted next to the same capacity on a new platform. Most of the remaining questions are mechanical, and a vendor can answer them in writing.

Whether the platform can take more capacity as it stands

Expansion is inexpensive when new capacity joins the existing pool and inherits the layout already in use. It stops being inexpensive the moment the platform answers the request by creating a second pool, a second appliance or a second deduplication domain. The quote can look identical in both cases, since it is priced per unit of capacity, but the operational result is not.

The questions are about scope. Does deduplication span the whole platform or stop at each pool, so a second pool starts its index from nothing. Does free space pool across new capacity, or is it reported per shelf, so a large job can use only part of what the dashboard shows as free.

Two limits are easy to miss because they are not about disks. The metadata layer can reach a ceiling before the drive bays do, and a platform whose nodes cannot run the release that supports larger drives is already at its limit whatever the bay count says.

The related question is whether old and new hardware can live in one namespace. Where nodes of different generations and drive sizes join the same pool, expansion and replacement stop being opposites: new nodes join, old ones retire one at a time, and the backup software keeps pointing at one target. Where they cannot, old and new are two systems, and that is a different project.

SignalWhat it usually meansWhat to check
Expansion is quoted as a second pool or applianceThe platform cannot grow the existing data layoutWhether deduplication, free space and lock settings span both
Added capacity needs a release the nodes cannot runThe hardware is past the boundary for new featuresThe last release the current nodes support
Rebuild time has grown with each expansionThe protection scheme is stretched across larger drivesMeasured rebuild time at current occupancy, not the installation figure
The renewal is offered for a shorter term than beforeEnd of support for this generation is being plannedPublished end of sale and end of support dates for the exact model
Free space is reported per shelf rather than per poolHeadroom is fragmented and only partly usableHow much of the free space one large job can actually consume
Migration is quoted as a service by the new vendorThe move needs effort the team does not haveWhat it assumes about source read speed and locked objects

What replacement costs in migration and parallel running

Replacement carries two costs that appear in neither quote. The first is reading everything off the old platform once, at the speed of a system that is full, several years old, and still serving nightly jobs while it drains. Where the backup software moves the data rather than the storage layer, it is a copy job competing for production backup windows.

The second is parallel running. Both platforms are racked, powered and monitored for the overlap, and retention usually sets its length, not copy speed. A third option is often cheaper: stand the new platform up, point new jobs at it, and let the old one age out as its restore points expire. That removes migration effort at the price of a longer overlap, and it is worth modeling wherever older restore points are unlikely ever to be restored.

Where the backup software mediates the move, the questions worth asking before a repository migration mostly concern what happens to existing chains, metadata and immutability flags on arrival.

How immutability locks set the timeline

Object Lock turns the decommissioning date from a choice into a calculation. Under compliance mode a retention date written to an object cannot be shortened by anyone, so the old platform stays readable until the last such date passes. Under governance mode a specifically privileged identity can lift the retention, an administrative act to be recorded rather than a routine migration step.

Copying locked data does not shorten the timeline. The copy is written with locks applied on arrival, starting a fresh retention period, while the source objects stay locked until their original dates pass. Both platforms then hold retained copies of the same data, and both consume capacity for it.

The practical output is one date: the latest lock expiry across every bucket on the old system. Reading it out of the platform rather than inferring it from job settings is worth the effort, since the lock actually applied is often longer than the retention typed into the job.

Where ARTESCA fits

ARTESCA is object storage software that runs on infrastructure the customer owns, used as a backup target in the range of roughly 50 TB to 8.5 PB. Capacity is added by adding nodes to the running system, and the interface presented to the backup software is the S3 API with Object Lock for immutability, so an expansion is a storage-side operation rather than a reconfiguration of backup jobs.

The same property affects the replace side. Where both the old and the new target speak S3, a change of platform is a change of endpoint, credentials and bucket, and the question becomes how the data moves rather than whether the protocol changes. Where a single namespace has to span multiple sites or exceed that range, RING covers the same S3 and Object Lock semantics at larger scale.

What to decide before the next renewal quote

Four facts turn this from a negotiation into an assessment, and all four can be written down now rather than in the two weeks a quote is valid. The last release the installed hardware can run, with its support end date. The latest lock expiry, read from the storage. Measured read throughput at current occupancy. And a written answer on whether an expansion joins the existing pool or creates a second one.

Those belong with the rest of the platform record, refreshed on a fixed annual date rather than on the vendor's schedule. The lock expiry in particular moves quietly as retention settings change, and it can make a sound replacement plan impossible to execute in the quarter it was planned for. Once a year, re-read the four, re-price expansion against new capacity in the same unit, and note what changed. That record also keeps routine storage maintenance windows predictable between refreshes.

Try ARTESCA free

Immutable object storage that scales from 20TB to petabytes. Deploy a working cluster in under an hour.

Start a free test drive