"Should we move to the cloud" is the wrong opening question, because it has no general answer. It has a specific answer per workload, and working it out takes about two weeks.
The four questions
1. What does this cost to run over five years, in each option?
Not the monthly figure against the purchase price. The full comparison: hardware plus power plus cooling plus rack plus the refresh in year four plus the administration hours, against consumption plus egress plus support plus the administration hours.
Do it properly and the result often surprises people, in both directions.
2. How long can it be down, and how much data can you lose?
Disaster recovery is where cloud usually wins decisively. Building real geographic redundancy on your own hardware means a second site, and that changes the arithmetic completely.
If your recovery tolerance is measured in days, on-premise may be fine. If it is hours, cost the second site properly before you conclude local is cheaper.
3. Does anything dictate where the data lives?
Regulation, a customer contract, a tender requirement. If the answer is yes, it is a constraint, not a preference, and it decides the question. Sometimes that is true for one workload and not the others, which is how a hybrid design happens on purpose instead of by accident.
4. Who operates it in eighteen months?
The hidden one. Cloud shifts work from hardware maintenance to platform engineering and cost management. Those are different skills, and if nobody in the building has them the running cost is higher than the invoice suggests.
Why sizing calculators mislead
Every vendor calculator asks for your requirements and assumes you know them. Almost nobody does, so the numbers get padded "for safety", and you pay for that padding every month, forever.
Measure instead. Two weeks of CPU, memory, IOPS and network data across a representative period, including month-end and your busiest day. Size against what you observe, with headroom agreed explicitly rather than smuggled in.
We have never done this exercise without it changing the specification.
Where each option wins
Cloud wins on spiky or seasonal load, on geographic distribution, on disaster recovery, on anything growing unpredictably, and on new projects where the requirement is still moving.
On-premise wins on steady predictable load running around the clock for five years, on very high data volumes where egress charges accumulate, on latency-sensitive workloads next to the people or machines using them, and where regulation requires it.
Hybrid wins more often than either camp admits: regulated or latency-critical data local, everything else in the cloud, with the link between them designed rather than improvised.
If your cloud bill is already out of control
This is one of the most common reasons we get called, and the causes are predictable: over-provisioned instances sized by guess, test environments nobody switched off, storage on the wrong tier, no reserved capacity or savings plan, and snapshots accumulating since 2022.
A cost review typically pays for itself inside a quarter. It is also the least glamorous work we do, which is presumably why it stays undone.
The version of this advice we would give a friend
Do not migrate for its own sake. Do not stay put for its own sake either.
Take your five most important workloads, answer the four questions for each, and you will probably end up with three in the cloud and two where they are. That is not a compromise. It is usually just what the numbers say.



