Skip to content
Innoveta Tech, fast forward with tech
All insights

Cloud

The real drivers on a cloud bill, ranked

Nobody asks for a cost review because they read something about FinOps. They ask because the bill grew faster than the business did and no one can name the workload responsible. Here is where the money usually is, and which fixes are worth doing first.

5 min readInnoveta Tech

Key takeaways

  • A runaway cloud bill is usually a design decision made under time pressure, charged to you again every month since.
  • Six causes account for most of the surprise: oversized compute, idle environments, untiered storage, mis-shaped commitment, egress, and missing tags.
  • Right-sizing against real utilisation is usually the single largest line item recovered, and one of the least disruptive to fix.
  • Tagging returns nothing on its own. It is what makes every other saving arguable, so it goes first.
  • Rank findings by annual saving against disruption to fix. Two of the six are architecture conversations, not switches.

Cloud cost reviews almost never start with a strategy document. They start with an invoice that has grown faster than the business has, and a finance conversation nobody in the technology team can answer confidently. The question on the table is usually some version of: what is this, and who asked for it?

The honest answer is that a runaway bill is rarely a pricing problem. Pricing is public and broadly comparable across the three major providers. What you are paying for is a set of design decisions, most of them made quickly, some of them made by people who have since left, and all of them charged to you again every month since.

Across the environments we are asked to review, the same handful of causes account for most of the surprise. They are worth knowing in advance, and worth fixing in a particular order.

Why a bill you cannot attribute cannot be argued with

Before any of the six causes, there is a step that returns nothing on its own and makes everything else possible. Spend that cannot be attributed to a team, a product or a client cannot be questioned by anyone, because there is nobody in the room whose budget it is.

This is why tagging is a governance control before it is a reporting one. Untagged spend has no owner, no defender and no challenger, so it survives every cost review by default. A cost baseline broken down by workload and owner, rather than by account, changes the conversation from "the bill is too high" to "this workload costs this much, and here is the person who can say whether that is reasonable".

The rule we work to

A runaway cloud bill is rarely a pricing problem. It is a design decision somebody made under time pressure two years ago, charged to you again every month since.

The six causes, ranked by what they return

Ranking matters because the fixes are not equally disruptive. Two of the six are a switch and a maintenance window. Two are a commercial decision. One is an architecture change that needs a proper design conversation. Sequencing them badly is how a cost programme stalls after the easy wins.

CauseWhat it typically returnsDisruption to fix
Compute provisioned for a peak that never comesUsually the largest single recoveryLow. A resize inside a maintenance window.
Environments nobody switched offImmediate, and repeats every monthLow, once somebody owns the decision to stop them.
Storage that never tiersCompounds. It grows every month it is leftLow. A lifecycle rule, once retention is agreed.
Commitment bought at the wrong shapeSignificant, and easy to get wrong twiceMedium. You are committing for one to three years.
Egress and cross-region trafficVariable. Occasionally it is the whole problemHigh. It is an architecture decision, not a setting.
No tags, so no accountabilityNothing directly. Everything indirectlyLow effort, high negotiation.
The usual causes of surprise on a cloud invoice, and what each costs to fix

The three that pay for the review

The first three are unglamorous, and they are where the money is. None of them require an architecture change and none of them are contested once the numbers are visible.

Compute sized for a launch-week estimate and never revisited is the most common finding of all. The estimate was reasonable when it was made. What did not happen is anyone going back six months later, when real utilisation data existed, and asking whether the guess had held.

Test, staging and proof-of-concept environments running at full rate through nights and weekends are the second. These are rarely defended once identified, because the people who provisioned them frequently no longer work there. The fix is a schedule, and the hard part is finding someone willing to say the environment can stop.

Storage is the quiet one. Snapshots, logs and backups accumulate on hot storage for years because no lifecycle rule was ever written, and the monthly increment is too small to trigger anyone's attention. It is the only cause on the list that gets worse purely through the passage of time.

Ask when the sizing was last checked

In most environments, the answer is at provisioning. Utilisation data has existed ever since and nobody has had a reason to open it. That single question tends to find more money than any tooling purchase.

The two that need an architecture conversation

Reserved capacity and savings plans are a bet on your own architecture staying roughly where it is. Bought well, they are the cheapest saving available on stable workloads. Bought badly, they lock you to instance families you have since moved away from, and you pay for the commitment and the replacement at the same time.

So commitment is worth buying only where you can defend the forecast, which usually means the workloads that have not changed shape in a year. Buying commitment across an environment you are about to modernise is one of the few cost decisions that can leave you worse off than doing nothing.

Egress is different again. Data moving between regions or out to the internet is almost always the visible symptom of a design decision: a service placed in one region talking to a store in another, a backup route nobody costed, an integration pulling a full dataset when it needed a delta. You cannot configure your way out of it. The invoice line is real, but the fix lives in the architecture.

ONE ENVIRONMENT, DESIGNED IN LAYERSOperateMonitoring & alertingPatchingCost reviewRunCompute & storageData platformPipelinesFoundationIdentity & accessNetworkBackup & recoverySkip the foundation and you pay for it in the operate layer, every month.
Foundation, workloads, operations. Cost problems usually trace back to the layer that was retrofitted rather than designed.

What to do with the list

A cost and resilience review runs two to three weeks and stands alone, without a migration attached. What makes the output usable is not the length of the findings list. It is that every finding carries the annual figure and the disruption alongside it, so the decision about what to act on belongs to you rather than to whoever wrote the report.

  1. 01Tag first, even though it saves nothing. Without attribution the rest of the list has no owner.
  2. 02Take the three low-disruption causes in the same change window. They fund the work that follows.
  3. 03Model commitment against the workloads you are confident will still exist in a year, and only those.
  4. 04Treat egress as an architecture item with a design conversation attached, not a line to be optimised.
  5. 05Put a monthly cost and capacity review in place with a named person. Everything on this list grows back otherwise.

That last point is the one most often skipped. Every cause on this list is a slow accumulation rather than a single event, which means a one-off clean-up buys you eighteen months and then the same conversation.

The short version

Get the bill attributable before you try to reduce it. Take the boring, low-disruption savings first, because they are the largest and the least contested. Treat commitment and egress as decisions rather than settings. Then put a monthly review in place, because a cloud bill is a thing that grows back.

Common questions

What causes a cloud bill to grow without usage growing?

Most commonly compute provisioned for a peak that never arrived and never revisited, non-production environments left running outside business hours, and storage with no lifecycle rule, so snapshots, logs and backups accumulate on hot storage for years. None of these track business growth. They track decisions nobody has revisited.

Can you do cloud cost optimisation without a migration?

Yes. A FinOps review covering right-sizing, commitment shape, storage tiering, idle environments and stale spend runs as a standalone two-to-three week engagement, and each finding comes with the annual saving attached so you can rank what is worth acting on.

Which cloud cost fix returns the most?

Right-sizing compute against real utilisation is usually the single largest recovery, and one of the least disruptive to make. It is also the one nobody schedules, because the original sizing was a reasonable estimate at the time and there is no event that prompts anyone to check it later.

Should we buy reserved instances or savings plans?

Only on workloads whose shape you can defend for the length of the commitment, which usually means the ones that have not changed in a year. Buying commitment across an environment you are about to modernise can leave you paying for the commitment and its replacement at once.

Why does tagging matter if it does not save anything?

Because spend that cannot be attributed to a team, product or client cannot be questioned by anyone. Tagging is a governance control before it is a reporting one. It is what turns "the bill is too high" into a specific workload with a specific owner who can say whether the cost is reasonable.

Start with a conversation, not a proposal.

Tell us the problem. We’ll tell you honestly whether it’s worth solving, and what solving it would take.

We use cookies for analytics and to improve this site. See our Privacy Policy.