Cloud Cost Governance for Custom Applications: FinOps Basics for Engineering Teams
Cloud bills are shaped by engineering decisions: how services are sized, how often they call each other, how long logs are kept, whether test environments run overnight. That is why cost governance for custom applications cannot live only in the finance team. It needs a shared practice that gives engineers visibility into what their software costs and gives the business a way to make trade-offs on purpose. That practice is usually called FinOps.
This guide covers FinOps basics for teams that build and run their own applications in the cloud: the principles and phases defined by the FinOps Foundation, how to allocate costs with tags, how budgets and alerts really behave, and how to build cost awareness into everyday engineering work.
What FinOps is (and isn't)
FinOps is not a cost-cutting project. The FinOps Foundation frames it around six principles:
- Teams need to collaborate.
- Business value drives technology decisions.
- Everyone takes ownership for their technology usage.
- FinOps data should be accessible, timely and accurate.
- FinOps should be enabled centrally.
- Take advantage of the variable cost model of the cloud.
Two of these are particularly important for custom software teams. "Everyone takes ownership" means the team that builds a service is accountable for what it costs to run, just as it is for its reliability. "Business value drives technology decisions" means the goal is not the lowest bill but the best value: spending more on a service that earns more can be the right decision.
The three phases: Inform, Optimize, Operate
The Foundation describes FinOps as three iterative phases:
- Inform — examine cost, usage and efficiency data; allocate costs and make them visible.
- Optimize — identify opportunities to improve efficiency and value, covering both usage optimisation (consuming less) and rate optimisation (paying less per unit).
- Operate — implement the changes identified and measure the results.
These are cycles, not stages you finish. Practitioners move through them repeatedly, and different teams can be in different phases at the same time. For a team just starting, almost all the early value is in Inform: you cannot optimise costs you cannot attribute.
Inform: make every cost belong to someone
Design a tagging schema before you need it
Tags (or labels) are key–value metadata attached to cloud resources. AWS's tagging best practices whitepaper positions a consistent tagging strategy as the basis for filtering resources, monitoring cost and usage, and managing the environment. A minimal schema for custom applications usually includes:
application— which product or service the resource belongs to.environment— production, staging, development, test.owner— the team accountable for the resource.cost-center— where the spend is reported financially.
Keep the list short and the allowed values controlled. A tag that is optional or free text will be inconsistent within weeks. Never put sensitive information in tags; AWS explicitly advises against it.
Know the platform's rules
Each provider has details that catch teams out. On AWS, for example, the cost allocation tags documentation explains that both AWS-generated and user-defined tags must be activated before they appear in Cost Explorer or cost allocation reports, and that tags can take up to 24 hours to appear in the billing console. The same documentation notes that cost allocation reports include both tagged and untagged resources, which is useful: the untagged share is your first governance metric.
Enforce tags where resources are created
Retrofitting tags onto a live estate is painful. For custom applications, the most reliable approach is to define tags in infrastructure-as-code modules so that every resource inherits them, and to fail deployments that omit required tags. Pair that with a regular report of untagged spend and a named person responsible for driving it down.
Decide how shared costs are split
Some costs belong to no single application: networking, shared clusters, observability platforms, support plans. Agree a simple, documented rule — for example, split proportionally to each application's direct spend, or hold shared platform costs centrally — and apply it consistently. The exact rule matters less than having one that everyone understands.
Move toward unit costs
Total spend going up can be good news if the business is growing. Unit costs — cost per order processed, per active customer, per API call, per report generated — tell you whether efficiency is improving. Pick one or two units that the business already tracks and publish them alongside total spend for each application.
Budgets and alerts: know what they do
Budgets are one of the first controls teams set up, and one of the most misunderstood. Google Cloud's budget documentation is explicit that an alerts-only budget does not automatically cap usage or spending. Budgets notify; they do not stop resources. Google's budgets can alert on actual costs or on forecasted costs to the end of the period, with default thresholds at 50%, 90% and 100% of the budget amount.
Practical guidance:
- Set budgets per application and environment, not just per account, so alerts reach the team that can act.
- Use forecasted-cost alerts to get warning before the month is lost, not after.
- Route alerts to a team channel with an owner, not a shared finance inbox.
- If you need hard limits, implement them deliberately (for example, automation that scales down non-production environments), and understand the availability risk before applying them to production.
Optimize: where custom applications usually waste money
The FinOps Foundation separates usage optimisation from rate optimisation. For teams that own their code, usage is where most of the control sits. Common areas to review:
- Idle non-production environments. Development and test environments that run around the clock but are used during working hours.
- Over-provisioned compute. Instance or container sizes chosen at launch and never revisited against actual utilisation.
- Orphaned resources. Unattached volumes, old snapshots, unused load balancers and forgotten experiments.
- Storage without lifecycle rules. Data kept in the most expensive tier indefinitely.
- Logging and telemetry volume. Verbose debug logging left on in production, or retention periods longer than anyone needs.
- Data transfer. Chatty service-to-service calls across zones or regions, or large responses sent where a smaller payload would do.
Rate optimisation — commitment-based discounts and negotiated pricing — typically comes after usage is under control and predictable, and is usually coordinated centrally, in line with the principle that FinOps should be enabled centrally.
Operate: build cost into the engineering workflow
Lasting results come from making cost a normal engineering signal rather than a monthly surprise:
- Design reviews include an expected cost profile for new services and notable changes.
- Dashboards show each team its own spend and unit cost, next to its reliability metrics.
- Anomaly reviews happen weekly for significant unexpected changes, with a short note on cause.
- Backlogs carry cost-improvement items prioritised against feature work, with the expected saving stated.
- Results are measured after each change, closing the Operate loop and feeding the next Inform cycle.
A practical starting plan
- Agree the tagging schema and shared-cost rule, and activate cost allocation tags where your provider requires it.
- Add required tags to infrastructure-as-code modules and block untagged deployments.
- Create per-application budgets with forecasted alerts routed to owning teams.
- Publish a simple dashboard: spend by application and environment, untagged share, one unit cost.
- Run a first usage review of non-production schedules, orphaned resources and log retention.
- Add a cost section to design reviews and start a weekly anomaly check.
If you are planning a migration, put steps 1–3 in place before the first workload moves; our guide to choosing a migration strategy for custom applications explains why replatforming decisions have a large effect on ongoing cost. Integration design matters too: batch-heavy, chatty integrations can drive both cost and throttling, as covered in our API-first ERP integration guide. For help setting this up for your applications, see our cloud solutions.
Frequently asked questions
Does setting a cloud budget stop spending when it is reached?
Generally, no. Google Cloud's documentation states that alerts-only budgets do not cap usage or spending. Budgets send notifications; stopping resources requires separate automation.
Who should own cloud costs — finance or engineering?
Both. The FinOps principles call for teams to collaborate, for everyone to take ownership of their usage, and for FinOps to be enabled centrally. Engineering teams own the usage of their services; a central function provides data, tooling, policies and rate optimisation.
What should we measure first?
Spend by application and environment, and the share of spend that is untagged or unallocated. Once allocation is reliable, add one unit cost that reflects business value.
Planning a Migration or Integration?
Tell us what you're trying to solve. We'll help you choose the right approach for your applications.
Contact Us →