Last month I watched a startup burn through $8,000 in AWS credits in three weeks because someone spun up a db.r5.24xlarge instance for “testing” and forgot about it. The database had exactly zero connections. This wasn’t stupidity or carelessness. This was what happens when cloud cost optimization becomes an afterthought instead of a design principle.

The dirty secret of cloud infrastructure is that the default settings are optimized for vendor revenue, not your budget. Every “easy deploy” button, every auto-scaling policy, every managed service comes with assumptions that probably don’t match your actual workload. Let’s audit those assumptions and fix them.

Right-Size Your Instances (Because Nobody Needs 96 Cores for a Blog)

The first place your money disappears is instance sizing. AWS, GCP, and Azure all make it ridiculously easy to over-provision compute resources. Click a few buttons and suddenly you’re running a machine with more RAM than most data centers had in 2005. The problem isn’t that these resources exist. The problem is that the pricing model pushes you to guess high and optimize never.

Start with your actual metrics, not your fears. Install monitoring on every instance and collect at least two weeks of CPU, memory, and network utilization data. You’ll discover that most web applications spend 95% of their time using less than 20% of available resources. That t3.large running at 8% CPU utilization? It could be a t3.small saving you $30 per month. Multiply that across dozens of instances and you’re looking at real money.

Here’s the weird part: smaller instances often perform better. A properly configured t3.medium with burstable CPU credits will outperform an oversized t3.xlarge under typical web traffic patterns. The bigger instance creates lazy code that assumes infinite resources. The smaller instance forces efficient design patterns that actually scale better long-term.

Storage Classes Are Not Created Equal (Stop Paying Ferrari Prices for Bicycle Trips)

Storage optimization is where I see the most ridiculous waste. Teams default to high-performance SSD storage for everything because it’s the path of least resistance. Then they store log files, backups, and static assets on the same tier as their production database. This is like parking a Lamborghini in your grocery store parking spot every day.

Amazon S3 alone has eight different storage classes with dramatically different pricing. Standard storage costs $0.023 per GB per month. Glacier Deep Archive costs $0.00099 per GB per month. That’s a 23x price difference. The catch? Deep Archive has retrieval times measured in hours, not milliseconds. But for backups older than 90 days, who cares about retrieval time?

Set up lifecycle policies that automatically move data to cheaper storage classes based on access patterns. Files that haven’t been accessed in 30 days move to Infrequent Access. Files untouched for 90 days go to Glacier. Files older than a year go to Deep Archive. Configure this once and watch your storage costs drop by 60-80% over the following year without changing a single line of application code.

Reserved Instances Are Loans Disguised as Discounts

Cloud providers love selling reserved instances because they get cash upfront and you get locked into their ecosystem. The marketing pitch sounds great: save 30-70% compared to on-demand pricing. The reality is more complicated. Reserved instances are basically loans where you pay interest in the form of reduced flexibility.

The math works if you can accurately predict your usage patterns 12-36 months in advance. But how many engineering teams can honestly claim that level of foresight? I’ve seen companies buy three-year reserved instances for projects that got cancelled six months later. Those reserved instances become very expensive paperweights.

A smarter approach is spot instances for development and batch workloads. Spot instances offer the same discount as reserved instances but without the long-term commitment. Yes, they can be terminated with two minutes notice. But if you’re running stateless applications or fault-tolerant batch jobs, spot termination becomes a feature, not a bug. It forces you to build resilient systems that handle failures gracefully.

Serverless Isn’t Automatically Cheaper (Math Doesn’t Lie, but Marketing Does)

The serverless marketing machine wants you to believe that Lambda functions and similar services automatically optimize costs by charging only for execution time. This is true for workloads with unpredictable traffic patterns and long idle periods. It’s catastrophically expensive for steady-state applications with consistent load.

Lambda charges $0.0000166667 per GB-second of execution time. That sounds tiny until you do the math for a function that runs continuously. A 1GB Lambda function running 24/7 costs about $43 per month. A t3.small instance with 2GB RAM costs $16.79 per month and gives you twice the memory plus persistent storage. The break-even point is roughly 30% utilization.

Serverless works great for event-driven architectures, batch processing, and applications with sporadic traffic. It’s expensive overkill for APIs that handle steady traffic or background workers that run continuously. The key is matching the pricing model to your usage pattern, not defaulting to whatever technology seems most modern.

Monitor Everything, Trust Nothing

Cost optimization without monitoring is just guessing with extra steps. Set up billing alerts that trigger before you hit budget thresholds, not after. Configure automated reports that break down costs by service, project, and team. Most importantly, make cost visibility part of your deployment pipeline.

Tag every resource with owner, project, and environment metadata. This isn’t bureaucracy. This is the only way to identify which team spun up that mystery EC2 instance that’s been running for six months with no traffic. Create cost allocation reports that show each team their monthly spend. Nothing motivates optimization like making the costs visible to the people creating them.

Consider tools like Infracost that estimate infrastructure costs directly from Terraform plans. Build cost estimation into your code review process. Make cost impact as visible as security vulnerabilities or performance regressions. The goal isn’t to eliminate all spending, but to make every dollar intentional.

Your cloud bill reflects your infrastructure decisions six months ago. Every instance you spin up, every storage bucket you create, every managed service you enable is a monthly commitment. The question isn’t whether you can afford the current bill. The question is whether you’re getting enough value to justify paying it every month for the next year. What assumptions are hiding in your infrastructure that deserve a skeptical second look?