Skip to main content
The fastest way to reduce an AWS bill is usually not a complex pricing model. It is removing resources that should not exist and shrinking resources that are larger than the workload needs. Updated October 9, 2026. Start with waste, then improve architecture, then consider commitments. A discount on unused capacity is still waste.

1. Delete idle resources

Look for resources with cost but no useful traffic or owner:
  • Old load balancers and NAT gateways
  • Unattached EBS volumes and obsolete snapshots
  • Databases created for abandoned environments
  • Elastic IP addresses no longer attached to active resources
  • CloudWatch log groups with indefinite retention
  • Preview and staging environments that nobody uses
Require an owner, environment, and lifecycle tag. A resource without an owner should enter a review queue, not run forever.

2. Right-size compute and databases

Compare provisioned CPU, memory, storage, and database capacity with actual usage over a representative period. Include peaks, deployments, and failure recovery. For ECS, reduce task CPU and memory gradually, then watch throttling, out-of-memory exits, latency, and queue depth. For databases, check connections, CPU, memory pressure, I/O, and recovery requirements before changing size. Do not right-size from average CPU alone. A service may be limited by memory, network, disk, or latency during short peaks.

3. Turn off what does not need to run

Development and test environments often run all night and all weekend. Schedule them when startup time and team workflow allow it. Scale workers to demand, remove minimum capacity where the queue can tolerate cold starts, and expire preview environments automatically. Stateful services need a deliberate shutdown and recovery plan, so do not treat them like disposable compute.

4. Reduce expensive data paths

Data movement can cost more than storage or compute.
  • Keep chatty services in the same region.
  • Avoid accidental cross-Availability Zone traffic.
  • Use CloudFront for cacheable internet content.
  • Review NAT gateway data processing.
  • Use VPC endpoints only when their hourly and traffic costs fit.
  • Compress large responses and avoid repeated downloads.
Map the complete request path before changing it. A cheaper service can become more expensive when traffic crosses several priced boundaries.

5. Use interruptible capacity safely

Spot capacity can lower compute cost for workloads that tolerate interruption. Good candidates include queue workers, parallel batch jobs, CI tasks, media processing, and stateless services with enough baseline capacity elsewhere. Use a mixed strategy for customer-facing services. Keep critical baseline capacity on On-Demand and add Spot for flexible capacity. Handle termination signals, retry work safely, and spread across supported sizes or capacity pools.

6. Commit after the baseline is stable

Savings Plans and Reserved Instances can reduce rates for predictable usage. They also create a commitment. Measure a stable baseline first. Commit only the portion you expect to keep using through product changes and migrations. Leave uncertain or burst capacity flexible.
Do not purchase a long commitment to hide an oversized architecture. Right-size first, then apply the discount to the remaining baseline.

7. Make cost visible beside engineering decisions

A monthly total is too late and too broad. Show teams the daily cost, run rate, owner, and workload unit that produced it. Review large changes after deployments and architecture changes. Springwinter places cost information beside the resource that generated it. See cost visibility. Combine that view with AWS budgets and billing reports because the AWS bill remains authoritative.

Use this order

1

Remove

Delete unused resources and expire temporary environments.
2

Resize

Match capacity to measured demand and recovery requirements.
3

Reshape

Reduce data transfer, storage growth, and noisy telemetry.
4

Schedule

Stop non-production resources when nobody needs them.
5

Discount

Use Spot and commitments only where the workload can support them.
Cost optimization works when it becomes ordinary engineering. Make one measurable change, verify reliability, and continue. Large one-time cost projects tend to fade; a small recurring review keeps the bill aligned with the product.

Frequently asked questions

Delete idle resources first, then right-size compute and databases using measured utilization. Review NAT gateway traffic, log ingestion, unattached storage, old snapshots, abandoned load balancers, and unused development environments. Apply Spot and commitments only after the remaining baseline is understood.
Savings Plans discount eligible compute usage up to a committed hourly amount. They do not reduce every service, storage, data transfer, NAT gateway, or observability charge. Purchase commitments only for a stable baseline after removing waste and right-sizing resources.
Use Spot for interruptible, retryable workloads such as queue workers, batch processing, CI, media jobs, and flexible service capacity. Keep critical baseline capacity on On-Demand when needed, handle interruption signals, make jobs idempotent, and diversify supported capacity pools.

Sources and further reading

Last modified on October 8, 2026