The Real Cost of a Petabyte

“A petabyte costs about $23 a month in S3.”

You may hear this (you may not if you have a social life). It’s not really true, in the sense that nobody actually pays $23 a month for a petabyte and gets anything useful out of it.

The headline storage price is a few percent of the actual bill, and the rest of the bill is the part the marketing slides don’t list. Worth working through honestly, because the petabyte-of-data shape comes up a lot in enterprise data conversations and the cost intuitions are mostly wrong.

The Headline Number

S3 Standard, the default tier, costs roughly $0.023 per GB per month for the first 50 TB, dropping to $0.022 and then $0.021 as you scale past 500 TB. A petabyte (1,024 TB, or 1,048,576 GB) lands at about $22,600 a month once you blend those tiers — call it $24,000 at the flat headline rate — which works out to roughly $270,000 a year. That’s real money, but not catastrophic (if your revenue streams justify it). Just remember it’s storage only: request and egress charges (moving a full petabyte off S3 runs close to $90,000 on its own) can easily eclipse the storage line depending on how often you touch the data.

Other tiers are cheaper: Infrequent Access about half that, Glacier Deep Archive about a quarter of a cent per GB. If you can put your petabyte on Glacier, you’re looking at $30K a year, which sounds like the “a petabyte costs $23 a month” quote some bloggers paraphrase. The catch is that you can’t do anything with Glacier data without paying restore costs, so it’s a deep archive, not a working store.

What Else is in the Bill

The honest costs that nobody puts on the slide:

Requests. S3 charges per API call. PUT, GET, LIST, DELETE all cost something. For a petabyte broken into millions of Parquet files queried by Spark, the request cost is non-trivial. A heavy query workload can run into thousands of dollars a month just in GET requests.

Egress. Moving data out of S3 to anywhere else costs $0.05-0.09 per GB depending on destination. If you read your petabyte once a month for analytics across regions, that’s $50-90K a month in egress alone. More on this in next week’s post.

Replication. Most production deployments want cross-region replication for DR. Doubles the storage cost and adds replication transfer charges.

Versioning. S3 versioning means you have multiple copies of objects when they change. Useful, but it multiplies the storage footprint. Most teams underestimate by 2-3x.

Snapshots and backups. Backup tools that snapshot S3 to another bucket or another cloud add more storage cost. Often forgotten.

Compute. Storage without compute is a museum. You need to read it. EMR, Athena, Glue, Spark on EC2, your warehouse’s external table reader – all of these have their own costs that often dwarf the storage bill.

The Warehouse Multiplier

If you’re not querying directly from S3 but loading into Snowflake or BigQuery, you pay twice. Once for the S3 storage, again for the warehouse storage (Snowflake at ~$23/TB/month for compressed data, BigQuery similar). On top of that, the compute to query is metered by the warehouse, separately from the storage.

This goes away largely with Iceberg tables acting as external tables. But not everyone is doing that.

For a petabyte of working data, you might pay $20K/month in S3 + $20K/month in Snowflake storage + $50-150K/month in Snowflake compute, depending on workload intensity. Storage is the smallest line.

The Compression Discount

The actual logical data volume people talk about is usually bigger than the physical storage. Parquet compresses 5-10x for typical analytical data. So a “petabyte of data” in business terms is often 100-200 TB on disk. This works in your favour for storage cost and against you for compute (the queries still process the logical data).

When teams quote “we have 5 PB of data,” ask whether that’s logical or physical. The number can vary by 10x depending on which they mean.

The Cloud Cost Shape

A realistic per-petabyte annual cost for working analytics data on AWS, rough orders of magnitude:

  • S3 storage: $250K
  • Requests and listings: $20K
  • Cross-region replication storage: $250K
  • Replication transfer: $100K
  • Egress to compute: highly variable, $50K-500K
  • Compute (EMR / Athena / Glue / Snowflake): $500K-2M
  • Backup and snapshots: $50-100K

Total: $1.2M to $3M a year for a petabyte of working analytics data. Plus or minus depending on workload intensity.

The $23-a-month framing is off by about three orders of magnitude.

What This Changes

Several practical implications.

  • Don’t blindly trust “cheap storage means keep everything.” The storage is cheap; everything around the storage isn’t.
  • Second, the compute decision matters more than the storage decision. Query engine choice (Athena vs Snowflake vs Databricks), pricing model, and workload shape drive the bill.
  • Third, architecture decisions like “keep raw data in the lake, materialise selectively in the warehouse” have real financial consequences. Hold everything in the warehouse and you pay warehouse-tier storage forever; hold everything in the lake and you pay compute every time you query.
  • Fourth, the FinOps function for data is genuinely needed. The bill is large enough and the levers are sufficiently non-obvious that having someone whose job is to optimise it pays for itself, often within months.

Discover more from Data Lingua. Where Data Engineering Meets Agentic Business Strategy

Subscribe now to keep reading and get access to the full archive.

Continue reading