Skip to content

Blog

AWS Clean Rooms Synthetic Data (re:Invent 2025)

AWS Clean Rooms ML now generates synthetic data for privacy-safe model training and cross-cloud collaboration. Here's how it works.

Forged Concepts
  • aws
  • application-modernization
  • aws-msp
  • cloud-strategy
  • cost-optimization
  • devops
AWS re:Invent 2025 blog post cover image

Two companies want to train a model on their combined customer data. Both have lawyers. Neither can hand over raw records. That standoff has killed more joint analytics projects than any technical limitation, and at re:Invent 2025 AWS shipped something aimed squarely at it.

AWS Clean Rooms ML can now generate privacy-enhancing synthetic datasets for machine learning. Organizations and their partners get data that preserves the statistical characteristics of the original, then train regression and classification models on it, all without exposing real, identifiable records. The use cases that used to die in legal review are back on the table.

For context, re:Invent is Amazon’s largest annual cloud conference, where AWS announces new services and sets the direction for the year ahead. This one came out of that batch.

What AWS Announced

Clean Rooms ML now creates synthetic training datasets that mimic the patterns and distributions of real data. The key detail is what the training code can actually see: only the synthetic version, never the original records. That cuts the risk of a model memorizing or leaking sensitive information, and it makes collaborative ML across organizational boundaries a lot less terrifying for everyone who has to sign off on it.

Partners can jointly generate ML-ready datasets for uses such as:

  • Campaign optimization
  • Fraud detection
  • Medical or scientific research
  • Cross-brand customer analytics
  • Joint product or promotion planning

AWS gave the example of an airline and a hotel brand that want to sharpen customer targeting without trading raw data. Synthetic data lets the two collaborate without either one exposing sensitive consumer information.

Why Synthetic Data Matters

Keeps the data useful while protecting identities

Synthetic datasets hold onto the statistical structure of the original while de-identifying individuals. You get data that still trains a useful model, with less risk of exposing personal information and a smaller chance the model quietly memorizes specific records.

Lets teams train across organizations

Sometimes the data carries sensitive personal information. Sometimes it legally cannot leave the environment it was created in. Either way, partners can still run ML initiatives together, because the synthetic dataset becomes a safe intermediary when regulatory restrictions block direct sharing.

Opens up use cases that privacy used to block

Travel, healthcare, finance, advertising. The industries that live and die on regulated or sensitive data are exactly the ones that have been stuck. They can now build joint ML models without throwing data governance out the window.

Expanded Support for Multiple Clouds and Data Sources

Synthetic data was not the only announcement. Clean Rooms now supports Snowflake and Amazon Athena as data sources inside Clean Rooms collaborations.

That matters because partner data no longer has to live in Amazon S3 for the collaboration to work. There is no copying, no migration. You skip:

  • Compliance risk
  • Pipeline maintenance overhead
  • Stale or outdated datasets
  • Costs associated with data movement

In other words, collaboration with zero extract, transform, and load (zero-ETL).

Example Use Case

Picture an advertiser with data in Amazon S3 and a media publisher with data in Snowflake. They can run an audience overlap analysis without sharing raw data or standing up ETL pipelines. No source data from external locations is permanently stored in AWS Clean Rooms, and anything temporarily read during analysis is deleted once the query completes.

How Clean Rooms Uses Privacy Controls

The enforcement layer here is analysis rules, which govern how data can be queried and what outputs are allowed. Data owners decide:

  • Which columns are accessible
  • What types of queries are allowed
  • Whether additional analyses are permitted
  • Whether queries require manual review

This is the part that earns trust. The data owner keeps precise control over how their dataset gets used in a collaboration, instead of hoping a partner behaves.

Partner and Customer Feedback

Early adopters pointed to:

  • Privacy-enhanced collaboration across multiple cloud environments
  • Improved data interoperability between partners
  • Safer joint analytics and personalization workflows

Organizations including Kinective Media by United Airlines and Snowflake voiced support for the cross-cloud and privacy-preserving features.

Why This Announcement Matters

Put the two pieces together, synthetic data generation plus multi-cloud sources, and the practical wins are clear. Organizations can:

  • Use fresher and more complete datasets
  • Reduce compliance and data-handling burdens
  • Avoid moving or exposing sensitive records
  • Build more accurate joint ML models
  • Scale partnerships across diverse technology stacks

If your team has a collaborative analytics or ML idea that keeps stalling on privacy boundaries, this update changes what is feasible.

We help teams design and implement secure data collaboration and privacy-safe ML on AWS. If that is on your roadmap, contact us and we will work through the right approach with you.

Official Sources

AWS Clean Rooms Documentation: https://docs.aws.amazon.com/clean-rooms/

AWS Blog: https://aws.amazon.com/blogs/aws/aws-clean-rooms-now-supports-multiple-clouds-and-data-sources/

Ready when you are

Need senior AWS expertise without building a full internal team?

Forged Concepts helps growing companies improve AWS performance, control cloud costs, modernize infrastructure, and build with confidence. If your team needs stronger cloud architecture, better operations, or a clearer path forward on AWS, let's talk.