Computational fluid dynamics data catalogue

Open CFD datasets for engineering machine-learning research

The catalogue documents four simulation datasets covering automotive external aerodynamics and high-lift aircraft flows. Each record summarises the geometry, numerical method, available data products, licence and associated publication.

  • 4datasets in the catalogue
  • 2application domains
  • Up to 500Mcells in an individual case
Abstract computational fluid dynamics streamlines in blue, teal and amber
AutomotiveAhmed · Windsor · DrivAer
AerospaceHiLiftAeroML · 1,800 cases

Catalogue

Datasets and simulation scope

The collection ranges from parameterised bluff bodies to a complete high-lift aircraft configuration, with different numerical methods and data volumes.

Ahmed body diagram showing the separated vortical structures represented in AhmedML Automotive

Ahmed body

AhmedML

AhmedML provides 500 parametric Ahmed-body CFD cases and official train, validation and test splits for data-efficiency and out-of-distribution evaluation.

Scale
500 geometries
Method
Hybrid RANS–LES
Size
About 2 TB
Parametric Windsor body geometry used to create the WindsorML dataset Automotive

Windsor body

WindsorML

WindsorML provides 355 Windsor-body WMLES cases and official train, validation and test splits for data-efficiency and out-of-distribution evaluation.

Scale
355 geometries
Method
WMLES
Size
About 8 TB
DrivAer vehicle parameters varied across the DrivAerML design space Automotive

DrivAer notchback

DrivAerML

DrivAerML is a 500-geometry road-car CFD dataset with official train, validation and test partitions for data-efficiency and out-of-distribution evaluation.

Scale
500-geometry design space
Method
Scale-resolving CFD
Size
About 31 TB
Wall-shear-stress rendering of a HiLiftAeroML CRM-HL aircraft case Aerospace

NASA CRM-HL

HiLiftAeroML

HiLiftAeroML comprises 1,800 WMLES cases of NASA CRM-HL aircraft variants for aerodynamic machine-learning research.

Scale
1,800 cases
Method
Explicit WMLES
Size
About 66.9 TB stored

About the catalogue

Scope and documentation

This site collates dataset-level information needed to assess suitability for a research task. It complements, rather than replaces, the repository documentation and source publications linked from each record.

Data products

Records distinguish geometry, integrated coefficients, surface fields, volume fields and derived data products.

Methods and provenance

Simulation method, solver, publication, licence, contributors and documented limitations are reported together.

Access considerations

File-selection examples and dry-run commands are provided because the complete repositories range from approximately 2 TB to 66.9 TB.

Data access

Recommended sequence for initial access

Because the repositories are multi-terabyte, file lists and transfer sizes should normally be inspected before data are downloaded.

  1. 1

    Select

    Compare domain, geometry, numerical method and available data products.

  2. 2

    Estimate

    Use a selective Hugging Face command with --dry-run enabled.

  3. 3

    Transfer

    Retrieve the required tables, geometry or field data after reviewing the estimate.

Read the data access guide

Release record

Dataset releases and publications

View all updates

Official WindsorML train, validation and test splits are now available. The release defines eight deterministic benchmark regimes: an approximately 80/10/10 random baseline, three nested data-efficiency subsets, and four out-of-distribution evaluations based on geometry, drag and image-derived wake structure. Download the JSON manifest, read the methodology, or browse the complete reproducibility files on Hugging Face. Partition sizes, provenance and a usage example are included in the WindsorML dataset record.

Official AhmedML train, validation and test splits are now available. The release defines eight deterministic benchmark regimes: a random baseline, three nested data-efficiency subsets, and four out-of-distribution evaluations based on geometry, drag and image-derived wake structure. Download the JSON manifest, read the methodology, or browse the complete reproducibility files on Hugging Face. Partition sizes and a usage example are included in the AhmedML dataset record.

Official DrivAerML train, validation and test splits are now available. The release provides eight deterministic benchmark regimes for a random public baseline, nested data-efficiency studies, geometry extrapolation, drag-regime extrapolation and rear-separation evaluation. Download the JSON manifest, read the methodology, or browse the complete reproducibility files on Hugging Face. The partition sizes and usage example are documented in the DrivAerML dataset record.

Catalogue maintenance

Corrections and related publications

Contact the maintainers to report a catalogue error, repository change, derived dataset or publication that uses these data.

Contact the maintainers