Virtual Embryo Challenge
Submit

One embryo resource, two modalities

Roughly a million cells across eleven developmental time points from E6.75 to E12.5. Early gastrulation through cardiac progenitor emergence, heart-tube formation, looping, and later morphogenesis.
The four parts of the resource: embryo collections from E6.75 to E12.5; paired single-cell measurements of RNA and chromatin accessibility; coronal sectioning decoded into 3D MERFISH; and conditional knockout embryos profiled the same way.

Embryo collections across eleven stages (A), single-cell measurement (B), 3D MERFISH from serial sections (C), and the conditional knock-outs (D).

Download

The training stages are released to registered participants. Validation and test targets come back as a score rather than as answers, until the final phase releases the validation answers for every task.

To look at these stages before downloading them. Cells in 3D, coloured by type, with gene expression. They are in the atlas: browse the challenge training data ↗

More data comes later

This is not all the data you will get. At the final test phase the validation stages are released too, with their answers, so they become training material rather than something you can only probe through the board. By then ranking has moved to the hidden test split, so releasing them costs nothing and gives every entrant more to train on for the round that decides the prizes. The held-out heart stages become training input at the same point.

Gene panel, per board

A submission’s var_names is compared to the board’s panel element by element: same genes, same order. The panel is not the same on every board, and it is not simply “the panel the target file was measured on”.

It is the ordered intersection of the genes present in every stage that board actually loads: the training stages visible before the target, and the target itself. Where a stage was measured on fewer genes, the intersection shrinks. In the embryo setting, Casp4 and Pnliprp1 were not measured at E6.75 or E7.25, so those two are absent from every embryo board. Including the one whose target, E7.5, does carry them. A prediction cannot be scored against a gene the reference stage never observed.

BoardGenesCellsPanel
T1:val32,2851,000-5,118T1__val.genes.txt ↓
T2:embryo:val_interp498583-5,000T2__embryo__val_interp.genes.txt ↓
T2:heart:val_extrap5001,000-25,179T2__heart__val_extrap.genes.txt ↓
T2:heart:val_interp5001,000-17,616T2__heart__val_interp.genes.txt ↓
T3:gata45001,000-7,449T3__gata4.genes.txt ↓

Each file is one gene name per line, in the order the scorer expects. Take the panel from here rather than from a data file, no released stage is guaranteed to match a board, and a submission whose genes are right but ordered differently is rejected just the same. Machine-readable index, including cell limits and the required obsm keys: index.json.

How many cells to submit

Anywhere in the range above. Cell count is not a scored quantity on any task, and there is no advantage to either end of it. The number of cells in a submission is a sample size. How much evidence you give the metrics to estimate from, and nothing else. It is not a prediction of how large an embryo is, and it is not compared to anything.

The upper bound is operational. A Task 1 answer densifies to n_cells × 32,285 float32 on the scoring host, so the cap is what keeps one submission inside its memory. It is set at three times the larger of the board's truth half and its reference stage, with a floor of 5,000. Sized so that the floor baseline, which resubmits an observed stage verbatim, can satisfy it. The lower bound is statistical: every primary metric estimates a distribution, and below it the estimate is dominated by its own sampling noise rather than by your model.

Going high buys nothing. The three heaviest metrics take their own subsample whatever you send: the unbiased MMD draws 2,000 cells, the energy distance and the variogram 1,500 each, so cells beyond roughly 2,000 are discarded before they are measured. The remaining terms are pseudobulk and proportion statistics, where more cells reduce noise slightly and change nothing systematically. A few thousand cells is a good submission at every board.

One thing that surprises people, because it makes the numbers above look far smaller than the files you downloaded: the scorer subsamples every stage to 10% before comparing anything. E9.5 as released holds 17,057 cells; the scorer's reference stage is 1,706. The target is downsampled the same way and then split in half, one half to score against and one to estimate the ceiling. So truth_cells and ref_cells in index.json are counts of our working copy, not of an embryo and not of the release. Your submission is not expected to mirror them.

What the Task 1 dissection covers

The Task 1 stages are a heart-centred dissection, not whole embryos and not isolated hearts. How much surrounding tissue comes with the heart differs by stage, because what can be dissected cleanly changes as the embryo grows. Populations that are plainly not cardiac are therefore present at some stages and absent at others: Neural Tube is 4.6% of E8.5 and 0% of E9.5, and Paraxial Mesoderm and extra-embryonic mesoderm behave the same way.

This matters for how you read the task. The metrics compare your submitted population to the target population, so the target carries both the biology and the dissection. A model is not being asked to predict our dissection protocol, but it is being scored against a sample that reflects it.

So that the observation process is visible rather than something each team reconstructs, the cell-type composition of the released stages is published: t1_composition.json. Released stages only. Nothing there describes E10.5 or E12.5.

One caution when you read it. The annotation vocabulary is not harmonised across stages: Endocardium at E9.5 appears as Endocardium-1 and Endocardium-2 later, and a label missing from a stage may be present under a different name rather than biologically absent. Comparing label sets across stages will overstate how much has changed. We are working on a harmonised mapping.

Modalities

Single-cell
Task 1
Single-cell RNA

Real dissociated cells profiled across the whole transcriptome, 32,285 genes rather than a panel. Cell-type labels ship with the training stages. This is what Task 1 predicts.

Spatial
Tasks 2 and 3
3D MERFISH, 500-gene panel

A whole heart or whole embryo is sliced continuously into serial 2D spatial-transcriptomic sections, which are then reassembled into a 3D volume. Every cell therefore carries both a 3D position and measured RNA, across several developmental stages, with cell-type labels. Coordinates are per-embryo local and are not registered across time points. The conditional knock-outs (Mab21l2, Gata4 and β-catenin, each with a matched wild type) are profiled the same way.

Task 1: single-cell RNA

Real dissociated single cells, whole transcriptome, not the MERFISH panel.

E7.75
unused
E8.5
train
E9.5
train
E10.5
val
E12.5
test
Train Validation Test Not in split
Stage / conditionRoleNotesSize
E8.5trainE8.5_RNA.h5ad571 MB
E9.5trainE9.5_RNA.h5ad590 MB
E10.5valvalidation target, answers withheld647 MB
E12.5testhidden test target360 MB
E7.75unusednot released — held out as the Task-2 embryo test stage

.h5ad · log1p-normalised .X · 32,285 genes · obs["celltype"] · no spatial coordinates

E7.75 is not released. A single-cell stage exists at E7.75, but E7.75 is the hidden test target for Task 2’s embryo interpolation, so no measured data from that stage is distributed for any task. It is shown here only so the gap in the series is accounted for. There is no E9.25 stage in this single-cell release, so training uses the two real stages that exist before the target.

File format

AnnData .h5ad, one file per stage. Entries marked tbc are fixed with the data release; the keys and dtypes below are already pinned by the loaders.

Train
E8.5_RNA.h5ad 571 MB
n_obs × 32,285tbc
Train
E9.5_RNA.h5ad 590 MB
n_obs × 32,285tbc
Validation
never distributed. Submit and the server scores it
n_obs × 32,285tbc
Test
inputs released at the final test phase, without labels
n_obs × 32,285tbc
KeyTypeWhat it holds
.Xfloat32 [n_obs, 32285]Log1p-normalised expression over the whole transcriptome. Submit sparse or dense. It changes nothing, because every metric runs on a PCA of a couple of thousand subsampled cells and the linear algebra behind that wants a dense array. The conversion happens after subsampling, so it is small.
obs["celltype"]categoricalCell-type label, released with the training stages. Never part of a submission: the scorer types every prediction with its own frozen classifier.
obs indextbcstringCell barcode. Not matched between prediction and target; the metrics are distributional, so submissions need not preserve cell identity or even cell count.
var indexstring [32285]Gene symbols in the fixed Task-1 order. A submission must carry exactly these names in exactly this order, or pass --allow-reorder and let the scorer reindex.
  • No spatial coordinates anywhere in Task 1. Dissociated cells have no position to predict.
  • Task 1 is the only task where training on external public single-cell data is allowed, provided the source is disclosed with the submission.

Task 2: 3D MERFISH (embryo setting)

The same 3D MERFISH assay applied to the whole embryo rather than the heart, over the gastrulation window.

E6.75
train
E7.25
train
E7.5
val
E7.75
test
E8
train
Train Validation Test
Stage / conditionRoleNotes
E6.75trainearliest released stage
E7.25trainlast stage before the held-out pair
E7.5valinterpolation validation, answers withheld
E7.75testinterpolation test, hidden
E8.0trainthe stage after the pair, both targets are bracketed

.h5ad · log-normalised .X · 500-gene MERFISH panel · obs["celltype"] · obsm["spatial_3D"] (per-embryo local frame, not cross-timepoint registered)

The scope follows what can be sectioned. Through gastrulation the whole embryo is small enough to slice end to end and reassemble, so these stages cover the entire embryo; by the stages in the heart setting it is far larger, and profiling is focused on the organ of interest instead. Both held-out stages here sit strictly inside the training range, E7.5 and E7.75 fall between E7.25 and E8.0, so the embryo setting is entirely interpolation, with no extrapolation axis. It is also the only setting that tests interpolation: heart trains and validates that question but has no stage left to hold back for it, so E7.75 is where interpolation is judged in the final phase. It is still the harder of the two: it spans gastrulation, where composition turns over fastest, and the stages are packed far more tightly in time than the heart series. Scored separately from the heart setting. Per-file sizes are not in the handoff document, so they are omitted here rather than guessed.

File format

AnnData .h5ad, one file per stage, per setting. Entries marked tbc are fixed with the data release; the keys and dtypes below are already pinned by the loaders.

Train
E<stage>.h5ad
n_obs × 500tbc
Validation
never distributed. Submit and the server scores it
n_obs × 500tbc
Test
inputs released at the final test phase, without labels
n_obs × 500tbc
KeyTypeWhat it holds
.Xfloat32 [n_obs, 500]Log-normalised expression on the 500-gene MERFISH panel: a measured panel, not a transcriptome. Must be finite and non-negative; a negative entry usually means the matrix was centred somewhere upstream.
obs["celltype"]categoricalCell-type label, released with the training stages. Not submitted, and not read from a submission.
obs indextbcstringCell identifier. Not matched across stages. There is no cell-level correspondence between time points in the released data.
var indexstring [500]The 500 panel genes, in panel order. Identical across every Task-2 and Task-3 file.
obsm["spatial_3D"]float32 [n_obs, 3]Per-cell x y z. The frame is per-embryo local and is not registered across time points, so no shared coordinate system holds between stages. Columns beyond the third are ignored.
  • The two settings, heart and embryo, share this schema exactly and differ only in which stages they load and which are held out.
  • A submission carries both channels: .X and obsm["spatial_3D"]. A file with expression but no coordinates is rejected before scoring.

Task 2: 3D MERFISH (heart setting)

Continuous serial 2D sections reassembled into a 3D volume, so every cell carries a position as well as its RNA.

E8.25
train
E8.5
val
E8.75
train
E9.5
train
E10.5
val
E12.5
test
Train Validation Test
Stage / conditionRoleNotesSize
E8.25trainships as E8.25_late.h5ad225 MB
E8.5valinterpolation validation, answers withheld85 MB
E8.75trainalso the Task 3 wild-type reference84 MB
E9.5train4D MERFISH release164 MB
E10.5valextrapolation validation, answers withheld484 MB
E12.5testextrapolation test, hidden629 MB

Identical to the embryo setting above.

The heart setting holds stages out on two axes at once. E8.5 sits strictly between the training stages E8.25 and E8.75, so recovering it is interpolation; E10.5 and E12.5 sit past the last training stage, so reaching them is extrapolation. The two questions are scored separately and never averaged. Heart trains and validates interpolation but never tests it: the interval E8.25-E8.75 contains just E8.5, so there is no second stage to hold back. Interpolation is tested in the embryo setting instead, on the hidden E7.75, which is where to look if you want to know how that skill will be judged. What this setting carries into the final phase is extrapolation, on the hidden E12.5. In the final phase every heart stage is training input.

The file format is the same as Task 2: 3D MERFISH (embryo setting) above: same container, same keys, same dtypes.

Task 3: conditional knock-outs

3D MERFISH from conditionally knocked-out embryos with matched wild-type controls at the same stage. Both knockouts released here are Mesp1-Cre driven, so the deletion is restricted to the mesodermal lineage: Mesp1-Cre; Gata4 F/F; Gata6 F/+ for the Gata4 condition, and Mesp1-Cre; β-catenin F/F for β-catenin.

E8.75
Gata4 KO
β-catenin KO
WT
E9.5
Mab21l2 KO
WT
Train Validation Test Reference
Stage / conditionRoleNotesSize
Mab21l2 KO @ E9.5trainclear phenotype, very specific gene458 MB
Gata4 KO @ E8.75valMesp1-Cre; Gata4 F/F; Gata6 F/+, two replicates; answers released at the final phase374 + 409 MB
β-catenin KO @ E8.75testMesp1-Cre; β-catenin F/F, two replicates, broadly expressed and the hardest300 + 381 MB
WT @ E9.5referencematched wild-type control (shared with Task 2)164 MB
WT @ E8.75referencematched wild-type control84 MB

Identical to Task 2.

Training and validation use knockouts of very specific genes. The hidden test gene is broadly expressed, and its effect correspondingly diffuse.

File format

AnnData .h5ad, one file per condition; knock-outs ship as two replicates. Entries marked tbc are fixed with the data release; the keys and dtypes below are already pinned by the loaders.

Train
Mab21l2 KO @ E9.5 458 MB
n_obs × 500tbc
Validation
Gata4 KO @ E8.75. 2 replicates 374 + 409 MB
n_obs × 500tbc
Test
inputs released at the final test phase, without labels
n_obs × 500tbc
Reference
matched WT @ E8.75 and E9.5 84 + 164 MB
n_obs × 500tbc
KeyTypeWhat it holds
.Xfloat32 [n_obs, 500]Same 500-gene panel and normalisation as Task 2. The prediction is the mutant embryo, not the difference from wild type.
obs["celltype"]categoricalAs Task 2. Released for the wild-type reference and the training knock-out.
obs["condition"]tbccategoricalGenotype of the embryo the cell came from, the knocked-out gene, or wild type.
var indexstring [500]The same 500 panel genes as Task 2, in the same order.
obsm["spatial_3D"]float32 [n_obs, 3]Per-cell x y z. The frame is per-embryo local and is not registered across time points, so no shared coordinate system holds between stages. Columns beyond the third are ignored.
  • Every differential-expression metric is computed against the matched wild type at the same stage, never against a preceding stage, so the wild-type reference is part of the input, not an extra.
  • Knock-outs ship as two biological replicates. A submission is scored against each of them, and the difference between those two scores is published alongside the result: it is the same prediction measured against two embryos of the same genotype, so it is a direct read of how much of any gap between two entrants is biology and how much is noise. Which replicate the leaderboard rank uses is fixed by the scorer, not chosen by the entrant.

Where the data comes from

The 3D MERFISH resource, the conditional knock-out embryos and the single-cell Multiome data were all generated by Qingquan Zhang and Neil Chi at UC San Diego. That covers the data behind all three tasks. The benchmark is built on that resource; the tasks, splits and scoring are the competition's.

Competition datasets are provided for research and educational purposes. Any use, reuse, publication, presentation, or redistribution of the data outside the Challenge must appropriately acknowledge the Virtual Embryo Challenge and credit Dr. Neil Chi’s group at the University of California, San Diego (UCSD), which generated the dataset. As the dataset has not yet been published, any use of the data in a publication or public presentation must receive prior approval from Dr. Neil Chi and the data-generating team. Any publication using the dataset must also appropriately credit Dr. Qingquan Zhang and Dr. Neil Chi.

Access and reproducibility

Training data is released for method development. Validation and test ground truth are withheld while they can still affect the board: a validation submission comes back as a leaderboard score, not the answers. At the start of the final phase the validation answers are released for every task. ranking has moved to the hidden test split by then, and the test leaderboard opens. Task 1 additionally permits training on external public single-cell data; only the evaluation stages are fixed, provided any external source is disclosed with the submission.

You may train on external data, public or your own, and for Task 1 this is expressly encouraged. What you may not use is measured data from a held-out stage or genotype, by any route: not directly, not through a model pre-trained on it, and not through a public dataset that contains it. Those are E10.5 and E12.5 for Task 1; E7.5 and E7.75 in the embryo setting and E8.5, E10.5 and E12.5 in the heart setting for Task 2; and the Gata4 and β-catenin knockouts at E8.75 for Task 3. Nearness counts, and the stage label is not the test: external data is treated as a held-out stage when it sits closer to that stage than to the nearest stage we released for the same task, the midpoint being the boundary. Training on an external E9.45 embryo while E9.5 is held out is using the answer under a different name, and so is a dataset labelled by somite count or Theiler stage that lands in the same window. For the Task 1 extrapolation targets the window is stated outright rather than derived, and it relaxes the midpoint rule rather than following it: external data is excluded from after E9.5 up to and including E13.5, E10.5 and E12.5 are excluded absolutely, and anything after E13.5 may be used provided its source and stages are disclosed. The two ends of that window differ because one of them is ours to publish and the other is not. Data measured at exactly E9.5 is permitted, since E9.5 is a released training stage and data at a stage everyone already holds tells nobody anything they were not given. Data measured at exactly E13.5 is not, since E13.5 is the far edge of the protected window rather than the first point outside it, and it sits one day from the test target. For interpolation targets the boundary is the released stage on either side of the held-out one rather than a midpoint: external data may not be used if it stages strictly between the two released stages that bracket a held-out one. In Task 2 that means heart data strictly between E8.25 and E8.75, and embryo data strictly between E7.25 and E8.0. The bracketing stages themselves are permitted, being released training stages. The same applies to genotype: data from the held-out knockouts, from another allele of the same gene at a comparable stage, or from a perturbation that phenocopies them, is treated as the held-out condition. Every external source must be disclosed with the submission; an undisclosed one is a violation whether or not it changed the result, because the disclosure is what makes this checkable.

Each task's loader in the starter kit reads a documented path list, and DATA_SOURCES.md records which file every task reads, its verified size, and its status. Including two dependencies of the richer reference baselines that are not yet on shared storage. The floor and the simple baseline need neither.

A related resource

This site also hosts the Virtual Embryo atlas. It is not the benchmark: the challenge is scored against the releases listed above. but it covers the same organism over the same window, and it is browsable without downloading anything. Useful for orienting yourself in mouse development, and for Task 1, where external public single-cell data may be trained on as long as it is disclosed with the submission.

With one limit that applies to every external source, this atlas included: no measured data from a held-out stage or condition may be used, not for training, not for fine-tuning, not transferred in under another name. The atlas spans the same window as the benchmark, so some of what it contains sits at target time points; using it there is using the answer, however it is routed. The held-out stages are named for each task above, so what is off limits is knowable before you start.