A virtual cell world model (VCWM) represents a persistent cellular state and simulates its evolution under genetic, chemical, environmental, and other biological interventions.
A VCWM seeks to model the underlying cellular system, allowing interventions to propagate through an evolving state and generate coherent molecular, structural, interactional, and morphological outcomes over time.
□ DeepSCENIC: transfer learning from sequence-to-function models enables causal gene regulatory network inference
deepSCENIC learns hierarchical TF→region→gene regulatory cascades by integrating scRNA-seq, scATAC-seq, and DNA sequence data. It combines a VAE-based multimodal framework with Enformer-derived sequence embeddings to infer cell-type-specific gene regulatory networks.
DeepSCENIC recovers TF-region (TF-RE) interactions with high fidelity, and captures de novo TF binding motifs without prior position weight matrices knowledge.
DeepSCENIC model can be used for in-silico perturbation by propagating the effect of TF expression changes or sequence variants in the input sequences. DeepSCENIC infers GRNs by finetuning S2F models or DNA language models with single-cell RNA and ATAC multiome atlases.
□ OrchAlign: Orchestrated Multimodal Alignment for Gene Expression Prediction
OrchAlign, the first multimodal alignment framework for gene expression prediction that explicitly establishes fine-grained cross-modal alignment between DNA sequences and epigenomic signals while preserving intra-modal structural organization, enabling principled multimodal fusion.
OrchAlign disentangles each modality into shared and specific representations. It then performs focused cross-modal alignment on the shared subspace within a key regulatory region, producing aligned representations with token-level correspondences.
□ Deciphering the gene-centric spatial dynamics across tissue and time with GeneSPOT
Gene SPatial Optimal Transport (GeneSPOT), a computational framework that models the relational spatial organization of genes. GeneSPOT positions genes within a context-specific Gene Constellation Embedding Space.
GeneSPOT first represents the tissue as a spatial-aware cell graph, with cells as nodes and spatial proximity encoded by edges. For each gene, its expression across cells is normalized into a probability distribution over the graph, representing its spatial expression pattern.
GeneSPOT then applies structure-preserving OT, using pairwise graph geodesic distances as transport costs, to quantify the minimum effort required to reconfigure one gene’s spatial distribution into that of another while respecting the underlying tissue structure.
□ Contrastive Poincaré Maps: uncovering developmental lineages from single-cell data with scalable hyperbolic embeddings
Contrastive Poincaré Maps (CPM) replaces the precomputed similarity matrix and manifold-metric estimation with a learned, contrastive encoder that maps raw expression directly into the Poincaré disk.
CPM processes the cells x genes count matrix in mini-batches, encoding original and corrupted views via a trainable projection replacing PCA and a hyperbolic block onto the 2D Poincaré disk. InfoNCE draws similar cells closer and pushes dissimilar cells apart in hyperbolic space.
□ DECIPHER integrates disentangled representation learning and prototype-based cell-type deconvolution across molecular modalities
DECIPHER (DECoupled Invariant Prototypes for Heterogeneous Expression Reconstruction), an end-to-end representation-learning framework for cell-type deconvolution that can be applied across multiple molecular modalities.
DECIPHER learns a domain-constant representation for deconvolution and a domain-specific representation to capture domain variation. DECIPHER estimates cell-type proportions by combining nonlinear representation learning and differentiable non-negative least-squares optimization.
□ scURL: Uncertainty-Sharpened Representation Learning for Single-cell Multi-omics Clustering
scURL, an uncertainty-sharpened representation learning framework for multi-omics clustering. It introduces representation uncertainty and establishes a connection to clustering generalization risk, enabling the derivation of cell-wise omics weights for uncertainty-aware fusion.
scURL further introduces a multi-granular calibration module (MCM) that calibrates omics-specific representations from both cluster-level semantic and local-level structural perspectives before fusion.
□ PerturbBridge: Conditional Latent Schrödinger Bridge for Single-Cell Perturbation Prediction
PerturbBridge, a conditional latent Schrödinger Bridge framework that reformulates stochastic population transport over high-dimensional, sparse gene-expression profiles as bridge learning in a compact cell latent space.
PerturbBridge learns a compact representation based on the B-TCVAE paradigm, augmented with an interpolation-consistency regularizer.
Beyond endpoint reconstruction, this regularizer encourages consistency between decoded latent interpolations and their expression-space counterparts, reducing geometric distortion during latent transport.
Recent virtual cell models, such as GEARS, CellFlow, STATE, TxPert, STACK, X-Cell, and MORPH, are trained on large-scale scRNA-seq or CRISPR-based Perturb-seq atlases, accept one or a population of cells as input, and predict the transcriptional impact of a given perturbation.
Because perturbations are experimentally assigned, these datasets can provide causal information about population-level responses within the cellular contexts they cover. However, the measurements are destructive, so they do not reveal how individual cells evolve under perturbation.
Virtual cells, therefore, often carry the population of cells as the state, without requiring an identified cell-to-cell coupling. When this predicted output population can serve as a valid starting state for further iteration, these models are aspirationally cellular world models.
The useful counterfactual is which variants we would test without AVI. Getting to causal variants with fewer experiments would be real scientific progress, even without a calibrated probability of pathogenicity.
□ DeepMind’s new genome ‘atlas’ charts effects of all nine billion human gene mutations
LucaCell, a nucleic acid sequence-centric foundation model for single-cell analysis. Departing from traditional identifier-based approaches, LucaCell represents genes through pre-trained mRNA sequence embeddings rather than relying solely on static gene annotations.
LucaCell maps each gene into a high-dimensional latent space based on sequence information rather than discrete labels, enabling sequence-based representation across species and data types without requiring manual gene mapping.
□ FORGE: Flow Orchestrated Regulatory Genomics Engine: A Configurable Nextflow Pipeline for End-to-End snMultiome Analysis
FORGE (Flow Orchestrated Regulatory Genomics Engine), a configurable workflow that automates standalone snRNA-seq and snATAC-seq analysis, integrates the pair through complementary linear and nonlinear latent-variable models, and carries them through regulatory-network inference.
A central design principle is evidence convergence: candidate regulatory relationships can be evaluated across independent layers including expression, accessibility, motif activity, footprinting, co-accessibility, domains of regulatory chromatin (DORC), and eRegulons.
□ scLDM: a conditional diffusion framework for single-cell perturbation prediction
scLDM (Single-cell Latent Diffusion Model) compresses high-dimensional gene expression into a compact latent space via a variational autoencoder, followed by a conditional diffusion process to generate post-perturbation states.
scLDM utilizes optimal transport (OT) to infer trajectory couplings, constructing a paired training dataset that links control states to their perturbed counterparts in the latent space.
scLDM leverages these OT-aligned pairs to train a conditional diffusion model that predicts the post-perturbation latent representation. This generation process is guided by the pre-perturbation latent state and specific conditions, such as cell type and perturbation type.
scLDM, a novel fully transformer-based VAE architecture for exchangeable data that uses a single set of fixed-size, permutation-invariant latent variables.
scLDM utilizes a Multi-head Cross-Attention Block (MCAB) that serves two purposes: It acts as a permutation-invariant pooling operator in the encoder, and functions as a permutation-equivariant unpooling operator in the decoder.
scLDM replaces the standard Gaussian prior with a latent diffusion model trained with the flow matching loss and linear interpolants using the Scalable Interpolant Transformers formulation, and a denoiser parameterized by Diffusion Transformers.
□ FlowMap: Geometry Dynamics Consistent Embedding of RNA Velocity for Interpretable Cellular Trajectories
FlowMap simultaneously reconstructs a smooth low-dimensional manifold of gene expression and constrains RNA velocity to follow its local geometry, producing coherent and denoised representations of cellular trajectories.
FlowMap recovers interpretable dynamical patterns, including continuous progressions, branching events, cyclic behaviors, and stable-like states.
FlowMap proceeds through four steps: phase-distance construction, spline-based manifold fitting, tangent-space velocity projection, and Jacobian-based velocity mapping.
FlowMap projects noisy velocity vectors onto the local tangent space of the learned manifold, producing geometry-consistent embedded dynamics. FlowMap constructs a velocity-aware embedding and refines it into a smooth manifold representation with coherent vector-field structre.
□ Pop-Corn: Predicting Perturbation Phenotype Effects Across Single-Cell and Spatial Contexts
Pop-Corn reformulates perturbation prediction as the direct prediction of cell-type composition. It aggregates control single-cell states with perturbation representations to predict population composition under unseen perturbations, without reconstructing gene expression.
Pop-Corn addresses dissociated screens, in which tissue architecture is lost. In intact tissues, however, perturbation phenotypes may also depend on the local cellular context.
Pop-Corn enables context-aware prediction of niche composition and using attention to generate hypotheses about neighbourhood-associated perturbation effects. It provides a direct route from control-cell profiles and perturbation identity to predicted cell-state proportions.
□ SpiderNet: A meta-interaction basis for cell-cell communication in tissues
SpiderNet learns a compact basis of cell-cell meta-interactions (MIs) from ST data. Each MI is a recurrent, directed communication module linking sender-side regulators to receiver-side targets through LR pairs, with activity quantified over ordered neighboring cell pairs.
SpiderNet learns this basis by reconstructing cell-pair LR co-expression and gene expression, so these latent MI activities explain intercellular LR signaling and sender- and receiver-side transcriptional variation, as a separate component captures cell-type-intrinsic expression.
□ AtlasOT - The Fused Unbalanced Gromov-Wasserstein for Multimodal Integration of Disease Atlases
AtlasOT (Atlas alignment with Optimal Transport), an optimal-transport framework based on the fused unbalanced Gromov-Wasserstein (FUGW) formulation for multi-modal integration.
AtlasOT jointly models a shared cross-modality feature space and modality-specific geometric structures while restricting transport to within-sample cell pairs. It relaxes strict mass conservation of the balanced OT-optimization to accommodate unbalanced cell numbers.
□ SPECTRA: predicting cellular perturbation responses with Graph Learning over Gene Regulatory Networks
SPECTRA — SPEctral CRISPR Transcriptome Regulatory Autoencoder — a graph-based model that treats CRISPR perturbations as interpretable, localized signals injected into a Gene Regulatory Network and propagated through directed graph neural networks.
SPECTRA combines a variational control-cell encoder and a directed heterophilic graph decoder, integrating single-cell expression, pretrained transcriptomic context, and prior regulatory topology into a single node-level representation.
□ stEDGE enables edge-guided multiscale reconstruction of hierarchical spatial domains and transition interfaces in spatial transcriptomics
stEDGE is the first unified framework in which local boundary structure directly guides fine-grained spatial domain reconstruction and is carried forward into subsequent multiscale hierarchical organization.
stEDGE first estimates local boundary signals and uses them to guide boundary-constrained expansion from stable low-boundary cores.
stEDGE employs a domain transition index (DTI) that quantifies domain-level transition propensity and characterizes the continuum from stable compartments to transition-prone or structurally mixed spatial states.
□ scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository
scBase Count is the first comprehensive single-cell database built by directly mining all publicly accessible 10x Genomics scRNA-seq data from the SRA and applying a standardized processing pipeline to improve data harmonization.
□ MESIC: Mapping Gene Expression to an Interpretable Semantic Space
A semantic gene embedding has dimensions with no direct biological meaning. MESIC (Mapping Expression to Semantic space with Interpretable Components) makes the dimensions themselves interpretable.
MESIC compresses the semantic gene embeddings into a small number of components, each concentrated on a small set of genes, so a component is defined by those genes and explained by their summaries and annotations.
□ DELPHAI predicts heterogeneous perturbation responses with learned single-cell fitness
DELPHAI (Deep Explainable Perturbation Heterogeneity-Aware Inference) jointly trains a fitness network and an optimal transport network without biological priors and applies fitness-gated transport with direct gene-space retrieval during inference.
DELPHAI ranks first in predicting differentially expressed genes, while revealing which cell lineages a perturbation depletes. First, normalised gene-expression matrices are reduced to a low-dimensional embedding by PCA.
□ VCCV: conservative transcriptomic corroboration for measurement prioritization of computational drug–target hypotheses
Virtual-to-Cellular Corroboration for Validation (VCCV), a model-agnostic posterior-triage layer for pre-trained DTI models. VCCV updates calibrated DTI working weights with context-aligned perturbational evidence using a near-identity affine map and exact covariance transport.
VCCV introduces an empirical warning branch for abstention and converts residual ambiguity into compact follow-up gene panels, using a submodular objective for deep near-ties. Each query is assigned one of three actionable states: a target-resolved hypothesis, an abstention, or a prioritized panel.
□ scMGFGRN: Multi-model biological and sequence information fusion for gene regulatory network inference from single-cell transcriptomics
scMGFGRN, a multi-model deep learning framework that integrates single-cell transcriptomic profiles with gene functional hierarchical relation-ships, gene sequences by leveraging denoising auto-encoders, graph attention feature extraction and pertained DNA language model.
scMGFGRN captures multi-source dependencies within multi-model biological knowledge, while its gated multi-head attention module effectively identifies informative regulatory signatures and integrate complementary features from different sources to predict accurate gene regulatory networks.
□ EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models
EvSpark, a speculative decoding system for Evo2 that verifies draft blocks in parallel and restores all classes of inference state by selecting retained intermediate states, without replay. A Hidden-state-conditioned drafter proposes each block in one parallel forward pass.
Frozen Evo2 supplies intermediate features and offline teacher supervision; the 7billion configuration extracts layer 27. A projected window of up to 64 target state forms the K/V prefix of each of two causal drafter layers.
□ UniWave-2: A Hybrid Model for Nucleic Acid Waveform Feature Extraction Enhanced by Fourier and Wavelet Transforms
UniWave-2 incorporates three key advancements, integrating hydrophobicity into the encoding scheme and enhancing waveform resolution through Fourier transform, coupled with wavelet transform to precisely capture both local details and global periodic patterns.
UniWave-2 introduces a lightweight multi-scale GRU model that facilitates cross-dimensional feature interactions, captures bidirectional temporal dependencies, and integrates long- and short-range patterns with spatial positional awareness, thereby yielding rich feature representations.
□ RNASeek: A Cross-Phyla Generative Foundation Model for Multipurpose RNA Modeling and Reinforcement Learning-Based Design
RNASeek, a 1.6-billion-parameter generative foundation model built on a DeepSeek architecture and trained on a cross-phyla transcriptomic corpus for RNA sequence representation and generation.
RNASeek captures species-specific transcript features and intron-exon boundaries in a zero-shot setting. Natural-language tokens enable flexible conditional prediction and sequence design using a unified backbone.
RNASeek uses a unified token vocabulary spanning RNA nucleotide sequences, genomic annotation markup, RNA secondary structure representations, and ~12,000 natural language tokens, providing a unified representation for sequence, structure, annotation, and task specifications.
□ scACORN: Context-engineered agent orchestration of specialized small language models for single-cell transcriptomic interpretation
scACORN is an agentic specialist-orchestrator that offers an alternative to monolithic single-cell language models. It combines specialized small language models with context-engineered agent orchestration for their selection and composition at inference time.
scACORN builds each expert in two stages: domain-aligned contrastive adaptation fits a pretrained cell-to-text backbone to the transcriptomic geometry of a dataset, and geometry-preserving specialization learns question-conditioned biological completions without eroding geometry.
□ IECP: iterative equilibration of cell-type expression profiles improves accuracy of reference-free deconvolution
GEPMC-Loc, a dynamic gated ensemble framework that combines pretrained language models with attention-enhanced multi-scale convolutional feature learning. GEPMC-Loc employs three parallel feature extraction branches.
□ GTVelo: RNA velocity inference based on graph transformer
GTVelo, a graph transformer-based neural ordinary differential equation model with three key innovations. GTVelo introduces a dual-branch graph transformer architecture to encode spliced and unspliced RNAs into a latent space.
GTVelo incorporates a multi-origin mechanism helps mitigate the constraint of a single initial state and supports more flexible inference of multi-branch trajectories.
GTVelo also supports the integration of multi-omics data, comprehensively characterizing cellular dynamic mechanisms from multiple dimensions, providing an enhanced analytical tool for single-cell dynamics research.
□ PHACTn enables training-free, context-independent inference of nucleotide variant tolerance across the genome
PHACTn (Phylogeny-Aware Computing of Tolerance for nucleotide variants) is a training-free, parameter-minimal method that infers nucleotide variant tolerability by traversing the mammalian phylogenetic tree.
PHACTn allocates a weight to each node that is inversely proportional to its phylogenetic distance from the query sequence.
□ SEAHORSE: A Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments
SEAHORSE (Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments), a discovery engine that exhaustively precomputes all pairwise associations across heterogeneous data types and presents them as a searchable association landscape.
□ LIGER2: Scalable Single-Cell Integration with On-Disk Datasets
LIGER (Linked Inference of Genomic Experimental Relationships), introduced an integrative non-negative matrix factorization (iNMF) framework to find metagene loading in cells, shared gene loading in metagenes, and dataset-specific gene loading in metagenes.
LIGER2 is a highly-optimized parallel factorization solution with on-demand loading from disk. It uses a new downstream embedding alignment method significantly improved performance in conserving biological variation while still aligning corresponding cell types across datasets.
□ PARNET: A CLIP-SEQ-BASED FOUNDATION MODEL FOR RNA SEQUENCE REPRESENTATION LEARNING
Parnet is a multi-task foundation model trained end-to-end on 223 eCLIP-seq experiments spanning 150 RPs to predict base-resolution RBP binding profiles directly from RNA sequence.
This CLIP-seq pretraining strategy departs fundamentally from the masked-language-modeling paradigm, anchoring learned RNA representations directly in measured protein-RNA interactions rather than sequence statistics.
Frozen Parnet embeddings, without task-specific fine-tuning, match or exceed the performance of both task-specific tools, as well as larger self-supervised RNA and genomic language models across diverse downstream tasks.
□ dicast: a machine learning method for accurate structural variant detection from short-read sequencing data
dicast, a machine-learning method that scores SV calls from short-read data using alignment and genomic-context features.
dicast is trained on a new multi-technology ground truth built from nine samples, with extensive manual curation. It outperforms existing short-read callers and consensus approaches, recovering substantially more true positives at high precision.
□ POME: Graph-based embeddings for partially observed mixed-type data
POME (partially observed mixed-type data embeddings) learns variable semantics based on a graph representation of the input data that enables the joint modeling of variables with heterogeneous data types and encodes missing values as absent edges within the graph structure.
POME is a versatile embedding method for POM data which relies on the key insight that partially observed discrete data can be represented very naturally as a non-complete bipartite graph between samples and variable values.
□ Signature Distance: Generalizing Energy Statistics
Signature Distance (SD), a statistical distance that compares empirical distributions through the mean absolute difference of their sorted pointwise distance profiles. SD is a structural generalization of energy distance and matches its quadratic pairwise-distance cost.
SD detects density changes with greater sensitivity than energy distance in the tested scale-contraction scenarios. Per-point mean-distance and signature-profile landscapes reveal the geometric mechanisms behind their different penalties.
Linearly interpolated biological samples that receive no increased penalty from energy distance are penalized by SD. It provides a direct differentiable potential energy for model-free Langevin data expansion, with a bootstrap resampling protocol to assess the stopping epoch.
□ DiffGSP: reversing mRNA diffusion to unlock high-fidelity spatial transcriptomics
DiffGSP, a physics-informed framework that integrates Fick's law with graph signal processing to explicitly model diffusion-induced distortions and recover the underlying spatial gene expression landscape.
DiffSP integrates inverse diffusion reconstruction with spectral filtering across both spatial and gene domains in a unified graph spectral framework, enabling joint recovery of pre-diffusion expression signals and suppression of technical noise, thereby improving stability.
□ AthenaBIO: BUILDING A TRANSLATIONAL RESEARCH ORGANIZATION TO DECODE FEMALE BIOLOGY IN THE AGE OF AI
AutoRNA, a variational autoencoder-based model that learns sequence-conditioned structural priors in the form of inter-nucleotide distance matrices. The model was trained on RNA structures obtained from the Protein Data Bank and restricted to sequences up to 64 nucleotides.
Predicted distance matrices were converted into 3D coordinates using multidimensional scaling, followed by template-based assembly and molecular dynamics refinement.
□ TriGer: Structured cross-omics interaction discovery with a triple-graph model
TriGer (Triple-Graph Multi-omics Interaction Model), an integrative framework that treats the primary unit of signal as a structured cross-omics module rather than an isolated association.
TriGer starts from the premise that biologically meaningful interactions are often expressed as coordinated many-to-many relationships, in which groups of features in one molecular layer align with groups in another while also retaining coherent organization within each layer.
TriGer combines evidence from cross-layer association and within-layer dependency in a single graph-based representation, allowing weak but concerted signals to accumulate at the module level rather than being evaluated only as disconnected pairs.
□ Delphy: Scalable near-real-time Bayesian phylogenetics for outbreaks
Delphy, an exact reformulation of Bayesian phylogenetics designed to transform its speed, scalability and accessibility while retaining Bayesian state-of-the-art accuracy.
Delphy's central data structure, an explicit mutation-annotated tree, takes advantage of the high sequence similarity of large-scale epidemic datasets for efficient tree exploration and convergence.
Delphy samples models from the posterior distribution using a standard Markov-chain Monte Carlo (MCMC) algorithm, generating a sequence of correlated samples by repeatedly proposing small changes, or moves, that are probabilistically accepted.
□ Cell-Lens: Advancing LLM Reasoning at Single-Cell Resolution
Cell-Lens, a training-free structured representation of the measured cell. It writes the cell as typed blocks, from paired measurements to marker-supported programs.
CellSpectrum, a benchmark spanning eight sources, seven tasks, and paired modalities. It extends evaluation beyond the limited modalities and task types covered by existing benchmarks.
□ Multimodal profiling of atherosclerosis: Protocol and pilot data for the AtherOMICS biobank
AtherOMICS, a bio-banking study that integrates deep phenotyping of human atherosclerotic plaques with peripheral blood biobanking, multimodal in vivo noninvasive imaging, and clinical data collection.
□ RosaSeed: Faster and Accurate Short Read Alignment Using a Configurable Seeding Strategy
RosaSeed replaces the BWA-MEM2 seeding kernel and directly supplies candidate seeds to the existing BWA-MEM2 chaining and alignment extension pipeline, producing standard SAM output suitable for downstream variant analysis.
RosaSeed accelerates seed generation using an s-base reference genome, a compressed occurrence array, and a k-mer jump table that finds all alignment positions for a short read segment that contains exactly k nucleotides.
□ RecAlign: Faster and accurate recombination-aware sequence-to-graph aligner
RecAlign is a sequence-to-graph aligner written in Rust. Differently from most aligners, RecAlign is an exact approach that exploits the A* algorithm for computing an optimal alignment between a string and a variation graph.
RecAlign can allow multiple recombinations in the alignment in a controlled (i.e., non heuristic) way - in other words, it can perform optimal alignment to path not included in the input graphs.
A great reminder that there is no single cellular clock. Our sampling intervals determine which processes we can resolve and which get blurred into an apparently stable state.
□ DisenTE: Sparse Pattern-Context Modeling for Interpretable Translation-Efficiency Matrix Completion
DisenTE, a sequence-conditioned neural model that combines separate sequence and context branches. DisenTE uses separate sequence and context branches together with a sparse low-rank pattern-context channel.
The channel is organized as Co-Translational Modules (CTMs). Each CTM combines a sequence-derived activation, a shared effect direction, and a per-context deployment weight. The resulting dictionary links sequence-derived activations to their fitted context weights.
□ EpiZoo: a DNA sequence-aware foundation model for cross-species single-cell epigenomics
EpiZoo converts million-dimensional single-cell epigenomic profiles from diverse species into compact cell sentences that integrate DNA-encoded regulatory information, sequence-independent epigenomic context and accessibility-based importance.
Built around a mixture-of-experts transformer and containing 2.6 billion parameters in total, EpiZoo is pretrained on our manually curated multi-species Omni-scATAC corpus of approximately 20.9 million cells to learn regulatory programs across species.
□ EpiZoo: a DNA sequence-aware foundation model for cross-species single-cell epigenomics
EpiZoo converts million-dimensional single-cell epigenomic profiles from diverse species into compact cell sentences that integrate DNA-encoded regulatory information, sequence-independent epigenomic context and accessibility-based importance.
Built around a mixture-of-experts transformer and containing 2.6 billion parameters in total, EpiZoo is pretrained on our manually curated multi-species Omni-scATAC corpus of approximately 20.9 million cells to learn regulatory programs across species.