Josué Ortega Caro

Swartz Postdoctoral Fellow · Department of Neuroscience, Yale University · Cardin Lab

I develop machine learning methods to understand how neural circuits give rise to perception and learning: generative models that discover latent regimes in longitudinal recordings, interpretable transformers for cortex-wide imaging, and foundation models for brain activity.

Portrait of Josué Ortega Caro

About

I am a Swartz Postdoctoral Fellow in the Department of Neuroscience at Yale University, working with Jess Cardin. From 2022 to 2025 I was a Wu Tsai Postdoctoral Fellow at Yale’s Wu Tsai Institute, co-mentored by Jess Cardin and David van Dijk.

I received my PhD in Quantitative and Computational Biosciences from Baylor College of Medicine, advised by Ankit Patel, and my B.Sc. in Biology from Universidad Peruana Cayetano Heredia in Lima, Peru. My work combines deep learning, dynamical systems, and neuroscience to build models that predict neural activity and are interpretable enough to generate hypotheses about how the brain computes.

I also mentor undergraduate researchers through the Research Experience for Peruvian Undergraduates (REPU) program.

News

Research

Understanding brain computation through machine learning

Three connected threads: modeling how neural dynamics change with learning, building predictive models of brain and biological data, and understanding the inductive biases that shape perception in brains and machines.

02

Predictive models of brain and biological dynamics

Self-supervised foundation models and operator-learning architectures that learn from large collections of recordings and transfer to new subjects, tasks, and modalities.

03

Perception, recurrence & inductive bias

Which architectural choices make networks robust and brain-like? I study how locality, convolution, and recurrence shape the information that vision models rely on.

Projects

Selected projects

Project pages include interactive schematics of each method, results, code, and citation details.

FLUX overview: unpaired snapshots from a dynamical system, a router over mixture-of-experts velocity fields, and three regimes discovered on a learned manifold.
NeurIPS 2026

FLUX: Longitudinal flow matching with mixture of experts

Flow matching for unpaired longitudinal snapshots along a learned data geometry. A mixture-of-experts velocity field routes each transport step to one expert, discovering regime switches without labels in Lorenz dynamics, cortex-wide calcium imaging across learning, and embryoid-body differentiation.

Explore the project
Cortex-wide calcium (jRCaMP1b) and acetylcholine (ACh3.0) maps for hit and miss trials early and late in learning.
bioRxiv 2025

PRISMt: Structured motifs in multimodal transformers

Maps a transformer’s predictions back onto region, time, and modality and decomposes them into motifs faithful to what the model uses. In cortical calcium and acetylcholine imaging, cholinergic motifs reorganized toward frontal cortex with learning.

Explore the project
BrainLM overview: fMRI recordings from 41,986 individuals train a foundation model used for zero-shot inference and fine-tuning tasks.
ICLR 2024

BrainLM: A foundation model for brain activity

Trained on 6,700 hours of fMRI with self-supervised masked prediction. BrainLM predicts clinical variables, forecasts future brain states, and recovers functional networks zero-shot.

Explore the project
Neural integral equations: an integral-equation solver iterated to convergence, trained against data, with attention-based (ANIE) or Monte Carlo (NIE) integration.
Nature Machine Intelligence 2024

Neural integral equations

Learning unknown integral operators from data: a principled framework for modeling spatiotemporal dynamics in physical and biological systems, with an attention-based solver (ANIE).

Read the paper
Neural manifold matching between model features and neural responses to natural images.
PLOS Comp. Biol. 2023Frontiers 2024

Robustness & frequency bias in vision models

Robust object-recognition models rely on low-frequency information, and localized convolutions create an implicit bias toward high-frequency adversarial examples.

Read the paper
Publications

Selected publications

All publications