Selected Publications

See Google Scholar for a full list.

2026

  1. Test-Time Scaling Makes Overtraining Compute-Optimal
    Nicholas Roberts, Sungjun Cho, Zhiqi Gao, and 7 more authors
    arXiv, 2026
  2. Harbor: A framework for evaluating and optimizing agents and models in container environments
    Harbor Framework Team
    In preparation, 2026
  3. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
    The Terminal-Bench Team
    ICLR, 2026
  4. A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
    Xavier Gonzalez*, E. Kelly Buchanan*, Hyun Dong Lee, and 5 more authors
    TMLR, 2026

2025

  1. Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification
    Jon Saad-Falcon*, E. Kelly Buchanan*, Mayee F. Chen*, and 9 more authors
    NeurIPS, 2025
  2. Batik: Behavior discovery, interpretation and annotation directly from raw video using video LLMs
    Aditya Nair, Rohan Kolhe, Nestor Coria, and 8 more authors
    Under Review, Nature Methods, 2025
  3. Extracting task-relevant preserved dynamics from contrastive aligned neural recordings
    Yiqi Jiang, Kaiwen Sheng, Yujia Gao, and 8 more authors
    NeurIPS, 2025 (Spotlight Presentation)
  4. Brain-wide representations of prior information in mouse decision-making
    Charles Findling, Félix Hubert, and International Brain Laboratory
    Nature, 2025
  5. Reproducibility of in vivo electrophysiological measurements in mice
    International Brain Laboratory
    eLife, 2025
  6. An Architecture Search Framework for Inference-Time Techniques
    Jon Saad-Falcon, Adrian Gamarra Lafuente, Shlok Natarajan, and 8 more authors
    ICML, 2025

2024

  1. Pathologies of Predictive Diversity in Deep Ensembles
    Taiga Abe, E. Kelly Buchanan, Geoff Pleiss, and 1 more author
    TMLR, 2024 (Featured Certification)
  2. Brain-to-Text Benchmark ’24: Lessons Learned
    Francis R. Willett, Jingyuan Li, Trung Le, and 13 more authors
    arXiv, 2024

2023

  1. The Effects of Ensembling on Long-Tailed Data
    E. Kelly Buchanan, Geoff Pleiss, Yixin Wang, and 1 more author
    Heavy Tails in ML Workshop, NeurIPS, 2023

2022

  1. Deep Ensembles Work, But Are They Necessary?
    Taiga Abe*, E. Kelly Buchanan*, Geoff Pleiss, and 2 more authors
    NeurIPS, 2022
  2. The Best Deep Ensembles Sacrifice Predictive Diversity
    Taiga Abe*, E. Kelly Buchanan*, Geoff Pleiss, and 1 more author
    NeurIPS I Can’t Believe It’s Not Better Workshop, 2022 (Entropic Award for Most Surprising Negative Result)
  3. Neuroscience Cloud Analysis As a Service: An open-source platform for scalable, reproducible data analysis
    Taiga Abe, Ian Kinsella, Shreya Saxena, and 8 more authors
    Neuron, 2022
  4. Plex: Towards Reliability using Pretrained Large Model Extensions
    Dustin Tran, Jeremiah Liu, Michael W. Dusenberry, and 23 more authors
    ICML Pre-training Workshop, 2022 (Contributed Talk, 5.9% of accepted papers)

2020

  1. Deep Graph Pose: a semi-supervised deep graphical model for improved animal pose tracking
    Anqi Wu*, E. Kelly Buchanan*, Matthew Whiteway, and 14 more authors
    NeurIPS, 2020

2018

  1. Penalized matrix decomposition for denoising, compression, and improved demixing of functional imaging data
    E. Kelly Buchanan*, Ian Kinsella*, Ding Zhou*, and 8 more authors
    BioRxiv, 2018