Research

Upcoming

RL to SFT crossover

RL to SFT Crossover in Micro-Models

In preparation

SFT vs SDFT catastrophic forgetting

SFT vs SDFT: An Analysis of Catastrophic Forgetting

In preparation

Selected Papers

Data Science Agents

Recovering Wasted Compute in Autoresearch Agents

Au Kwok Chun, Abhigyan Acherjee, Amrutha Rao, Zaiqian Chen, Kazem Meidani, C. Bayan Bruss, Micah Goldblum

Conference on Language Modeling (COLM), 2026

A slew of recent works have developed agents for solving data science tasks end-to-end. Such agents have inspired large industry investment, motivated by their potential to automate time-consuming human labor and customize machine learning solutions for specialized applications. In this paper, we identify common failure modes of these agents when applied to tabular datasets: buggy code generation, absent hyperparameter tuning, inefficient usage of compute budget, and insufficient exploration. We further explore patches for these failure modes to enable efficient and effective solution generation. In particular, we find that prompt- and control-level enhancements, a debug consultant that shares discovered runtime constraints across all branches of the search tree, and refined tree-search algorithms successfully mitigate failure modes. Our results indicate that significant gains in data science agent performance are achievable through better agentic design alone, independent of the underlying language model's capabilities.
Human Rights Benchmarking

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

Savannah Thais, Wm. Matthew Kennedy, Abhigyan Acherjee, Matilda Wysocki, Malcolm Langford, Caitlin Kraft Buchman

AI for Law Workshop, International Conference on Machine Learning (ICML), 2026

Large language models (LLMs) increasingly mediate decisions that affect what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench — the first expert-validated, scenario-based benchmark for human rights reasoning. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work (substituting P, "proposing remedies," for C, "legal conclusion," yielding IRAP) to structure our evaluation heuristics. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real-world human rights issues and annotated by human rights lawyers and professionals across the world. Ultimately, we find that frontier model accuracy scores range considerably across legal reasoning tasks (0.339–0.577), which strongly implies that HumRightsBench is a capable instrument for advancing this emerging subfield of AI evaluations science at a critical time.
Anomaly Detection in Astrophysics

Anomaly Detection in Astrophysics — VAE Separates Interacting Binary Stars from Normal Red Giants

Abhigyan Acherjee, Savannah Thais, J. L. Sokoloski

Machine Learning and the Physical Sciences Workshop, Neural Information Processing Systems (NeurIPS), 2025

Symbiotic binary stars, in which a white dwarf accretes from a red-giant companion, can be difficult to detect yet are key to understanding binary stellar evolution and supernova progenitors. We present a novel application of a variational autoencoder (VAE) applied to optical and infrared colors of red giants measured in the SkyMapper, 2MASS, and AllWISE surveys. By learning the latent structure of red giants, the VAE flags anomalies, some of which align with known symbiotic binaries, successfully recovering most above the 90th percentile in anomaly scores. This analysis demonstrates the efficacy of unsupervised anomaly detection for uncovering hidden interacting binaries — and possibly other populations of astronomical objects with small training sets — in forthcoming large surveys such as Rubin Observatory's LSST.
Sex Education Policy Adherence

From Policy to Practice: Quantifying Sex Education Policy Adherence for Improved Policy Impact Analysis

Abhigyan Acherjee, Savannah Thais, Amaya Kejriwal

Computational Social Science Society of the Americas (CSSSA), 2025

Sex and HIV education is believed to play a critical role in shaping adolescent health outcomes, yet there is vast disagreement across the United States about what should be taught and how. This has resulted in a patchwork of state-level policies that differ in scope, content, and emphasis. It is therefore critical to characterize the impact that different sex and HIV education policy designs have on student behavior. Many current causal modeling techniques rely on the assumption of ignorability, which is often unmet in real-world policy impact analysis scenarios. In particular, a key confounder for these analyses is policy implementation. In this paper, we propose a method to measure how closely schools in states adhere to their own sex and HIV education policies. Focusing on data from 2020, we leverage the biennial CDC School Health Profiles Survey, which provides insights into what curricula are being taught in schools across participating states. We then align these reported practices with policy data from the Sexuality Information and Education Council of the United States (SIECUS). By comparing state-level educational practices with existing policies, we calculate a quantitative adherence score for each state. This score addresses a critical measurement gap in policy evaluation and serves as a foundational metric for causal analyses examining the relationship between education policies and social outcomes.