Anany Kotawala

Hi! I’m Anany, a senior at Princeton University studying Financial Engineering, with minors in computer science, optimization, and cognitive science.
These days I spend most of my time on reinforcement learning, strategic games, and markets, and what they share: decisions made fast, under uncertainty, against other decision-makers. I love problems where reasoning and coordination either snap into place or fall apart.
Currently:
- Cooperative multi-agent RL in SocialJax: adapting distributed-systems scalability models (Universal Scalability Law) to learned-team performance across sizes 2–10, comparing how RL agents scale against human teams on the same environments.
- Connect 4 as a clean RL testbed for genuine strategic structure versus local pattern-matching: training self-play DQN and tabular Q-learning agents, then probing their value functions across 63 diagnostic positions to localize where policies and humans transition between local heuristics and global structure.
- Lucky CoT: developing a step-level faithfulness metric and diagnostic benchmark for chain-of-thought reasoning, quantifying when frontier models do the multi-step work versus when they land on the right answer by accident.
Always up for a conversation, especially with people thinking from angles I haven’t. Reach me at akotawala[at]princeton.edu.
Papers
- GENSTRAT: Toward a Science of Strategic Reasoning in Large Language ModelsVartan Shadarevian, Kia Ghods, Alex Kenich, Anany KotawalaUnder Review · 2026 [arXiv]
- NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation ModelsAnany Kotawala2nd Workshop on the Impact of Memorization on Trustworthy Foundation Models (MemFM) at ICML · 2026 [arXiv]
- Resolution Diagnostics for Paired LLM EvaluationAnany KotawalaAccepted, ICML 2026 Workshop on Hypothesis Testing · Seoul, South Korea [arXiv]
- Your Benchmark Score Is Not a Measurement: Capability Does Not Determine Cue-RobustnessAnany KotawalaAI Measurement Science (AIMS) @ COLM 2026 [OpenReview]
- What Black-Box Benchmarks Miss: Open-Ended Generation Under-Measures Memorization in Foundation ModelsAnany KotawalaScientific Understanding of Foundation Models (Sci-FM) @ COLM 2026 [OpenReview]
- Trusted by Default, Doubted When Official: Source Framing Gates Whether Context Overrides Memorized KnowledgeAnany KotawalaContext Beyond the Window (CBW) @ COLM 2026 [OpenReview]
- Arbitrage-Free Forecasts from Language Models via Coherence ProjectionAnany KotawalaForecasting as a New Frontier of Intelligence @ ICML 2026 [OpenReview]
Academic Service
- Reviewer, The Impact of Memorization on Trustworthy Foundation Models @ ICML 2026
- Reviewer, Philosophy Meets Machine Learning @ ICML 2026