About Me

I am an Assistant Professor in the NLP group at IT University of Copenhagen.

I study what language models learn from their data and whether our evaluations can tell, with a focus on low-resource languages.

Data and evaluation. Pretraining corpora are filtered by classifiers that score how a page reads, not whether it is true. In an audit of FineWeb-Edu, about four percent of a random sample was misinformation (audit). In a re-evaluation of a tabular language model, its reported accuracy was explained by majority-class baselines, contaminated test sets and instruction tuning without any tabular data (ICML’26). I also study synthetic pretraining data made by rephrasing web text, in particular how the choice of rephrasing model trades off faithfulness against diversity, and what that does to models pretrained on it (study, G-Vendi).

Low-resource languages. With AI4Bharat I worked on multilingual language models, including romanization (RomanSetu, ACL’24, RomanLens, ACL’25), instruction following in Indic languages (IndicIFEval) and reasoning across languages (The Reasoning Lingua Franca, EACL’26). Current work is on Danish.

Applied work. I lead two funded projects that apply language models to maritime documents, and I build Interleaver, an open audio course for learning a language, whose content is generated and checked with language models. Both are described below.

I did my PhD at the University of Edinburgh on planning for long-document generation, which received the SICSA Best Dissertation in Scotland award, and I interned with the Summarization team at Google Research, London. Recent output is on Google Scholar and Hugging Face.

If you are interested in research collaboration, write to me at rapu@itu.dk.

I am also open to advisory work for companies interested in the work above.

Projects

Maritime Copilot: Maritime Safety Assistant via Retrieval-Augmented Generation. Principal Investigator, funded by Orient’s Fond (2026–2027). Maritime accident investigation reports, such as those from DMAIB and MAIB, are long, unstructured PDFs. This project builds a retrieval-augmented system for querying them, to extract actionable insights and promote safer maritime operations.

ACCENT: NLP for Nautical Charts and Marine Data. Principal Investigator, funded by Den Danske Maritime Fond (2026–2027). Extracts structured information from Notices to Mariners to draft nautical chart corrections for expert approval. This improves the efficiency of the chart correction process and ensures terminological consistency. The project is in collaboration with Geodatastyrelsen, the Danish Geodata Agency.

Interleaver. A course for audio-first language learning at interleaver.org. The story follows a character, Maya, through her first year in a new city, in Danish, German, Hindi, Marathi, Sanskrit and Malayalam. The first 28 weeks of a 52-week design are released. Source at github.com/ratishsp/interleaver.

News

  • 25 Jul 2026: Released Interleaver, an open course for audio-first language learning, currently covering Danish, German, Hindi, Marathi, Sanskrit and Malayalam. - Listen - Code
  • 13 Jun 2026: A study on the G-Vendi diversity metric for evaluating synthetic pretraining data. - Details
  • 30 Apr 2026: Our paper on re-examining the evaluation of Tabular Language Models was accepted to ICML 2026. - Paper
  • 24 Apr 2026: Gave a talk 'A Tale of Two Audits' at the Pioneer Centre for AI (P1), Copenhagen. - Slides
  • 3 Apr 2026: Released the FineWeb-Edu Misinformation Audit, a dataset auditing the prevalence and types of misinformation in the FineWeb-Edu pretraining corpus. - Dataset
  • 18 Mar 2026: A quality study of rephrased web data used as synthetic pretraining data. - Details
  • 25 Feb 2026: Preprint introducing IndicIFEval, a benchmark for verifiable instruction-following evaluation across 14 Indic languages. - Preprint
  • 3 Feb 2026: Preprint evaluating generalization claims in Tabular Language Models. - Preprint
  • 5 Dec 2025: Received a grant from Den Danske Maritime Fond for the ACCENT project, which develops automated methods for maritime chart correction preparation and English translation. - Funding Announcement
  • 24 Nov 2025: Received a grant from the Orient Fund for the Maritime Safety Assistant project.
  • 28 Oct 2025: Preprint on RiddleBench, a benchmark for evaluating complex reasoning and puzzle-solving in LLMs. - Preprint
  • 23 Oct 2025: Preprint on multilingual reasoning in LLMs, revealing the Lost-in-Translation failure mode when reasoning in English. - Preprint
  • 29 Aug 2025: Chimera accepted to Transactions of Machine Learning Research (TMLR). It proposes a unified state-space model for sequences, graphs, and images. - Paper
  • 12 Jun 2025: Paper on self-pretraining for genome modeling accepted to ICML 2025 Workshop on Generative AI for Biology - Paper
  • 16 May 2025: RomanLens to appear in Findings of ACL 2025 - Paper
  • 11 Feb 2025: RomanLens paper on latent romanization in multilingual LLMs - Paper
  • 25 Sep 2024: Paper on vocabulary expansion and initialization strategies for LLMs accepted to CoNLL 2024 - Paper
  • 13 Jun 2024: VerityMath paper accepted to AI4Math workshop at ICML 2024 - Paper
  • 16 May 2024: Two papers accepted to ACL: RomanSetu and a paper on Indic MT Eval - RomanSetu Preprint - Indic MT Eval Preprint
  • 25 Jan 2024: Introducing Airavata, Hindi Instruction-tuned LLM - Blog
  • 24 Jan 2024: RomanSetu for unlocking multilingual capabilities of Large Language Models via Romanization - Preprint
  • 21 Nov 2023: IndicTrans2 is accepted to Transactions of Machine Learning Research (TMLR) - Preprint
  • 13 Nov 2023: VerityMath for applying unit consistency check for math problem solving - Preprint
  • 9 Oct 2023: Two papers accepted to EMNLP. DecoMT is accepted to Main and CTQScorer to Findings - DecoMT Preprint - CTQScorer Preprint

Selected Publications

For my latest publications, please visit my Google Scholar profile.