Vikram Natarajan

Vikram Natarajan

AI Researcher & Engineer

I'm Vikram, an engineer and AI safety researcher based in New York. I currently work at Vanta building AI agent evaluations and infrastructure for security and compliance. I used to work at Nuro on self-driving car routing features, and at Salesforce before that as a data scientist on internal ML tools.

My research focuses on interpretability and deception detection. Most recently, building more targeted deception detectors for LLMs by leveraging a human-interpretably taxonomy of deception, which led to a paper at ICML 2026.

When I'm not in front of a computer, I enjoy learning languages (I am fluent in seven), playing squash, improvising on the piano, and long walks through new cities!

One Probe Won't Catch Them All: Towards Targeted Deception Detection

Vikram Natarajan, Devina Jain, Shivam Arora, Satvik Golechha, Joseph Bloom

ICML 2026

Uses a deception taxonomy from the psychology literature to show that deception is heterogenous in LLMs and needs targeted detectors

Mechanistic Decomposition of Sentence Representations

Matthieu Tehenan, Vikram Natarajan, Jonathan Michala, Milton Lin, Juri Opitz

Preprint, 2025

Introduces a method to mechanistically decompose sentence embeddings into interpretable components via dictionary learning on token-level representations.