Ongoing Work
- MS Thesis(IIT Kanpur): Building mechanistic tools to analyze feature formation during reverse denoising in diffusion models. I am applying the Schrödinger bridging framework to trace multi-step trajectories from noise to structured output and locate where early generation choices lock in.
- ARENA 9.0 Fellow: Acceped in ARENA 9.0 cohort, will start the program at LISA(London Initiative for Safe AI, Shoreditch) Offices from 5th October.
- Ongoing Project(Model Preferences or Aversion to Difficulty): We know Models have personas which manifest themselves as their values, they in turn affect model choices/outputs in almost every context. Models thus prefer to do certain tasks(for ex:coding over writing poetry etc) over the others when given a choice, I aim to investigate if this can really be attributed to model preferences/values or is it just averison to difficulty, where models just chose the task which they perceive to be easier.
Research
Evaluation Awareness in Large Language Models
Supervisor: Maheep Chaudhary. Mapped how evaluation awareness scales across compact architectures (Gemma 3, Phi-3, Llama-3 8B) using CoT analysis, representation probing, and Integrated Gradients attribution. Built a dual-pathway mitigation pipeline which uses prompt sanitization plus activation counter-steering, reaching a 70.58% average behavioral flip rate across 200 high-confidence prompts.
Constitutional AI Widens Secret Loyalty of LLMs
Research Funded By: Bluedot Impact. The work started as a project for Bluedot Technical AI Safety Project Sprint, after getting the initial traction, Bluedot Impact graciously supported my attempts to generalise the results for larger model.Recent work has shown that it is possible to instill Secret Loyalties in models, which trigger them to output responses favoring a certain principal(ex: a politician, a corporate entity etc) whom they are loyal to. We show that Constitutional AI (CAI) can widen a model’s Narrow Loyalty (i.e. trigerring in very specific contexts) making it adaptable to conversational context. We also report that a CAI fine-tuned model’s ability to dodge Black-Box audits remains at par with the Narrow model.
VEIL: Mechanistic Upstream Guardrails for Generative Biology
Trained an L1-regularized linear probe on ProteinMPNN decoder latents to flag structural markers of pathogenic risk, isolating causal features independent of sequence length. Surfaced polysemanticity where safety vectors conflated structural threat with structural rigidity, validated the working on out-of-distribution molecules like Ricin.
Investigated simplicity bias via custom GradCAM and PGD attacks, mechanistically proving ”lazy” models index on localized color while robust models learn global geometric curvature.Trained a Sparse Autoencoder (SAE) on activations to disentangle features.
Replication Studies & Data Attribution
Replicated foundational mechanistic interpretability findings focusing on circuit discovery and linear representation hypotheses, adversarial training and simplicity bias.
Also worked on
Encrypted neural network inference, noise budget management
SSD object detection, ResNet-50 / ResNet-18
C++ stochastic algorithms, continuous-time pricing
Drone telemetry segmentation
Contact
Reach out about interpretability research, fellowships, or collaboration.