Technical AI Safety — BS-MS Physics — IISER Mohali
Navraj Singh

Navraj Singh

Hi, I'm Navraj. I like maths, study foundational deep learning, and use both to make AI safer. Currently, that means tracing feature trajectories through diffusion models via Schrödinger bridges, and probing residual streams for evaluation awareness and hidden loyalties.

Ongoing Work

Ongoing Work

  • MS Thesis(IIT Kanpur): Building mechanistic tools to analyze feature formation during reverse denoising in diffusion models. I am applying the Schrödinger bridging framework to trace multi-step trajectories from noise to structured output and locate where early generation choices lock in.
  • ARENA 9.0 Fellow: Acceped in ARENA 9.0 cohort, will start the program at LISA(London Initiative for Safe AI, Shoreditch) Offices from 5th October.
  • Ongoing Project(Model Preferences or Aversion to Difficulty): We know Models have personas which manifest themselves as their values, they in turn affect model choices/outputs in almost every context. Models thus prefer to do certain tasks(for ex:coding over writing poetry etc) over the others when given a choice, I aim to investigate if this can really be attributed to model preferences/values or is it just averison to difficulty, where models just chose the task which they perceive to be easier.
Selected Work

Research

2026
arXiv

Evaluation Awareness in Large Language Models

Supervisor: Maheep Chaudhary. Mapped how evaluation awareness scales across compact architectures (Gemma 3, Phi-3, Llama-3 8B) using CoT analysis, representation probing, and Integrated Gradients attribution. Built a dual-pathway mitigation pipeline which uses prompt sanitization plus activation counter-steering, reaching a 70.58% average behavioral flip rate across 200 high-confidence prompts.

2026
LessWrong

Constitutional AI Widens Secret Loyalty of LLMs

Research Funded By: Bluedot Impact. The work started as a project for Bluedot Technical AI Safety Project Sprint, after getting the initial traction, Bluedot Impact graciously supported my attempts to generalise the results for larger model.Recent work has shown that it is possible to instill Secret Loyalties in models, which trigger them to output responses favoring a certain principal(ex: a politician, a corporate entity etc) whom they are loyal to. We show that Constitutional AI (CAI) can widen a model’s Narrow Loyalty (i.e. trigerring in very specific contexts) making it adaptable to conversational context. We also report that a CAI fine-tuned model’s ability to dodge Black-Box audits remains at par with the Narrow model.

2026
Apart Research

VEIL: Mechanistic Upstream Guardrails for Generative Biology

Trained an L1-regularized linear probe on ProteinMPNN decoder latents to flag structural markers of pathogenic risk, isolating causal features independent of sequence length. Surfaced polysemanticity where safety vectors conflated structural threat with structural rigidity, validated the working on out-of-distribution molecules like Ricin.

2025
Independent

Simplicity Bias in CNNs

Investigated simplicity bias via custom GradCAM and PGD attacks, mechanistically proving ”lazy” models index on localized color while robust models learn global geometric curvature.Trained a Sparse Autoencoder (SAE) on activations to disentangle features.

2025
Independent

Replication Studies & Data Attribution

Replicated foundational mechanistic interpretability findings focusing on circuit discovery and linear representation hypotheses, adversarial training and simplicity bias.

Other Technical Work

Also worked on

Fully Homomorphic Encryption in AI (CKKS)
Encrypted neural network inference, noise budget management
Valid Circuit Detection, Hexa-board PCB
SSD object detection, ResNet-50 / ResNet-18
Computational Option Pricing
C++ stochastic algorithms, continuous-time pricing
Grain Size Distribution via Mask-RCNN
Drone telemetry segmentation
Get in Touch

Contact

Reach out about interpretability research, fellowships, or collaboration.