AK

Research Engineer — AI Evaluations & Agentic Safety

I design empirical evaluations for agentic AI systems and study behavioural measurement, adversarial failure modes, memory integrity, and evaluation robustness.

My work follows a practical research arc: systems engineering → agent memory → adversarial benchmarking → evaluation methodology → agentic AI safety. I write about experiments, measurement choices, failure analysis, and the evidence needed to make claims about agent behaviour.

arXiv preprint · 2026 Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking An empirical study of memory poisoning, content screening, provenance-weighted retrieval, and evaluation robustness in the reported LongMemEval setup. Read the publication →

Engineering evidence for agentic AI

My systems-engineering background supplies the reliability discipline; work on Aegis Memory supplies the experimental substrate. Quantify Labs is the organisational provenance behind that work—not the centre of the research identity.