I design empirical evaluations for agentic AI systems and study behavioural measurement, adversarial failure modes, memory integrity, and evaluation robustness.
My work follows a practical research arc: systems engineering → agent memory → adversarial benchmarking → evaluation methodology → agentic AI safety. I write about experiments, measurement choices, failure analysis, and the evidence needed to make claims about agent behaviour.
Writing
-
Evaluating the Evaluator: What a Broken Scorer Taught Me About LLM Evals
A small learning experiment showing how perfect aggregate scores can give way to scorer failures—and why transcript inspection still matters.
-
Scope-Aware Memory Access Control for Multi-Agent Systems
A systems case study in preserving memory boundaries: the architecture, threat model, tests, and failure conditions behind three-tier access control for multi-agent systems.
Research
arXiv preprint · 2026 Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking An empirical study of memory poisoning, content screening, provenance-weighted retrieval, and evaluation robustness in the reported LongMemEval setup. Read the publication →-
Benchmarking agent memory without grading the benchmarkPublication condition: after the controlled memory-type comparison and evaluator checks are run.
-
Measuring behaviour in stateful agent systemsPublication condition: after the hybrid scorer experiment is run.
-
Testing evaluation sensitivity to context policyPublication condition: after retrieval, context-budget, ordering, and repeat-run variants are tested.
-
Validating multi-agent simulation outcomesPublication condition: after construct-validity checks and behavioural baselines are evaluated.
-
Testing adversarial failures at agent–tool boundariesPublication condition: after prompt-injection, trust-boundary, and integrity test cases are run reproducibly.
Projects
Engineering evidence for agentic AI
My systems-engineering background supplies the reliability discipline; work on Aegis Memory supplies the experimental substrate. Quantify Labs is the organisational provenance behind that work—not the centre of the research identity.