About
My current work focuses on the empirical evaluation of agentic AI systems under adversarial or ambiguous conditions. I build experiments that examine how agents behave when instructions, evidence, or operating contexts are incomplete, conflicting, or deliberately manipulated.
Earlier, I built Aegis Memory, an open-source testbed for persistent memory in agent systems. Benchmarking the quality of stored and retrieved memories led to questions about memory poisoning, provenance, adversarial retrieval, and the methods used to evaluate those effects.
I authored the arXiv preprint Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking, which reports an empirical evaluation of that research thread.
I am extending this work into broader behavioural and agent evaluations using Inspect AI, with an emphasis on clearly defined tasks, appropriate scorers, and reproducible analysis.
Current focus
- LLM and agent evaluation design
- Inspect AI evaluation development
- Scorer validity and measurement quality
- Adversarial agent behaviour
- Memory poisoning and provenance
- Python research engineering
Education
I hold an MSc in Big Data Analytics from Sheffield Hallam University.