Trustworthy AI for Software and Agentic Systems
Core Vision
Building scientific and engineering foundations for AI-generated software
and autonomous agents that can be specified, measured, assured, and held
accountable.
My research investigates when AI-generated software and autonomous agent
actions can be trusted, when they should be constrained or corrected, and
when they must be rejected.
The research programme follows four connected stages:
Specify → Measure → Assure → Account.
Pillar 1: Semantic Specification and Intent Alignment
This pillar examines how human intentions, software requirements, and
organisational policies can be translated into verifiable behavioural
specifications.
- Behavioural specifications
- User intent modelling
- Permission semantics
- Requirement reconstruction
- Ambiguity and conflict detection
Core Research Question
How can we determine what constitutes correct and authorised behaviour for
AI-generated software and autonomous agents?
Pillar 2: Evidence-Based Evaluation and Benchmarking
This pillar develops rigorous methods for identifying hidden failures,
unsafe behaviours, permission violations, and unintended consequences.
- Executable benchmarks
- Real-world datasets
- Failure taxonomies
- Dynamic and longitudinal evaluation
- Security, reliability, and cost-aware metrics
Core Research Question
How can we reliably measure whether AI-generated software and agent actions
are correct, secure, robust, and consistent with user intentions?
Selected Related Publications
- Characterizing Large Language Model Agentic Workflows: A Study on the n8n Ecosystem
- No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT
- ChatGPT vs Search-Based Software Testing: An Exploratory Study
- Measuring and Mitigating Bias in Code Generated by Large Language Models
- An Empirical Study on Low-Code Programming Using Traditional and LLM Support
- Lost in Translation, Found in Summary: Summary-Driven Supervision for Cross-Language Code Retrieval
Pillar 3: Runtime Assurance and Intervention
This pillar develops independent assurance mechanisms that verify, constrain,
repair, or safely stop AI-generated actions before harmful consequences occur.
- Software testing
- Program analysis
- Runtime monitoring
- Sandboxing and isolation
- Intent-aware policy enforcement
- Safe fallback, rollback, and repair
Core Research Question
How can untrusted software outputs and autonomous actions be prevented,
constrained, corrected, or safely degraded during execution?
Selected Related Publications
- MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
- LLM-CompDroid: Repairing Configuration Compatibility Bugs in Android Apps with Large Language Models and Dynamic Analysis
Pillar 4: Traceability, Accountability and Human Control
This pillar investigates what evidence is required for people to understand,
question, restrict, reverse, and hold AI systems accountable for their actions.
- Software and action provenance
- Auditable execution traces
- Causal evidence of agent decisions
- Consequence explanations
- Meaningful consent and refusal
- Undo, appeal, and accountability mechanisms
Core Research Question
How can AI systems provide sufficient evidence for users and organisations
to make informed trust decisions and retain meaningful control?
Application Streams
AI-Generated Software
This stream studies when AI-generated code, patches, tests, configurations,
and workflows should be accepted, tested, repaired, or rejected.
- Functional correctness and requirement conformance
- Security and privacy vulnerabilities
- Regression and unintended side effects
- Test adequacy and oracle generation
- Dependency and software supply-chain risks
- Code provenance and validation evidence
Tool-Using and Agentic Systems
This stream studies how autonomous actions can remain aligned with user
intent, permission boundaries, organisational policies, and acceptable
consequences.
- Intent and permission semantics
- Tool and workflow security
- Prompt injection and untrusted inputs
- Long-horizon and multi-agent failures
- Permission delegation and laundering
- Runtime enforcement and execution auditing
Research Outcomes
The programme aims to produce reusable scientific and engineering assets,
including:
- Formal and empirical models of trustworthy behaviour
- Failure taxonomies and annotated datasets
- Executable benchmarks and evaluation infrastructure
- Runtime monitors, policy engines, and secure sandboxes
- Standardised provenance and audit formats
- User-facing mechanisms for consent, control, and accountability