Trustworthy AI for Software and Agentic Systems

Core Vision

Building scientific and engineering foundations for AI-generated software and autonomous agents that can be specified, measured, assured, and held accountable.

My research investigates when AI-generated software and autonomous agent actions can be trusted, when they should be constrained or corrected, and when they must be rejected.

The research programme follows four connected stages: Specify → Measure → Assure → Account.

Pillar 1: Semantic Specification and Intent Alignment

This pillar examines how human intentions, software requirements, and organisational policies can be translated into verifiable behavioural specifications.

  • Behavioural specifications
  • User intent modelling
  • Permission semantics
  • Requirement reconstruction
  • Ambiguity and conflict detection

Core Research Question

How can we determine what constitutes correct and authorised behaviour for AI-generated software and autonomous agents?

Pillar 2: Evidence-Based Evaluation and Benchmarking

This pillar develops rigorous methods for identifying hidden failures, unsafe behaviours, permission violations, and unintended consequences.

  • Executable benchmarks
  • Real-world datasets
  • Failure taxonomies
  • Dynamic and longitudinal evaluation
  • Security, reliability, and cost-aware metrics

Core Research Question

How can we reliably measure whether AI-generated software and agent actions are correct, secure, robust, and consistent with user intentions?

Selected Related Publications

  • Characterizing Large Language Model Agentic Workflows: A Study on the n8n Ecosystem
  • No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT
  • ChatGPT vs Search-Based Software Testing: An Exploratory Study
  • Measuring and Mitigating Bias in Code Generated by Large Language Models
  • An Empirical Study on Low-Code Programming Using Traditional and LLM Support
  • Lost in Translation, Found in Summary: Summary-Driven Supervision for Cross-Language Code Retrieval

Pillar 3: Runtime Assurance and Intervention

This pillar develops independent assurance mechanisms that verify, constrain, repair, or safely stop AI-generated actions before harmful consequences occur.

  • Software testing
  • Program analysis
  • Runtime monitoring
  • Sandboxing and isolation
  • Intent-aware policy enforcement
  • Safe fallback, rollback, and repair

Core Research Question

How can untrusted software outputs and autonomous actions be prevented, constrained, corrected, or safely degraded during execution?

Selected Related Publications

  • MCP-SandboxScan: WASM-based Secure Execution and Runtime Analysis for MCP Tools
  • LLM-CompDroid: Repairing Configuration Compatibility Bugs in Android Apps with Large Language Models and Dynamic Analysis

Pillar 4: Traceability, Accountability and Human Control

This pillar investigates what evidence is required for people to understand, question, restrict, reverse, and hold AI systems accountable for their actions.

  • Software and action provenance
  • Auditable execution traces
  • Causal evidence of agent decisions
  • Consequence explanations
  • Meaningful consent and refusal
  • Undo, appeal, and accountability mechanisms

Core Research Question

How can AI systems provide sufficient evidence for users and organisations to make informed trust decisions and retain meaningful control?

Application Streams

AI-Generated Software

This stream studies when AI-generated code, patches, tests, configurations, and workflows should be accepted, tested, repaired, or rejected.

  • Functional correctness and requirement conformance
  • Security and privacy vulnerabilities
  • Regression and unintended side effects
  • Test adequacy and oracle generation
  • Dependency and software supply-chain risks
  • Code provenance and validation evidence

Tool-Using and Agentic Systems

This stream studies how autonomous actions can remain aligned with user intent, permission boundaries, organisational policies, and acceptable consequences.

  • Intent and permission semantics
  • Tool and workflow security
  • Prompt injection and untrusted inputs
  • Long-horizon and multi-agent failures
  • Permission delegation and laundering
  • Runtime enforcement and execution auditing

Research Outcomes

The programme aims to produce reusable scientific and engineering assets, including:

  • Formal and empirical models of trustworthy behaviour
  • Failure taxonomies and annotated datasets
  • Executable benchmarks and evaluation infrastructure
  • Runtime monitors, policy engines, and secure sandboxes
  • Standardised provenance and audit formats
  • User-facing mechanisms for consent, control, and accountability