Archie Marshall

AI Engineer · Founder of Marshall AI Engineering

I build production AI systems for industries where getting it wrong isn’t an option.

See my work Get in touch

Siemens PLC GB&I Inventor of the Year 2022

What I solve

Unstructured to structured

Your experts have years of knowledge trapped in spreadsheets, PDFs, and emails. I build systems that extract it, structure it, and make it queryable, without losing the nuance that makes it valuable.

Reduced 17,000 de-identified veterinary research records to 586 relevant ones using a local model. No data left the building.

Eval frameworks that prove it works

Most AI demos look impressive. Most AI demos break in production. I build evaluation frameworks with ground truth, scored accuracy targets, and provenance, so you know your system works before your client finds out it doesn’t.

Open-source eval framework with nine scorers: exact match, JSON schema with partial credit, LLM judge, cascade, RAG faithfulness, plus scorer-reliability and model-robustness analysis.

Sovereign and secure AI

When your data can’t leave your infrastructure - PII, defence, regulated industries - I deploy local models with the capability the task actually needs. No data exfiltration. No third-party dependencies. Full audit trail.

De-identified veterinary research records processed entirely on a local model. Zero API calls. Full traceability.

Teams that actually deliver

AI projects fail because of people, not technology. I embed with your team, set up requirements traceability, evaluation pipelines, and governance structures, then build the thing.

Requirements traceability from contract clause to acceptance criterion to scored evaluation question, with evaluation pipelines and governance in place before the build starts.

Tech I have shipped with

Hover or tap any item to see what I have actually done with it.

Languages

  • Production services, FastAPI, Pydantic, asyncio, pandas.
  • Full-stack web applications, Next.js frontends.
  • Postgres, analytical queries, data pipelines.
  • Neo4j graph queries, knowledge graph design.

AI and LLMs

  • Tool-calling agents, structured outputs, raw API integration.
  • Claude models in client delivery work.
  • Agent loops and tool orchestration.
  • Managed model access inside regulated deployments.
  • Local models for sovereign and sensitive-data environments.
  • Agentic workflows, stateful tool routing.
  • Hybrid retrieval: semantic, keyword and graph traversal.
  • Vector search, similarity scoring, document retrieval.
  • Neo4j and SurrealDB ontology design, plan and context stores.
  • Model Context Protocol tool integration.
  • Screenshot and document text extraction feeding LLM analysis.

Evaluation

  • Scored question sets, agreed accuracy thresholds.
  • Rubric scoring, cascade tiers, judge bias awareness.
  • Grounding checks and context sufficiency.
  • Scorer reliability versus model degradation.
  • LLM observability and tracing.
  • Agent step tracing and tool call inspection.

Infrastructure

  • Knowledge graphs, ontology design, relationship modelling.
  • Structured data, analytical queries, deterministic engines.
  • EC2, S3, Bedrock, data sovereignty.
  • Containerised services, multi-service orchestration.
  • Infrastructure defined as code.
  • CI/CD pipelines, automated testing, deployment.

Process and governance

  • Contract clause to acceptance criterion to scored question.
  • Safety-critical development within product acceptance programmes.
  • ISMS preparation, Vanta integration, pen-test readiness.
  • Product ownership, sprint planning, backlog management.

Domains

  • Siemens Mobility: signalling, ETCS, safety-critical systems (SIL2).
  • Aston Martin: test management and programme automation, via Novecom Digital.
  • Customer-facing guided design tooling.
  • De-identified research triage, local LLM deployment.
  • Language QA automation, UI screenshot analysis.

Case studies

Email reply automation - SimplyGrow

Overnight enquiries drafted and waiting in the team’s drafts folder by 8am.
Problem
Team spending hours every morning reading and drafting email responses to client enquiries.
What I built
AI system that reads overnight emails, drafts responses, and places them in the drafts folder by 8am. Team reviews and sends - not the other way around.
Result
Client reported 3.1× faster operations. Team shifted from writing emails to reviewing them.
Tech
Python, OpenAI API, email integration

Language QA tool - MeepleCorp

Every UI screen screenshotted and language-checked before release.
Problem
Multi-language game with hundreds of UI screens. Manual checking for language errors was slow and error-prone.
What I built
Screenshot every UI page, detect all visible text, check language correctness, flag issues with location and suggested fix.
Result
Comprehensive language coverage across all game screens. Errors caught before release.
Tech
Computer vision, OCR, LLM-based language analysis

Veterinary research triage - academic study

17,000 de-identified records triaged to 586, entirely on-premises.
Problem
An academic study needed 17,000 veterinary case records reviewed for a specific condition. Manual review was impractical, and the records could not be sent to a third-party API.
What I built
A local LLM pipeline that processed every record on-premises, identified the 586 relevant cases, and recorded a reasoning trail for each classification. Records were de-identified before they reached me.
Result
17,000 → 586 cases. No data left the local machine. Every classification auditable.
Tech
Local LLM (Ollama), Python, structured output parsing

Guided jewellery design - Von Ara

A customer-facing design tool, still in daily use over a year on.
Problem
High-end jewellery clients struggle to articulate what they want. Designers spend hours in consultation trying to understand personal meaning and aesthetic preferences.
What I built
AI-guided design system that extracts personal meaning from conversation, fuses it with the Von Ara house style, and generates design directions the client can react to.
Result
Clients arrive at a design direction faster and with deeper personal connection to the piece. Over a year later the tool remains in daily customer use and under active maintenance.
Tech
LLM, prompt engineering, generative design integration

Dead Reckoning - LangChain × SurrealDB hackathon winner

Ask a Python codebase architecture questions in plain English.
Problem
Hackathon challenge: build something useful with LangChain and SurrealDB in 48 hours.
What I built
A code-navigation agent that parses a Python codebase into a SurrealDB knowledge graph, then answers architecture questions in natural language via a LangGraph agent: hybrid semantic and keyword search, version diffing, and resumable crash-safe ingestion. Built with Júlia Sala-Bayo at the LangChain × SurrealDB London Hackathon.
Result
Won the hackathon. Open-source on GitHub.
Tech
LangGraph, SurrealDB, Python, LangSmith
Source
github.com/atwmarshall/dead-reckoning

Evals framework - open source

An LLM evaluation framework built from first principles, with nine scorers.
Problem
Existing eval frameworks are heavy, opaque, and don’t let you understand what is actually being measured. Most also measure the model without ever checking whether the scorer itself is reliable.
What I built
Nine scorers: exact match, normalised match, regex, multi-regex, JSON schema with partial credit, LLM judge, cascade, RAG faithfulness, and dataset-level context sufficiency, plus two analyses most frameworks leave out. Sensitivity measures how far the scorer’s output moves when the input is rephrased. Robustness measures how far the model degrades under adversarial input. Three-model separation is enforced in code, because using one model to both generate and score variations measures self-consistency, not reliability.
Result
Open source. 262 tests, CI on GitHub Actions.
Tech
Python, Ollama, no external eval dependencies
Source
github.com/atwmarshall/evals-from-scratch

Patents and credentials

Patents

  • EP4059805 - Real-time computer vision-based track monitoring
  • EP4059804 - Moving vehicle warning system
  • EP4144614 - Trespassing deterrence
  • GB2622872 - Trackside worker warning method

Recognition

  • Siemens PLC GB&I Inventor of the Year 2022
  • 4 patents and 5 further invention disclosures (2020–2022)
  • Recognised for inventing a new product range in railway safety

Education

  • MEng (Hons) Mechanical Engineering with Mechatronics - First Class
  • University of Southampton (2015–2019)

Professional

  • Incorporated Engineer (IEng) - IET
  • SAFe Agile Product Management, SAFe Product Owner
  • IPMA Level D / APM PMQ

What clients say

Our operations are running 3.1 times faster now.
SimplyGrow Max WarkoczynskiDeputy Executive Director, SimplyGrow
Archie's expertise in AI integration saved us from investing in the wrong technology. His approach delivered exactly what we needed.
MeepleCorp Matthew HilsonManaging Director, MeepleCorp
Archie built us a custom AI design tool that our customers use every day to create personalised jewellery. Over a year later, he still maintains it and it's become central to how we work. What stood out was his ability to understand our business first, then find the right technical solution - not the other way around.
Lara, Managing Director and Founder of Von Ara LaraManaging Director & Founder, Von Ara

Domain experience

  • Siemens MobilityRailway signalling, ETCS, safety-critical
  • Aston MartinSupercar programme, test management, via Novecom Digital
  • TRLFatal car crashes, autonomous vehicles
  • BMTSubmarines, batteries

And independent engagements across veterinary research, luxury retail and games.

Archie Marshall

About

I’m Archie Marshall. I build AI systems for industries where failure has consequences. Seven years of engineering across Siemens Mobility, the Aston Martin supercar programme via Novecom Digital, and enterprise AI delivery.

I don’t build demos. I build systems that pass eval, hold up under security review, and still work on a Tuesday morning when nobody’s watching.

Four patents, Inventor of the Year, First Class MEng. Based in Bristol, working internationally.