Skip to content

AI development · Agent integration · Workflow implementation

Most AI projects stop at the demo. I build the ones that keep running.

ControlStackAI is the engineering practice of Matthew Mangano. I design, build, and operate AI systems that hold real context, take real actions inside real tools, and leave a record you can review afterward — for teams that need the work finished, not another pilot.

What I actually do
Write the code, wire the tools, run the evaluations, and stay on until it works in production.
Who this is for
Engineering teams and technical founders who already tried the obvious AI tools and hit the wall where the real work starts.
Where I come from
Aerospace — 6-DOF flight simulation, control systems, and verification work where “probably correct” was never an acceptable answer.

Proof of work

An operating system run by its agent. Built here, running daily.

This is not a concept for something I hope to build. It is my own working environment: one spoken request becomes retrieved context, a plan, delegated execution across specialist agents, and a consequential action that stops for human approval and leaves a receipt. The same architecture is what I build for clients.

1:48 · Captioned, so it plays with the sound off.
  • Voice-first

    The request is spoken. No prompt-chaining, no copy-paste between tabs.

  • Persistent memory

    Decisions and constraints survive the session, so the system stops re-asking what it was already told.

  • Verified delegation

    Specialist agents do the narrow work and report back with evidence, into one continuous thread.

  • Review before action

    Anything consequential freezes, asks for approval on the exact action, and writes a receipt after.

Project case study · Agent Harness

Contract in. Independent verification. Signed evidence out.

I built a provider-independent runtime for agent work where the executor never grades its own homework. It runs against a definition of done, survives interruption, verifies the result separately, and leaves a tamper-evident receipt.

Explore the architecture
  1. 01

    Define the contract

    Intent, checks, constraints, budget.

  2. 02

    Execute durably

    Local, API, or delegated agent.

  3. 03

    Verify separately

    Checks plus adversarial review.

  4. 04

    Keep the proof

    Signed verdict and evidence.

Services

Three ways engagements usually start.

Different entry points, same underlying job: get a model doing useful work inside systems that already have owners, constraints, and consequences.

01 — Build

AI application development

Products with a model inside, built like software rather than like a prototype that got promoted. Retrieval that returns the right thing, structured output you can parse, evaluations that catch regressions, and cost and latency treated as requirements.

  • · Agent and RAG applications end to end
  • · Evaluation harnesses and regression suites
  • · Model routing, caching, and cost control
  • · Deployment inside your infrastructure boundary

02 — Integrate

Agent harness integration

Coding agents and assistants are only as useful as their reach. I connect them to the repositories, tools, and services your team actually uses, with permissions that are scoped deliberately instead of granted by default.

  • · Harness setup and multi-agent orchestration
  • · MCP servers and custom tool interfaces
  • · Chat, issue tracker, and CI integration
  • · Scoped permissions, approval gates, audit trails

03 — Automate

Workflow implementation

Take a process a person runs by hand every week and make it a supervised one. Map the real workflow — including the exceptions nobody documented — automate the load-bearing parts, and put a human at the decisions that deserve one.

  • · Workflow mapping and opportunity ranking
  • · Document, report, and analysis pipelines
  • · Engineering toolchain automation, including MATLAB and Simulink
  • · Human checkpoints where the stakes justify them

Method

Context, plan, execution, proof.

The same four beats the film walks through are how the engagement itself runs. AI work fails in predictable places: nobody wrote down what the system is allowed to assume, and nobody defined what "working" means before building it.

  1. 01

    Map the real work

    Watch the process as it is actually performed, not as the documentation describes it. Name the constraints, the data boundaries, and the steps that carry consequences.

  2. 02

    Build the thin path

    One workflow, end to end, in production, early. A narrow system people use beats a broad one they evaluate.

  3. 03

    Put the gates in

    Scoped permissions, approval on consequential actions, and an evaluation suite that fails loudly. Autonomy gets extended once it is earned, not assumed on day one.

  4. 04

    Hand over something operable

    Documented, tested, and runnable by your team. You should be able to change it without calling me — and call me because you want to, not because you are stuck.

Stack posture

Model-agnostic and harness-agnostic, on purpose.

The frontier moves every few months. Anything built tightly around one vendor's assumptions becomes a migration project the moment the pricing or the capability changes. I build the boundary first — your workflows and evaluations on one side, the model and harness on the other — so swapping either is a decision rather than a rewrite.

In practice that means working across the current generation of coding agents and agent runtimes, commercial and local inference, and the tool protocols that connect them — and being honest with you about which parts of that are stable enough to depend on.

Who you are hiring

An engineer who ships, not an agency that staffs.

I am an aerospace engineer. Before AI systems, my work was aircraft 6-DOF simulation, flight dynamics, control law development, and the verification that has to hold up when something flies. That background is the reason I build AI the way I do: define what correct means, instrument it, and never confuse a convincing output for a validated one.

I also run the thing in the film. My own infrastructure is a fleet of specialist agents on a declaratively managed Linux host, with persistent memory, scheduled autonomous work, and approval gates on anything that leaves the machine. Every failure mode I would otherwise discover on your project, I have already hit on mine.

You work with me directly. No account manager, no handoff to a junior team after the pitch.

More about the practice

Background

Domain
Aerospace — flight dynamics, control systems, avionics, simulation and verification
Engineering tooling
MATLAB and Simulink, model-based design, automated test and reporting pipelines
AI systems
Multi-agent orchestration, retrieval and memory, tool protocols, evaluation harnesses
Infrastructure
Declarative Linux, containers, edge and cloud deployment, local inference

Engagements

Start small enough to be worth trying.

Most people should start with the audit. It is fixed-fee, it ends with something useful whether or not we work together again, and it means neither of us is guessing at scope when the build starts.

Start here

Fixed fee

Workflow audit

One to two weeks

I map how the work is done now, find where a model realistically helps and where it does not, and rank the opportunities by value against difficulty.

You get: a workflow map, a ranked opportunity list with honest effort estimates, a recommended architecture, and a build plan you can execute with me or without me.

Build

Scoped

Implementation sprint

Four to eight weeks

One workflow taken all the way into production, with the evaluation suite, the permissions, and the approval gates that make it safe to leave running.

You get: a working system in your infrastructure, source and tests in your repositories, runbook documentation, and a handover session with the people who will operate it.

Operate

Monthly

Ongoing AI engineering

Retained, rolling

For teams running agentic systems in production who want someone accountable when the model, the pricing, or the workflow changes underneath them.

You get: continuous extension and maintenance, evaluation and cost monitoring, model migrations handled, and an engineer who already knows the system.

Pricing is quoted per engagement after a short scoping conversation — there is no useful honest number before knowing what the work is. Tell me the workflow and I will tell you whether it is worth automating at all.

Contact

Describe the workflow you want to stop doing by hand.

The most useful first message is a concrete one: what the process is, who runs it today, how often, and what breaks when it goes wrong. I reply personally, usually within a business day.

Prefer email? [email protected]