Test & Control · AI Force Multiplication · Huntsville, AL

Proof

Most consulting proof is a logo wall you have no way to check. Here is what you can verify before you hire me: four technical reports published in full, no email required. Two on running AI in a commercial business, two on running it inside a CUI boundary. Every one of them prices the exercise instead of gesturing at it, cites its sources, and ends by separating what is measured from what is my opinion. Read them and judge the work before you ever speak to me.

Commercial

Paying for AI

The three ways to buy AI, the four discounts that stack, and five ways to bill it back to the teams that spent it. It covers what token billing actually charges you for — including the tool definitions that tax every call in an agentic workflow and never show up as a line item — then works the arithmetic on caching, batching and model right-sizing, showing a nightly job go from $340 a month to under $20 without changing its output.

The middle of the report is the part most companies postpone: five concrete options for attributing spend to teams, from provider-native projects through cloud tagging to a full gateway with virtual keys, each with its setup effort and its ceiling stated plainly. It says which one is right at which size, including the two cases where the answer is to do nothing.

It is equally plain about the evidence for AI productivity. It cites the 55.8 percent Copilot result and the METR trial that found experienced developers 19 percent slower on their own codebases, and explains why those two findings do not actually contradict each other.

Download the report — PDF, 8 pages

An LLM Gateway in AWS GovCloud

One controlled door in front of every model. The full LiteLLM control surface, meaning budgets at seven levels at once and the parameter names to set them with. A four-tier menu you can publish internally so you stop answering the access question one email at a time. The thirteen NIST SP 800-171 requirements a gateway materially contributes to, named individually, and the ninety-seven it does not touch.

It contains a paragraph of SSP language you can lift, and it is honest about the economics: at forty engineers the gateway costs more than the inference it optimises, and it is justified by the audit trail rather than by token savings. It also names the single failure mode that kills these deployments — the temporary direct Bedrock exception that outlives its reason.

The architecture is the same one a commercial business needs, with the compliance mapping added. It sits in both tracks for that reason.

Download the report — PDF, 9 pages

Defense and CUI

Running AI in AWS GovCloud

What is actually available in GovCloud, what is actually authorized, and why those are two different lists published in two different places on two different schedules. The frontier models are callable in us-gov-west-1 today; the models AWS has announced as FedRAMP High and IL4/IL5 approved are an older and partly different set. Calling an available but unannounced model with CUI in the prompt is an assessment finding that is hard to explain afterwards.

It documents the region asymmetry — eight Bedrock models in West against two in East — which means West to East is not an AI failover no matter what your continuity plan says. It covers the commercial AWS account you cannot avoid, because model enablement in GovCloud requires accepting the EULA in a commercial region first, and gives the five-step handling that turns that from a finding into a footnote.

It also carries the FedRAMP Consolidated Rules 2026 change from Low, Moderate and High to Classes A through D, and what that does to DFARS 7012 language that still says Moderate.

Download the report — PDF, 10 pages

Running AI in Azure Government

Five regions, two of which have the AI. Microsoft Foundry is in usgovvirginia and usgovarizona only — not US Gov Texas, and not either DoD region. That is a fact worth establishing in week one rather than discovering in month four.

The report documents something genuinely non-obvious: in Azure Government the deployment type changes which models exist. Data Zone Standard carries the full catalogue. Plain Standard has no GPT-5.6, no GPT-5.1 and no o3-mini, so a team that provisions Standard by habit concludes Azure Government is behind when they simply picked the wrong deployment type.

It also covers the GPT-4.1 context window, which is 128,000 tokens in government rather than the advertised 1,047,576, the documented tool-definition failure that returns a 500 telling you nothing useful, and the fact that model retirement dates in Azure Government diverge from commercial in both directions.

Download the report — PDF, 8 pages

Deliverables and standards

Migrating LabVIEW and TestStand to Python, with AI assistance

A working paper on what actually breaks when you move a National Instruments test estate to Python: dataflow-to-sequential concurrency errors, hidden state, the real-time and FPGA code you should not port at all, DAQmx to nidaqmx timing traps, and why TestStand's process model, not the code, is the actual work.

It also covers the acceptance contract that makes a migration verifiable: golden data plus hardware-in-the-loop, with the migration log as the deliverable.

This is the level of specificity I bring to an engagement. If it reads like someone who has done it, that is because the scar tissue is real. Email me and I will send you the paper.

Ask for the NI-to-Python paper

What a Teardown deliverable looks like

Every Workflow Teardown ships a written workflow map, an hour-cost figure at your wrap rate, working source you own outright, and an honest list of what the automation does badly. Sanitized examples available on request. Ask and I will walk you through a real one.

Case studies, and the standard they are held to

Every completed engagement gets written up in the same format and published once the client approves it. That format is fixed, and it is worth reading before there is anything in it, because the last field is the one this industry leaves out:

  • The workflow. What the process was, and who performed it.
  • What it cost. Hours per cycle, priced at their wrap rate.
  • What shipped. The system, the source, and who owns it.
  • What it cost after. Measured on the same workflow, not estimated.
  • What it does badly. The honest limitations of what I handed over.

Numbers come from the same workflow measured twice, before and after. Not modelled, not projected, and not supplied by me alone. A case study with no limitations section is advertising, so mine carries one and the client reads it before it goes up.

First engagement · write-up in preparation, pending client approval.

Second engagement · write-up in preparation, pending client approval.

Third engagement · write-up in preparation, pending client approval.

If you would rather talk to a client than read about one, ask during the briefing and I will put you in touch directly.

Book the briefing and ask for a reference

Want to see it on your own workflow?

That is what the Teardown is for. Two weeks, $2,500, and you find out whether I deliver for the price of a decent laptop.

Start a conversation