Applied AI Systems Lab

Applied AI Systems Lab — working prototypes, visible architecture, honest limitations. Financial infrastructure, agentic systems, and human-centered AI, built with clarity, control, and measurable value.

Liquidity Intelligence

The question

Can AI identify meaningful intraday liquidity movement and emerging funding risk without burying treasury teams in alerts?

The prototype

A monitoring workflow that combines payment-flow data, thresholds, anomaly detection, and AI-generated explanations for material changes.

What it demonstrates

Event ingestionTime-series analysisAnomaly detectionLLM explanationAlert orchestrationHuman review

Executive relevance

Earlier visibility, reduced monitoring effort, faster investigation, and better prioritization of liquidity events.

Control point

AI explains and prioritizes activity; treasury retains the decision and escalation authority.

Functional prototype · Synthetic / sample data

See how it works
Architecture
Event ingestion → time-series store → threshold and anomaly detection → LLM explanation layer → alert orchestration → human review queue.
Models & tools
Statistical anomaly detection generates the signal; a frontier LLM is used only to explain and rank material changes — deliberately not as the detector.
Data sources
Synthetic and sample payment-flow data; no production or customer data.
Evaluation
Precision and recall measured against seeded anomalies, plus an explicit alert-volume budget so 'more alerts' never counts as success.
Human-in-the-loop
Every alert lands in an analyst review queue; escalation and decision authority stay with treasury.
Cost & usage
Token spend metered per alert; the LLM is invoked only on material changes, which keeps cost proportional to signal, not volume.
Known limitation
Alert quality depends heavily on establishing institution-specific seasonal baselines. The prototype demonstrates the workflow but is not presented as a production liquidity-risk model.
Lesson from building it
The hard problem was never detection — it was suppression. Most of the work went into not alerting.

Also in this lab: Tokenized deposit prototype · Digital check processing · Payments / ISO data model · Reporting standardization

Executive Workbench

The question

Can an executive delegate real multi-step work to agents without losing control of quality, cost, or permissions?

The prototype

A multi-agent workbench that routes tasks to specialized agents with shared context, scoped tool access, and an approval gate before anything leaves the building.

What it demonstrates

Multi-agent orchestrationModel routingPermissionsMemoryTool useObservabilityCost control

Executive relevance

Recurring analytical and drafting work gets delegated, while review, approval, and accountability stay with the executive.

Control point

Agents draft and assemble; the executive approves. Any outbound or irreversible action requires explicit human confirmation.

Working prototype · Internal use

See how it works
Architecture
Task router → specialized agents → shared context and memory store → scoped tool layer → approval gate → output.
Models & tools
Multi-model routing across Claude, Devin, Codex, and Gemini — the model is matched to the task rather than standardized for convenience.
Agent design
Narrow, single-purpose agents with explicit handoffs, rather than one general agent asked to do everything.
Evaluation
Task-level rubric plus human accept/reject rate tracked over time; a rejected draft is treated as an evaluation signal, not a failure to hide.
Human-in-the-loop
A hard approval gate before any outbound action. Permissions are scoped per agent, not granted globally.
Cost & usage
Per-task token and cost telemetry, so the workbench can be judged on cost-per-outcome rather than usage volume.
Known limitation
Agent reliability degrades on long-horizon, ambiguous tasks. The workbench compensates with checkpoints and approval gates rather than claiming autonomy it doesn't have.
Lesson from building it
Orchestration value came from constraining agents, not empowering them. Scope and handoffs mattered more than model choice.

Also in this lab: Orchestration dashboard · Executive CRM · CRO Agent · Executive & personal assistants · AI Operations Console

ReadinessIQ

The question

Can a parent get real visibility into what a child actually missed — and turn that into targeted practice — instead of a dashboard of vague scores?

The prototype

An adaptive reading and math platform that generates passages and problem sets, tracks item-level performance, recommends books by reading level, and gives parents review and targeted-session controls.

What it demonstrates

Adaptive content generationItem-level evaluationLevel-aware recommendationParental review gateGamification

Executive relevance

The same discipline enterprises need: generated content behind a review gate, measurable item-level outcomes, and a human who stays accountable for what gets used.

Control point

Parents review generated material and see every missed item. AI proposes; the adult decides what gets assigned.

In active family use · Iterating

See how it works
Architecture
Content generation → parental review queue → session delivery → item-level scoring → analytics and recommendation.
Models & tools
Built end to end with Claude Code; LLM generation for reading passages and problem sets.
Evaluation
Item-level correctness, category growth over time, and the review-gate rejection rate — how often generated content wasn't good enough to assign.
Human-in-the-loop
The review gate is the product. Nothing reaches a child unreviewed, including book recommendations.
Known limitation
Generated passages still need review for tone and level fit. The review gate exists precisely because generation isn't reliable enough to assign sight-unseen.
Lesson from building it
Better information changed the conversation more than better content did. 'Where did you get stuck?' works when you know the actual item.

Also in this lab: Chess Trainer (Lichess data + Stockfish analysis)

What I build with

Engineering capabilities tied to the prototypes that demonstrate them.

Capability Demonstrated through
Multi-agent orchestration Executive Workbench, CRO Agent
Model routing and observability AI Operations Console
Multimodal document processing Digital Check Processing
Time-series monitoring Liquidity Intelligence
Domain data modeling Payments / ISO Repository
RAG and enterprise knowledge Executive CRM, Reporting Standardization
Financial modeling and tokenomics Tokenized Deposits, AI Value Model
Human-in-the-loop controls Check Processing, Executive Workbench
Evaluation and quality gates ReadinessIQ, Engineering Value Model
Adaptive user experiences ReadinessIQ, Chess Trainer

The model behind the work

How this becomes enterprise value

The lab shows the engineering. The AI Engineering Value Model shows how capability becomes measurable enterprise value — capacity, constraints, and the allocation decision.

Let's talk.

Board service, advisory, speaking, or executive AI education — start a conversation.