Technical advisory
$250
per hour
2-hour minimum
Focused architecture, design review, technical due diligence, or troubleshooting.
Human-AI PartnershipELI Labz builds Intelligent Machine Systems (IMS) that augment Human Crews and AI agents with lifelike reasoning, memory, and collaboration.
ELI Agent Loop
Hover or tap a node to trace the agent loop.
Observe → Reason → Act → Verify → Remember → Collaborate
Most AI work that stalls does so the same six ways. Tick what is true for you this month and you get the first thing to fix, free, plus a link to send to whoever owns the workflow.
Your result
One tick is enough to name where the loop is stuck. Two or more usually means the record is missing rather than the model being weak.
Nothing is stored. The result lives in the page link, so you can send it to whoever owns the workflow.
Six integrated layers turn raw model capability into reliable, bounded, human-aligned digital work.
System-2 style reasoning, tool-assisted reflection, self-checking, and verification before action.
Short-term, working, long-term, episodic, and workflow memory that persists across tasks.
Text, image, audio, video, screen state, and interface awareness for situated understanding.
Agents that operate browsers, files, documents, spreadsheets, and dashboards with bounded action.
Approvals, handoffs, role awareness, escalation, preference learning, and cognitive offload.
Policy gates, confidence signals, audit traces, recovery loops, and human override.
Direct senior engineering for environments where a slide deck is not enough. Four published ways to start, from a focused advisory session to a standing retainer.
Technical advisory
$250
per hour
2-hour minimum
Focused architecture, design review, technical due diligence, or troubleshooting.
Deployment sprint
$18,000
starting at
Two weeks
Prove, integrate, or harden one high-value AI workflow.
Embedded FDE
$32,000
starting per month
Up to 80 hours
Senior engineering embedded with your crew for sustained delivery and adoption.
Retainer
$16,000
starting per month
Priority access
Standing senior availability when scope changes month to month.
Final scope and price are confirmed in writing before work begins. Travel, cloud usage, software licenses, and specialized compliance work are quoted separately.
Seven levels, from tooling that tidies itself up to a full intelligence explosion: an honest map of what our systems credibly do today.
Claim
AGI-level intelligence explosion
Your current status
No public system has conclusively demonstrated this
Claim
Recursive improvement of the research loop
Your current status
Partial: the system adjusts its own research approach. Any change to how that works is approved by a person first
Claim
Automated AI-research loop
Your current status
Partial: an automated research loop over models we already control. Actual training runs still need a person's go-ahead
Claim
The main model adapts how it works
Your current status
Partial: the system tunes how it behaves within limits set in advance. Its underlying model is untouched
Claim
It trains smaller, focused models
Your current status
Yes, within bounds: it trains smaller models locally, and only when the conditions we agreed beforehand are met
Claim
Code and policy changes, approved by a person
Your current status
Yes, within bounds: every change is checked, approved by a person, and can be undone
Claim
The tooling around the model improves itself
Your current status
Yes, credible
One named industry and one outcome we can point at, instead of a list of sectors we hope to serve.
Defense and government program operations
Aviation, seismic, weather, cyber, news, and sanctions feeds load into a single map-centered view, viewport by viewport, with each layer independently switchable. The advisory path is feature-flagged, written down as decision support only, and ends at a Human reviewer with an audit record rather than an automatic action.
This is our own published deployment, not a client engagement. No customer name, contract, or measured client result is claimed here.
See the page written for that operatorWhat was built
Third Eye
The week, in the operator's own units
Feeds per shift
Aviation, seismic, weather, cyber, news, sanctions, and the rest. Each lives in its own window with its own layout, and none of them wait for you.
Minutes to first brief
The brief is due before the picture is complete, so the gap between what you have seen and what you can stand behind is yours to close.
Name per recommendation
Every recommendation you forward carries your name, not the tool's. A wrong call is yours. A missed one is yours too.
Re-checks after hours
A feed you did not check is a gap you own, so the re-check follows you home and into the weekend.
Record per decision
When leadership asks who approved an action, the answer has to exist in a record, not in someone's memory.
Hours back per new tool
New tooling usually arrives as one more screen to watch. It only counts if it gives hours back.
Everything below sits on someone else's infrastructure and carries a date. Open it without asking us first.
Routing layer commit, 137722a
GitHub
A dated, public change by the ELI Labz account adding task routing and a Human handoff to a research loop.
Check itInstallable release with a benchmark runner
GitHub Releases
Built Windows and Linux artifacts and a repeatable evaluation runner published under a tag, not a screenshot.
Check itLive geospatial workspace
geospatialcommand.center
A self-hosted intelligence workspace with a documented Human review path is running and reachable today.
Check itGovernance and Human review design
GitHub docs
The approval boundary is written down in the repository, feature-flagged, and marked decision support only.
Check itCapability taxonomy with tests and CI
GitHub Actions
A machine-readable taxonomy is validated by a repeatable command and a workflow anyone can read.
Check itLabor market application in production
laboreconomics.dev
A full-stack analysis application is deployed and open for inspection, including its scoring and forecasts.
Check itWindows desktop agent, patent pending
cursorkeyboardagent.app
A shipped desktop product that moves the cursor, types, clicks, and reads the screen on real Windows apps, with its reasoning visible on every action.
Check itThe client workflow above is not on this list, and it cannot be.
The named client, the repository names, and the run records behind the twelve-run figure sit behind an NDA, so the time saved there is our own measurement and you should treat it that way. Everything listed here is different: it is on GitHub or on a public URL, it carries a date, and it can be opened without asking us first. That is the part of the evidence you do not have to take on trust.
Public work proves method and build capability. It does not prove what a client's numbers looked like, and no page here should be read as saying it does.
If the check above named your week, this is what the same exercise looks like when it runs all the way through. Read it as an example of the record, not as proof of your numbers.
Stage
Input state
What happens
Before anything changes, the system writes down exactly what it was looking at: the version of the work, the checks that were failing, and the rules it was allowed to operate under. Nothing moves until that record is fixed.
Stage
Policy decision
What happens
The system may only suggest a change that fits inside limits agreed in advance. Anything larger is stopped and sent to a person rather than guessed at.
Stage
Human gate
What happens
A person sees the change being proposed, the reason for it, and what the system expected would happen. Nothing is applied until someone says yes.
Stage
Post-action verification
What happens
The same checks run again once the change is in. If the result is worse than expected, the change is rolled back on its own and the whole sequence is kept on record.
Stage
Measured time saved
What happens
About four hours of manual checking, reviewing, and re-testing each cycle, down to roughly thirty minutes of Human oversight. Measured across the last 12 runs.
Where this sits on the claim ladder: Level 2 - yes, within bounds.
The client's name and the repository names are removed from this write-up. The record of each run, what was checked, and what came back can be reviewed under NDA.
Measured outcome, our own numbers
4 hours → 30 minutes
Four hours of manual checking, reviewing, and re-testing each cycle, now about thirty minutes of Human oversight, across the last 12 runs of the workflow above. That figure is ours and it cannot be published, which is why the checkable records are the ones listed just above this section.
Start small
Bring one workflow you are already running. In two hours you will know where it is safe to give an AI system more responsibility, where it is not, and what to fix first. You leave with a written summary you can hand straight to your Crew.
$250 per hour - 2-hour minimum
No budget to spend yet? Send one failing run and the check that should have caught it, and you get a one-paragraph read back on where it is stuck. That costs nothing and starts no engagement.
dmesa@e-l-i.net
Bounded
Verifiable digital work with human oversight where appropriate.
Evidence-first
Reasoning grounded in observation, tools, and verification.
Human + Agent
An operating model that augments people, not replaces them.
Collaborate on research, prototype an agent, or follow the builds. We work in the open.