Rhea AI Consulting takes LLM and ML features from "demo that works" to systems that are fast, accurate, and affordable. Senior, hands-on, in your codebase this week. The person you talk to is the person who builds it.
| total inference cost | −66% (≈$7–8K/mo) |
| api spend (cache + preproc) | −80% |
| hybrid retrieval p50 | 22 ms · 27 ms w/ rerank |
| retrieval throughput | 130 QPS |
| concurrent live streams | 50 → 500+ |
| external calls (on-prem) | 0 |
| uptime @ 50M tx/day | 99.99% |
Every engagement is something already done in production, priced so a CTO can say yes without a board meeting. Start with the audit; most clients do.
Line-by-line review of inference spend, token flow, caching, batching, and failure modes. You get a prioritized savings plan with projected dollar impact, and two quick wins implemented before the audit ends.
Production-grade retrieval: hybrid lexical + dense search, reranking, an eval harness, and citation verification, wired into your product. Latency and accuracy benchmarks before and after.
Document intelligence or inference on your own hardware with zero external calls, a tamper-evident audit trail, and runbooks your team can operate unattended.
Embedded with your team: architecture, code review, vendor decisions, hiring help, and hands-on delivery. A senior AI lead without the six-month search.
One intensive week producing an architecture decision record, a cost model, and a 90-day build plan for a new AI feature or platform.
No free pilots, no spec work, no slide decks. A short call, a two-page proposal, and a start date.
Bring your inference bill or your architecture diagram. You leave with at least one concrete observation whether or not we work together.
One recommended offer, one price, a scope with explicit exclusions, and the outcome we will measure against.
Deposit on signature, access sorted in the first day, code in your repo by the end of the first week.
Anyone can wire a vector database to a prompt. The hard parts are cost, correctness, privacy, and staying up under load.
Documented record of cutting inference spend by more than half while load went up tenfold: preprocessing pipelines, tiered caching, keyframe filtering, batch inference, and multi-step LLM cost optimization.
Hybrid BM25 + dense search fused with Reciprocal Rank Fusion, cross-encoder reranking, evals, and cite-only generation that marks unsupported claims instead of emitting them.
Real experience running AI with zero external calls on modest hardware, for clients whose data cannot leave the building. Vendor-neutral: local models, hosted APIs, or both.
Idempotency, backpressure, crash recovery, circuit breakers, and graceful degradation come standard, learned running 50M+ transactions a day.
Public work that shows how Rhea builds: engines and runtimes, written in Rust, Python, and Go.
A DuckDB extension implementing a full-service AI/ML engine inside the database.
lyraRustA symbolic computation engine inspired by the Wolfram Language.
whiskeyPythonDependency injection and IoC framework designed for Python AI applications.
sockdGoA container runtime optimized for serverless workloads, based on SOCK containers.
coralPythonActive Python systems work; see the repository for current scope.
more on GitHub→Sixty-plus repositories across ML, compilers, distributed systems, and tooling.
Parker architected an ML-powered streaming platform from zero to 4,000+ users and 10K+ monthly actives in under six months, then cut its inference costs by two-thirds while it scaled to 500+ concurrent streams. Before that: 50M+ transactions a day at PayPal at 99.99% availability, ML matching that became 51% of revenue at Qwick, and $3M a year saved through onboarding automation at State Farm.
He currently builds MinerAI, an on-prem document intelligence platform for professional-services firms: hybrid retrieval, cite-only drafting with source verification, and overnight indexing on a CPU-only workstation with no external calls. He holds a B.S. in Mathematics from the University of Arizona and served as a U.S. Army combat medic.
Email a sentence about what you're building and what's hurting: cost, quality, or scale. You'll get a reply within one business day and a concrete observation on the call whether or not we work together.