Applied AI engineering · Phoenix, AZ · Remote US

Production AI for startups that need to ship — and stop bleeding on inference.

Rhea AI Consulting takes LLM and ML features from "demo that works" to systems that are fast, accurate, and affordable. Senior, hands-on, in your codebase this week. The person you talk to is the person who builds it.

engagement_summarybefore → after
total inference cost−66% (≈$7–8K/mo)
api spend (cache + preproc)−80%
hybrid retrieval p5022 ms · 27 ms w/ rerank
retrieval throughput130 QPS
concurrent live streams50 → 500+
external calls (on-prem)0
uptime @ 50M tx/day99.99%
# figures from shipped work at TL;DR AI, MinerAI, PayPal
66%total inference cost reduction while scaling to 500+ concurrent streams
22 msp50 hybrid retrieval with reranking and citation verification
0external network calls for on-prem document AI on a 16 GB CPU box
99.99%availability on 50M+ daily transactions across 200+ countries
Services

Five offers. Fixed scope, fixed price.

Every engagement is something already done in production, priced so a CTO can say yes without a board meeting. Start with the audit; most clients do.

01

AI Cost & Reliability Audit

2 weeks · the front door

Line-by-line review of inference spend, token flow, caching, batching, and failure modes. You get a prioritized savings plan with projected dollar impact, and two quick wins implemented before the audit ends.

$15,000fixed fee
02

Retrieval & RAG Pilot

6–8 weeks

Production-grade retrieval: hybrid lexical + dense search, reranking, an eval harness, and citation verification, wired into your product. Latency and accuracy benchmarks before and after.

$60K–$90Kfixed fee
03

Private / On-Prem AI Deployment

8–12 weeks

Document intelligence or inference on your own hardware with zero external calls, a tamper-evident audit trail, and runbooks your team can operate unattended.

$75K–$120Kfixed fee
04

Fractional Head of AI Engineering

2 days/week · 3-month minimum

Embedded with your team: architecture, code review, vendor decisions, hiring help, and hands-on delivery. A senior AI lead without the six-month search.

$14,000per month
05

Architecture Design Sprint

1 week

One intensive week producing an architecture decision record, a cost model, and a 90-day build plan for a new AI feature or platform.

$9,500fixed fee
How it works

Discovery to kickoff in under three weeks.

No free pilots, no spec work, no slide decks. A short call, a two-page proposal, and a start date.

Step 1 · 30 minutes

Discovery call

Bring your inference bill or your architecture diagram. You leave with at least one concrete observation whether or not we work together.

Step 2 · within 48 hours

Two-page proposal

One recommended offer, one price, a scope with explicit exclusions, and the outcome we will measure against.

Step 3 · within 2 weeks

Kickoff

Deposit on signature, access sorted in the first day, code in your repo by the end of the first week.

Why Rhea

What most AI consultancies don't do.

Anyone can wire a vector database to a prompt. The hard parts are cost, correctness, privacy, and staying up under load.

Cost is a feature

Documented record of cutting inference spend by more than half while load went up tenfold: preprocessing pipelines, tiered caching, keyframe filtering, batch inference, and multi-step LLM cost optimization.

Retrieval done properly

Hybrid BM25 + dense search fused with Reciprocal Rank Fusion, cross-encoder reranking, evals, and cite-only generation that marks unsupported claims instead of emitting them.

Private by default

Real experience running AI with zero external calls on modest hardware, for clients whose data cannot leave the building. Vendor-neutral: local models, hosted APIs, or both.

Payments-grade reliability

Idempotency, backpressure, crash recovery, circuit breakers, and graceful degradation come standard, learned running 50M+ transactions a day.

Open source

Infrastructure, not just prompts.

Public work that shows how Rhea builds: engines and runtimes, written in Rust, Python, and Go.

About

Parker Lackey, founder and principal engineer.

Parker architected an ML-powered streaming platform from zero to 4,000+ users and 10K+ monthly actives in under six months, then cut its inference costs by two-thirds while it scaled to 500+ concurrent streams. Before that: 50M+ transactions a day at PayPal at 99.99% availability, ML matching that became 51% of revenue at Qwick, and $3M a year saved through onboarding automation at State Farm.

He currently builds MinerAI, an on-prem document intelligence platform for professional-services firms: hybrid retrieval, cite-only drafting with source verification, and overnight indexing on a CPU-only workstation with no external calls. He holds a B.S. in Mathematics from the University of Arizona and served as a U.S. Army combat medic.

  • 2026 –Founder & Principal Engineer, MinerAI · Founder, Rhea AI Consulting
  • 2025 –Lead Software Engineer & Team Lead, TL;DR AI
  • 2023 – 2025Senior Software Engineer, PayPal
  • 2022 – 2023Senior Software Engineer, Qwick
  • 2020 – 2022Software Developer, State Farm
  • 2012 – 2014Combat Medic, U.S. Army
Contact

Send the inference bill. Get a free 30-minute review.

Email a sentence about what you're building and what's hurting: cost, quality, or scale. You'll get a reply within one business day and a concrete observation on the call whether or not we work together.

[email protected]

A good fit looks like

  • Seed to Series B, 10–80 people, an AI feature customers already use
  • An inference bill growing faster than revenue
  • A RAG or search feature that is "mostly right"
  • Load, latency, or GPU-memory problems at scale
  • Data that cannot leave your own infrastructure
Copied