Anyone can ship with AI now.
Knowing if it’s good is the job.
I build products, then build the thing that grades them. Founder of Calibr, a live AI product with 200+ users.
01
Product
Two things I built, and the decisions behind them.
I built an AI product that is not allowed to lie, and then proved it.
25+
I started from the complaint, not the idea.
Interviews on a fixed script, eight tools torn down, a hundred public complaints coded into themes. Keyword tools were safe and useless. ChatGPT was useful and unsafe: it invented the metrics people then had to defend in a room.
9 / 13
So I made it worse on first impression, on purpose.
If a rewritten line contains a number the user never wrote, it does not ship. Next to ChatGPT, Calibr often looks less impressive. I took that trade, because a resume you cannot defend in an interview is a liability, not a feature.
67 → 93
Then I stopped trusting my own taste.
Judging quality by reading it is a mood, not a method. I wrote ten criteria and had a model score every run against them. It caught what I had been missing for weeks: quantification failing on technical roles while the average looked fine.
Read the full case
200+
Users, from the first 12. 500+ resumes processed.
2 hrs → 20 min
To tailor a resume to one company.
8
People on the team I lead across product, strategy and marketing.
The plan I was handed described a product reps would not have bought.
18 → 3
I cut the roadmap before writing any code.
A teardown of twelve competing tools, plus research on what reps complained about. Eighteen screens became the three features reps named unprompted. Waitlist enrollment rose 45%, and the product reached its first 200 reps.
Then I respecified the build as acceptance criteria per screen, not descriptions. A description can be satisfied several ways; a criterion cannot. Three weeks of estimate shipped in five days.
Read the full case
02
Analysis
Twice, the deliverable was a decision.
Neither of these ended in a feature. Both ended in a number someone had to bet on.
The client walked away from a $20M acquisition.
I rebuilt fifteen years of cohort economics and stress-tested three downside cases. The analysis exposed an estimated 40% erosion in projected returns. They did not proceed.
$20M
Acquisition the analysis argued against.
$2.5M
Ad budget reallocated on a separate case, after decomposing return on spend across 100+ SKUs.
Four workforce trackers, four different answers.
Fifteen leaders, and no two counted a head the same way. Getting them to agree on one definition was the actual work. The SQL that unified four systems came after, and the self-serve forecasting after that.
80,000
Employees covered by the forecasting model.
6 hrs → minutes
A manual process, automated end to end, across 10+ lines of business.
03
Leadership
Before any of it, two years standing between four countries.
I learned what a bad translation costs before I learned what a PRD was.
I led a 10-person operations team across 300+ cross-border missions, coordinating live communication between the US Army, UN Command and four other nations. When a line had to stay open, I was the person keeping it open. The unit named me its best squad leader, and my proposal won its development competition.
+40%
Operating log accuracy, after I overhauled the workflows behind it.
5+
Companies that adopted the real-time speaker-identifying transcriber I built.
300+
Cross-border missions run by the team I led.
Joshua Lee
easyhoon75@gmail.com