henny.aiGet in touch
bootstrapped ai studio · products and agents

henny.ai

henny.ai is a small, bootstrapped AI studio that builds products and agents. We’re building Bookmark Buddy, an iOS app that uses AI to organize, summarize, and resurface the links you save.

buildingBookmark Buddy for iOS · AI-powered link organizer◆betaBookmark Buddy · invite-only◆researchingeval harness for tool-using models◆deployingAI agents for teams◆fine-tuningsmall models for routing and repeatable steps◆buildingBookmark Buddy for iOS · AI-powered link organizer◆betaBookmark Buddy · invite-only◆researchingeval harness for tool-using models◆deployingAI agents for teams◆fine-tuningsmall models for routing and repeatable steps◆buildingBookmark Buddy for iOS · AI-powered link organizer◆betaBookmark Buddy · invite-only◆researchingeval harness for tool-using models◆deployingAI agents for teams◆fine-tuningsmall models for routing and repeatable steps◆
── what we’re building

Our own products, built with AI.

01 · ios appinvite-only beta

Bookmark Buddy

An iOS app that uses AI to organize, summarize, and resurface the links you save. Bookmark Buddy is in an invite-only beta.

Learn about Bookmark Buddy →
✓ organizes the links you save✓ summarizes what each link says✓ resurfaces links worth another lookpowered by AI
── agents for teams

We also design and deploy AI agents for teams.

The same approach we use for our own products, applied to your workflows: scoped tightly, evaluated against the real job, and wired into the tools your team already uses. We fine-tune custom models when an off-the-shelf one isn’t enough.

./henny · agent sandbox○ idle
›
try →
── harnesses we deploy

We meet you where the agent should live.

We work across models and agent harnesses and don’t pick one out of taste. We pick the one your work fits inside, set it up against your tools and data, and stay until it’s measurably useful.

OpenClaw
OpenClaw
open agent harness

A flexible open-source harness. We deploy it inside your VPC with the channels, tools, and memory your team actually uses.

best for
  • +self-host on your infra
  • +Slack / Telegram / email
  • +long-running ops agents
── how we work

Set up in weeks, not quarters. Trained on your work, not the internet’s.

01
Scope the work
Most agent failures are scope failures. We start by writing exactly what the agent does, refuses, and escalates, for your business and not in the abstract.
02
Set it up in your stack
We deploy inside the right harness (OpenClaw, Hermes, Perplexity, or whatever fits the job), wired to the tools, channels, and data your team already uses.
03
Fine-tune when it matters
When prompting plateaus, we train. SFT, DPO, distillation on your records. The leverage is enormous and underused: most teams never get here.
04
Evaluate against the real job
A good eval looks like a job description, not a benchmark. We replay your historical cases and ship to a number you trust.
eval-first

We define the job, failure modes, and review rubric before tuning or automation.

tool-native

Agents ship with real integrations: repos, inboxes, CRMs, queues, docs, observability.

model-flexible

Open weights when control matters. Frontier APIs when capability wins. Usually both.

handoff-ready

Every system leaves behind traces, runbooks, datasets, and tests your team can own.

── workbench

The work is less “chatbot”, more lab bench.

Henny systems are built as inspectable pipelines. You can see what the agent saw, which skill it chose, why it acted, and how it improved.

traceable/tunable/owned by your team
agent_build.pipeline02/5
compose
Break the agent into skills: retrieval, tool calls, critique, memory, handoff.
✓ artifact written✓ eval added✓ handoff documented
── engagements

Useful shapes of work.

Start with a contained build, then keep what earns trust. No transformation theater.

012–4 weeks
Agent prototype

Turn a messy workflow into a working agent with tools, traces, and a hard yes/no eval.

→ workflow decomposition
→ tool wiring
→ human review loop
023–6 weeks
Fine-tune sprint

Use your examples to train model behavior that prompting cannot reliably produce.

→ dataset curation
→ SFT/DPO recipes
→ model + eval report
03ongoing
Production build

Embed agentic systems into product or operations with monitoring, cost controls, and ownership transfer.

→ runtime architecture
→ observability
→ deployment + runbooks
── contact

Want a Bookmark Buddy invite, or have an agent problem?

Tell us what you’re working on in a paragraph. We read everything and reply to most.