Ulric
Book a call

Eugene, Oregon · one person, whole builds

Product line · live

Pippin.

The agentic harness builder. Sweet, nurturing, and quietly everywhere. Custom agent flows your team can actually trust.

Pippin

Agents your team can actually trust.

01

Grounded from day one

Retrieval over your real data, so agents quote facts instead of inventing them.

  • Retrieval over your real data
  • It hands back the page it used
  • Facts, not inventions

Pippin

Agents your team can actually trust.

02

Guardrails as architecture

Dedicated safety models screening every turn, topical fences, and honest refusals.

  • A safety model, wired in
  • Topical fences
  • An honest refusal when it should not answer

Pippin

Agents your team can actually trust.

03

Multi-model by default

Groq, then NVIDIA, then Cerebras, then Google Gemini; a four-provider ladder that falls through on any error, so one outage never takes the agent down.

  • Four providers, one ladder
  • Falls through on any error, not just a rate limit
  • The last rung is an alias that cannot be retired

Pippin

Agents your team can actually trust.

04

Agents that do work

Lead capture, CRM filing, dashboard copilots, workflow automation; harnesses, tools, and evaluation loops, not just chatbots.

  • Lead capture and CRM filing
  • A staff mode inside the dashboard
  • Evaluation loops, not just a chat window

The other product lines →

The breakdown

What Pippin does.

Pippin designs and ships production agent systems: the discipline behind this site's grounded assistant and the Nemotron Agent Lab.

Pippin is the discipline I use to build agents that do real work: the retrieval, the guardrails, the fallback wiring and the evaluation around a model, rather than the model itself. She is the assistant on this site and the resident agent in the lab, and she is the harness I ship inside Frida for clients.

Grounded from day one

Retrieval over your real data, so agents quote facts instead of inventing them.

Guardrails as architecture

Dedicated safety models screening every turn, topical fences, and honest refusals.

Multi-model by default

Groq, then NVIDIA, then Cerebras, then Google Gemini; a four-provider ladder that falls through on any error, so one outage never takes the agent down.

Agents that do work

Lead capture, CRM filing, dashboard copilots, workflow automation; harnesses, tools, and evaluation loops, not just chatbots.

Nurtured after launch

Transcripts reviewed, knowledge tended, behavior tuned as your business changes.

The inventory

What she ships with.

The harness as it runs on this site today. Every item is in the chat module, and the lab at /nemotron will explain its own pipeline if you ask it.

  • Hybrid retrievalEvery answer starts with the site's own published pages. Keyword scores from MySQL full-text search are blended with cosine similarity over NVIDIA embeddings (nemotron-3-embed-1b, 2,048 dimensions), so a query finds both the exact name and the right idea. If the embedding service is down, retrieval falls back to keywords alone rather than to nothing.
  • A safety model, wired inA dedicated model, Llama 3.1 Nemotron Safety Guard 8B, can screen every turn before anything is generated: anything not explicitly safe is refused, politely, with no model in the loop to argue with. It runs on the lab at /nemotron, where the models are open to the public.
  • A four-provider ladderGroq, then NVIDIA, then Cerebras, then Google Gemini, each with more than one model. The chain cascades on any per-model error, not only a rate limit, and it ends on a floating alias, gemini-flash-latest, so the last rung cannot be retired out from under it.
  • Answers that cite, and refuseEvery note carries the page it came from, and the reply hands those pages back as links under the answer. If the notes do not cover something, she says so and points to the free call rather than inventing a number, a client or a capability.
  • Lead capture into the CRMWhen a visitor is ready, the conversation becomes a lead in the CRM, and the same hooks fire that a form submission would.
  • Staff modeSigned-in users get a different assistant inside the dashboard: the in-house guide, with a live pulse of new leads and follow-ups, suggesting the highest-impact next step. Counts only, never a dump of personal data.
  • Voice, in the labThe lab takes speech in through Whisper on Groq and can speak its replies back. It also names the model it is running on, this turn, if you ask.
  • Guardrails as prompt and as codeShe cannot reveal passwords, keys or server details, because they are not in her notes: the assistant reads the knowledge base and the site's own pages, and nothing else is in front of her. The topical fence is a rule, not a hope.

Under the hood

How she is built, and why.

A model on its own writes text. Everything that makes an agent useful lives around it: what it may see, what it must check first, which tools it can reach, and what happens when a provider fails at two in the morning. I build that part, and I keep the model a swappable component, because the models get cheaper and better every quarter and the harness is what a client keeps.

Retrieval is hybrid on purpose. Keyword search is unbeatable at proper nouns, prices and part numbers; embeddings are unbeatable at meaning. Blending the two scores catches the question that names a thing and the question that describes it. And when the embedding API is unavailable, keywords alone still answer, which is the difference between a slow day and a dead assistant.

The provider ladder came from watching free tiers fail in every way they can: rate limits, retired models, a 500 at the worst moment. So the chain falls through on any error at all, runs several models per provider, and ends on an alias that Google keeps pointed at the current Flash model. The last rung, by construction, never answers 404.

The safety guard runs before generation, as its own model, because a guardrail written into the same prompt that is being attacked is not a guardrail. She is in beta: the transcripts are read, the knowledge is tended, and the behaviour is tuned by hand as the site changes.

The honest fit

Who she is for, and who she is not.

For

  • A business with a body of written material (pages, documents, policies, listings) and a stream of the same questions arriving every day, from customers or from staff.
  • A site where a visitor should be able to ask, get a straight answer with its source, and become a lead without a form standing in the way.
  • A team that wants a copilot inside its own dashboard, one that knows what is new this week and says what to do first.

Not for

  • Anything that has to act unsupervised on day one. Autonomy is earned by being right for a few weeks first.
  • A question the documents do not answer. She will say so and hand you to a person; if that is the wrong behaviour for your case, she is the wrong tool.
  • Records that cannot leave the building. Today the ladder runs on hosted providers. If your data has to stay on your own hardware, that is a consulting conversation about local models, not a Pippin install.

What it costs

How you get her.

Pippin arrives with the Relay, as the assistant trained on your business with guardrails you approve, and with the Consult, when the agent has real work to do. The care plan with the assistant keeps her retrained on what changed each month. The prices are published on one page.

Gentle with your team, protective of your customers. Ask about the beta. Fifteen minutes, no pitch deck, and you keep the plan either way.