Ella
Deciding what an AI agent is allowed to do
- Role
- Sole author of the PRD and the release acceptance criteria
- Shipped
- 13 August 2026
- Proves
- Agent behaviour design, permission scoping, monetisation, and a decision reversed on new information
“I’d rather ask ChatGPT what to sell than click through this.” Paraphrased from creator interviews, early 2026
At a glance
| Role | Product Manager. Sole author of the PRD and the release acceptance criteria. |
| Collaborators | Nelson, CTO (architecture, engineering delivery). |
| Timeline | PRD 18 June 2026 → acceptance criteria 19 June 2026 → shipped 13 August 2026 |
| What it is | An in-product AI agent that builds products with creators through conversation, and answers questions using their own account data |
| Status | Live. Bundled into the paid tier. Replaced the platform’s previous support assistant. |
| What I owned | What the agent may do, what it must ask permission for, what it must refuse, the interaction and design decisions, and the criteria that gate release |
| What I did not own | Architecture, model selection, infrastructure. Those were Nelson’s calls. |
Context
Nestuge is a creator-economy platform serving primarily African creators. Creators sell custom products, courses, events, consultation sessions and bundles; run paid community hubs; take payments across 135+ countries; and run affiliate programmes and email campaigns from the same account.
That breadth is the product’s advantage and its problem. Every capability a creator might want is present. Reaching any of them means understanding a builder with more options than most creators have opinions about.
The problem
Four signals, gathered separately, pointed at the same gap.
1. The support queue was mostly a knowledge queue. I sampled and counted inbound tickets. “How do I do X” questions dominated. These were not defects. They were creators who could not find a path through the product, being answered one at a time by humans.
2. Creators were asking what to sell. A recurring, unserved category of request: not “how do I build this” but “what should I build.” Nothing in the product addressed it, so it landed in support, where it did not belong either.
3. Creators were asking about their own performance, and what to do about it. Two questions arriving together. The first was a lookup, answerable entirely from data the platform already held, routed to a human who then had to go and find it. The second was harder and more valuable: having seen the number, how do I improve it, and what should I do differently. A dashboard answers the first question. Nobody was answering the second.
4. Creators were already using AI. Elsewhere. In interviews, creators described routinely opening general AI tools outside Nestuge to think through product ideas and structure, then coming back to the platform to do the data entry.
Signal 4 reframed the whole thing. The job was already being done. It was being done off-platform, badly matched to the creator’s actual catalogue and audience, and then manually transcribed back in. We were not deciding whether to add AI. We were deciding whether the AI a creator used to plan their business would know anything about their business.
Constraints
- The users are non-experts. Not “less technical.” Creators for whom this product is their income and who have no reason to have opinions about subscription billing cadence.
- The agent touches money. Payments, transactions, wallet balances, settlement timing.
- It was a beta. Nothing had been proven in production.
- There was no design system. The prototype was built without Figma or an agreed component set, which constrained how far the interaction quality could be specified rather than described.
- A support assistant already existed and was integrated with our customer support platform. Removing it risked breaking context and logging that CX depended on.
Decisions
Six forks. Each one is a place the product could have gone another way.
1. A named surface, not a support widget
At stake: whether creators would read this as a real capability or as help.
Ella got a dedicated navigation item opening a full-page surface, positioned between “My Purchases” and “Settings.” It also got a floating panel that persists across pages, which replaced the existing support widget rather than sitting alongside it.
The rejected model was a modal or a launcher: open it, use it, close it, reopen it later. That pattern makes every interaction a discrete errand. I wanted continuous use, where a creator can keep a conversation open while moving through the dashboard and watching a result land.
What I gave up: navigation real estate, and a clean line in the creator’s mental model between “support” and “building things.” Those are now the same door.
2. Review before commit, on everything
At stake: how much autonomy an agent gets in a product where actions have consequences.
Nothing Ella produces commits without the creator seeing it first. Not products, not emails, not the low-risk items. Anything customer-facing or irreversible requires explicit approval.
Three reasons, in order of weight. The creator keeps final say, so agency stays with the person whose business it is. Ella was in beta and nothing had earned trust yet. And on this platform there are very few genuinely low-stakes actions, because almost everything Ella produces is something a paying customer eventually sees.
What I gave up: speed, and the demo-friendly feeling of an agent that just does things.
Worth being clear about: this is a posture calibrated to the product’s maturity, not a permanent principle. As output quality becomes measurable, the right move is to relax it selectively, starting with the actions that are cheapest to undo. Holding it universally forever would be caution, not judgment.
3. Replacing the support assistant, having specified coexistence
At stake: whether to break a working integration to own the whole surface.
The PRD is explicit. Section 2, non-goals: replacing the existing customer support. Ella coexists with it. Section 8.6 spells out the coexistence.
The reason was not attachment. The existing assistant carried conversation context and was wired into our customer support platform. Breaking that integration meant breaking logging and continuity that the CX team relied on. Coexistence was the conservative call and, at the time I wrote it, the correct one.
Then we found a way to connect Ella directly to the same support platform. That removed the constraint the non-goal was built on. Owning the connection ourselves meant owning context building, logging, and a seamless path from conversation to ticket.
We replaced it. Before launch.
Why this one matters most to me: there is a document, dated and written by me, stating the opposite decision, with the reasoning that made it correct. When the reasoning stopped holding, the decision changed. A spec is a hypothesis with a date on it, not a commitment to be defended.
4. Ella can read financial data. It cannot touch it.
At stake: the most requested capability we deliberately did not build.
Ella looks up transactions, explains payment status, and can tell a creator why an international payment has not settled yet. It cannot issue refunds, trigger payouts, or amend receipts. There is no path to it. The boundary is enforced at the action layer, independent of anything the creator or the conversation says.
Two reasons. The first is obvious: keeping creators’ money out of reach of an autonomous system’s mistakes.
The second is the one I actually care about. Creators should perform financial actions themselves, because the act of doing it is what makes the ledger legible to them. If an agent moves money on a creator’s behalf and the wallet balance later looks wrong, the creator has no memory of the action to reason from. Every discrepancy becomes a support ticket. Active participation is not friction here; it is what makes the record trustworthy to the person who owns it.
What I gave up: the single most-requested thing creators would ask an agent to do.
5. Deferring the integration layer, then shipping it somewhere else
At stake: whether to let third parties build on this before we understood it.
Letting external tools connect into Nestuge accounts through Ella was deferred at MVP. Not rejected, deferred, with a stated reason: observe how the agent was actually used, and what data flowed through it, before opening it up.
It later shipped, but alongside the public API and webhooks in the paid tier rather than inside Ella. The demand was real. The right surface for it was not the one we originally imagined.
The point: “deferred” should mean sequenced, with a condition for revisiting. Deferrals without a condition are just rejections that nobody wants to defend.
6. Pricing: information is free, action is paid
At stake: how to monetise a capability that spans support and creation.
The PRD left this open. Section 10 lists it as TBD, because the paid tier’s own value proposition had not been settled yet, and pricing a feature into a bundle that has not been defined is guesswork.
Resolution: Ella is bundled into the paid tier. Free accounts get an Ella that answers questions and handles support. Paid accounts get the agentic half, the part that drafts and creates.
The line is clean enough to explain in a sentence, which is the test I use for a gate. Information is free. Action is paid. Support stays universal, which protects the deflection benefit for every creator, and the capability that saves creators the most time sits behind the tier.
The smaller decisions
| Decision | Reasoning |
|---|---|
| Selectable option chips, not free-text prompts | Reduces the effort of answering and removes the blank-page problem for creators who do not know the vocabulary |
| Ask only what is required to go live | Optional settings are not blocking and can be edited later; the goal is a reviewable draft, not a complete one |
| Explain the option, then recommend | A creator asked to choose between things they do not understand will either guess or stall |
| Voice input | See research section |
| No auto-redirect at handoff | Ella prompts the creator to go and review; it does not move them. Being relocated by software is disorienting |
| Saved as draft, never live | The handoff produces something safe to abandon |
| Persona detection with mid-conversation switching | A creator who arrives undecided and becomes decided should stop being asked discovery questions |
| “Created via Ella” tagging | Context survives the session, so later edits know where the item came from |
| Scope guardrail | Ella declines off-platform requests. It is not a free general assistant |
| Handoff carries a summary, not the transcript | See below |
One of these deserves more than a row. When Ella escalates to a human, it sends CX a summary of the creator’s problem rather than the full message history. It was a small call at spec time. In practice it changed how CX works: they open a ticket and read the problem, instead of reading a long conversation to find it. Fewer clarifying round-trips. The small interaction decisions are often the ones with operational consequences.
What shipped
13 August 2026. Both halves in the first release: agentic creation across all five product types, and supportive guidance with read-only account lookups. A dedicated full-page surface plus the persistent panel. Conversation history and multiple sessions. Human handoff. Voice input.
Gated by tier from day one. The previous support assistant was retired at launch.
The release was gated on a set of criteria I wrote as pass or fail, with a subset marked as hard gates: account scoping enforced at the data layer with zero cross-account exposure, the financial read-only boundary holding against direct instruction, no customer-facing or irreversible action committing without approval, and resistance to instructions embedded in inputs attempting to override any of the above.
How we knew it worked
Directional. Three weeks of production is not enough for settled numbers, and I would rather say that than imply otherwise.
Ticket volume fell. Measured by sampling and counting tickets across the three weeks before and the three weeks after release.
The residual queue changed shape, which matters more. What remains is largely what Ella is not scoped to see or do, for example creators reporting downtime during a scheduled maintenance window. The escalation boundary is behaving the way it was designed to. Volume falling tells you the agent is useful; the queue narrowing to out-of-scope issues tells you the scope guardrail is holding.
CX handling changed. Summarised handoffs replaced transcript-reading.
Creators named specific things, unprompted. In interviews after launch, creators reported Ella helping them understand their business performance, think through what to build, and produce products and emails. Two features came up by name without being asked about: voice input and the option picker.
We also ran an open trial on launch day where creators used Ella unsupervised and fed back. That produced a set of specific optimisation requests, which are documented and queued rather than absorbed silently.
Research and theory
I want to be precise about which research informed a decision at the time and which is a lens I applied afterward, because the difference is the whole value of citing anything.
Informed the decisions
Original AI UX research on orality → voice input. My own research into AI user experience found that speaking, rather than typing, has become a substantial part of how people actually use AI. Voice input does not appear in the PRD because the research postdated it. It was added before launch on that basis. Creators then named it unprompted in post-launch interviews, which is the loop closing: research, decision, validation.
Heuristic evaluation of mainstream generative AI interfaces → persona split, explain-then-recommend, option chips. Batool, A., & Hussain, W. (2025). Evaluating the usability and ethical implications of graphical user interfaces in generative AI systems. Computers, 14(10), 418. https://doi.org/10.3390/computers14100418
A heuristic and user-based evaluation of ChatGPT, Gemini and Claude against fourteen usability heuristics, with an ethics lens covering transparency, autonomy and error prevention. Three findings shaped Ella directly:
- Under adaptation to growth, participants criticised the one-size-fits-all pattern where a system offers experienced and new users identical guidance. Ella’s persona split is the direct answer: detect whether a creator arrives decided or undecided, and adapt.
- The guidance heuristic is defined as guiding the user to the next appropriate step and offering recommendations, rather than presenting choices and leaving them there. That is explain-then-recommend.
- Under error prevention, participants across all three products flagged missing confirmation for irreversible actions and absent consequence warnings. That is the case for approval gating, observed as a deficiency in the best tools on the market.
Limitation I should state rather than hide: the study’s twelve participants were research scientists at a national research organisation, so they were AI-literate professionals. My users are non-expert creators. If AI-literate experts hit those walls, non-experts hit them harder, but that is an inference I am drawing, not a finding the study makes.
Applied afterward, not at the time
Trust calibration. Araujo, T. (2026). Unpacking the dynamics of generative AI use in our daily lives: towards an integrative trust calibration framework. AI & SOCIETY. https://doi.org/10.1007/s00146-026-03279-0
Published 11 August 2026, two days before Ella shipped and roughly eight weeks after the decisions it explains. I did not have it when I decided. I am including it because it names what I was reaching for and, in one place, complicates it.
The framework concerns non-expert users calibrating trust in generative AI agents, and holds that the goal is trust matched to the system’s actual trustworthiness, avoiding both naive over-trust and blanket distrust. Read that way, generate-then-review is a calibration mechanism, not just a safety catch: it puts the creator in the position of evaluating output rather than receiving it, every time, which is how a person builds a working sense of what the agent is good at.
The paper also notes that an agent’s name is a system-level factor shaping how users assign trust. We named it “Ask Ella” to align with an existing campaign and to make the interaction feel like talking to someone. That was a marketing-inflected decision. It has trust consequences I had not thought about.
And it names something uncomfortable. Explanations can calibrate trust or produce over-trust. Ella explains its pricing reasoning, which I treated as unambiguously good. On this reading it is a live risk: a confident explanation of a price can make a creator more willing to accept a bad one. I do not have a resolution. It is the most useful thing I have read on this product since shipping it.
Neither
The financial read-only boundary and the decision to replace the support assistant were product judgment and operational reasoning. No literature behind either. Saying so seems better than finding some.
What I would do differently
The acceptance criteria are rigorous about whether Ella behaves correctly and silent on whether its output is good.
They will catch a chip that does not render, a scoping filter that leaks, an action that commits without approval. They will not catch Ella confidently recommending a price that is 40% too low. One criterion requires Ella to propose a price and explain its reasoning. Nothing anywhere defines what a good proposal is.
That is a real gap in a product whose outputs are non-deterministic, and it is more exposed given the over-trust risk above: an explained bad answer is worse than an unexplained one. I have since written the evaluation framework that should have been part of the original spec, covering what “good” means per output type, an offline evaluation set, review sampling, quality thresholds, a failure taxonomy, and operator controls.
Writing it after launch is the thing I would change. Not writing it would have been worse.
Artifacts
- Ella acceptance criteria (redacted) — release-gating criteria across navigation, conversation model, interaction fidelity, agentic creation, guardrails, data integrity and non-functional requirements
- Ella PRD (redacted) — problem, personas, scope, non-goals, open decisions
- Ella output evaluation framework — written post-launch, addressing the gap above
Attribution: I wrote the PRD and the acceptance criteria and owned the product and design decisions described here. Nelson (CTO) owned the architecture and engineering delivery, including how the account-scoping and financial boundaries are enforced in the system. Where this case study describes a technical guarantee, it describes what I asked engineering to guarantee, not how they built it.