Skip to main content
Each Jev request sees only the state you send it. What a customer told your app last month, such as their edition or their plan, is not in their next request unless your code puts it there. The memory system here does that with one extra Jev request per call and a Python dict. Our example routes support tickets. In August, a customer said they self-host on Windows Server. In September they write “Dashboards have been really slow since this morning”. Without memory, Jev sends that ticket to the cloud team instead of the self-hosted team. In the example below, adding memory sends all six later tickets to the expected teams. Routing each ticket alone gets two out of six correct. The simple memory system has two parts:
  • After each call, Jev sees the original request and its answer together. A Noul question, which we call the gate, asks whether the request states an important fact about the customer that should be remembered. If P(yes) is at least 0.5, the ticket and Jev’s answer are saved as a memory for that customer.
  • A wrapper function, ask_with_memory(), takes the same state and questions arguments as ask(), which is the TypeSafe Python SDK client.system_one method with the model pinned and the call cached, and returns the Jev response. Before each call, it adds the customer’s saved tickets and Jev’s answers to them to the state.
That wrapper is a basic agent harness. A harness is the code around a model that gives it extra capabilities such as tool calls and running in loops. This basic toy one only gives it memory in a plain Python dict with one list per customer ID. The wrapper makes two Jev requests, one after the other: the routing request, then the gate, which judges the ticket together with Jev’s answer. The wrapper returns the routing response whether or not anything is saved. Because ask_with_memory() takes the same arguments as ask() and returns the same response, you can swap one for the other without changing any other code. The wrapper finds the customer’s memories, adds them to the state and saves new ones itself. The example sends five earlier tickets from two customers through the wrapper, then routes three later messages from each customer twice: once with the memories it saved, once without. A table puts the two answers side by side for all six.

Setup

Set TYPESAFE_API_KEY. Every API call is cached in json_cache.json. Download it into the folder you run this code from to replay the published numbers instead of calling the API, or leave it out to run everything live. Numbers below came from jev-1.13.0 on 2026-09-28. The memory threshold of 0.5 is a starting point. You should evaluate it against your own examples.

Define the tickets

Two customers, five earlier tickets and three later messages, all written for this cookbook. C-1042 (Harbor Logistics) runs the self-hosted edition on Windows Server under an Enterprise contract. C-2210 (Fernwood Studio) is on the cloud edition’s self-serve Team plan. Three of the earlier tickets state a lasting fact: an edition, a plan, a contract. Two don’t: a phone at the airport, a public holiday. No later message mentions the customer’s setup, so without memory Jev has to guess it. Each ticket becomes a state with two parts: customer, who the customer is, and ticket, the ticket thread. The ticket_state() function builds a simplified ticket, like one from a helpdesk. In your app, the state is each ticket as your helpdesk sends it, and ask_with_memory() adds the customer’s saved tickets to it. The split into earlier tickets and later messages is only for this example.

Route a ticket

One Choice question picks the support team. The none_suitable option gives Jev an answer for when no team fits, and a ticket with too little context can end up there. The helper ask() is client.system_one with the model version fixed and every call cached. It takes a state and a dict of questions and returns the SDK’s own response object, so code that already reads client.system_one responses works with it unchanged.
Plain Jev sends C-1042’s slow dashboards to the cloud team. This ticket alone gives it no reason to do anything else.

Ask Jev whether the ticket is worth remembering

First, as_memory() pairs the original state with Jev’s answer: the memory the wrapper may save. The request key holds the state before memories were injected, and jev_answered maps each question to its answer. This keeps old memories from being nested inside every new memory. A second Jev call reads that pair. The gate asks whether the request states an important fact about the customer, such as their edition, operating system, plan or contract. Jev’s answer is there as context, but a fact that appears only in the answer does not count. Without that rule, a generic answer such as “General product questions” can outweigh a real fact in the request, and an answer that relied on an earlier memory can be saved again as if it were new.

Wrap the call so it remembers

The wrapper, ask_with_memory(), has the same arguments and return value as ask(). The memory store is a plain dict with one list per customer ID. The wrapper uses the ID in the state to select the customer’s memories.
Adding memory this way has four consequences for your app:
  • Every call makes two sequential requests, your call followed by the memory gate. The gate judges the ticket together with Jev’s answer, so it runs after the routing request. If your gate ignores the answer, it could instead be a second question in the routing request (see fan-out).
  • Memory only grows. Every saved memory goes into every later call for that customer. There is no size limit, no check for relevance and no expiry date here. The Next steps section has ideas for each.
  • Your questions never mention the memories. The memories go into the state under earlier_tickets_from_this_customer, and Jev reads the whole state. Nothing in ROUTE changes.
  • The state needs customer.id. Plain Jev also accepts a string as the state, but this wrapper needs the ID to find the customer’s list.

Send the earlier tickets and watch the answers change

Send the five earlier tickets through ask_with_memory(), in the order they arrived. The same customer sends two of the later messages again: slow dashboards and 20 more seats. That shows how each saved memory changes Jev’s answers. The repeated messages go through the wrapper too, so the gate judges them as well.
The three tickets that state an edition, a plan or a contract are saved at 0.91 to 0.92. T-102 is saved even though Jev routed it to the general team, because the gate counts the customer’s words, not the answer. The airport and public-holiday tickets score 0.07 and 0.10 and are not saved, so the answers after them do not change. Each saved memory changes the answers it is relevant to. For C-1042, the Windows ticket moves slow dashboards from the cloud team (0.80) to the Windows team (0.99). The Enterprise contract then settles seats at 0.99. For C-2210, the Team-plan ticket raises the cloud team from 0.75 to 0.91 and moves seats from the Enterprise desk (0.52) to billing (0.98). Adding the Enterprise memory lowers C-1042’s slow dashboards slightly, from 0.99 to 0.94, which is one reason to choose which memories to add (see Next steps). The gate saved none of the repeated messages.

Route the same messages for both customers

Each customer now sends the three later messages. Each message goes to plain ask() and to ask_with_memory(), with the same state and the same question.
With memory, all six messages reach the expected team, against two without. The same message now goes to a different team for each customer: slow dashboards go to the Windows team for C-1042 and stay with the cloud team for C-2210, and 20 more seats go to the Enterprise desk for C-1042 and to billing for C-2210. Where plain Jev was already right, memory raises the probability, from 0.62 to 0.99 for C-1042’s seats. The wrapper also judged the six later messages, and saved none: they state no important fact. Each row’s expected team follows from the customer’s setup in the earlier tickets: C-1042’s self-hosted edition and Enterprise contract, C-2210’s cloud Team plan.

Open it in the playground

The link holds the exact routing state used for C-1042’s first slow-dashboards call, including its memories at that point, plus the routing question.
Open C-1042’s ticket with its memories in the TypeSafe playground →

Next steps

  • Choose what your app should remember. Edit GATE to describe what is worth remembering for your own task, and test REMEMBER_AT against examples you have labelled.
  • Choose which memories to add. The wrapper adds every saved memory. A relevance question can select a smaller set once the list gets long.
  • Replace outdated memories. Ask whether a new memory replaces an older one, then remove the old entry in code when appropriate.
  • Keep memories between runs. Replace MEMORIES with a database keyed by customer ID. The in-memory dict disappears when the process exits.
  • Act on the route. Map each team to a handler in code, such as booking a call for Enterprise customers or sending a setup-specific article. See intent routing and confidence-based routing. Add your reply to the ticket thread as a support message, and the gate can remember what you did.

Appendix

Can a saved answer reinforce a mistake?

Yes. The gate ignores anything that appears only in Jev’s answer, so an answer can’t get a ticket saved by itself. Once a ticket is saved, though, its answer is saved with it, mistaken inference and all, and every later call for that customer reads it. A later answer can then repeat the mistake.

How should I choose the threshold?

The 0.5 in this wrapper is a starting value. To pick your own, label request-and-answer pairs from your traffic as worth remembering or not, and find the probability that best separates the two groups. Check it on pairs you didn’t use to pick it. Raise it if a misleading memory costs more than a missing one; lower it if the reverse. See Confidence and Confidence-Gated Routing for choosing thresholds and matching them to the risk of each action.