Classify in
milliseconds.
Handle clear messages locally. Send uncertain ones to your LLM. MemCat uses a small encoder and your examples to reduce the model calls spent choosing a category. Test the quality before switching.
1.8 msper message, medianLocal embedding + category decisionBANKING77 · Apple M4 Max · Node 22.22.3 · Warm CPU inference; excludes setup. See the measurement
54 wrong automatic routes in 3,080 test messages. Results for one banking policy; start in shadow mode.
TypeScript SDK · Your categories · Your existing model
770 messages.
Two ways to classify them.
BANKING77 · 77 categories · Qwen2.5 7B Q4_K_M · Apple M4 Max · local comparison
Message LLM Category
770 recorded classification requestsMessage MemCat Category or LLM
227 recorded fallback requestsFewer calls. Quality still needs work.
70.5% fewer LLM calls. Accuracy changed from 69.5% to 83.4%, below our preset 90% floor. This configuration did not pass.
Of 543 local decisions, 527 were correct. The actual Qwen fallback solved 115 of 227 deferred messages. The fallback accounts for most remaining errors.
| Measurement | Qwen on every message | MemCat + Qwen fallback |
|---|---|---|
| Correct / all messagesOverall accuracy | 535 / 77069.5% | 642 / 77083.4% |
| Wrong automatic labels | 218 | 123 |
| Sent for review | 17 | 5 |
| Failed or not run | 0 | 0 |
| LLM classification calls | 770 | 227 |
| Reported model tokensPrompt + generated | 2,103,931 | 620,541 |
| Total classification time | 299.67 s | 87.72 s |
Correct means the final label matched the labelled answer. Reviews and failures count against overall accuracy.
Sum of measured query times; setup and warmups reported separately. Qwen's prompt cache was active. Tokens are returned usage counts, not dollar savings. Both routes made real requests. Reported prompt tokens can include cached work. Dollar savings and total compute cost are unmeasured.
The decisions behind the totals.
- MessageThe same recorded input
- MemCatverify_source_of_funds26.8 ms recorded
- ResultCompared with the labelled answer
The live browser lab runs the same 77-category banking policy. Load the model on demand, inspect the matches, and download your result. Your text stays in the browser.
One layer.
Your existing stack.
Describe your categories and supply examples. An encoder turns text into a vector; MemCat compares category examples and applies similarity and margin thresholds. Your application keeps control.
- 01
Show it what belongs.
Provide labelled examples and an embedding adapter. Prepare category vectors once; compare new messages with them. Update examples as your categories change.
- 02
Set a threshold, then test it.
Use separate validation data to choose similarity and margin thresholds. Test on untouched messages. A cosine score is closeness, not the probability of being correct.
- 03
Watch before you switch.
Shadow mode observes MemCat without replacing your existing result. Log disagreements, failures and real usage. Route automatically only when your workflow’s quality target is met.
A category proposes a queue. It does not authorize a payment, refund or account change. Mixed-intent handling still has known failures.
Small enough to add.
Clear enough to inspect.
Example-based classification, an evaluator that counts mistakes, and shadow execution that leaves your current result in place. Run the banking demo locally with the ready-to-use Node starter, or bring an encoder for your own categories.
- Category examples, similarity and review decisions
- Coverage, precision, confusion matrices and raw receipts
- Bounded callbacks, cancellation and explicit failure states
@memcat/sdk v0.3.1. Private alpha · UNLICENSED · not published on npm. The core package has no runtime dependencies; model weights and an inference runtime are separate. The Node starter includes a pinned encoder adapter for the 77 banking categories. No API key needed; model assets download on first use.
import {
createClassifier,
evaluateClassifier,
createShadowClassifier,
} from "@memcat/sdk";
// Your encoder, categories and limits.
const classifier = await createClassifier(config);
const candidate = async (text) => {
const result = await classifier.classify(text);
if (result.status === "unavailable") {
throw new Error("Candidate unavailable");
}
return { label: result.category }; // null = review
};
const report = await evaluateClassifier({
...evaluationSet, classify: candidate,
});
const shadow = createShadowClassifier({
...shadowPolicy, candidate,
});
const { baseline, observation } = await shadow.run(
request, () => existingWorkflow(request),
);Integration sketch. Your application supplies the encoder, categories, evaluation set and current workflow. The package includes runnable examples and a pilot guide.
Fewer calls.
Only if quality holds.
The comparison above measures classification requests, returned token counts and time on one local workload. Your total cost also includes encoding, setup, reviews, retries and maintenance. Customer savings and paid-provider bills remain unmeasured.
Required to justify switching
Useful predictions. Visible limits.
The frozen banking policy made 54 wrong automatic decisions on 3,080 test messages. It also failed 3 of 14 authored stress cases, including a mixed-intent request. Shadow evaluation comes first; benchmark precision alone is not permission to automate consequential actions. Read the errors and methodology.
When a rule
is all you need.
The SDK also runs explicit rules. This older demo uses six case-sensitive prefixes on synthetic support text. Its speed measures string matching, not the semantic model above.
1,000 messages. 3.2 ms. Median local rule routing.
Six literal rules · Apple M4 Max · No embedding or model inferenceWhat did we measure?
100 measured batches after 20 warmups per size. Cache off. Includes validation, rule checks, callbacks and traces; excludes setup and rendering. This is a separate workload from the semantic benchmark.
| Short messages | Median | p95 |
|---|---|---|
| 100 | 0.32 ms | 0.56 ms |
| 1,000 | 3.18 ms | 3.46 ms |
| 10,000 | 33.21 ms | 35.88 ms |
Incoming messages
1,000Mixed requests, in their original order.
Clear next steps
A suggested queue for each match. Unfamiliar messages stay in review.
Run the SDK on real text, then inspect why each group was selected.
See the six rules behind this demo
"I was charged "Billing"Please cancel my "Billing"I cannot log in "Technical"The app crashes "Technical"Can I book a demo "Sales"I need pricing for "SalesThese rules check exact, case-sensitive prefixes. First match wins; everything else goes to review. They don’t understand paraphrases or mixed intent. Categories don’t authorize actions.
Real SDK execution, synthetic dataset. No tickets are sent and no accounts are changed. Timing covers local rule checks, callbacks, and traces; it excludes rendering and model inference. Your rules and inputs determine coverage.
Start with a message.
Keep the evidence.
Run the local classifier. Inspect the decision. Test your own examples.