Skip to content
MemCat
Classification before your LLMPrivate alpha

Classify in
milliseconds.

Handle clear messages locally. Send uncertain ones to your LLM. MemCat uses a small encoder and your examples to reduce the model calls spent choosing a category. Test the quality before switching.

1.8 msper message, medianLocal embedding + category decision

BANKING77 · Apple M4 Max · Node 22.22.3 · Warm CPU inference; excludes setup. See the measurement

97.5% accepted precision69.7% automatic coverage

54 wrong automatic routes in 3,080 test messages. Results for one banking policy; start in shadow mode.

TypeScript SDK · Your categories · Your existing model

Illustration · slowed to follow
Categories defined by examplesInspect every decisionEvaluate before routing
Measured local LLM comparison

770 messages.
Two ways to classify them.

BANKING77 · 77 categories · Qwen2.5 7B Q4_K_M · Apple M4 Max · local comparison

Qwen on every message

Message LLM Category

770 recorded classification requests
MemCat + Qwen fallback

Message MemCat Category or LLM

227 recorded fallback requests
Preset target not met

Fewer calls. Quality still needs work.

70.5% fewer LLM calls. Accuracy changed from 69.5% to 83.4%, below our preset 90% floor. This configuration did not pass.

Of 543 local decisions, 527 were correct. The actual Qwen fallback solved 115 of 227 deferred messages. The fallback accounts for most remaining errors.

Actual outcomes and model work on the same messages
MeasurementQwen on every messageMemCat + Qwen fallback
Correct / all messagesOverall accuracy535 / 77069.5%642 / 77083.4%
Wrong automatic labels218123
Sent for review175
Failed or not run00
LLM classification calls770227
Reported model tokensPrompt + generated2,103,931620,541
Total classification time299.67 s87.72 s

Correct means the final label matched the labelled answer. Reviews and failures count against overall accuracy.

Sum of measured query times; setup and warmups reported separately. Qwen's prompt cache was active. Tokens are returned usage counts, not dollar savings. Both routes made real requests. Reported prompt tokens can include cached work. Dollar savings and total compute cost are unmeasured.

Follow one message

The decisions behind the totals.

Recorded run · paced for explanation
How can I lookup where funds came from?
  1. MessageThe same recorded input
  2. MemCatverify_source_of_funds26.8 ms recorded
  3. ResultCompared with the labelled answer
Labelled answerverify_source_of_funds
Qwen on every message
verify_source_of_fundsMatches the labelled answer · 185.8 ms
MemCat + Qwen fallback
verify_source_of_fundsMatches the labelled answer · 26.8 ms
Case banking77-test-02022Inspect this receipt
Try your own message

The live browser lab runs the same 77-category banking policy. Load the model on demand, inspect the matches, and download your result. Your text stays in the browser.

One layer.
Your existing stack.

Describe your categories and supply examples. An encoder turns text into a vector; MemCat compares category examples and applies similarity and margin thresholds. Your application keeps control.

Your applicationA message + your category examples
MemCatEncoder → similarity → review threshold
Shadow mode firstCompare with your current result
Only after quality is acceptableSend accepted categories to a queue
Uncertain matchYour model or reviewer
  1. 01

    Show it what belongs.

    Provide labelled examples and an embedding adapter. Prepare category vectors once; compare new messages with them. Update examples as your categories change.

  2. 02

    Set a threshold, then test it.

    Use separate validation data to choose similarity and margin thresholds. Test on untouched messages. A cosine score is closeness, not the probability of being correct.

  3. 03

    Watch before you switch.

    Shadow mode observes MemCat without replacing your existing result. Log disagreements, failures and real usage. Route automatically only when your workflow’s quality target is met.

A category proposes a queue. It does not authorize a payment, refund or account change. Mixed-intent handling still has known failures.

Small enough to add.
Clear enough to inspect.

Example-based classification, an evaluator that counts mistakes, and shadow execution that leaves your current result in place. Run the banking demo locally with the ready-to-use Node starter, or bring an encoder for your own categories.

  • Category examples, similarity and review decisions
  • Coverage, precision, confusion matrices and raw receipts
  • Bounded callbacks, cancellation and explicit failure states

@memcat/sdk v0.3.1. Private alpha · UNLICENSED · not published on npm. The core package has no runtime dependencies; model weights and an inference runtime are separate. The Node starter includes a pinned encoder adapter for the 77 banking categories. No API key needed; model assets download on first use.

Evaluate. Observe. Then adopt.TypeScript
import {
  createClassifier,
  evaluateClassifier,
  createShadowClassifier,
} from "@memcat/sdk";

// Your encoder, categories and limits.
const classifier = await createClassifier(config);
const candidate = async (text) => {
  const result = await classifier.classify(text);
  if (result.status === "unavailable") {
    throw new Error("Candidate unavailable");
  }
  return { label: result.category }; // null = review
};

const report = await evaluateClassifier({
  ...evaluationSet, classify: candidate,
});
const shadow = createShadowClassifier({
  ...shadowPolicy, candidate,
});
const { baseline, observation } = await shadow.run(
  request, () => existingWorkflow(request),
);

Integration sketch. Your application supplies the encoder, categories, evaluation set and current workflow. The package includes runnable examples and a pilot guide.

Fewer calls.
Only if quality holds.

The comparison above measures classification requests, returned token counts and time on one local workload. Your total cost also includes encoding, setup, reviews, retries and maintenance. Customer savings and paid-provider bills remain unmeasured.

Required to justify switching

Encoder + fallback + retriesYour current workflow
See what is measured

Useful predictions. Visible limits.

The frozen banking policy made 54 wrong automatic decisions on 3,080 test messages. It also failed 3 of 14 authored stress cases, including a mixed-intent request. Shadow evaluation comes first; benchmark precision alone is not permission to automate consequential actions. Read the errors and methodology.

When a rule
is all you need.

The SDK also runs explicit rules. This older demo uses six case-sensitive prefixes on synthetic support text. Its speed measures string matching, not the semantic model above.

1,000 messages. 3.2 ms. Median local rule routing.

Six literal rules · Apple M4 Max · No embedding or model inference
What did we measure?

100 measured batches after 20 warmups per size. Cache off. Includes validation, rule checks, callbacks and traces; excludes setup and rendering. This is a separate workload from the semantic benchmark.

Short messagesMedianp95
1000.32 ms0.56 ms
1,0003.18 ms3.46 ms
10,00033.21 ms35.88 ms
Inspect the historical measurements
01

Incoming messages

1,000

Mixed requests, in their original order.

We need to change the owner of workspace #1.
I was charged for an extra seat on workspace #2. Could you check the invoice?
Can I book a demo of the enterprise features? Our enquiry reference is #3.
I was charged for an extra seat on workspace #4. Could you check the invoice?
Synthetic sample · seed 20261006 · 1,000 messages40 authored templates, sampled by seed. Not customer traffic or an accuracy score.
02 / Check
MemCat6 text rulesRuns locally
03

Clear next steps

A suggested queue for each match. Unfamiliar messages stay in review.

Run the SDK on real text, then inspect why each group was selected.

See the six rules behind this demo
"I was charged "Billing
"Please cancel my "Billing
"I cannot log in "Technical
"The app crashes "Technical
"Can I book a demo "Sales
"I need pricing for "Sales

These rules check exact, case-sensitive prefixes. First match wins; everything else goes to review. They don’t understand paraphrases or mixed intent. Categories don’t authorize actions.

Real SDK execution, synthetic dataset. No tickets are sent and no accounts are changed. Timing covers local rule checks, callbacks, and traces; it excludes rendering and model inference. Your rules and inputs determine coverage.

Start with a message.
Keep the evidence.

Run the local classifier. Inspect the decision. Test your own examples.

Open the classification lab