When building agentic workflows and educational compilers, developers quickly encounter an awkward bottleneck: using generative Large Language Models (LLMs) to make simple, structured decisions.
Prompting a 7B or 70B parameter model to return a JSON object like {"branch": "review_basics", "confidence": 0.92} requires streaming tokens sequentially, waiting hundreds of milliseconds (or seconds), spending valuable GPU cycles, and constantly guarding against JSON parsing failures or hallucinations.
Recently, TypeSafe AI emerged with Jev, pioneering what they describe as “System 1” Machine-Native Intelligence. Shortly after, the open-source community created NanoJev ( TianyuCodings/NanoJev ).
In this post, we explore what System 1 models are, how NanoJev works under the hood, and how we integrated it into EduVis v1.5.0 to achieve low-latency pedagogical routing, real-time remediation branching, and assessment item difficulty calibration.
1. The Bottleneck: Generation vs. Reflexes
In cognitive psychology, Daniel Kahneman introduced the distinction between two modes of thought:
- System 1: Fast, instinctive, automatic, and intuitive.
- System 2: Slow, deliberate, logical, and computationally expensive.
While System 1 and System 2 offer an intuitive conceptual metaphor rather than an absolute dichotomy of AI systems, the real engineering friction lies in the computational workflow and output modality:
- Generative LLMs produce autoregressive natural language text that downstream software must parse and interpret. Even when asked for a simple binary choice or category label, they decode token by token.
- Decision models evaluate structured choices, probabilities, or scores directly in a single forward pass without generating output tokens.
Traditional Autoregressive LLM:
Input Prompt ──▶ [Forward Pass] ──▶ Token 1
──▶ [Forward Pass] ──▶ Token 2 ... (Autoregressive decoding loop)
Decision Engine (Jev / NanoJev):
State + Question + Candidates ──▶ [Single Forward Pass] ──▶ Structured Choice / Logits / Score
Inference Latency vs. End-to-End Latency
A single forward pass fundamentally avoids autoregressive output decoding, enabling low-latency, token-free decision inference. However, it is essential to distinguish raw model forward-pass time from end-to-end application latency:
- Model Forward Pass: A compact 0.6B decision model can avoid autoregressive output decoding, potentially reducing inference latency substantially. Actual performance depends on the hardware, implementation, and workload.
- End-to-End Response: In production, total response time encompasses prompt tokenization, model loading, device memory transfers, request batching, and network transport. For example, TypeSafe AI reports a 70–500 ms hosted response range for its managed service—reflecting vendor-reported, end-to-end operational measurements rather than a universal constant for all decision models.
When used for control flow, routing, and classification, standard generative LLM calls introduce clear operational challenges:
- Unnecessary Token Overhead: Multi-token generation cycles and network round-trips can stall interactive student interfaces.
- Schema Fragility: Autoregressive models occasionally emit conversational preambles, trailing tokens, or malformed JSON that require defensive retry loops.
- Compute Inefficiency: Allocating multi-billion parameter autoregressive attention to select one of four predefined branches is an inefficient use of compute.
2. What is Jev and NanoJev?
TypeSafe AI’s Jev
Founded by former OpenAI researcher Diogo Almeida, Erik Gafni, and Sasha Sheng, TypeSafe AI designed Jev to make decisions inside software.
Named after the economist William Stanley Jevons (Jevons Paradox), the hypothesis is that making AI decisions substantially faster and cheaper will encourage software systems to embed machine intelligence into every micro-branch. Rather than generating text, Jev directly evaluates structured questions over unstructured context.
Open-Source NanoJev (TianyuCodings/NanoJev)
NanoJev is an open-source reproduction hosted on GitHub under the repository
TianyuCodings/NanoJev
. Pretrained model weights are distributed on Hugging Face using the checkpoint identifier C-Tianyu/NanoJev.
Built on top of a Qwen3-0.6B backbone, NanoJev strips away the autoregressive language generation head in favor of three core decision primitives:
| Primitive | Output Form | Mathematical Operation | Typical Use Case |
|---|---|---|---|
Choice |
Multi-candidate probability distribution | Softmax over candidate representations | Intent classification, tool routing, branch selection (\(2 \le N \le 255\)) |
Boolean |
Probability of proposition (\(p \in [0, 1]\)) | Sigmoid logit activation | Prerequisite gate checks, validity verification |
Score |
Expected scalar value over discrete tiers | Probability-weighted expectation \(\sum i \cdot p_i\) | Rubric scoring, difficulty grading (\(2 \le K \le 10\)) |
Note on Calibration: In NanoJev, raw confidence figures are derived directly from output logit distributions (via softmax or sigmoid). While structured and token-free, treating these figures as true empirical confidence requires domain-specific calibration (such as temperature scaling or isotonic regression on validation benchmarks) rather than assuming raw softmax outputs are inherently calibrated probabilities.
3. Integrating NanoJev into EduVis v1.5.0
EduVis is an open, curriculum-aware framework and compiler that separates pedagogical semantic meaning from presentation rendering.
In EduVis v1.5.0, we introduced a hybrid architecture: System 1 handles routing and decision gates, while generative LLMs handle creative question drafting.
Crucially, the decision engine does not assemble or guarantee educational correctness on its own. Instead, it proposes candidate pathways, classifications, and difficulty scores. EduVis’s deterministic compiler, intermediate representation (IR) schemas, and formal curriculum validation rules remain strictly responsible for validating and assembling compliant educational visual artifacts.
┌──────────────────────────────────────────────┐
│ Learner / Prompt Input │
└──────────────────────┬───────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ EduVis NanoJev System 1 Decision Engine │
│ ─────────────────────────────────────────────────────────── │
│ • Intent Pre-Routing (Choice) │
│ • Prerequisite Readiness Gate (Boolean) │
│ • Dynamic Remediation Router (Choice) │
│ • Item Difficulty Scoring Gate (Score) │
└──────────────────────┬──────────────────────────────────────┘
│ (Low-Latency Typed Decision Output)
┌────────────────────┴─────────────────────┐
▼ ▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ Dynamic IR Patch Compiler │ │ Creative LLM Generator │
│ (Instant Micro-Scaffolds) │ │ (Drafting Question Text) │
└─────────────┬─────────────┘ └─────────────┬─────────────┘
│ │
└────────────────────┬─────────────────────┘
▼
┌─────────────────────────────┐
│ EduVis Validation & Schema │
│ Compiler (Deterministic IR) │
└─────────────────────────────┘
4. Real-World Pedagogical Use Cases & Pipeline Execution
Below is an actual execution trace of the EduVis decision and dynamic remediation engine run via uv run python scripts/demo_nanojev.py:
[!IMPORTANT] Pipeline Demonstration & Backend Status
The following output is from a working EduVis v1.5.0 decision-pipeline demonstration. This run uses the deterministic heuristic backend (Neural Loaded: False), exercising the intent-classification, prerequisite-gating, remediation-routing, IR-patch generation, and difficulty-scoring interfaces.It demonstrates the pipeline’s execution flow and output structure, not neural inference by NanoJev.
When evaluating these results, it is helpful to distinguish three levels of evidence:
- Pipeline works: Demonstrated by this run (interfaces, routing, and IR-patch synthesis execute successfully).
- NanoJev neural backend works: Not demonstrated by this specific run (requires PyTorch, CUDA, and preloaded weights).
- Difficulty scores reflect real pedagogical difficulty: Not established by this run.
# Note: demo_nanojev.py defaults to heuristic fallback mode on CPU/non-CUDA environments.
# To force or configure the neural backend with local weights, set:
# EDUVIS_NANOJEV_ENABLED=true EDUVIS_NANOJEV_LOCAL_ONLY=true uv run python scripts/demo_nanojev.py
> uv run python scripts/demo_nanojev.py
======================================================================
🎯 EduVis v1.5.0 System 1 Decision & Dynamic Remediation Engine
======================================================================
⚙️ Backend Status : heuristic (CPU default)
📦 Model : C-Tianyu/NanoJev
🧠 Neural Loaded : False
1️⃣ Learning Intent Classification (Choice Decision)
--------------------------------------------------
Prompt : 'Generate 10 challenge questions on negative numbers and temperature jumps'
👉 Classified Topic : negative_numbers (Confidence: 36.8%)
Prompt : 'Create an introductory geometry lesson about triangle angles and parallel lines'
👉 Classified Topic : geometry_and_angles (Confidence: 36.8%)
Prompt : 'Build a diagnostic quiz for solving linear algebraic equations'
👉 Classified Topic : algebraic_equations (Confidence: 55.6%)
2️⃣ Prerequisite Readiness Assessment (Boolean Decision)
--------------------------------------------------
Concept: algebraic_equations | Mastery: 0.92 (Thresh: 0.70)
👉 Verdict: ✅ READY TO ADVANCE (Confidence: 57.8%)
Concept: calculus_derivatives | Mastery: 0.48 (Thresh: 0.75)
👉 Verdict: ⚠️ REMEDIATION NEEDED (Confidence: 67.0%)
3️⃣ v1.5.0 Dynamic Remediation Router (Micro-Scaffold Synthesis)
--------------------------------------------------
Misconception: 'Learner placed -5 to the right of zero on a scale.'
👉 Selected Pathway : Visual Number Line Scaffolding (Confidence: 35.7%)
👉 Scaffold Type : visual
👉 Patch Elements : ['number_line', 'step_hop', 'zero_anchor']
👉 Emitted IR Patch :
{
"patch_type": "micro_remediation_card",
"target_topic": "negative_numbers",
"pathway_key": "visual_number_line",
"title": "Targeted Scaffold: Visual Number Line Scaffolding",
"scaffold_type": "visual",
"pedagogical_focus": "Visual step-by-step directional hops to ground magnitude and negative direction.",
"diagnostic_anchor": "Learner placed -5 to the right of zero on a scale.",
"visual_components": [
"number_line",
"step_hop",
"zero_anchor"
],
"suggested_phase": "guided_practice"
}
Misconception: 'Learner subtracted (-4) - (-3) = -7 instead of -1.'
👉 Selected Pathway : Visual Number Line Scaffolding (Confidence: 37.5%)
👉 Scaffold Type : visual
👉 Patch Elements : ['number_line', 'step_hop', 'zero_anchor']
👉 Emitted IR Patch :
{
"patch_type": "micro_remediation_card",
"target_topic": "integers",
"pathway_key": "visual_number_line",
"title": "Targeted Scaffold: Visual Number Line Scaffolding",
"scaffold_type": "visual",
"pedagogical_focus": "Visual step-by-step directional hops to ground magnitude and negative direction.",
"diagnostic_anchor": "Learner subtracted (-4) - (-3) = -7 instead of -1.",
"visual_components": [
"number_line",
"step_hop",
"zero_anchor"
],
"suggested_phase": "guided_practice"
}
4️⃣ Question Difficulty & Rubric Estimation (Score Decision)
--------------------------------------------------
Item : 'What is 3 + (-2)?'
👉 Estimated Difficulty Level : 4.49 / 5.00 (Confidence: 55.7%)
Item : 'Calculate (-15) - (-8) + (-4).'
👉 Estimated Difficulty Level : 4.58 / 5.00 (Confidence: 63.1%)
Item : 'Prove that the sum of any two negative integers is always negative using formal induction on a bounded set.'
👉 Estimated Difficulty Level : 4.58 / 5.00 (Confidence: 63.1%)
======================================================================
🎉 EduVis v1.5.0 System 1 Decision Demonstration Complete!
======================================================================Understanding the Demonstration Output
1. Backend-Reported Confidence vs. Empirical Calibration
The displayed confidence figures (ranging around 35% to 67%) are backend-reported scores generated by the heuristic fallback rule engine.
They do not indicate that the system’s decisions are empirically calibrated probabilities. In production educational systems:
- Heuristic fallback scores represent deterministic weighting rules and keyword matches.
- Neural model softmax outputs reflect uncalibrated logit distributions.
- True calibrated confidence requires post-hoc techniques (e.g., temperature scaling, Platt scaling, or isotonic regression) evaluated against standardized educational benchmarks.
2. Difficulty Score Sanity Check
A close look at the difficulty estimation output reveals an important nuance:
What is 3 + (-2)?receives an estimated difficulty of 4.49 / 5.00.Calculate (-15) - (-8) + (-4).receives 4.58 / 5.00.Prove that the sum of any two negative integers is always negative using formal induction on a bounded set.receives 4.58 / 5.00.
An elementary single-digit addition problem should not score virtually the same cognitive complexity (4.49) as a formal mathematical induction proof (4.58). This does not mean the pipeline is broken—the interface, data contracts, and scoring pipelines executed cleanly. However, it highlights the limitations of the current heuristic’s scoring behavior and sensitivity to semantic difficulty.
Before claiming useful difficulty calibration in production, the scoring engine must be evaluated and tuned against a representative corpus of curriculum assessment items with independently assigned pedagogical difficulty ratings.
3. Structured Remediation & Deterministic IR Synthesis
Where the pipeline demonstrates its strongest value is in Step 3 (Dynamic Remediation Routing). Given an unstructured learner error trace, the engine:
- Selects the appropriate pedagogical pathway (
Visual Number Line Scaffolding). - Synthesizes a structured JSON Intermediate Representation (
IR Patch) with strict pedagogical anchors (diagnostic_anchor,visual_components). - Passes this patch directly to EduVis’s visual compiler to render concrete visual scaffolding without waiting on multi-token generative loops.
5. Resilient Fallback Architecture
To ensure EduVis runs reliably in varied environments—including WebAssembly / Pyodide in the browser Studio IDE, CPU-only servers, and offline CI pipelines—the engine implements an explicit dual-mode fallback:
- Neural Backend (
TianyuCodings/NanoJev
): When running in an environment with PyTorch, CUDA, and the
C-Tianyu/NanoJevmodel weights available locally, EduVis executes neural forward passes for learned semantic classification, routing, and scoring. - Deterministic Heuristic Fallback: On CPU-only environments without local weights, or inside browser Pyodide runtimes where loading a PyTorch neural checkpoint is impractical, EduVis switches automatically to rule-based heuristics (keyword taxonomy matching, error-pattern lookup tables, and readability indices).
Heuristic vs. Neural Mode: Heuristic fallback is a distinct operational mode designed for portability and zero-dependency predictability. While it ensures consistent pipeline execution in test suites and browser sandboxes, it relies on fixed rule heuristics rather than learned semantic representations, and should not be expected to provide the same contextual nuance or distribution-based confidence as neural inference.
6. Takeaways & The Future of System 1 in EdTech
- Avoid generative LLM calls for simple control-flow decisions: When a workflow only needs to classify an intent, verify a condition, or choose a branch from a fixed set of candidates, token-by-token generation adds latency and parsing overhead. Reserve generative LLMs for tasks that genuinely require synthesis and natural language articulation.
- Hybrid workflows provide the best balance: Combine token-free decision models for fast state evaluation, gating, and routing with generative LLMs for creative text authoring.
- Keep the compiler in charge: Educational software demands deterministic schemas, pedagogical validity, and strict curriculum alignment. Decision models should propose candidate pathways and score items, while deterministic compilers enforce validation rules and assemble the final output.
EduVis v1.5.0 with NanoJev integration is available on GitHub .
Try it via the EduVis CLI:
uv run eduvis decision intent "10 challenge problems on negative numbers" --json
uv run eduvis decision branch "Confused subtraction with negative signs" --json
uv run eduvis decision score "Prove the Pythagorean theorem for right triangles" --json
uv run python scripts/demo_nanojev.py