Reverse-engineered Jev-like model vinnylarouge / jevlike Public Notifications You must be signed in to change notification settings Fork 23 Star 230 mainBranchesTagsGo to fileCodeOpen more actions me
By Coderz Club · 2026-09-17 · Tags: git
Reverse-engineered Jev-like model
vinnylarouge / jevlike Public Notifications You must be signed in to change notification settings Fork 23 Star 230 mainBranchesTagsGo to fileCodeOpen more actions menuLatest commit History3 Commits3 CommitsFolders and filesNameNameLast commit messageLast commit datedocsdocs examplesexamples jevlikejevlike scriptsscripts teststests .gitignore.gitignore AGENTS.mdAGENTS.md LICENSELICENSE README.mdREADME.md pyproject.tomlpyproject.toml View all filesRepository files navigationJevlike Train a small model that chooses among a changing list of text options. A Jev-like model takes a piece of text and a list of N text options. It returns one probability for each option. It does this in one pass instead of writing an answer word by word. Jev is TypeSafe's commercial model for this kind of task. TypeSafe has not published its design. This repository is an independent starter model with the same input and output shape. Demo The same option-attention head can score controller buttons from image patches. This ten-second film joins two selected five-second windows: live deadly_corridor combat on the seven Doom buttons, then a chess controller walking to and playing moves with five keys. The diagram shows the tensors used for each decision. The Doom window came from the supplied joint checkpoint, which averaged 0.60 kills and -97.50 reward across its ten recorded episodes. The chess window came from the stronger chess-only checkpoint, which scored 4 wins, 46 draws and 0 losses in 50 sampled games against a random mover, but 0 wins, 2 draws and 48 losses against Stockfish level 0. The windows were selected for activity and are not typical-play or competence claims. Install the game extras and record a fresh 640 by 480 Doom trace from the released joint checkpoint: uv pip install -e '.[games]' python examples/doom/play.py examples/checkpoints/joint-imitation.pt --episodes 10 --game-seconds 35.3 --device cpu --capture-resolution 640x480 --output runs/doom.mp4 --trace runs/doom-trace.json Render the trace in the same visual layout. This writes a silent film because the author-owned soundtrack source is not part of the repository. (cd examples/film && npm install && npx playwright install chromium) examples/film/make-film.sh runs/doom-trace.json runs/doom-film.mp4 10 The release includes the Doom example, the chess example, the single-game checkpoints and the shared 12-option checkpoint. Both games import the visual scorer from jevlike.vision; there is no second model copy in either example. Architecture Each option becomes a query vector, which is a short list of numbers representing its text. The query assigns attention weights to the context tokens. Those weights make one context vector for that option. A shared dot product turns each option and context pair into one score. A softmax, which converts scores into probabilities that sum to one, runs across the options. The default encoder learns byte embeddings from scratch. An encoder is the part that turns text into vectors. The optional Hugging Face path uses a frozen pretrained encoder, whose existing weights stay fixed while the small scorer learns. Data format Use one JSON object per line: {"context":"The customer needs a refund.","options":["refund","sales","technical support"],"label":0} label is the zero-based index of the correct option. Each row may have a different number of options, with a minimum of two. Quickstart Run these commands from the repository root. They create local synthetic data, train on it, evaluate the saved model and score one new menu. uv venv source .venv/bin/activate uv pip install -e '.[dev]' jevlike-data synthetic --output data/synthetic jevlike-train data/synthetic/train.jsonl \ --validation data/synthetic/validation.jsonl \ --output runs/synthetic.pt jevlike-eval runs/synthetic.pt data/synthetic/test.jsonl jevlike-predict runs/synthetic.pt \ --context "Choose the exact badge amber badger. Badge: amber badger." \ --option "azure crane" \ --option "amber badger" \ --option "gold heron" The evaluation prints top-1 accuracy, which is the fraction of correct first choices. Top-3 accuracy is the fraction with the right answer among the three highest scores. Expected calibration error compares confidence with observed accuracy. The command also prints a shuffled-context control, which pairs each menu with the wrong context. A useful model should beat that control. Use your own data Export train, validation and test JSONL files in the format above. Keep all options that the model will see at prediction time in each row. Split related records together. For example, keep all records for one customer or one target page in one split. This prevents near-duplicates from leaking into the test set. Run jevlike-train with your train and validation files. Run jevlike-eval once on the held-out test file. Held-out means the file was never used for training or model selection. The default byte encoder truncates context to 192
vinnylarouge / jevlike Public Notifications You must be signed in to change notification settings Fork 23 Star 230 mainBranchesTagsGo to fileCodeOpen more actions menuLatest commit History3 Commits3 CommitsFolders and filesNameNameLast commit messageLast commit datedocsdocs examplesexamples jevlikejevlike scriptsscripts teststests .gitignore.gitignore AGENTS.mdAGENTS.md LICENSELICENSE README.mdREADME.md pyproject.tomlpyproject.toml View all filesRepository files navigationJevlike Train a small model that chooses among a changing list of text options. A Jev-like model takes a piece of text and a list of N text options. It returns one probability for each option. It does this in one pass instead of writing an answer word by word. Jev is TypeSafe's commercial model for this kind of task. TypeSafe has not published its design. This repository is an independent starter model with the same input and output shape. Demo The same option-attention head can score controller buttons from image patches. This ten-second film joins two selected five-second windows: live deadly_corridor combat on the seven Doom buttons, then a chess controller walking to and playing moves with five keys. The diagram shows the tensors used for each decision. The Doom window came from the supplied joint checkpoint, which averaged 0.60 kills and -97.50 reward across its ten recorded episodes. The chess window came from the stronger chess-only checkpoint, which scored 4 wins, 46 draws and 0 losses in 50 sampled games against a random mover, but 0 wins, 2 draws and 48 losses against Stockfish level 0. The windows were selected for activity and are not typical-play or competence claims. Install the game extras and record a fresh 640 by 480 Doom trace from the released joint checkpoint: uv pip install -e '.[games]' python examples/doom/play.py examples/checkpoints/joint-imitation.pt --episodes 10 --game-seconds 35.3 --device cpu --capture-resolution 640x480 --output runs/doom.mp4 --trace runs/doom-trace.json Render the trace in the same visual layout. This writes a silent film because the author-owned soundtrack source is not part of the repository. (cd examples/film && npm install && npx playwright install chromium) examples/film/make-film.sh runs/doom-trace.json runs/doom-film.mp4 10 The release includes the Doom example, the chess example, the single-game checkpoints and the shared 12-option checkpoint. Both games import the visual scorer from jevlike.vision; there is no second model copy in either example. Architecture Each option becomes a query vector, which is a short list of numbers representing its text. The query assigns attention weights to the context tokens. Those weights make one context vector for that option. A shared dot product turns each option and context pair into one score. A softmax, which converts scores into probabilities that sum to one, runs across the options. The default encoder learns byte embeddings from scratch. An encoder is the part that turns text into vectors. The optional Hugging Face path uses a frozen pretrained encoder, whose existing weights stay fixed while the small scorer learns. Data format Use one JSON object per line: {"context":"The customer needs a refund.","options":["refund","sales","technical support"],"label":0} label is the zero-based index of the correct option. Each row may have a different number of options, with a minimum of two. Quickstart Run these commands from the repository root. They create local synthetic data, train on it, evaluate the saved model and score one new menu. uv venv source .venv/bin/activate uv pip install -e '.[dev]' jevlike-data synthetic --output data/synthetic jevlike-train data/synthetic/train.jsonl \ --validation data/synthetic/validation.jsonl \ --output runs/synthetic.pt jevlike-eval runs/synthetic.pt data/synthetic/test.jsonl jevlike-predict runs/synthetic.pt \ --context "Choose the exact badge amber badger. Badge: amber badger." \ --option "azure crane" \ --option "amber badger" \ --option "gold heron" The evaluation prints top-1 accuracy, which is the fraction of correct first choices. Top-3 accuracy is the fraction with the right answer among the three highest scores. Expected calibration error compares confidence with observed accuracy. The command also prints a shuffled-context control, which pairs each menu with the wrong context. A useful model should beat that control. Use your own data Export train, validation and test JSONL files in the format above. Keep all options that the model will see at prediction time in each row. Split related records together. For example, keep all records for one customer or one target page in one split. This prevents near-duplicates from leaking into the test set. Run jevlike-train with your train and validation files. Run jevlike-eval once on the held-out test file. Held-out means the file was never used for training or model selection. The default byte encoder truncates context to 192