General
Broad Chinese-language decision tasks across eight categories.
Weights coming soon
CHINESE-LANGUAGE DECISION MODELS / RESEARCH RELEASE 2026
Bringing System One Model
to Chinese-Language Tasks
An encoder-based model family for general, medical, legal and financial decisions. Calibrated candidate probabilities in a single forward pass.
Code and data resources linked · Model weights planned for open release
System One models offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. We introduce Chinese-Jev, a compact model for Chinese-language decision-making across general tasks and the medical, legal and financial domains. A unified data-processing protocol converts heterogeneous annotations into probability targets over candidate options. We first train the model on 10 million general-purpose decisions, then independently fine-tune medical, legal and financial specialists. All models share a lightweight encoder-only architecture and the same decision interface.
We introduce Chinese-Jev Bench (CJ-Bench), comprising 307,900 held-out decisions, to evaluate accuracy, calibration and latency. Chinese-Jev General achieves 69.20% accuracy and 3.78% expected calibration error on the general subset, with a median latency of 14 ms on one H200 GPU at batch size one. An INT8 browser prototype also enables local smartphone inference at approximately one second per decision on the tested iPhone 15 Pro configuration.
A general checkpoint followed by three independent domain fine-tunes. All variants use the same decision interface and approximately 322 million parameters.
Broad Chinese-language decision tasks across eight categories.
Weights coming soonMedical-domain specialist fine-tuned from the general checkpoint.
Weights coming soonLegal-domain specialist fine-tuned from the general checkpoint.
Weights coming soonFinancial-domain specialist fine-tuned from the general checkpoint.
Weights coming soonResults on general Chinese tasks and three specialized domains.
100,000 decisions · 8 task categories · approximately balanced decision types
Click a metric to sort| # | Model | Execution |
|---|
Latency: local models are measured on one H200 GPU, at batch size one after warm-up. Values are median end-to-end times. Jev uses a hosted API and includes network round-trip time.
Metrics. ACC and ECE are computed on single-label examples, including hard score targets. Soft targets are excluded. ECE uses 15 equal-width bins, native model probabilities and no additional target-dataset calibration. The component sizes above include all decisions; exact metric-eligible counts are not supplied in the paper.
Models. Qwen3.5-2B is evaluated in non-thinking mode by scoring candidate labels. Open-Jev refers to Zefan Cai’s Qwen-based decision model. SemIF reads decision probabilities from pretrained language models’ option-token logits. Laya Multilingual uses a compact multilingual encoder. Jev is accessed through its hosted API.
Chinese-Jev General is the first-stage model. Medical, Legal and Finance use three independently fine-tuned specialists with the same architecture and decision interface.
Selected Chinese inputs and their recorded predictions.
These examples show recorded outputs, not live inference. Reference labels and probabilities are retained from the evaluation records. Selected examples do not represent overall benchmark accuracy. Source and license information appears with each example; third-party data keeps its original terms.
The paper reports an INT8 browser prototype that performs tokenization and inference locally on a mobile device. An interactive demo is planned.
Tested configuration: iPhone 15 Pro · iOS 26.7 · INT8 · Browser CPU / WASM.
Approximately 1 s per decision after the model is loaded; initial download time is excluded.
Chinese-Jev is trained on 10 million general Chinese decisions. Medical, Legal and Finance specialists are independently fine-tuned from the general checkpoint. Each model has approximately 322 million parameters.
| Domain | Training decisions | Benchmark decisions |
|---|---|---|
| General | 10,000,000 | 100,000 |
| Medical | 3,000,000 | 200,000 |
| Legal | 24,000 | 4,300 |
| Finance | 24,000 | 3,600 |
Annotations are converted into candidate-probability targets for choice, noul and score. General training covers eight categories: aspect mention; aspect status and sentiment; category and intent; reading and reasoning; review rating; semantic matching; sentiment and emotion; textual inference and stance.
The training, development, calibration and test partitions are separate. The General benchmark is selected from a larger held-out test pool. Counts refer to decisions, not necessarily independent source records. Source datasets retain their original license terms.
The repository includes the dataset construction, fine-tuning, calibration and evaluation pipeline. Start with your own compatible model bundle and decision dataset.
Prepare → train → calibrate → evaluate. Includes the decision format, configuration examples and documentation.
Available on GitHub02 · EVALUATIONHeld-out benchmark spanning general, medical, legal and financial decision tasks.
View on Hugging Face03 · TRAINING DATADataset release and documentation for the Chinese-Jev data construction workflow.
View on Hugging FaceUse the CLI to prepare a typed-decision dataset, fine-tune a bundle, then calibrate and evaluate it. The repository supplies code and format documentation; trained Chinese-Jev weights are planned for a later release.
Read the quickstart ↗Chinese-Jev is an independent community project and is not affiliated with or endorsed by TypeSafe AI. The repository does not include its proprietary model, weights or training data. Original code is released under Apache-2.0; external models and datasets keep their own licenses.
If this work supports your research, cite the paper.
@misc{wang2026chinesejev,
title = {Chinese-Jev: Bringing System One Model to Chinese-Language Tasks},
author = {Zexiao Wang and Zihao Zhang and Xudong Wang and Pan Wang and Ziyi Ye and Haoyu Zhao and Zuxuan Wu and Shuicheng Yan},
year = {2026},
eprint = {2609.36965},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.36965}
}