EVIDENCE · QUALITY AND PERFORMANCE
Benchmarks
Morphology quality and execution cost are measured as separate workloads. Every quality comparison includes raw and contract-adjusted results. Initialization, throughput, latency, and memory retain their own units.
These measurements are evidence for fixed inputs and settings. A high Canonical score does not establish recall for typos, spacing errors, or new repositories. Compare both precision and recall for smart and any in the Robust results below, then check fixed queries on your own documents. See the source reports for the sample scope and limitations of real-document evaluations.
Evaluation scope
Canonical contains 500 positive and 500 negative sentences manually reviewed for standard Korean. kfind, Kiwi, Lindera, MeCab-ko, and KOMORAN perform the same lemma, POS, and sentence task.
Robust contains 250 positive and 250 negative natural sentences with real errors. Its results are not combined with Canonical, and the robustness setting, error location, and denominator are recorded separately.
The query matrix applies multiple positive and negative queries to each sentence. It measures partial recovery and within-sentence false positives. A versioned product-contract registry defines its reviewed expectations.
Raw and contract-adjusted metrics
Raw metrics preserve the source corpus gold. TP, FP, TN, FN, precision, recall, and F1 are reported without modification.
Contract-adjusted metrics apply a registry fixed before product execution to the same predictions. It may reclassify semantically indistinguishable homographs, source-aligned internal components, and gold-span errors. Only nonstandard input outside the product contract may be excluded. Unsupported or expensive grammar remains in the denominator.
A contract-adjusted confusion matrix adds a superscript c to each raw abbreviation: TPᶜ, FPᶜ, TNᶜ, and FNᶜ. A dataset without contract reviews keeps its raw expectations and records a reviewed-case count of zero.
Canonical quality and performance
Confusion matrix
| Profile or product | Raw TP / TN / FP / FN | TPᶜ / TNᶜ / FPᶜ / FNᶜ |
|---|---|---|
| kfind embedded · any | 486 / 493 / 7 / 14 | 486 / 493 / 7 / 12 |
| kfind embedded · smart | 461 / 500 / 0 / 39 | 461 / 500 / 0 / 37 |
| kfind full POS · any | 497 / 493 / 7 / 3 | 497 / 493 / 7 / 1 |
| kfind full POS · smart | 498 / 500 / 0 / 2 | 498 / 500 / 0 / 0 |
| Kiwi | 418 / 500 / 0 / 82 | 418 / 500 / 0 / 80 |
| Lindera | 377 / 500 / 0 / 123 | 377 / 500 / 0 / 121 |
| MeCab-ko | 390 / 500 / 0 / 110 | 390 / 500 / 0 / 108 |
| KOMORAN | 393 / 500 / 0 / 107 | 393 / 500 / 0 / 105 |
Same-workload performance
Fresh processes run one warm-up followed by five measurements. Quality and cost remain separate.
| Profile or product | Initialization | cases/s | p95 | Peak RSS |
|---|---|---|---|---|
| kfind embedded · any | 0.0020 s | 52,504.5 | 0.0512 ms | 5.5 MiB |
| kfind embedded · smart | 0.0446 s | 37,824.3 | 0.0567 ms | 41.3 MiB |
| kfind full POS · any | 0.0356 s | 41,755.7 | 0.0731 ms | 21.0 MiB |
| kfind full POS · smart | 0.0773 s | 25,162.4 | 0.1105 ms | 56.6 MiB |
| Kiwi | 1.5310 s | 1,745.6 | 0.9617 ms | 528.5 MiB |
| Lindera | 0.0273 s | 24,199.6 | 0.0697 ms | 180.5 MiB |
| MeCab-ko | 0.0003 s | 10,946.6 | 0.1489 ms | 102.8 MiB |
| KOMORAN | 1.0870 s | 1,569.8 | 1.1011 ms | 756.1 MiB |
Query-matrix quality and performance
The query matrix selects up to three lemma-POS-span queries that should match in one source sentence and pairs each with a same-POS query that should not match. These values are aggregated per query and remain separate from the Canonical regression baseline.
Confusion matrix
| Profile or product | Raw TP / TN / FP / FN | TPᶜ / TNᶜ / FPᶜ / FNᶜ |
|---|---|---|
| kfind embedded · any | 1268 / 1272 / 24 / 28 | 1270 / 1272 / 21 / 26 |
| kfind embedded · smart | 1197 / 1292 / 4 / 99 | 1201 / 1293 / 0 / 95 |
| kfind full POS · any | 1291 / 1272 / 24 / 5 | 1293 / 1272 / 21 / 3 |
| kfind full POS · smart | 1292 / 1292 / 4 / 4 | 1296 / 1293 / 0 / 0 |
| Kiwi | 1108 / 1296 / 0 / 188 | 1107 / 1293 / 0 / 189 |
| Lindera | 994 / 1296 / 0 / 302 | 993 / 1293 / 0 / 303 |
| MeCab-ko | 1036 / 1296 / 0 / 260 | 1035 / 1293 / 0 / 261 |
| KOMORAN | 1064 / 1296 / 0 / 232 | 1063 / 1293 / 0 / 233 |
Same-workload performance
Fresh processes run one warm-up followed by five measurements. Quality and cost remain separate.
| Profile or product | Initialization | cases/s | p95 | Peak RSS |
|---|---|---|---|---|
| kfind embedded · any | 0.0020 s | 52,001.7 | 0.0508 ms | 8.4 MiB |
| kfind embedded · smart | 0.0425 s | 39,565 | 0.0539 ms | 44.0 MiB |
| kfind full POS · any | 0.0348 s | 45,002.9 | 0.0595 ms | 21.7 MiB |
| kfind full POS · smart | 0.0775 s | 25,472.5 | 0.1007 ms | 57.4 MiB |
| Kiwi | 1.4731 s | 1,701.5 | 0.9763 ms | 532.8 MiB |
| Lindera | 0.0282 s | 23,457.6 | 0.0735 ms | 201.3 MiB |
| MeCab-ko | 0.0003 s | 11,348.2 | 0.1404 ms | 103.9 MiB |
| KOMORAN | 1.0911 s | 1,837.9 | 0.9018 ms | 868.7 MiB |
Robust quality and performance
Confusion matrix
| Profile or product | Raw TP / TN / FP / FN | TPᶜ / TNᶜ / FPᶜ / FNᶜ |
|---|---|---|
| kfind embedded · any | 230 / 244 / 6 / 20 | 230 / 244 / 6 / 20 |
| kfind embedded · smart | 182 / 249 / 1 / 68 | 182 / 249 / 1 / 68 |
| kfind full POS · any | 232 / 244 / 6 / 18 | 232 / 244 / 6 / 18 |
| kfind full POS · smart | 201 / 249 / 1 / 49 | 201 / 249 / 1 / 49 |
| Kiwi | 213 / 250 / 0 / 37 | 213 / 250 / 0 / 37 |
| Lindera | 208 / 250 / 0 / 42 | 208 / 250 / 0 / 42 |
| MeCab-ko | 206 / 250 / 0 / 44 | 206 / 250 / 0 / 44 |
| KOMORAN | 205 / 250 / 0 / 45 | 205 / 250 / 0 / 45 |
Same-workload performance
Fresh processes run one warm-up followed by five measurements. Quality and cost remain separate.
| Profile or product | Initialization | cases/s | p95 | Peak RSS |
|---|---|---|---|---|
| kfind embedded · any | 0.0020 s | 54,253 | 0.0486 ms | 4.7 MiB |
| kfind embedded · smart | 0.0438 s | 38,596 | 0.0545 ms | 40.4 MiB |
| kfind full POS · any | 0.0362 s | 43,249.8 | 0.0798 ms | 20.8 MiB |
| kfind full POS · smart | 0.0797 s | 23,698.2 | 0.1203 ms | 56.5 MiB |
| Kiwi | 1.8788 s | 1,713.4 | 1.3710 ms | 527.5 MiB |
| Lindera | 0.0293 s | 24,956.2 | 0.1007 ms | 165.4 MiB |
| MeCab-ko | 0.0003 s | 11,454.7 | 0.2169 ms | 95.4 MiB |
| KOMORAN | 1.1644 s | 1,329.2 | 1.8062 ms | 522.4 MiB |
Morphology queries and regex baselines
Seven full-POS morphology queries with explicit POS run once with any and once with smart. The results are compared with a regex that manually enumerates inflected surfaces and another that lists only short stems. Each query has eight positives and eight negatives, for 112 cases.
This constructed fixture diagnoses coverage and boundary trade-offs for the same queries. It is neither a held-out quality benchmark nor a general ranking of Korean search quality.
Contract-adjusted values apply expectations fixed before execution to each method’s same predictions. The false-positive difference between any and smart remains visible; adjusted results never replace raw errors.
Confusion matrix
| Search strategy | Raw TP / TN / FP / FN | TPᶜ / TNᶜ / FPᶜ / FNᶜ |
|---|---|---|
| kfind full POS · any | 56 / 47 / 9 / 0 | 62 / 47 / 3 / 0 |
| kfind full POS · smart | 56 / 50 / 6 / 0 | 62 / 50 / 0 / 0 |
| Enumerated-surface regex | 50 / 48 / 8 / 6 | 55 / 47 / 3 / 7 |
| Short-stem regex | 46 / 22 / 34 / 10 | 52 / 22 / 28 / 10 |
Seven-query batch time
| Search strategy | Median | Minimum | Maximum | p95 | Effective throughput |
|---|---|---|---|---|---|
| kfind full POS · any | 491.54 ms | 488.00 ms | 507.89 ms | 507.89 ms | 192.8 MiB/s |
| kfind full POS · smart | 2311.31 ms | 2283.46 ms | 2421.43 ms | 2421.43 ms | 41 MiB/s |
| rg · enumerated surfaces | 64.94 ms | 63.03 ms | 73.39 ms | 73.39 ms | 1,458.9 MiB/s |
| grep · enumerated surfaces | 5266.18 ms | 5235.50 ms | 5341.44 ms | 5341.44 ms | 18 MiB/s |
| rg · short stems | 67.63 ms | 63.23 ms | 73.38 ms | 73.38 ms | 1,400.9 MiB/s |
| grep · short stems | 1786.89 ms | 1779.30 ms | 1989.68 ms | 1989.68 ms | 53 MiB/s |
d28fa2790b5592710291c5a90220698065ff8999 · 40e3c1fcfc7b803fe48121d8e73e7e7dde2a61de70f90213890e58b43bec5021
Benchmark contract and reports
Source evidence
The approved snapshot records the source report’s Git revision and SHA-256. The report contains the environment, fixture checksums, tool versions, and case-level failures.
d28fa2790b5592710291c5a90220698065ff8999 · 3258d172ce7098f9434150fb24d9a59fbedc73ec4dd0b84b17ca24062a68af28