Benchmark
A vertical model is only credible if it is measured on a yardstick its maker does not grade for itself. So before flinq is a number, it is a benchmark.
BauSatz
Section titled “BauSatz”BauSatz is an open, neutral benchmark for the specialist language of construction. It starts with German, built over 647 real GAEB project files and spanning 26 tasks across 8 families, and the roadmap extends it across European languages. It exercises the work that actually matters in AEC: retrieval, matching, classification and regression on bidding and BIM text.
The point is the neutrality: the test set is held out and no single vendor grades its own work on it. That lets any embedding model be compared on construction language on the same terms, which is what makes a result on it meaningful rather than self-graded. It is a credibility instrument, not a marketing chart, and it is still being built.
Scores during the pilot
Section titled “Scores during the pilot”Two pilot models — flinq-pilot-otter and flinq-pilot-wombat — are currently
available side by side behind the API (see Models). Per-model
BauSatz scores for the pilot ids will be published once the pilot lineup is
final; until then, the most meaningful benchmark is your own data — both models
are one model parameter apart, so comparing them on your corpus is a
five-minute exercise.
Reading these figures honestly
Section titled “Reading these figures honestly”- BauSatz is still in development; scores are a Borda count over the task families, and confidence intervals are still being finalized, so we do not claim to beat any specific competitor in a statistical sense.
- The benchmark’s aim is broader than one model: an open, multi-language yardstick for construction language in Europe. flinq’s deepest-trained language is German.
- Semantic textual similarity (STS) is not part of how we describe flinq’s capability. It is noisy across all models on this domain and is not a claim we make.
The point to keep: construction-specialized models measured on an open benchmark rather than a self-graded one.