Skip to content

Benchmark

A vertical model is only credible if it is measured on a yardstick its maker does not grade for itself. So before flinq is a number, it is a benchmark.

BauSatz is an open, neutral benchmark for the specialist language of construction. It starts with German, built over 647 real GAEB project files and spanning 26 tasks across 8 families, and the roadmap extends it across European languages. It exercises the work that actually matters in AEC: retrieval, matching, classification and regression on bidding and BIM text.

The point is the neutrality: the test set is held out and no single vendor grades its own work on it. That lets any embedding model be compared on construction language on the same terms, which is what makes a result on it meaningful rather than self-graded. It is a credibility instrument, not a marketing chart, and it is still being built.

Two pilot models — flinq-pilot-otter and flinq-pilot-wombat — are currently available side by side behind the API (see Models). Per-model BauSatz scores for the pilot ids will be published once the pilot lineup is final; until then, the most meaningful benchmark is your own data — both models are one model parameter apart, so comparing them on your corpus is a five-minute exercise.

  • BauSatz is still in development; scores are a Borda count over the task families, and confidence intervals are still being finalized, so we do not claim to beat any specific competitor in a statistical sense.
  • The benchmark’s aim is broader than one model: an open, multi-language yardstick for construction language in Europe. flinq’s deepest-trained language is German.
  • Semantic textual similarity (STS) is not part of how we describe flinq’s capability. It is noisy across all models on this domain and is not a claim we make.

The point to keep: construction-specialized models measured on an open benchmark rather than a self-graded one.