Back
Legal & Contract Analysis

Three Domain-Specific Extraction Databases Built on One Shared Contract Engine

How Pfactorial Technologies fine-tuned one shared extraction architecture into three purpose-built databases for lease, service fee and royalty agreement terms - and tracked each one's real maturity honestly.

August 21, 2026
Share
ENGAGEMENT SNAPSHOT

Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 1
Figure 1 - Key figures from this engagement, at a glance.
EXECUTIVE SUMMARY
Building on a cleaned SEC/SEDAR contract corpus, our client needed structured databases for three distinct agreement categories - lease terms, service fees, and royalty rates - each with its own fields, terminology and extraction challenges.
Building three separate extraction systems from scratch would have tripled the engineering effort. But treating all three as identical problems would have ignored real differences: lease and royalty terms follow relatively consistent patterns, while service fee terms vary enough that even defining the extraction criteria required more up-front analysis.
Pfactorial fine-tuned one shared extraction architecture - QA, NER and classification models - three times, once per agreement category, and documented each database's actual maturity level explicitly rather than presenting all three as equally finished.
Why this engagement is representative This engagement is a good example of Pfactorial's engineering honesty in client deliverables: rather than presenting three databases as uniformly complete, the documentation states plainly which ones are production-stable, which are worked examples, and which still have open questions - because that distinction is exactly what a client needs to plan around.
THE CHALLENGE
Extending one extraction engine across three different agreement categories surfaced real differences in how tractable each category actually was.

1. Building three systems from scratch doesn't scale

Lease, service fee and royalty agreements each need their own target fields and extraction logic, but building each as an independent system would have multiplied engineering effort with no shared foundation.

2. Not every category is equally pattern-consistent

Lease and royalty terms follow relatively predictable structures; service fee terminology varies enough that even defining a reliable extraction pattern required substantially more upfront analysis.

3. Labeling effort at scale is a real resourcing question

With roughly 200,000 service agreements in scope and an estimated minimum of 5 minutes per agreement for adequate labeling depth, getting to a trained model for just one category was a multi-month resourcing commitment - one that needed to be surfaced honestly, not glossed over.
The real brief Not “build three extraction databases” but “build one reusable extraction foundation and apply it honestly to three categories with genuinely different levels of tractability - without pretending they're equally finished.”
THE SOLUTION
Pfactorial fine-tuned the same three-model architecture - QA for fee/rate fields, NER for parties and dates, classification for categorical fields - three times, once per agreement category, on top of the shared cleaned corpus.
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 2
Figure 2 - One extraction engine, fine-tuned three ways for three different agreement categories.

Architectural principles

  • Shared model types, domain-specific tuning - The same QA, NER and classification model types are used across all three databases, each fine-tuned on category-specific data rather than built as three separate architectures.
  • Real sample clauses before generalized patterns - For Lease, every field's extraction pattern started from real sample clauses before being generalized into keyword or regex patterns - grounding the pattern in actual document language.
  • Stability confirmed explicitly, not assumed - Royalty's keyword regex was reviewed and re-confirmed stable at a later checkpoint, with that confirmation recorded rather than left implicit.
  • Open questions documented, not hidden - Service Fee's extraction criteria being unresolved is stated as a known gap with an explicit resourcing estimate, not smoothed over as a minor detail.
CAPABILITIES DELIVERED
Capabilities differ meaningfully by database, reflecting each category's actual maturity at time of review.
CAPABILITY
WHAT IT DOES
Lease Rate Database
Category, rent amount, payment frequency, area, address, and start/end dates - with tested regex against real sample clauses.
Royalty Rate Database
Royalty rate, base, milestone type, exclusivity, license rights, territory and parties - with a confirmed, stable, misspelling-tolerant keyword regex.
Service Fee Database (planned)
Target fields and model architecture defined; extraction keyword criteria identified as an open task requiring further labeling investment.
Quarterly Lease updates
New SEC filing batches automatically filtered and checked against the existing Lease Structure field.
OCR-tolerant royalty matching
A fuzzy regex pattern tolerant of common misspellings and OCR errors (royatly, royally, roaylty, and similar variants).
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 3
Figure 3 - Extraction maturity, stated plainly rather than flattened into one status.
Design note Royalty's keyword regex includes deliberate misspelling and OCR-tolerance variants - a level of robustness not documented for any other single-term keyword set in this project, reflecting how much real-world text variation that specific field encounters.
ENGINEERING FOR SCALE AND RELIABILITY
Five decisions kept the actual state of each database visible, rather than letting documentation drift ahead of reality.

Per-database maturity notes

Each database's section states its own maturity level plainly - Lease as a worked example, Royalty as confirmed stable, Service Fee as still in the planning stage - rather than presenting a single unified status.

Explicit labeling-cost estimation

The Service Fee labeling estimate was revised once real per-agreement labeling time was understood - from an initial 1-minute assumption to a realistic 5-minute minimum - and the revised, larger cost was surfaced rather than kept at the more favorable initial number.

Known gaps recorded, not silently closed

A two-quarter update backlog for the Lease Rate Database, and an unresolved rentals-as-percentage-of-sales edge case, are both documented as open items rather than omitted.

Image-agreement limitation acknowledged

At least one image-based royalty agreement was identified as unprocessable until OCR completes and re-enters the standard pipeline - a shared limitation across all three databases, explicitly noted rather than treated as an isolated edge case.

Cross-database comparison as a deliverable

A side-by-side maturity comparison across all three databases was produced specifically to make the resourcing gap visible at a glance, rather than requiring a reader to infer it separately from each section.
DELIVERY APPROACH
The engagement extended the shared extraction foundation to each category in sequence, documenting real progress as it happened.
1. Shared architecture extension - adapting the QA/NER/classification model pattern for fee, rate and party field types across all three categories.
2. Lease Rate Database - field extraction pipeline built and tested against real sample clauses, with quarterly update automation.
3. Royalty Rate Database - keyword regex development, refinement, and later-checkpoint stability confirmation.
4. Service Fee Database scoping - target field and model architecture definition, plus labeling feasibility analysis.
5. Cross-database maturity documentation - side-by-side comparison and known-gaps reporting across all three databases.
RESULTS AND IMPACT

Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 4
Figure 4 - Key outcomes from this engagement.
Two of the three domain databases - Lease and Royalty - reached working, tested extraction status on top of the shared engine, while Service Fee's real resourcing requirement was surfaced clearly rather than left as an underestimated afterthought.
The honest maturity tracking gave the client an accurate basis for prioritizing further investment, rather than a status report that implied uniform completeness across all three databases.

What it enabled commercially

The client got two production-ready extraction databases plus a clear, resourced picture of what completing the third would actually require - letting them make an informed investment decision instead of discovering the Service Fee gap after committing further budget.
WHY PFACTORIAL
This engagement reflects a Pfactorial value that shows up repeatedly in client-facing documentation: distinguishing what's actually finished from what's still in progress, even when a less careful summary would present everything as done.
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 5
Figure 5 - Service lines this engagement draws on.
Engagement enquiries Pfactorial Technologies works with legal-tech and contract-analytics teams extending a shared extraction foundation across multiple document categories. If you're evaluating a multi-category contract extraction initiative, we're happy to give you an honest read on scope and risk before anyone commits to anything. · pfactorial.ai
APPENDIX A - TECHNOLOGY STACK
The technology stack underpinning the system, grouped by the layer it serves.
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 6
© 2026 Pfactorial Technologies. Client identity and product-specific implementation detail are withheld or generalized; no client data, credentials, source code, or infrastructure detail is included in this document.

Result and Analysis

ENGAGEMENT SNAPSHOT

How Pfactorial Technologies fine-tuned one shared extraction architecture into three purpose-built databases for lease, service fee and royalty agreement terms - and tracked each one's real maturity honestly.

Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 1
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 2
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 3
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 4
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 5
Pfactorial_Case_Study_Material_Contracts_Domain_Databases image 6