
Back
Healthcare & Clinical
Generating Realistic Synthetic Datasets Without Exposing Real Customer Data
A live, natural-language-driven tool that produces industry-specific structured and unstructured datasets for testing, AI training, analytics, and demonstrations, without using any real customer or patient information.
August 21, 2026
Share
ENGAGEMENT SNAPSHOT

Figure 1 - Key figures from this engagement, at a glance.
EXECUTIVE SUMMARY
Organizations regularly need realistic datasets for software testing, AI model development, analytics, and product demonstrations, but using real customer or patient data for those purposes carries privacy risk and, in regulated industries, direct compliance exposure. Real data is also often limited, sensitive, expensive, or simply unavailable at the volume a team actually needs.
Pfactorial Technologies built and deployed Synthora, a synthetic data generation tool available as a live application. A user describes the type of data they need in natural language, and the platform proposes an appropriate schema or content structure, which can be reviewed and refined before the dataset is generated - producing production-ready output in minutes rather than requiring a bespoke data-generation script per request.
The tool spans both structured data, for databases, spreadsheets, and analytics, and unstructured data, such as clinical notes, customer conversations, contracts, emails, and reports, across industry domains including healthcare, finance, retail, legal, human resources, and manufacturing.
Why this engagement is representative This engagement demonstrates Pfactorial's approach to privacy-preserving data infrastructure: making synthetic data generation accessible through a natural-language interface rather than a bespoke script per request, while keeping every generated dataset free of real personally identifiable information by design.
THE CHALLENGE
Teams that need realistic datasets for testing, training, or demonstration purposes run into the same set of problems whether the underlying use case is a QA suite, a machine learning model, or a sales demo.
1. Testing against live customer data is risky
Using real customer data to test applications before production creates unnecessary exposure; a synthetic alternative needs to be realistic enough to catch real issues without carrying that risk.
2. Real-world data for AI training is often limited or sensitive
Building representative datasets for AI models is difficult when real-world data is limited, sensitive, expensive to license, or simply unavailable at the volume a model needs.
3. Regulatory exposure follows real PII wherever it's used
Any workflow that touches personally identifiable information carries regulatory exposure under frameworks such as HIPAA and GDPR, which a synthetic dataset with no real customer or patient information avoids entirely.
4. Demonstrations need realistic data, not empty screens
Showing a product convincingly requires realistic business scenarios rather than empty screens or manually assembled sample records, which are time-consuming to build and maintain by hand.
The real brief Not "generate some sample rows" but "produce production-ready, industry-realistic datasets in minutes, from a natural-language description, without touching real customer or patient data."
THE SOLUTION
Pfactorial built Synthora around a five-step workflow that takes a user from an industry and format selection to a reviewed, exported dataset, without requiring the user to write a data-generation script or define a schema by hand.

Figure 1 - The five-step workflow: industry, format, natural-language requirements, generation configuration, and export.
Architectural principles
- Natural language, not a schema editor, is the starting point - a user describes the dataset they need in plain language, and the platform proposes an appropriate schema or content structure automatically, which can then be reviewed and refined before generation.
- Industry context shapes the output - selecting a business domain - healthcare, finance, retail, legal, human resources, manufacturing, and others - lets the generated data reflect industry-specific terminology and realistic business scenarios rather than generic placeholder content.
- Structured and unstructured data are both first-class outputs - the platform generates structured data for databases, spreadsheets, and analytics, as well as unstructured data such as clinical notes, customer conversations, contracts, emails, and reports, from the same workflow.
- Real-world imperfection is a configurable option, not an afterthought - generation options include deliberately introducing missing values, duplicates, formatting inconsistencies, and typographical errors, so test data can reflect real-world data quality conditions rather than only a clean ideal case.
CAPABILITIES DELIVERED
The platform's capabilities span the full workflow from describing a dataset to exporting it in a ready-to-use format.
CAPABILITY | WHAT IT DOES |
|---|---|
Industry Selection | Scopes generated data to a specific business domain so terminology and scenarios reflect that industry. |
Structured & Unstructured Generation | Produces database- and spreadsheet-ready structured data, or free-form unstructured content such as notes, conversations, and contracts. |
Natural-Language Requirement Capture | Proposes a schema or content structure from a plain-language description, reviewable and refinable before generation runs. |
Configurable Generation Parameters | Controls record count, geographic region and language, and deliberate data-quality variation to match real-world conditions. |
Quality Review Before Export | Lets users preview generated data and review quality metrics prior to finalizing a dataset. |
Multi-Format Export | Exports generated datasets as CSV, JSON, SQL, or ZIP packages for immediate use downstream. |

Figure 2 - Why organizations use synthetic data, and what the platform is designed to replace.
Design note The source material describing this engagement is a short overview of a live, deployed application rather than a detailed technical specification; this case study is drafted conservatively from what that overview states, without extending into implementation detail the source does not cover.
ENGINEERING FOR SCALE AND RELIABILITY
The platform's design choices, as described, center on making synthetic data generation accessible without sacrificing industry realism or configurability.
Schema proposal reduces setup effort
rather than requiring a user to define a schema manually, the platform proposes one from a natural-language description, with review and refinement built into the workflow before generation.
Region and language are explicit generation parameters
geographic region and language are configured directly as part of dataset generation, rather than being fixed to a single default locale.
Data imperfection is generated deliberately
missing values, duplicates, formatting inconsistencies, and typographical errors are offered as configurable characteristics, letting a generated dataset match real-world data quality rather than only representing a clean, idealized case.
Review happens before export, not after
the workflow includes a preview and quality-metrics review step prior to export, rather than only surfacing dataset quality after a user has already integrated the output downstream.
Output formats match common downstream needs
CSV, JSON, SQL, and ZIP package export options are supported directly, so generated data can move into a database, an application, or an analytics tool without an extra conversion step.
DELIVERY APPROACH
As described in the source material, the engagement delivered a live, deployed application built around the five-step generation workflow.
1. Define the industry and format model - scoped the supported business domains and the structured/unstructured format distinction that shapes every downstream generation request.
2. Build natural-language requirement capture - implemented the flow that turns a plain-language description into a proposed, reviewable schema or content structure.
3. Build configurable generation options - implemented controls for record count, region, language, and data-quality variation parameters.
4. Build review and export - implemented dataset preview, quality-metric review, and export to CSV, JSON, SQL, and ZIP formats.
5. Deploy as a live application - shipped the platform as a live, accessible application rather than an internal-only prototype.
RESULTS AND IMPACT

Figure - Key outcomes from this engagement.
The source material describes Synthora as a live application rather than a benchmarked or audited deployment, so this case study does not represent measured accuracy, adoption, or performance figures beyond what the overview itself states.
What the overview does establish is a complete, working generation workflow - industry selection through to multi-format export - spanning both structured and unstructured data across at least six named industry domains, with realistic data-quality variation available as a configuration option rather than only clean, idealized output.
What it enabled commercially
By producing production-ready datasets in minutes from a natural-language description, the platform is positioned to reduce teams' dependency on real customer or patient data for testing, AI training, analytics, and demonstrations - directly supporting privacy and compliance requirements such as HIPAA and GDPR by design, rather than as an added control.
WHY PFACTORIAL
This engagement draws on Pfactorial's data and pipeline infrastructure capability: building accessible, natural-language-driven tools that remove real personally identifiable information from testing, training, and demonstration workflows entirely, rather than trying to de-identify it after the fact.

Figure - Service lines this engagement draws on.
Engagement enquiries Pfactorial Technologies works with teams whose testing, AI training, or demo data pipeline still depends on real customer or patient information. If you are evaluating whether synthetic data generation is a fit for your workflow, we are happy to give you an honest read on scope, cost, and risk before anyone commits to anything. · pfactorial.ai
APPENDIX A - TECHNOLOGY STACK
The technology stack underpinning the system, grouped by the layer it serves.

Result and Analysis
ENGAGEMENT SNAPSHOT
A live, natural-language-driven tool that produces industry-specific structured and unstructured datasets for testing, AI training, analytics, and demonstrations, without using any real customer or patient information.
CASE STUDIES
You might also like...

Computer VisionML Infra, Classifiers & RL
Aug 21, 20266 min readRead

Conversational AI & ChatbotsML Infra, Classifiers & RL
A Layered Safety Pipeline for a Healthcare Patient Companion
Aug 21, 20268 min readRead

ML Infra, Classifiers & RL
A Machine-Learning Screening Model for Pulmonary Hypertension from Routine Medical Records
Aug 21, 20266 min readRead

Speech & Audio PipelinesML Infra, Classifiers & RL
A Verified, Multi-Source News Platform With AI-Anchor Narration and Hourly Refresh
Aug 21, 20267 min readRead

Data Scraping & Aggregation
Automating Clinical Report Generation From Multi-Source Patient Files
Aug 21, 20269 min readRead

Automotive & Vehicle
A Browser-Native Neural Network Simulation That Learns to Drive From Experience
Aug 21, 20267 min readRead





