
Back
Healthcare & Clinical
Identifying Disease-Driving Gene Expression Pathways in Pulmonary Arterial Hypertension
How Pfactorial Technologies built an RNA-seq analysis pipeline that pinpoints which genes and biological pathways are dysregulated in PAH patients versus healthy controls.
August 21, 2026
Share
ENGAGEMENT SNAPSHOT

Figure 1 - Key figures from this engagement, at a glance.
EXECUTIVE SUMMARY
Pulmonary Arterial Hypertension (PAH) is a progressive disease driven by molecular processes that aren't fully understood, and our client wanted to identify which genes and biological pathways actually drive its progression using RNA sequencing data from PAH patients and healthy controls.
Raw RNA-seq data is noisy before it's useful - sequencing depth, GC-content bias and batch effects all distort expression counts unless corrected for - and a list of individually up- or down-regulated genes says little about the biological mechanisms actually at work.
Pfactorial built a pipeline that quantifies, normalizes and validates the RNA-seq data before running differential expression analysis, then connects the resulting gene list to biological pathways through gene set enrichment analysis across four curated reference databases.
Why this engagement is representative This engagement shows Pfactorial's ability to take a raw sequencing dataset all the way to biologically interpretable findings, with the statistical rigor at every stage that a research or therapeutic-development team can actually trust and act on.
THE CHALLENGE
Turning raw sequencing reads into trustworthy biological insight required solving four problems before any conclusion could be drawn.
1. Raw RNA-seq data is noisy before it's useful
Sequencing depth, GC-content bias and batch effects all distort raw expression counts, and any of them left uncorrected can masquerade as a real biological signal.
2. Individual gene changes don't explain disease mechanism
A list of up- and down-regulated genes on its own says little about what biological processes are actually driving PAH - the genes need to be connected to pathways to be actionable.
3. Confounders can masquerade as disease signal
Age and smoking status differ between PAH patients and controls independent of disease biology; without adjustment, those differences could be mistaken for PAH-driven expression changes.
4. Subtle expression changes still matter biologically
A significance threshold that only catches large fold-changes would miss modest but biologically important shifts in expression - the analysis needed to stay sensitive to those.
The real brief Not “run a differential expression tool” but “build a pipeline that goes from raw sequencing reads to biologically interpretable pathways, with statistical rigor trustworthy at every step.”
THE SOLUTION
Pfactorial built the pipeline around QC and normalization first, a statistical model suited to RNA-seq's specific noise characteristics, and pathway-level enrichment analysis to connect individual genes to disease mechanism.

Figure 2 - From raw sequencing reads to validated pathway-level findings.
Architectural principles
- QC and normalization before anything else - Transcript-level quantification (via Salmon), gene-level aggregation, and low-expression filtering all happen before differential expression is attempted, so downstream results rest on a clean foundation.
- Variance decomposition catches batch effects early - Principal Variance Component Analysis identifies which technical or biological factors actually drive variability in the dataset, before that variability is mistaken for disease signal.
- A statistical model built for RNA-seq's quirks - limma with voom transformation, rather than a generic linear model, accounts for the mean-variance relationship specific to sequencing count data.
- From genes to pathways - Gene set enrichment analysis connects individual differentially expressed genes to the biological pathways they collectively point to, using curated reference sets spanning core processes, established pathway databases, regulatory motifs and functional classifications.
CAPABILITIES DELIVERED
Each deliverable builds on a validated prior stage, from raw counts to ranked biological pathways.
CAPABILITY | WHAT IT DOES |
|---|---|
QC'd expression matrix | Normalized RNA-seq data for 87 PAH patients and 81 healthy controls. |
Differential expression analysis | FDR-corrected significance testing via limma with voom transformation. |
Clustering diagnostics | PCA and t-SNE validation confirming sample grouping reflects disease, not technical artifact. |
Gene set enrichment analysis | Pathway enrichment across four curated reference gene-set databases. |
Gene-level visualization | Volcano, MA and boxplot visualizations for key findings such as MAOA. |
Documented statistical thresholds | FDR < 0.05 with no arbitrary fold-change cutoff, applied consistently throughout. |

Figure 3 - Four curated databases triangulate which pathway findings are robust.
Design note No fold-change cutoff was applied on top of the FDR threshold - a deliberate choice to keep subtle but biologically important expression changes visible, rather than filtering them out for looking small.
ENGINEERING FOR SCALE AND RELIABILITY
Six methodological decisions keep the pipeline's findings defensible from raw reads through to pathway-level conclusions.
Bias-corrected transcript quantification
Salmon-based quantification accounts for transcript length, GC-content bias and sequencing depth variation, improving the reliability of expression estimates before any downstream analysis.
Variance decomposition before conclusions
PVCA-driven analysis identifies which metadata variables contribute most to variance in the dataset, separating technical batch effects from true biological variability.
Clustering diagnostics as a sanity check
PCA and t-SNE clustering were run before differential analysis to confirm sample groupings reflected disease status rather than a technical artifact.
RNA-seq-appropriate statistical modeling
limma with voom transformation was used specifically because it accounts for how RNA-seq data's variability changes with expression level, unlike a generic linear model.
Consistent FDR-adjusted thresholds
A p<0.05 FDR-adjusted significance threshold was applied consistently across all differential expression results, avoiding threshold-shopping.
Multi-database pathway triangulation
Gene set enrichment was run across four independent curated databases, so pathway findings that appear across multiple sources carry more confidence than a single-database hit.
DELIVERY APPROACH
The engagement moved from raw sequencing data to biological interpretation in clearly validated stages.
1. RNA-seq preprocessing & QC - transcript quantification via Salmon and gene-level aggregation of expression data.
2. Normalization & batch-effect assessment - PVCA-based variance decomposition and PCA/t-SNE clustering diagnostics.
3. Differential expression modeling - limma with voom transformation to identify significantly differentially expressed genes.
4. Gene-level visualization & validation - volcano, MA and boxplot visualizations to interpret and communicate key findings.
5. Gene set enrichment analysis - pathway enrichment across HALLMARK, curated, regulatory and Gene Ontology gene sets.
6. Biological interpretation & reporting - synthesis of upregulated and downregulated pathway findings into actionable research hypotheses.
RESULTS AND IMPACT

Figure 4 - Key outcomes from this engagement.
Upregulated pathways tied to vascular remodeling, inflammation and stress response, and downregulated pathways tied to microRNA regulation and immune function, gave the client's research team a prioritized set of pathway-level hypotheses for follow-up validation.
Gene-specific findings, including differential expression of MAOA, were flagged as candidates worth further investigation given their established links to vascular function and oxidative stress.
What it enabled commercially
The engagement turned a raw RNA-seq dataset into a ranked, statistically defensible set of pathway-level hypotheses the client's research or therapeutic-development team can prioritize for follow-up, rather than a large undifferentiated gene list.
WHY PFACTORIAL
This engagement reflects Pfactorial's bioinformatics service line: rigorous, reproducible analysis pipelines that turn raw genomic data into pathway-level findings a research team can act on with confidence.

Figure 5 - Service lines this engagement draws on.
Engagement enquiries Pfactorial Technologies works with life-sciences teams that need raw sequencing data turned into validated, interpretable biological findings. If you're evaluating an RNA-seq or gene-expression analysis initiative, we're happy to give you an honest read on scope and risk before anyone commits to anything. · pfactorial.ai
APPENDIX A - TECHNOLOGY STACK
The technology stack underpinning the system, grouped by the layer it serves.

© 2026 Pfactorial Technologies. Client identity and product-specific implementation detail are withheld or generalized; no client data, credentials, source code, or infrastructure detail is included in this document.
Result and Analysis
ENGAGEMENT SNAPSHOT
How Pfactorial Technologies built an RNA-seq analysis pipeline that pinpoints which genes and biological pathways are dysregulated in PAH patients versus healthy controls.
CASE STUDIES
You might also like...

Computer VisionML Infra, Classifiers & RL
Aug 21, 20266 min readRead

Legal & Contract Analysis
A HIPAA-Compliant De-Identification Pipeline for Multi-Format Clinical Data
Aug 21, 20266 min readRead

Conversational AI & ChatbotsML Infra, Classifiers & RL
A Layered Safety Pipeline for a Healthcare Patient Companion
Aug 21, 20268 min readRead

ML Infra, Classifiers & RL
A Machine-Learning Screening Model for Pulmonary Hypertension from Routine Medical Records
Aug 21, 20266 min readRead

OCR & Document Extraction
A Multi-Tool Extraction Pipeline for Structured Clinical Variables from Unstructured Notes
Aug 21, 20266 min readRead

Recruiting & HR TechRAG & Semantic Search
A Natural-Language Candidate Search Platform That Replaces Boolean Query Building
Aug 21, 20267 min readRead





