From Analytical Chemistry to Predictive Toxicology: A Review of Data Pipelines Linking Environmental Measurement to Health Prediction
DOI:
https://doi.org/10.21590/Keywords:
polychlorinated biphenyls, GC-ECD, GC-MS, data pipeline, predictive toxicology, FAIR principles, provenance, exposome, oxidative stress, congener analysis, reproducibility, uncertainty quantificationAbstract
Predictions of environmentally induced disease depend on a chain that begins with instrument signals and ends with a consequential interpretation. Errors often arise at the joints between stages, where units, censoring decisions, quality flags, provenance, and uncertainty are lost. This critical integrative review treats the environmental-health data pipeline as the object of study. It organizes the chain into five stages: measurement and quantification, quality control and harmonization, feature engineering and integration, predictive modeling, and interpretation and decision support. Polychlorinated biphenyl analysis by GC-ECD and GC-MS provides a worked example of how calibration, recovery correction, blank subtraction, detection-limit handling, panel mismatch, and congener aggregation propagate into downstream predictions. Evidence is assessed according to whether it preserves measurement meaning, provenance, uncertainty, and decision validity across stage boundaries. The review does not claim systematic coverage or estimate pooled effects. It uses direct environmental-health evidence for toxicological claims and clearly separates lessons drawn by analogy from diagnostic laboratories, data governance, and asset-integrity engineering. The analysis identifies four requirements for trustworthy pipelines: machine-actionable provenance, explicit interface contracts, versioned and tested workflows, and end-to-end uncertainty propagation. It further argues that reproducibility should be evaluated at every joint rather than inferred from agreement in final model outputs. The practical research agenda is to conduct multi-laboratory pipeline challenges using shared materials and raw files, then benchmark reproducibility, transportability, auditability, calibration, and decision consequences. The central claim is that strengthening the connective infrastructure of the pipeline will often improve predictive trustworthiness more than incremental refinement of the final algorithm.


