A set of metabolic markers for assessing perfluorooctane sulfonate exposure levels and applications thereof

By constructing a discriminant model based on metabolic biomarkers, the pollution risk problem in PFOS exposure detection was solved, achieving highly sensitive and specific PFOS exposure assessment and providing an environmentally friendly detection method.

CN122409923APending Publication Date: 2026-07-17HARBIN METANOTITIA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510663832.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Current PFOS exposure detection relies on PFOS standards, which poses a risk of contamination. A more environmentally friendly detection method is needed to assess the exposure level of perfluorooctane sulfonate (PFOS).

Method used

By constructing a discriminant model to differentiate between high- and low-exposure populations to PFOS, and using a set of metabolic biomarkers including 2,3-dihydroxybutyric acid, erythritol, mannose, triglycerides 58:7, lysophosphatidylethanolamine 18:2, and ceramide d42:1, a machine learning support vector machine model was constructed to screen out 11 differential metabolites for high-sensitivity and high-specificity prediction of PFOS exposure risk.

Benefits of technology

The constructed discriminant model has high discrimination accuracy, with an area under the curve (AUC) above 0.828. Both sensitivity and specificity are at a high level, and it can accurately distinguish between high- and low-exposure populations to PFOS.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122409923A_ABST
    Figure CN122409923A_ABST
Patent Text Reader

Abstract

This invention provides a set of metabolic biomarkers for assessing perfluorooctane sulfonate (PFOS) exposure levels and their applications, belonging to the field of organic pollutant poisoning detection technology. Utilizing plasma sample metabolomics data and machine learning algorithms, this invention successfully constructs a discriminant model capable of accurately distinguishing between high PFOS exposure (above 20 ng / mL) and low exposure (not exceeding 20 ng / mL). This method is simple to operate, non-invasive, and the sample collection process is quick and easy, making it particularly suitable for large-scale PFOS high exposure risk screening in populations. It eliminates the need for PFOS standards, avoiding potential contamination risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of organic pollutant poisoning detection technology, specifically involving a set of metabolic biomarkers for assessing perfluorooctane sulfonic acid (PFOS) exposure levels and their applications. Background Technology

[0002] Perfluorooctane sulfonate (PFOS) is a typical and widely distributed persistent organic pollutant with high chemical stability and bioaccumulation. Due to its excellent chemical stability and surface activity, PFOS is widely used in industrial production and consumer products, such as fire-fighting foams, textile coatings, and food packaging materials. Studies have shown that PFOS exhibits reproductive toxicity, developmental toxicity, neurotoxicity, and immunotoxicity; long-term exposure may lead to various diseases such as liver damage, thyroid dysfunction, and immunosuppression. Furthermore, PFOS is difficult to excrete from the body and eventually accumulates in the blood, liver, kidneys, and brain of humans and animals. Currently, the detection of PFOS exposure mainly relies on PFOS concentration measurement, but this method requires the use of PFOS standards, posing a potential pollution risk. Therefore, a more environmentally friendly detection method is urgently needed. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a set of metabolic biomarkers for assessing perfluorooctane sulfonate (PFOS) exposure levels without the need for PFOS standards. By constructing a discriminant model that distinguishes between high- and low-PFOS exposure populations, the invention can predict PFOS exposure risk with high sensitivity and strong specificity, thus aiding in health risk assessment and early intervention.

[0004] This invention provides a set of metabolic biomarkers for assessing perfluorooctane sulfonic acid (PFOS) exposure levels, comprising at least three of the following compounds: 2,3-dihydroxybutyric acid, erythritol, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminohexanoic acid.

[0005] Preferably, when the metabolic marker comprises three compounds, it includes at least one of the following: a first composition formed by erythric acid, triglycerides 58:7 (18:1 / 18:1 / 22:5) and 2-aminoadipic acid; a second composition formed by 2,3-dihydroxybutyric acid, mannose and lysophosphatidylethanolamine 18:2 (0:0 / 18:2); a third composition formed by ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4) and phosphatidylethanolamine 36:3 (18:1 / 18:2); and a fourth composition formed by triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid and 2-aminoadipic acid.

[0006] Preferably, when the metabolic marker comprises five compounds, it includes at least one of the following: a fifth composition formed from erythric acid, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), and 2-aminoadipic acid; a sixth composition formed from 2,3-dihydroxybutyric acid, mannose, triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4); and a seventh composition formed from erythric acid, mannose, triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminoadipic acid.

[0007] Preferably, when the metabolic marker comprises eight compounds, it includes at least one of the following: an eighth composition formed from erythritol, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminoadipic acid; 2,3-dihydroxybutyric acid, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), and ceramide A ninth composition consisting of d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2) and hippuric acid, and a tenth composition consisting of mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4) and 2-aminoadipic acid.

[0008] This invention provides the application of the aforementioned metabolic biomarkers in constructing a discriminant model to distinguish between high and low exposure to perfluorooctane sulfonate (PFOS), wherein the concentration of PFOS high exposure is above 20 ng / mL, and the concentration of PFOS low exposure is not higher than 20 ng / mL.

[0009] Preferably, the discrimination model is constructed using a machine learning support vector machine.

[0010] Preferably, when constructing the discrimination model to distinguish between high and low exposure to perfluorooctane sulfonic acid, the number of random loop iterations is not less than 1000.

[0011] This invention provides the application of a reagent for detecting the metabolic marker in the preparation of a kit for screening high exposure risk of perfluorooctane sulfonate (PFOS) in a population, wherein the concentration of PFOS in the population with high exposure is above 20 ng / mL.

[0012] Preferably, the reagents include methanol, acetonitrile, water, acetic acid and methyl tert-butyl ether of mass spectrometry grade purity and formic acid of chromatographic grade purity;

[0013] The instruments used for screening the high risk of perfluorooctane sulfonic acid exposure in the population are liquid chromatography and mass spectrometry systems.

[0014] This invention provides a method for screening the aforementioned metabolic biomarkers, comprising the following steps:

[0015] The sample to be tested was pretreated to obtain an organic phase test solution and an aqueous phase test solution;

[0016] The organic phase and aqueous phase test solutions were loaded and analyzed by liquid chromatography-mass spectrometry to obtain raw mass spectrometry data.

[0017] After converting the raw mass spectrometry data into a data matrix, the data is optimized to obtain homogenized data.

[0018] The homogenized data was analyzed to identify metabolites;

[0019] The metabolites were subjected to regression analysis using the least absolute contraction and selection algorithm, and 5-fold cross-validation was used to screen out the metabolic biomarkers.

[0020] This invention provides a set of metabolic biomarkers for assessing perfluorooctane sulfonate (PFOS) exposure levels, including at least three of the following compounds: 2,3-dihydroxybutyric acid, erythritol, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminoadipic acid. This invention utilizes metabolomics data from plasma samples, combined with LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis, to screen for 11 differentially expressed metabolites. A discriminant model was constructed based on some or all of the differentially expressed metabolites to distinguish between individuals with high and low PFOS exposure. Results showed that the constructed discriminant model had an area under the curve (AUC) above 0.828, with a sensitivity above 0.737 and a specificity above 0.783, indicating high discriminant power. Furthermore, the discriminant model also demonstrated accurate judgment in the validation group. Attached Figure Description

[0021] Figure 1 The results of multivariate ROC curve analysis for 11 important markers distinguishing HE vs LE in the modeling group;

[0022] Figure 2 To validate the results of multivariate ROC curve analysis of 11 key markers that distinguish HE vs LE in the validation group. Detailed Implementation

[0023] This invention provides a set of metabolic biomarkers for assessing perfluorooctane sulfonic acid (PFOS) exposure levels, comprising at least three of the following compounds: 2,3-dihydroxybutyric acid, erythritol, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminohexanoic acid.

[0024] In this invention, when the metabolic marker comprises three compounds, it preferably comprises at least one of the following: a first composition formed by erythric acid, triglycerides 58:7 (18:1 / 18:1 / 22:5) and 2-aminoadipic acid; a second composition formed by 2,3-dihydroxybutyric acid, mannose and lysophosphatidylethanolamine 18:2 (0:0 / 18:2); a third composition formed by ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4) and phosphatidylethanolamine 36:3 (18:1 / 18:2); and a fourth composition formed by triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid and 2-aminoadipic acid. In one embodiment of the present invention, taking the first composition as an example, a discriminant model was constructed to distinguish between high and low exposure to perfluorooctane sulfonate (PFOS). In this embodiment, based on the modeling group data, the AUC of the discriminant model constructed with the three metabolites was 0.828, with sensitivity = 0.737 and specificity = 0.826; simultaneously, based on the validation group data, the AUC of the discriminant model constructed with the three metabolites was 0.827, with sensitivity = 0.846 and specificity = 0.700. This indicates that the constructed model has high discriminant accuracy.

[0025] In this invention, when the metabolic marker comprises five compounds, it preferably includes at least one of the following: a fifth composition formed from erythric acid, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), and 2-aminoadipic acid; a sixth composition formed from 2,3-dihydroxybutyric acid, mannose, triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4); and a seventh composition formed from erythric acid, mannose, triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminoadipic acid. In this embodiment of the invention, taking the fifth composition as an example, a discriminant model was constructed to distinguish between high and low exposure to perfluorooctane sulfonate (PFOS). Based on the modeling group data, the AUC of the discriminant model constructed from the five metabolites was 0.867, with sensitivity = 0.800 and specificity = 0.783. Simultaneously, based on the validation group data, the AUC of the discriminant model constructed from the five metabolites was 0.850, with sensitivity = 0.923 and specificity = 0.833. This indicates that the constructed model has high discriminant accuracy.

[0026] In this invention, when the metabolic marker comprises eight compounds, it preferably includes at least one of the following: an eighth composition formed from erythritol, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminohexanoic acid; 2,3-dihydroxybutyric acid, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminohexanoic acid; and 2,3-dihydroxybutyric acid, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), ceramide d42:1 (d18:1 / 24:0), ceramide d42:1 (22:4 / 18:2 / 20:4), ceramide d42:1 (22:4 / 18:2 / 22:4), ceramide d42:1 (22:4 / 18:2 / 22:4), ceramide d42:1 A ninth composition consisting of amide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2) and hippuric acid, and a tenth composition consisting of mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4) and 2-aminoadipic acid. In this embodiment of the invention, taking the eighth composition as an example, a discriminant model was constructed to distinguish between high and low exposure to perfluorooctane sulfonate (PFOS). Based on the modeling group data, the AUC of the discriminant model constructed from the eight metabolites was 0.848, with sensitivity = 0.750 and specificity = 0.818. Simultaneously, based on the validation group data, the AUC of the discriminant model constructed from the eight metabolites was 0.850, with sensitivity = 0.808 and specificity = 0.833. This indicates that the constructed model has high discriminant accuracy.

[0027] In this invention, when the metabolic marker comprises 11 compounds, it preferably includes an eleventh composition formed from 2,3-dihydroxybutyric acid, erythric acid, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminohexanoic acid. In this embodiment of the invention, based on the modeling group data, the discriminant model constructed from 11 metabolites had an AUC of 0.857, a sensitivity of 0.850, and a specificity of 0.783; simultaneously, based on the validation group data, the discriminant model constructed from the 11 metabolites had an AUC of 0.851, a sensitivity of 0.846, and a specificity of 0.733. This indicates that the constructed model has high discriminant accuracy.

[0028] In this invention, the 58:7 of the triglyceride 58:7 (18:1 / 18:1 / 22:5) indicates that the total number of carbon atoms in all fatty acid chains of the triglyceride molecule is 58 and the total number of unsaturated double bonds is 7; 18:1 / 18:1 / 22:5 indicates that it is composed of three fatty acid chains: one fatty acid chain containing 18 carbon atoms and 1 unsaturated double bond (18:1), one fatty acid chain containing 18 carbon atoms and 1 unsaturated double bond (18:1), and one fatty acid chain containing 22 carbon atoms and 5 unsaturated double bonds (22:5). The 18:2 in lysophosphatidylethanolamine 18:2 (0:0 / 18:2) indicates that the total number of carbon atoms in all fatty acid chains of the lysophosphatidylethanolamine molecule is 18, and the total number of unsaturated double bonds is 2; 0:0 / 18:2 indicates that it is composed of two fatty acid chains, one of which is a fatty acid-free chain (0:0), and the other of which is a fatty acid chain containing 18 carbon atoms and 2 unsaturated double bonds (18:2). In ceramide d42:1 (d18:1 / 24:0), d42:1 indicates that the total number of carbon atoms in all fatty acid chains of the ceramide molecule is 42, and the total number of unsaturated double bonds is 1 (d indicates dihydrosphingolipid structure); d18:1 / 24:0 indicates that it is composed of two fatty acid chains, one containing 18 carbon atoms and 1 unsaturated double bond (18:1), where "d" indicates that the sphingosine base of this chain is dihydrosphingosine, and the other containing 24 carbon atoms and no unsaturated double bond (24:0). The 60:10 in the triglyceride 60:10 (22:4 / 18:2 / 20:4) indicates that the total number of carbon atoms in all fatty acid chains of the triglyceride molecule is 60, and the total number of unsaturated double bonds is 10. The 22:4 / 18:2 / 20:4 indicates that it is composed of three fatty acid chains: one fatty acid chain with 22 carbon atoms and 4 unsaturated double bonds (22:4), one fatty acid chain with 18 carbon atoms and 2 unsaturated double bonds (18:2), and one fatty acid chain with 20 carbon atoms and 4 unsaturated double bonds (20:4). The 36:3 in phosphatidylethanolamine 36:3 (18:1 / 18:2) indicates that the total number of carbon atoms in all fatty acid chains of the phosphatidylethanolamine molecule is 36, and the total number of unsaturated double bonds is 3; 18:1 / 18:2 indicates that it is composed of two fatty acid chains, one containing 18 carbon atoms and 1 unsaturated double bond (18:1), and the other containing 18 carbon atoms and 2 unsaturated double bonds (18:2).The 56:6 in the triglyceride 56:6 (14:0 / 20:2 / 22:4) indicates that the total number of carbon atoms in all fatty acid chains of the triglyceride molecule is 56, and the total number of unsaturated double bonds is 6; 14:0 / 20:2 / 22:4 indicates that it is composed of three fatty acid chains: one fatty acid chain with 14 carbon atoms and no unsaturated double bonds (14:0), one fatty acid chain with 20 carbon atoms and 2 unsaturated double bonds (20:2), and one fatty acid chain with 22 carbon atoms and 4 unsaturated double bonds (22:4).

[0029] This invention provides a method for screening the aforementioned metabolic biomarkers, comprising the following steps:

[0030] The sample to be tested was pretreated to obtain an organic phase test solution and an aqueous phase test solution;

[0031] The organic phase and aqueous phase test solutions were loaded and analyzed by liquid chromatography-mass spectrometry to obtain raw mass spectrometry data.

[0032] The original mass spectrometry data is converted into a data matrix and then optimized to obtain homogenized data.

[0033] The homogenized data was analyzed to identify metabolites;

[0034] The metabolites were subjected to regression analysis using the least absolute contraction and selection algorithm, and 5-fold cross-validation was used to screen out the metabolic biomarkers.

[0035] The present invention pretreats the sample to be tested to obtain an organic phase test solution and an aqueous phase test solution.

[0036] In this invention, the type of sample to be tested preferably includes blood or plasma. The pretreatment method preferably includes treatment with a first extract and a second extract. The first extract preferably includes methyl tert-butyl ether and methanol. The volume ratio of methyl tert-butyl ether to methanol is preferably 3:1. The second extract preferably includes methanol and water. The volume ratio of methanol to water is preferably 3:1. The volume ratio of the sample to be tested, the first extract, and the second extract is preferably 1:10:5. The pretreatment method preferably involves sequentially performing ultrasonication, settling, and layering. The ultrasonic power is preferably 300-350W, which can be 320W; the frequency is preferably 30-40kHz, which can be 35kHz; and the ultrasonic time is preferably 15 minutes. The settling time is preferably 15 minutes. The layering method is preferably centrifugation. The centrifugation speed is preferably 10000-14000g, which can be 12700g; and the centrifugation time is preferably 5-30 minutes, which can be 5-10 minutes. After the separation, the upper layer is the organic phase test solution. The lower layer is treated with ice-cold methanol to remove proteins, and the supernatant is collected, the water is removed, and then mixed with water. After ultrasonic treatment, the supernatant is collected to obtain the aqueous phase test solution.

[0037] After obtaining the organic phase and aqueous phase test solutions, the organic phase and aqueous phase test solutions are loaded onto a liquid chromatography-mass spectrometry (LC-MS) instrument for detection to obtain raw mass spectrometry data.

[0038] In this invention, the organic phase test solution is preferably Waters ACQUTTY. Separation was performed using a BEH C8 1.7μm 2.1×100mm column; the aqueous phase analyte was detected using a Waters ACQUTTY column. Separation was performed using an HSS T3 1.8μm 2.1×100mm column. The preferred liquid chromatography-mass spectrometry (LC-MS) system was an ACQUITY UPLCI-Class liquid chromatography system (Waters) and a Q-Exactive mass spectrometry system (Thermo Fisher Scientific). During LC detection, the mobile phase parameters for the organic analyte were as follows: mobile phase A was an aqueous solution containing 0.1% acetic acid and 0.1% ammonium acetate; mobile phase B was an acetonitrile-isopropanol (7:3, v / v) solution containing 0.1% acetic acid and 0.1% ammonium acetate. The elution gradient was as follows: 55%-89% mobile phase B for 0-12 minutes, and 100% mobile phase B for 12-19.5 minutes. The mobile phase parameters of the aqueous test solution are as follows: mobile phase A is an aqueous solution containing 0.1% formic acid; mobile phase B is an acetonitrile solution containing 0.1% formic acid. The separation and elution gradient is as follows: 0-13 minutes is 1%-70% mobile phase B, and 13-18 minutes is 99% mobile phase B. For mass spectrometry detection, Full MS and Full MS / dd-MS2 (each with both positive and negative modes) are preferred for acquisition. The parameters used for Q Exactive are as follows: Full MS mode has a resolution of 70,000, a scan range of 100-1500 m / z, an Automatic Gain Control (AGC) of 3E+6, and a Maximum IT of 200 ms; in Full MS / dd-MS2 mode, the resolution of the secondary mass spectrometer is 17,500, the quadrupole window is 1.5 m / z, the AGC is 1E+5, the maximum ion implantation time is 50 ms, and the Higher Energy Collisional Dissociation (HCD) is 30 eV.

[0039] After obtaining the raw mass spectrometry data, this invention transforms the raw mass spectrometry data into a data matrix and then optimizes the data to obtain homogenized data.

[0040] In this invention, before conversion, the raw mass spectrometry data is preferably extracted into FeatureXML format files to reduce the dimensionality of the raw mass spectrometry data and improve the signal-to-noise ratio. The conversion method preferably utilizes the peak alignment algorithm of OpenMS software to correct and align the retention times of the peak format data between samples, thereby converting the mass spectrometry data into a data matrix. The data optimization method preferably includes matching and filtering isotope peaks in the data matrix, replacing abnormal data (0, negative values, background noise, etc.) with empty values, discarding peaks with a detection rate <80% from all obtained feature peaks, filling the median value of feature peaks with a detection rate >80%, adding 5% random noise (following a standard normal distribution), and then using NormalizationAutoencoder (NormAE) for homogenization to remove systematic errors such as batch effects, thereby reducing the differences in metabolite concentrations between samples and making the data distribution more symmetrical.

[0041] After obtaining the homogenized data, the present invention analyzes and identifies metabolites from the homogenized data.

[0042] In this invention, the preferred method for analysis is to use conventional software to analyze the raw data, obtaining the spectral information of the primary precursor ion (MS1) and secondary fragment ion (MS2) of the compound. This information is then matched with the spectral information of primary and secondary metabolites in a public database to qualitatively identify the metabolites. The spectral information includes the mass-to-charge ratio (m / z) of the primary mass spectrometer and the fragments of the secondary ion. The preferred public databases include the Human Metabolite Database (HMDB, www.hmdb.ca), the Metabolomics Database (Metlin, metlin.scripps.edu), the Mass Spectrometry Database (www.massbank.jp), and the Lipid Map Database (Lipidmap, www.lipidmaps.org). After qualitative analysis, the metabolites are validated based on the retention time, MS1, and MS2 mass spectrometry information of standards separated under the same conditions. A total of 491 metabolites were obtained through screening.

[0043] After obtaining the metabolites, the present invention uses the least absolute contraction and selection algorithm for regression analysis of the metabolites and employs 5-fold cross-validation to screen and obtain the metabolic biomarkers.

[0044] In this invention, the preferred screening method is to select the metabolites with the smallest error, alpha, which is 0.0685, and retain the metabolites whose regression coefficients in the corresponding models are non-zero.

[0045] This invention provides the application of the aforementioned metabolic biomarkers in constructing a discriminant model to distinguish between high and low exposure to perfluorooctane sulfonate (PFOS), wherein the concentration of PFOS high exposure is above 20 ng / mL, and the concentration of PFOS low exposure is not higher than 20 ng / mL.

[0046] In this invention, the discriminant model is preferably constructed using a machine learning support vector machine. When constructing the discriminant model to distinguish between high and low exposure to perfluorooctane sulfonate (PFOS), the number of random loop iterations is preferably not less than 1000.

[0047] In this invention, the cutoff value for high and low exposure to perfluorooctane sulfonate (PFOS) is 20 ng / mL. According to the Tolerable Weekly Intake (TWI) of PFOS set by agencies such as the U.S. Environmental Protection Agency (EPA) and the European Food Safety Authority (EFSA), it is 8–13 ng / kg body weight / week, and the recommended health counseling level for drinking water is 20 ng / mL. Meanwhile, the concentration range of PFOS in human blood varies in different regions and populations; the concentration range in the blood of Chinese adults is 1.47–32.0 ng / mL, and 20 ng / mL is generally considered a threshold for high and low exposure.

[0048] This invention provides the application of a reagent for detecting the metabolic marker in the preparation of a kit for screening high exposure risk of perfluorooctane sulfonate (PFOS) in a population, wherein the concentration of PFOS in the population with high exposure is above 20 ng / mL.

[0049] In this invention, the reagents preferably include methanol, acetonitrile, water, acetic acid, and methyl tert-butyl ether of mass spectrometry grade purity, and formic acid of chromatographic grade purity. The instruments used for screening the high exposure risk of perfluorooctane sulfonic acid in the population are preferably a liquid chromatography system and a mass spectrometry system.

[0050] In this invention, the method for screening high PFOS exposure risk in a population uses at least three metabolic biomarkers to construct a discriminant model. The reagents are used to detect the levels of the corresponding metabolic biomarkers in the sample. The detected metabolite data are then substituted into the discriminant model. If the levels are greater than or equal to a threshold, the patient is identified as having high PFOS exposure; if the levels are less than the threshold, the patient is identified as having low PFOS exposure. Experimental results show that the detection results of the metabolites can accurately distinguish between patients with high and low PFOS exposure.

[0051] The following detailed description, in conjunction with embodiments, of a set of metabolic biomarkers for assessing perfluorooctane sulfonate (PFOS) exposure levels provided by the present invention and their applications, but these should not be construed as limiting the scope of protection of the present invention.

[0052] Example 1

[0053] A set of screening methods for metabolic biomarkers to assess perfluorooctane sulfonate (PFOS) exposure levels

[0054] 1. Subject Information

[0055] 1) Inclusion criteria:

[0056] Participants must meet all of the following inclusion criteria to be eligible to participate in this study:

[0057] (1) Males or females aged 18 years or older;

[0058] (2) Read and fully understand the information, sign the informed consent form, and be able to provide a blood sample for metabolomics testing;

[0059] (3) The patient reported no obvious clinical symptoms;

[0060] (4) The physical examination results are within the normal range and no obvious abnormalities were found;

[0061] (5) The test results are outside the normal range or there are abnormalities but the doctor judges that they have no clinical significance (the physical examination items include: LDCT, abdominal color Doppler ultrasound, 4 tumor markers, blood pressure, etc.).

[0062] 2) Exclusion criteria:

[0063] Subjects who meet any of the following exclusion criteria are ineligible to participate in this study:

[0064] (1) During pregnancy or lactation;

[0065] (2) History of long-term medication use within the past 6 months;

[0066] (3) History of mental illness, drug abuse, or drug dependence;

[0067] (4) Those who have undergone surgery within the past 3 months, or those who plan to undergo surgery in the near future;

[0068] (5) Those who have donated blood or experienced massive bleeding (>450mL) within the past 3 months, or who have received blood transfusions or used blood products.

[0069] (6) Subjects with a history of common chronic diseases (hypertension, diabetes, coronary heart disease), cancer treatment, and major surgery.

[0070] 3) Subject information

[0071] This invention collected plasma samples from 226 subjects across two medical centers. Based on a PFOS concentration >20 ng / mL as the high exposure threshold, the subjects were divided into a high-exposure (HE) group (n=105) and a low-exposure (LE) group (n=121). Specifically, the plasma samples used for the modeling group consisted of 79 individuals in the HE group and 91 individuals in the LE group; the plasma samples used for the validation group consisted of 26 individuals in the HE group and 30 individuals in the LE group (Table 1).

[0072] Table 1 Subject Information

[0073] High PFOS exposure group (HE) Low PFOS exposure group (LE) Number of people in the modeling team 79 91 Number of people in the verification group 26 30 total 105 121

[0074] 2. Plasma metabolite detection

[0075] 1) Test reagents:

[0076] Methanol, acetonitrile, water, acetic acid, methyl tert-butyl ether of mass spectrometry grade, and formic acid of chromatographic (HPLC) grade were all purchased from Sigma-Aldrich, USA.

[0077] 2) Sample preparation:

[0078] Take 100 μL of plasma and place it in 1000 μL of pre-cooled (methyl tert-butyl ether: methanol, volume ratio 3:1) solution. Vortex to mix the extracted blood sample and obtain the sample extract. Add 500 μL of (methanol: water, volume ratio 3:1) solution to the sample extract, sonicate, let stand, vortex and centrifuge to separate the layers. The upper layer is the organic phase and the lower layer is the aqueous phase.

[0079] Organic phase: After the sample is separated into layers, take 500 μL of the upper organic phase into a centrifuge tube, dry it, add 200 μL of (acetonitrile:isopropanol, volume ratio 3:1), and incubate at room temperature for 15 minutes; after incubation, vortex the centrifuge tube, sonicate for 5 minutes, and then centrifuge at room temperature for 5 minutes (12000 rpm); take 180 μL of the supernatant from the centrifuge tube into a 2 mL glass vial, which is the organic phase test solution, and perform LC-MS detection.

[0080] Aqueous phase: After sample separation, transfer the lower 400 μL aqueous phase to a centrifuge tube and add 1100 μL of ice-cold methanol to precipitate proteins. After protein precipitation, centrifuge the tube and transfer 1000 μL of the supernatant to a new centrifuge tube, then dry overnight. Add 200 μL of water to the dried centrifuge tube and incubate at room temperature for 15 minutes. After incubation, vortex the mixture, sonicate for 5 minutes, and then centrifuge at room temperature for 5 minutes (12000 rpm). Transfer 180 μL of the supernatant from the centrifuge tube to a 2 mL glass vial as the aqueous phase test solution. Analyze using LC-MS.

[0081] 3) Detection of small molecule metabolites:

[0082] Organic phase using Waters ACQUTTY BEH C8 1.7μm 2.1×100mm column, with Waters ACQUTTY used for the aqueous phase. Small molecule separation was performed using an HSS T3 1.8μm 2.1×100mm column; both liquid chromatography and mass spectrometry used an ACQUITYUPLC I-Class liquid chromatography system (Waters) and a Q-Exactive mass spectrometry system (Thermo Fisher Scientific).

[0083] The detection parameters of the liquid chromatography system are as follows:

[0084] The parameters of the mobile phase of the organic phase test solution are as follows: Mobile phase A is an aqueous solution containing 0.1% acetic acid and 0.1% ammonium acetate; Mobile phase B is an acetonitrile-isopropanol (7:3 v / v) solution containing 0.1% acetic acid and 0.1% ammonium acetate. The separation and elution gradient is as follows: 55%-89% mobile phase B for 0-12 minutes, and 100% mobile phase B for 12-19.5 minutes.

[0085] The mobile phase parameters of the aqueous test solution are as follows: mobile phase A is an aqueous solution containing 0.1% formic acid; mobile phase B is an acetonitrile solution containing 0.1% formic acid. The separation and elution gradient is as follows: 0-13 minutes is 1%-70% mobile phase B, and 13-18 minutes is 99% mobile phase B.

[0086] The mass spectrometry parameters are as follows:

[0087] Mass spectrometry data were acquired using Full MS and Full MS / dd-MS2 (each with both positive and negative modes). The parameters used by QExactive were as follows: Full MS mode had a resolution of 70,000, a scan range of 100-1500 m / z, an Automatic Gain Control (AGC) of 3E+6, and a Maximum IT of 200 ms; in Full MS / dd-MS2 mode, the resolution of the secondary mass spectrometer was 17,500, the quadrupole window was 1.5 m / z, the AGC was 1E+5, the maximum ion implantation time was 50 ms, and the Higher Energy Collisional Dissociation (HCD) was 30 eV.

[0088] 3. Metabolomics data preprocessing and metabolite identification

[0089] 1) Metabolomics data processing:

[0090] (1) Extract peaks from the RAW format files of the mass spectrometer and convert them into FeatureXML format files to reduce the dimensionality of the original mass spectrometry data and improve the signal-to-noise ratio;

[0091] (2) Using the peak alignment algorithm of OpenMS software, the retention time of the extracted peak format data is corrected and aligned between samples, thereby converting the mass spectrometry data into a data matrix;

[0092] (3) Match and filter the isotope peaks in the data matrix obtained in step 2, and then replace the abnormal data (0, negative values, background noise, etc.) with missing values;

[0093] (4) Remove the feature peaks with a detection rate of <80% from all the feature peaks obtained in step 3, fill the median value of the feature peaks with a detection rate of >80%, and add 5% random noise (following a standard normal distribution).

[0094] (5) In order to reduce the difference in metabolite concentrations between samples and make the data distribution more symmetrical, the NormalizationAutoencoder (NormAE) was used for normalization to remove systematic errors such as batch effects.

[0095] 2) Identification of metabolites:

[0096] After analyzing the raw data using software, the spectral information of the primary precursor ion (MS1) and secondary fragment ion (MS2) of the compound is obtained. This information, such as the mass-to-charge ratio (m / z) of the primary mass spectrometer and the fragment ion data, is matched with the spectral information of primary and secondary metabolites in public databases to qualitatively identify the metabolites. Commonly used metabolite databases include the Human Metabolite Database (HMDB, www.hmdb.ca), the Metabolomics Database (Metlin, metlin.scripps.edu), the Mass Spectrometry Database (www.massbank.jp), and the Lipid Map Database (Lipidmap, www.lipidmaps.org). Metabolites identified based on these databases are then finally validated using retention times, MS1, and MS2 mass spectrometry data obtained from separation of standards under the same chromatographic column and mass spectrometry conditions. The criteria for metabolite identification are a retention time difference within 0.1 min and a theoretical and measured molecular weight difference of less than 10 ppm.

[0097] 4. Screening of metabolic biomarkers

[0098] Metabolite detection was performed on the above samples, and a total of 491 metabolites were obtained after annotation. LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis was performed on the data of the modeling group. The mean error corresponding to each regularization parameter alpha was calculated using 5-fold cross-validation. The optimal alpha with the smallest error was found to be 0.0685. Metabolites with non-zero regression coefficients in their corresponding models were retained. Finally, 11 differential metabolites were selected (Table 2) as important metabolic markers to distinguish between the HE group and the LE group.

[0099] Table 2. 11 Important Metabolic Markers for Differentiating HE from LE

[0100] serial number Logo - English Logo - Chinese HMDBID 1 2,3-Dihydroxybutanoic acid 2,3-Dihydroxybutyric acid HMDB0245394 2 Erythronicacid Erythritol HMDB0000613 3 Mannose Mannose HMDB0000169 4 TAG58:7(18:1 / 18:1 / 22:5) Triglycerides 58:7 (18:1 / 18:1 / 22:5) HMDB0010468 5 LysoPE18:2(0:0 / 18:2) Lysophosphatidylethanolamine 18:2 (0:0 / 18:2) HMDB0011477 6 Cerd42:1(d18:1 / 24:0) Ceramide d42:1 (d18:1 / 24:0) HMDB0004956 7 TAG60:10(22:4 / 18:2 / 20:4) Triglycerides 60:10 (22:4 / 18:2 / 20:4) HMDB0054779 8 PE36:3 (18:1 / 18:2) Phosphatidylethanolamine 36:3 (18:1 / 18:2) HMDB0009027 9 TAG56:6(14:0 / 20:2 / 22:4) Triglycerides 56:6 (14:0 / 20:2 / 22:4) HMDB0042592 10 Hippuricacid hippuric acid HMDB0000714 11 2-Aminoadipicacid 2-Aminohexanoic acid HMDB0302754

[0101] Example 2

[0102] (1) Method for constructing a discriminant model to distinguish between high PFOS exposure (HE) and low PFOS exposure (LE)

[0103] To verify the discriminative effect of the 11 selected biomarkers in distinguishing between HE and LE, multivariate ROC curve analysis was performed on these 11 biomarkers in the modeling group. Three-quarters of the sample data from the HE and LE groups in the modeling group were randomly used as the training set and one-quarter as the test set for training. The model was then used to randomly iterate 1000 times using a support vector machine (SVM). By statistically analyzing the average accuracy of the final model, a discriminative model for distinguishing between high and low PFOS exposure was constructed.

[0104] ROC curves are a method for studying the relationship between model sensitivity and specificity. Sensitivity is plotted on the ordinate, and 1-specificity on the x-axis. The evaluation criterion is the area under the curve (AUC). An AUC greater than 0.5, and closer to 1, indicates better model performance and better discrimination. An AUC less than 0.5 indicates poor model accuracy. ROC classification prediction models, in addition to common parameters such as the receiver operating characteristic (ROC) curve and AUC, also include sensitivity and specificity.

[0105] Sensitivity calculation is shown in Formula I:

[0106]

[0107] Specificity is calculated using Formula II:

[0108]

[0109] Among them, TP (True Positive): True positive, the number of samples that are actually positive but were correctly predicted as positive;

[0110] TN (True Negative): The number of samples that are actually negative but were correctly predicted as negative.

[0111] FP (False Positive): The number of samples that are actually negative but are incorrectly predicted as positive.

[0112] FN (False Negative): The number of samples that are actually positive but are incorrectly predicted as negative.

[0113] The results are as follows Figure 1 As shown, AUC = 0.857 (sensitivity = 0.850, specificity = 0.783), indicating that the constructed model has high discriminative power.

[0114] (2) Validation of the discriminant model used to distinguish between high and low PFOS exposure

[0115] To further validate the discriminant model for distinguishing between high and low PFOS exposure groups based on the modeling group data, validation group data was used to validate the model. Multivariate ROC curve analysis was performed to evaluate the model's independent validation performance on unknown datasets outside the modeling group dataset. After the validation group samples were fed into the model constructed by the modeling group, the probability value was output for each sample based on the detection data of 11 important metabolic markers distinguishing between LE and HE. Using the probability value of each sample as the discrimination threshold, a confusion matrix (including true positive, true negative, false positive, and false negative) was obtained. Sensitivity and specificity could be calculated using formulas. A point could be marked on the ROC analysis graph with sensitivity on the ordinate and 1-specificity on the abscissa. Similarly, when the probability value of each sample was used as the discrimination threshold, multiple different points were obtained in the ROC analysis graph. Connecting these points would produce an ROC curve. Figure 2 Among them, the point with the best sensitivity and specificity is selected, and the discrimination threshold at this time is 0.4925.

[0116] As shown in Table 3, the confusion matrix results indicate that, based on the 11 metabolic biomarkers, the prediction model, with a discrimination threshold of 0.4925, resulted in 22 out of 26 subjects with high PFOS exposure being correctly identified as having high PFOS exposure, and 4 being incorrectly identified as having low PFOS exposure. Conversely, among 30 subjects with low PFOS exposure, 22 were correctly identified, and 8 were incorrectly identified as having high PFOS exposure. The ROC analysis results of the discrimination model in the validation group are shown below. Figure 2 As shown, sensitivity and specificity were calculated based on the confusion matrix results, with an AUC of 0.851 (sensitivity = 0.846, specificity = 0.733). These results indicate that the established discriminant model for distinguishing between high and low PFOS exposure groups also demonstrates good discriminant performance in the validation group.

[0117] Table 3. Confusion matrix of the discrimination model for distinguishing between high and low PFOS exposure.

[0118] Types of diseases PFOS high exposure PFOS low exposure High PFOS exposure, N=26 22(TP) 4(FN) Low PFOS exposure, N=30 8(FP) 22(TN)

[0119] Example 3

[0120] A discriminant model was constructed using a combination of three metabolic biomarkers: erythritol, triglycerides (58:7, 18:1 / 18:1 / 22:5), and 2-aminoadipic acid, and ROC curve analysis was performed, following the method described in Example 2. When using all three metabolic biomarkers, the AUC was 0.828 (sensitivity = 0.737, specificity = 0.826), demonstrating stable discriminant ability.

[0121] Validation was also conducted. The discriminant model constructed using the three metabolic biomarkers (erythritol, triglycerides in a 58:7 ratio (18:1 / 18:1 / 22:5), and 2-aminoadipic acid) achieved an AUC of 0.827 (sensitivity = 0.846, specificity = 0.700) in the validation group. These results indicate that our established prediction model also exhibits good discriminative performance in the validation group.

[0122] Example 4

[0123] A discriminant model was constructed using five metabolic biomarkers: erythritol, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), and 2-aminoadipic acid. ROC curve analysis was performed using the same method as described in Example 2. When using five metabolic biomarkers, the AUC was 0.867 (sensitivity = 0.800, specificity = 0.783).

[0124] In the discriminant model constructed using the above 5 metabolic biomarkers, the AUC was 0.850 in the validation group (sensitivity = 0.923, specificity = 0.833).

[0125] Example 5

[0126] Eight metabolic biomarkers were used: erythritol, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), triglycerides 56:6 (14:0 / 20:2 / 22:4), and a combination of hippuric acid and 2-aminoadipic acid. A discriminant model was constructed and AUC analysis was performed, following the method described in Example 2. The results showed that when using eight metabolic biomarkers, the AUC was 0.848 (sensitivity = 0.750, specificity = 0.818).

[0127] Furthermore, the constructed discriminant model was validated in the validation group, and multivariate ROC curve analysis was performed. The results showed that the above eight metabolic markers, namely erythritol, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and the combination of 2-aminoadipic acid, had an AUC of 0.850 in the validation group (sensitivity = 0.808, specificity = 0.833).

[0128] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A set of metabolic biomarkers for assessing perfluorooctane sulfonate (PFOS) exposure levels, characterized in that, It includes at least three of the following compounds: 2,3-dihydroxybutyric acid, erythritol, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminohexanoic acid.

2. The metabolic biomarker for assessing perfluorooctane sulfonate exposure levels according to claim 1, characterized in that, When the metabolic marker comprises three compounds, it includes at least one of the following: a first composition formed by erythric acid, triglycerides 58:7 (18:1 / 18:1 / 22:5) and 2-aminoadipic acid; a second composition formed by 2,3-dihydroxybutyric acid, mannose and lysophosphatidylethanolamine 18:2 (0:0 / 18:2); a third composition formed by ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4) and phosphatidylethanolamine 36:3 (18:1 / 18:2); and a fourth composition formed by triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid and 2-aminoadipic acid.

3. The metabolic biomarker for assessing perfluorooctane sulfonate exposure levels according to claim 1, characterized in that, When the metabolic marker comprises five compounds, it includes at least one of the following: a fifth composition consisting of erythritol, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), and 2-aminoadipic acid; a sixth composition consisting of 2,3-dihydroxybutyric acid, mannose, triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4); and a seventh composition consisting of erythritol, mannose, triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminoadipic acid.

4. The metabolic biomarker for assessing perfluorooctane sulfonate exposure levels according to claim 1, characterized in that, When the metabolic marker comprises eight compounds, it includes at least one of the following: an eighth composition formed from erythritol, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), triglycerides 56:6 (14:0 / 20:2 / 22:4), hippuric acid, and 2-aminohexanoic acid; 2,3-dihydroxybutyric acid, mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), ... A ninth composition consisting of 2:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2) and hippuric acid, and a tenth composition consisting of mannose, triglycerides 58:7 (18:1 / 18:1 / 22:5), lysophosphatidylethanolamine 18:2 (0:0 / 18:2), ceramide d42:1 (d18:1 / 24:0), triglycerides 60:10 (22:4 / 18:2 / 20:4), phosphatidylethanolamine 36:3 (18:1 / 18:2), triglycerides 56:6 (14:0 / 20:2 / 22:4) and 2-aminoadipic acid.

5. The application of the metabolic marker according to any one of claims 1 to 4 in constructing a discriminant model to distinguish between high and low exposure to perfluorooctane sulfonate (PFOS), wherein the concentration of PFOS high exposure is above 20 ng / mL and the concentration of PFOS low exposure is not higher than 20 ng / mL.

6. The application according to claim 5, characterized in that, The discriminant model is constructed using machine learning support vector machines.

7. The application according to claim 5, characterized in that, When constructing the discriminant model to distinguish between high and low exposure to perfluorooctane sulfonic acid, the number of random loop iterations is no less than 1000.

8. The use of a reagent for detecting the metabolic markers of any one of claims 1 to 4 in the preparation of a kit for screening high exposure risk of perfluorooctane sulfonate in a population, wherein the concentration of high exposure to perfluorooctane sulfonate in the population is 20 ng / mL or higher.

9. The application according to claim 8, characterized in that, The reagents include methanol, acetonitrile, water, acetic acid and methyl tert-butyl ether of mass spectrometry grade purity and formic acid of chromatographic grade purity; The instruments used for screening the high risk of perfluorooctane sulfonic acid exposure in the population are liquid chromatography and mass spectrometry systems.

10. A method for screening metabolic biomarkers according to any one of claims 1 to 4, characterized in that, Includes the following steps: The sample to be tested was pretreated to obtain an organic phase test solution and an aqueous phase test solution; The organic phase and aqueous phase test solutions were loaded and analyzed by liquid chromatography-mass spectrometry to obtain raw mass spectrometry data. After converting the raw mass spectrometry data into a data matrix, the data is optimized to obtain homogenized data. The homogenized data was analyzed to identify metabolites; The metabolites were subjected to regression analysis using the least absolute contraction and selection algorithm, and 5-fold cross-validation was used to screen out the metabolic biomarkers.