Metabolic marker combination for predicting pfos values in human blood and use thereof

By screening metabolic biomarker combinations through metabolomics and combining them with liquid chromatography and mass spectrometry, a multiple linear regression model was constructed, which solved the problem of accurately predicting PFOS levels in human blood, enabling early detection and intervention for high-risk groups and reducing the harm of PFOS.

CN122409931APending Publication Date: 2026-07-17HARBIN METANOTITIA INC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN METANOTITIA INC
Filing Date
2025-05-28
Publication Date
2026-07-17

Smart Images

  • Figure CN122409931A_ABST
    Figure CN122409931A_ABST
Patent Text Reader

Abstract

This invention provides a method for screening metabolic biomarkers and a combination of metabolic biomarkers for predicting perfluorooctane sulfonate (PFOS) levels in human blood. The invention provides a set of metabolic biomarkers for predicting PFOS levels in the blood. By detecting this combination of metabolic biomarkers in blood samples, the PFOS exposure status of patients can be predicted. Based on plasma metabolomics data, a highly accurate multiple linear regression model for predicting PFOS exposure values ​​is constructed. This method is convenient, economical, and facilitates sample acquisition, preventing further environmental harm from the use of PFOS standards. It is suitable for large-scale screening and long-term follow-up of high-risk populations, enabling timely intervention to reduce the harm caused by high PFOS exposure. This method has high accuracy and can accurately predict PFOS, providing important assistance in avoiding harm to the human body from PFOS exposure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of metabolomics analysis, specifically to a combination of metabolic biomarkers for predicting PFOS levels in human blood and their applications. Background Technology

[0002] The molecular formula of perfluorooctane sulfonate (PFOS) is C8F. 17 SO3 is a perfluorinated compound of octane sulfonic acid, which is also one of the most widely used perfluorinated compounds. As one of the most important chemical products of the 20th century, PFOS possesses both oleophobic and hydrophobic properties and is chemically very stable. Therefore, it is widely used in the production of nearly a thousand products, including antifouling agents, coatings, pesticides, pharmaceuticals, mining products, and paper food packaging materials and non-stick cookware that come into close contact with people's lives. However, due to its stability, PFOS is difficult to degrade and migrates with environmental media such as the atmosphere and water, causing global pollution. It can enter the human body through various exposure routes, circulating throughout the body via the bloodstream, causing organ damage, including hepatotoxicity, nephrotoxicity, neurotoxicity, and cardiovascular toxicity. Furthermore, as PFOS accumulates in the body, its toxic effects increase with age.

[0003] As a persistent organic pollutant, PFOS has been increasingly reported to cause harm to the human body in studies. For example, PFOS can cause endocrine disruption and metabolic effects, leading to thyroid dysfunction or abnormal lipid metabolism; immune system suppression; reproductive and developmental toxicity; liver damage; cardiovascular disease; and neurotoxicity. This demonstrates that PFOS accumulates in the body over a long period, producing toxic effects and inducing diseases. Therefore, there is an urgent need to develop a set of metabolic biomarkers to predict PFOS exposure levels in human blood, especially for high-risk groups such as pregnant women, children, occupationally exposed individuals, and residents of polluted areas. Accurate prediction of PFOS levels in the body can allow for timely detection of the harmful effects of high PFOS exposure and the implementation of appropriate interventions to minimize the damage caused by high PFOS exposure.

[0004] Metabolomics, as an important branch of systems biology, utilizes high-throughput analysis of small molecule metabolites (<1500 Da) in organisms and has been widely applied in disease mechanism research, precision medicine, and environmental toxicology assessment. Therefore, this patent, based on metabolomics data from blood samples, screens a set of metabolic biomarkers that can be used to predict PFOS levels, providing a basis for studying the damage caused by PFOS exposure to the human body, enabling early detection and intervention. Summary of the Invention

[0005] The technical problem to be solved by this invention is to provide a set of metabolic biomarkers for predicting PFOS levels in the blood. By detecting this set of metabolic biomarkers in blood samples, the PFOS exposure of patients can be predicted. This method has high accuracy and can accurately predict PFOS, providing important help in avoiding harm to the human body from PFOS exposure.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: This invention discloses a method for screening metabolic markers to predict perfluorooctane sulfonate (PFOS) levels in human blood, comprising the following steps: 1) Collect blood samples with different PFOS levels, and prepare organic and aqueous phases for each sample; 2) Metabolomics data were acquired using liquid chromatography and mass spectrometry for the organic and aqueous phases. The organic phase was constructed using a Waters ACQUTTY UPLC® BEH C8 1.7µm 2.1*100mm column, and the aqueous phase was constructed using a Waters ACQUTTY UPLC® HSS T3 1.8µm 2.1*100mm column. 3) Collect mass spectrometry data, convert the mass spectrometry data into a data matrix, and then perform qualitative analysis of metabolites after matching with commonly used databases; 4) Perform Lasso (Least Absolute Shrinkage and Selection Operator) analysis on the data of the modeling group samples to screen out metabolic biomarkers.

[0007] Preferably, the PFOS content is 1.5-90 ng / mL.

[0008] Preferably, the commonly used databases are human metabolite databases, metabolomics databases, mass spectrometry databases, and lipid metabolite databases.

[0009] This invention discloses a combination of metabolic markers, which includes: triglycerides 60:10, triglycerides 55:6, triglycerides 58:7, erythritol acid, phosphatidylcholine 34:2, (E)-ferulate ethyl ester, triglycerides 56:6, triglycerides 56:5, 2-hydroxybutyric acid, and L-seryl-L-leucine.

[0010] This invention discloses a metabolic marker composition comprising: triglycerides 60:10, triglycerides 55:6, triglycerides 58:7, erythritol acid, phosphatidylcholine 34:2, (E)-ferulate ethyl ester, triglycerides 56:6, triglycerides 56:5, 2-hydroxybutyric acid, and L-seryl-L-leucine.

[0011] This invention discloses a kit for predicting the value of perfluorooctane sulfonate (PFOS) in human blood, the kit comprising the aforementioned combination of metabolic markers.

[0012] Preferably, the kit also includes quality control products and / or standards.

[0013] Preferably, the kit also includes an instruction manual.

[0014] Preferably, the instruction manual includes instructions for constructing a multiple linear regression model using the Elastic Net Regression function.

[0015] Preferably, the Elastic Net Regression function is: .

[0016] Preferably, the kit uses the Elastic Net Regression function to construct a multiple linear regression model to analyze the perfluorooctane sulfonate (PFOS) value in the blood sample.

[0017] Preferably, the construction of the multiple linear regression model includes grid search and cross-validation using 5-fold cross-validation.

[0018] This invention discloses the use of the aforementioned combination of metabolic markers in the preparation of reagents and / or kits for predicting perfluorooctane sulfonic acid (PFOS) levels in human blood.

[0019] This invention discloses the use of the metabolic marker composition in the preparation of reagents and / or kits for predicting perfluorooctane sulfonic acid (PFOS) levels in human blood.

[0020] Preferably, the Elastic Net Regression function is used to construct a multiple linear regression model.

[0021] Preferably, model building includes grid search and cross-validation using 5-fold cross-validation.

[0022] Preferably, it also includes the operation of testing and evaluating the multiple linear regression model.

[0023] Preferably, mean absolute error (MAE), root mean squared error (RMSE), and coefficient of determination (R²) are used. 2 The Pearson correlation coefficient (Pearson R) and P-value were used for testing and evaluation.

[0024] Compared with existing technologies, this invention constructs a highly accurate multiple linear regression model for predicting PFOS exposure based on plasma metabolomics data. This method is convenient, economical, and easy to obtain samples, preventing further environmental harm caused by the use of PFOS standards. It is also suitable for large-scale screening and long-term follow-up of high-risk populations, and allows for timely intervention to reduce the harm caused by high PFOS exposure to the human body. Attached Figure Description

[0025] Figure 1 Linear correlation analysis between the predicted PFOS values ​​and the actual PFOS values ​​of subjects in the modeling and validation groups. Detailed Implementation

[0026] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0027] Example 1 Subject Sample Information 1. Subject Information 1) Sample inclusion criteria: Participants must meet all of the following inclusion criteria to be eligible to participate in this study: (1) Males or females aged ≥18 years; (2) Read and fully understand the information, sign the informed consent form, and be able to provide a blood sample for metabolomics testing; (3) The patient reported no obvious clinical symptoms; (4) The physical examination results are within the normal range and no obvious abnormalities are found; (5) The examination results are outside the normal range or are abnormal, but the doctor judges that they are not clinically significant (the physical examination items include: low-dose spiral CT (abbreviated as LDCT), abdominal color Doppler ultrasound, 4 tumor markers, blood pressure, etc.).

[0028] 2) Sample exclusion criteria: Subjects who meet any of the following exclusion criteria are ineligible to participate in this study: (1) During pregnancy or lactation; (2) Emergency room visit or resuscitation required; (3) History of blood transfusion within 7 days prior to sampling; (4) People who have received organ transplants or have previously received non-autologous (allogeneic) bone marrow or stem cell transplants; (5) History of malignant tumor within 5 years or any anti-tumor treatment before sampling; (6) Simultaneous co-occurrence of multiple primary malignant tumors.

[0029] 1) Subject information This study collected plasma samples from 225 subjects across two medical centers, including 169 samples in the modeling group and 56 samples in the validation group (Table 1). The actual PFOS values ​​in the samples were obtained using the external standard quantification method. A series of standards with known concentrations were prepared, and the mass spectrometry response values ​​were obtained using High-Performance Liquid Chromatography-Tandem Mass Spectrometer (HPLC-MS / MS). A standard curve was then constructed between the known concentrations and the mass spectrometry response values. The test sample was injected into the chromatographic mass spectrometry system, and the response value was recorded. Using the standard curve equation, the PFOS concentration in the sample was accurately calculated.

[0030] Table 1. Subject Information Modeling Group Verification group Number of people 169 56 Gender (Male / Female) 97 / 72 31 / 25 PFOS value (ng / ml, range) 1.57-87.93 2.93-59.89 Example 2: Detection and Data Analysis of Plasma Metabolites 1. Plasma metabolite detection 1) Reagents: Methanol, acetonitrile, water, acetic acid, and isopropanol of mass spectrometry grade purity, and formic acid, ammonium acetate, and methyl tert-butyl ether of chromatographic (HPLC) grade purity were purchased from Sigma-Aldrich, USA.

[0031] 2) Sample preparation: Take 100 μL of plasma and place it in 1000 μL of pre-cooled (methyl tert-butyl ether: methanol, volume ratio 3:1) solution. Vortex to mix the extracted blood sample and obtain the sample extract. Add 500 μL of (methanol: water, volume ratio 3:1) solution to the sample extract, sonicate, let stand, vortex and centrifuge to separate the layers. The upper layer is the organic phase and the lower layer is the aqueous phase.

[0032] - Organic phase: After the sample is separated into layers, take 500 μL of the upper organic phase into a centrifuge tube, dry it, add 200 μL of (acetonitrile:isopropanol, volume ratio 3:1), and incubate at room temperature for 15 minutes; after incubation, vortex the centrifuge tube, sonicate for 5 minutes, and then centrifuge at room temperature for 5 minutes (12000 rpm); take 180 µL of the supernatant from the centrifuge tube into a 2 mL glass vial, which is the organic phase test solution, and perform LC-MS detection.

[0033] - Aqueous phase: After the sample is separated into layers, take 400 μL of the lower aqueous phase into a centrifuge tube and add 1100 μL of ice-cold methanol to precipitate the protein. After the protein is precipitated, centrifuge the tube and transfer 1000 μL of the supernatant to a new centrifuge tube and dry it overnight. Add 200 μL of water to the dried centrifuge tube and incubate at room temperature for 15 minutes. After incubation, vortex the centrifuge tube, sonicate for 5 minutes, and then centrifuge at room temperature for 5 minutes (12000 rpm). Take 180 μL of the supernatant from the centrifuge tube into a 2 mL glass vial as the aqueous phase test solution and perform LC-MS analysis.

[0034] 3) Detection of small molecule metabolites: For small molecule separation, the organic phase was separated using a Waters ACQUTTY UPLC® BEH C8 1.7µm 2.1*100mm column, and the aqueous phase was separated using a Waters ACQUTTY UPLC® HSS T3 1.8µm 2.1*100mm column. The liquid chromatography and mass spectrometry systems used were the ACQUITY UPLC I-Class liquid chromatography system (Waters) and the Q-Exactive mass spectrometry system (Thermo Fisher Scientific).

[0035] The mobile phase parameters are as follows: Organic phase analyte mobile phase parameters – Mobile phase A is an aqueous solution containing 0.1% acetic acid and 0.1% ammonium acetate; Mobile phase B is an acetonitrile-isopropanol (7:3 v / v) solution containing 0.1% acetic acid and 0.1% ammonium acetate. The separation elution gradient is as follows: 55%-89% mobile phase B for 0-12 minutes, and 100% mobile phase B for 12-19.5 minutes.

[0036] Mobile phase parameters for the aqueous test solution: Mobile phase A is an aqueous solution containing 0.1% formic acid; Mobile phase B is an acetonitrile solution containing 0.1% formic acid. The separation and elution gradient is as follows: 0-13 minutes is 1%-70% mobile phase B, and 13-18 minutes is 99% mobile phase B.

[0037] The mass spectrometry parameters are as follows: Mass spectrometry data were acquired using Full MS and Full MS / dd-MS2 (each with both positive and negative modes). The parameters used by QExactive were as follows: Full MS mode had a resolution of 70,000, a scan range of 100-1500 m / z, an AGC of 3E+6, and a maximum IT of 200 ms; in Full MS / dd-MS2 mode, the resolution of the secondary mass spectrometer was 17,500, the quadrupole window was 1.5 m / z, the AGC was 1E+5, the maximum ion implantation time was 50 ms, and the HCD relative collision energy was 30 eV.

[0038] 2. Metabolomics data preprocessing and metabolite identification 1) Metabolomics data processing: (1) Extract peaks from the RAW format file of the mass spectrometer and convert it into FeatureXML format file to reduce the dimensionality of the original mass spectrometry data and improve the signal-to-noise ratio; (2) Use the peak alignment algorithm of OpenMS software to correct and align the retention time of the extracted peak format data between samples, thereby converting the mass spectrometry data into a data matrix; (3) Match and filter the isotope peaks in the data matrix obtained in step 2, and replace abnormal data (0, negative values, background noise, etc.) with missing values; (4) Remove the peaks with a detection rate of <80% from all the characteristic peaks obtained in step 3, fill the median value of the characteristic peak with the peaks with a detection rate of >80%, and add 5% random noise (following a standard normal distribution); (5) In order to reduce the difference in metabolite concentration between samples and make the data distribution more symmetrical, use Normalization Autoencoder (NormAE) for normalization processing to remove systematic errors such as batch effects.

[0039] 2) Identification of metabolites: After analyzing the raw data using software, the spectral information of the primary precursor ion (MS1) and secondary fragment ion (MS2) of the compound is obtained. This information, such as the mass-to-charge ratio (m / z) of the primary mass spectrometer and the fragment ion data, is matched with the spectral information of primary and secondary metabolites in public databases to qualitatively identify the metabolites. Commonly used metabolite databases include the Human Metabolite Database (HMDB, www.hmdb.ca), the Metabolomics Database (Metlin, metlin.scripps.edu), the Mass Spectrometry Database (www.massbank.jp), and the Lipid Map Database (Lipidmap, www.lipidmaps.org). Metabolites identified based on these databases are then finally validated using retention times, MS1, and MS2 mass spectrometry data obtained when separated from standards under the same chromatographic column and mass spectrometry conditions. The criteria for metabolite identification are a retention time difference within 0.1 min and a theoretical and measured molecular weight difference of less than 10 ppm.

[0040] 3. Data Analysis 1) Screening of metabolic biomarkers for predicting PFOS values To screen metabolic biomarkers associated with PFOS values, firstly, LASSO (Least Absolute Shrinkage and Selection Operator) analysis was performed on the modeling group data, including metabolomics data and corresponding PFOS values ​​for each sample. Through 5-fold cross-validation, the average error under different regularization parameters α was calculated, and the optimal regularization parameter α was determined to be 0.0547, which minimized the model error. Under this parameter, metabolites with non-zero regression coefficients were selected, while metabolites with zero coefficients were removed to eliminate noise variables. This resulted in 10 metabolites being selected as important metabolic biomarkers for predicting PFOS values.

[0041] Table 2. 10 Important Metabolic Markers for Predicting PFOS Values Logo - English Logo - Chinese 1 TAG 60:10 Triglycerides 60:10 2 TAG 55:6 Triglycerides 55:6 3 TAG 58:7 Triglycerides 58:7 4 Erythronic acid erythritol 5 PC 34:2 Phosphatidylcholine 34:2 6 (E)-Ethyl ferulate Ethyl (E)-feruloate 7 TAG 56:6 Triglycerides 56:6 8 TAG 56:5 Triglycerides 56:5 9 2-Hydroxybutanoic acid 2-Hydroxybutyric acid 10 L-Seryl-L-leucine L-seryl-L-leucine Example 3: Construction and Evaluation of a Multiple Linear Regression Model Using the Elastic Net Regression function, a multiple linear regression model was constructed for the above 10 metabolic biomarkers of the modeling group subjects. The independent variable was the mass spectrometry peak intensity of the 10 metabolites, and the dependent variable was the actual PFOS detection value. A multiple linear regression model between the metabolites and the actual PFOS detection value was constructed.

[0042] Elastic network regression is a multiple linear regression model that combines Ridge Regression and Lasso Regression (Least Absolute Shrinkage and Selection Operator Regression). Both Ridge and Lasso regressions are regularized versions of linear regression models that handle collinear data and perform variable selection. They reduce model complexity and avoid overfitting by adding a regularization term to the loss function. The regularization term typically consists of two parts: L1 regularization (Lasso): summing the absolute values ​​of the model coefficients, which can cause some coefficients to become zero, thus achieving variable selection; and L2 regularization (Ridge Regression): summing the squares of the model coefficients, which can shrink the values ​​of the coefficients but not set them to zero. The elastic network regression model combines these two regularization methods into a regularization term. The overall loss function of elastic network regression consists of two parts: mean squared error (MSE) and the regularization term, as follows:

[0043] J(W): Loss function, which is the objective function of elastic network regression, used to measure the difference between the model's predicted value and the actual value, as well as the complexity of the model; w: Weight vector, representing the parameters in the model; n: Number of samples; y i The actual value of the i-th sample, i.e., the value of the target variable; w T The transpose of the weight vector W; x i : The feature vector of the i-th sample, i.e., the value of the input variable; The L1 norm of the weight vector W is the sum of the absolute values ​​of all weights. The square of the L2 norm of the weight vector W, which is the sum of the squares of all weights.

[0044] λ is the total regularization strength (≥0). The larger the value, the stronger the regularization, and the more significantly the model coefficients are compressed, preventing overfitting. α (also known as l1_ratio) controls the ratio of L1 and L2 regularization terms (0≤α≤1). When α=0, ElasticNet degenerates into ridge regression; when α=1, it degenerates into Lasso regression. The advantage of the ElasticNet model is that it combines the variable selection ability of Lasso with the stability of ridge regression, enabling it to handle datasets with highly correlated features, and it is very effective in selecting important predictor variables and improving the model's generalization ability.

[0045] In model building, firstly, based on grid search, 5-fold cross-validation is used in the modeling group sample data. This involves randomly dividing the modeling group data into 5 equal subsets. In each iteration, one subset serves as the validation set, and the remaining 4 subsets are combined as the training set. This process is repeated 5 times to ensure that each subset serves as a validation set once. By calculating the MSE of each set of parameters (α and λ) on the validation set in the 5 training iterations, the average MSE of the 5 iterations is calculated. The set of parameters with the smallest MSE is selected as the optimal hyperparameters of Elastic Net and used to build the multiple linear regression model.

[0046] For the multiple linear regression model constructed above, several parameters are typically tested and evaluated to assess the model's performance (Table 3). These include Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and Coefficient of Determination (R²). 2 The Pearson correlation coefficient (Pearson R) and p-value are as follows: , is the average of the absolute values ​​of the differences between the predicted and actual values; the smaller the value, the better. in, The predicted value of the i-th sample; : The actual value of the i-th sample; Σ: The summation symbol, representing the summation of the absolute errors over all samples; n: The number of samples; RMSE = √MSE, which is the square root of MSE. MSE is the average of the squares of the differences between the predicted and actual values; the smaller the value, the better.

[0047] R 2 The coefficient of determination represents the proportion of the variance of the target variable explained by the model. It usually takes the range of 0-1, with values ​​closer to 1 being better, but values ​​>0.7 are generally considered good. Pearson R, the Pearson correlation coefficient, represents the linear correlation between predicted and actual values. Its value ranges from 0 to 1, with values ​​closer to ±1 being better, but values ​​>0.9 are generally considered good. The p-value tests the significance of the hypothesis that "correlation coefficient = 0".

[0048] Next, the mass spectrometry peak intensities of 10 metabolites from the validation group samples were input into the constructed multiple linear regression model to further validate the model's performance (Table 3). The results showed that the Pearson R value in the linear regression was greater than 0.9 and close to 1 in both the modeling and validation groups, indicating a strong linear correlation between the two variables and accurate prediction results. Figure 1 ).

[0049] Table 3. Parameter performance of the multiple linear regression model in the modeling and validation groups. parameter Modeling Group Verification group MAE 4.5091 3.7549 RMSE 5.4921 4.7276 <![CDATA[R 2 ]]> 0.8883 0.8939 Pearson R 0.9484 0.9628 P-value <1E(-5) <1E(-5) The present invention has been illustrated through the above embodiments, but the present invention is not limited to the above process steps, that is, it does not mean that the present invention must rely on the above process steps to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent substitutions of the raw materials used in the present invention, additions of auxiliary components, and selection of specific methods, etc., all fall within the protection scope and disclosure scope of the present invention.

Claims

1. A method for screening metabolic markers to predict perfluorooctane sulfonate (PFOS) levels in human blood, characterized in that, Includes the following steps: 1) Collect blood samples with different PFOS levels, and prepare organic and aqueous phases for each sample; 2) Metabolomics data were acquired using liquid chromatography and mass spectrometry for the organic and aqueous phases. The organic phase was constructed using a Waters ACQUTTY UPLC® BEH C8 1.7µm 2.1*100mm column, and the aqueous phase was constructed using a Waters ACQUTTY UPLC® HSS T3 1.8µm 2.1*100mm column. 3) Collect mass spectrometry data, convert the mass spectrometry data into a data matrix, and then perform qualitative analysis of metabolites after matching with commonly used databases; 4) Perform Lasso (Least Absolute Shrinkage and Selection Operator) analysis on the data of the modeling group samples to screen out metabolic biomarkers.

2. The screening method according to claim 1, characterized in that, The PFOS content is 1.5-90 ng / mL.

3. The screening method according to claim 1, characterized in that, The commonly used databases include human metabolite databases, metabolomics databases, mass spectrometry databases, and lipid metabolite databases.

4. The combination of metabolic biomarkers obtained by the screening method according to claim 1, characterized in that, The combination of metabolic markers includes: triglycerides 60:10, triglycerides 55:6, triglycerides 58:7, erythritol, phosphatidylcholine 34:2, (E)-feruloic acid ethyl ester, triglycerides 56:6, triglycerides 56:5, 2-hydroxybutyric acid, and L-seryl-L-leucine.

5. A kit for predicting perfluorooctane sulfonic acid (PFOS) levels in human blood, characterized in that, The kit comprises the combination of metabolic biomarkers as described in claim 4.

6. The reagent kit according to claim 5, characterized in that, The kit also includes quality control materials and / or standards.

7. Use of the combination of metabolic markers according to claim 4 in the preparation of reagents and / or kits for predicting perfluorooctane sulfonic acid (PFOS) levels in human blood.

8. The use according to claim 7, characterized in that, A multiple linear regression model is constructed using the Elastic NetRegression function.

9. The use according to claim 8, characterized in that, Model building includes grid search and cross-validation using 5-fold cross-validation.

10. The use according to claim 9, characterized in that, It also includes the testing and evaluation of the multiple linear regression model.