Mixing pollutant key element screening method based on gender difference

By combining machine learning and X-chromosome-related gene methods, a regression model was constructed to screen out key elements, solving the problem of gender differences not taken into account in existing technologies and achieving more efficient and accurate screening of pollutant elements.

CN120636610APending Publication Date: 2025-09-12SHANXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510599374.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-11
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing pollutant element screening methods do not integrate gender differences, resulting in screening bias, unscientific priority division of key elements, low screening efficiency and reliance on manual experience, making it difficult to ensure the objectivity and accuracy of the results.

Method used

A machine learning model was used in combination with X chromosome-related genes and physiological indicators to construct multiple regression models. The regression model with the best performance was selected through mean square error and determination coefficient evaluation, and multivariate correlation analysis was performed to identify key elements.

Benefits of technology

It improves the accuracy and efficiency of screening for key elements of pollutants, reduces labor costs, enhances the predictive ability and adaptability of the model, and provides more scientific screening results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636610A_ABST
    Figure CN120636610A_ABST
Patent Text Reader

Abstract

A sex difference-based mixed pollutant key element screening method belongs to the technical field of environmental health monitoring and biological analysis, solves the technical problems of strong subjectivity, low precision and low efficiency of the existing screening method, and comprises the following steps: S1, obtaining biological factor data; s2, acquiring in-vivo pollutant element data; s3, regression model screening: screening out a regression model with the best performance in pollutant key element screening; and S4, screening key elements. According to the method, a multi-dimensional fusion model is created by combining X chromosome related genes, typical physiological indexes and pollution element data in an organism; on the basis, through an automatic model trained by machine learning, manual intervention is reduced, the manpower and material resource cost is effectively reduced, the objectivity and efficiency of the screening process are improved, and according to the innovation, complex data are automatically processed through data analysis, and the accuracy and consistency of operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of environmental health monitoring and bioanalysis, and specifically relates to a method for screening key elements of mixed pollutants based on gender differences. Background Art

[0002] Numerous studies have confirmed that male and female individuals exhibit biological differences in metabolism, accumulation, and damage responses after exposure to environmental pollutants. Although a wide variety of methods are currently available for screening pollutant elements, they have the following shortcomings:

[0003] 1. Screening bias caused by missing data on sex differences: Existing element screening models do not integrate biological sex variables (especially X-chromosome-related biomarkers), making it impossible to interpret toxicity differences caused by differences in epigenetic regulation of pollutants in male and female organisms, resulting in inaccurate toxicity assessments of high-risk elements.

[0004] 2. Unscientific prioritization of key elements: Broad-spectrum element testing lacks screening logic for gender-sensitive elements, making it difficult to prioritize elements for control based on gender differences in toxicity.

[0005] 3. Insufficient screening efficiency and interpretability: Traditional models rely on manual experience to set element screening weights and do not introduce machine learning models and biological indicators, resulting in time-consuming models and opaque decision-making logic.

[0006] As can be seen from this, existing technologies largely rely on manual judgment and existing experience. Furthermore, traditional methods for screening key elements in mixed pollutants often ignore sample sex or only indirectly infer gender influences through physiological indicators (weight), making it difficult to guarantee the objectivity and accuracy of screening results. Furthermore, manual screening for key elements in pollutants consumes considerable time and manpower, resulting in low screening efficiency and making it difficult to cope with large-scale screening tasks.

[0007] In summary, a more scientific method for screening key elements of pollutants is urgently needed to identify and track the key elements of pollutants that cause gender differences. By combining X-chromosome-related genes, relevant parameters of typical physiological indicators, and data on pollutant elements in organisms to build a model, the screening of key pollutant elements can be effectively achieved. Summary of the Invention

[0008] The main purpose of the present invention is to overcome the deficiencies in the prior art and solve the technical problems of the existing screening methods, such as strong subjectivity, low precision and low efficiency. The present invention provides a method for screening key elements of mixed pollutants based on gender differences.

[0009] The present invention is achieved through the following technical solutions:

[0010] The method for screening key elements of mixed pollutants based on gender differences includes the following steps:

[0011] S1. Obtain biological factor data;

[0012] S2. Obtaining data on pollutant elements in the body;

[0013] S3. Regression model screening: Screen out the regression model with the best performance in screening key elements of pollutants;

[0014] S3-1. Regression model screening: Use several machine learning models, combine 80% of the training data and 20% of the test data, and construct multiple models to characterize different biological indicators; the several machine learning models include: linear regression, ridge regression, lasso regression, elasticnet regression, support vector regression, k-nearest neighbor regression, decision tree regression, gradient boosting regression, XGBoost regression, LightGBM regression, or multilayer perceptron regression.

[0015] S3-2. Using mean square error (MSE) and coefficient of determination (R) 2 As an indicator for evaluating the performance of regression models;

[0016] The mean square error (MSE) is used to measure the gap between the predicted value and the true value. The mean square error (MSE) is:

[0017]

[0018] In formula (1), y i is the true value, is the predicted value, n is the number of samples;

[0019] Coefficient of determination R 2 The coefficient of determination R is used to measure the degree of fit of the model to the data. 2 for:

[0020]

[0021] In formula (2), ss res is the residual sum of squares, ss tot is the total sum of squares;

[0022] Residual sum of squares ss res The calculation formula is:

[0023]

[0024] In formula (3), y i is the true value, is the predicted value, n is the number of samples;

[0025] Total sum of squares sstot The calculation formula is:

[0026]

[0027] In formula (4), y i is the true value, is the mean of the observed values, and n is the number of samples;

[0028] S3-3. Calculate the average of each model

[0029]

[0030] S3-4. Average of each model Sorting and selecting the regression model with the best performance for model building;

[0031] S4. Key element screening: Use the optimal regression model selected in step S3 to conduct multivariate correlation analysis on important indicator parameters in the key element screening of mixed pollutants to identify elements that are highly correlated with the target variables, and further screen out the key elements with the highest weights to complete the key element screening of mixed pollutants based on gender differences.

[0032] The important indicator parameters of the present invention mainly include: pollutant element factors in organisms (such as heavy metal factors, etc.) and biological factors after pollutant exposure (such as behavioral, physiological indicators, functional indicators, etc.). The present invention combines toxicology and machine learning to establish a clear-structured and hierarchical screening system for key pollutant elements, which can comprehensively analyze and evaluate the core pollution sources in mixed pollutants, improve the objectivity and fairness of the screening results, and provide reliable support for decision makers. Moreover, the present invention effectively improves the accuracy and reliability of source analysis, and has important application value, especially in complex pollution systems.

[0033] Furthermore, in step S1, the biological factor data is at least one of an epigenetic index, a functional index or a physiological response index.

[0034] Furthermore, epigenetic indicators include DNA methylation levels and histone modification levels of X-chromosome-linked genes; functional indicators include the mRNA abundance and protein activity of sex-differentially expressed genes (such as XIST and UTX) and function-related biomarkers; physiological response indicators include lung function and behavioral physiology indicators.

[0035] Furthermore, steps S1 and S2 are used to obtain important indicator parameters in the screening of key elements of mixed pollutants, and the correlation data of the pollutant elements in the biological samples in steps S1 and S2 are determined by biological experiments and mass spectrometry technology (such as ICP-MS / GC-MS).

[0036] The beneficial effects of the present invention are:

[0037] First, the present invention combines X-chromosome-linked genes, typical physiological indicators, and data on pollutant elements within organisms, effectively improving screening accuracy through the integration and modeling of multivariate information. By fusing data from different fields, it can provide more accurate analysis results for different types of organisms, greatly enhancing the model's predictive capabilities. Furthermore, the biological factor data used offers a variety of options, allowing for flexible adjustments based on actual needs, further enhancing adaptability and customizability.

[0038] Secondly, in terms of cost, this invention utilizes machine learning training models, reducing reliance on manual intervention, significantly reducing both human and material costs while improving the objectivity and efficiency of the screening process. Traditional methods may be subject to manual bias, while machine learning methods provide more accurate and reliable data processing capabilities.

[0039] Finally, the present invention has greater environmental adaptability and can provide efficient solutions in a variety of fields. Whether in biomedicine, environmental monitoring, or other related industries, the present invention can be flexibly applied to solve complex problems that are difficult to handle with traditional methods, further expanding its application potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a flowchart of the present invention;

[0041] Figure 2 The control group and PM 2.5 Comparison of lung function indicators of male and female individuals in the exposure group; Figure 2 (A) is a comparison chart of PIF (peak inspiratory flow) content, Figure 2 (B) is a comparison chart of PEF (peak expiratory flow) content, Figure 2 (C) is a comparison chart of F (respiratory frequency) content, Figure 2 (D) is a comparison chart of mv (minute ventilation) content, Figure 2 (E) is a comparison chart of EF50 (expiratory flow rate corresponding to 50% of the exhaled air volume). Figure 2 (F) is a comparison chart of Penh (enhanced expiratory pause) content;

[0042] Figure 3 This is a comparison chart of genes related to lung injury; among them, Figure 3 (A) is a comparison chart of the expression levels of Ccsp (club cell secretory protein), Figure 3 (B) is a comparison chart of the expression levels of Csf2rb (colony stimulating factor 2 receptor subunit Beta), Figure 3 (C) is a comparison chart of Spa (surfactant protein A) expression levels, Figure 3 (D) is a comparison chart of Spb (surfactant protein B) expression levels, Figure 3 (E) is a comparison chart of Spc (surfactant protein C) expression levels;

[0043] Figure 4 The control group and PM 2.5 Comparison of X chromosome linked gene expression levels in male and female individuals in the exposure group; Figure 4 (A) is a comparison chart of Kdm6a (lysine demethylase 6A) expression levels, Figure 4 (B) is a comparison chart of Kdm5c (lysine demethylase 5C) expression levels, Figure 4 (C) is a comparison of the expression levels of Ace 2 (angiotensin-converting enzyme 2), Eif2s3 (eukaryotic translation initiation factor 2 subunit γ), and Ddx3x (DEAD-box helicase 3, X-linked);

[0044] Figure 5 The control group and PM 2.5 Comparison of the contents of metal elements such as calcium, sodium, potassium, iron, and magnesium in the lung tissues of male and female individuals in the exposure group;

[0045] Figure 6 In the specific implementation, several regression models R are used 2 Comparison value results;

[0046] Figure 7 These are the results of key elements of mixed pollutants screened out in a specific implementation method. DETAILED DESCRIPTION

[0047] The present invention will be described in further detail below with reference to the accompanying drawings and examples.

[0048] like Figure 1 The method for screening key elements of mixed pollutants based on gender differences shown includes the following steps:

[0049] S1. Obtain biological factor data. In this embodiment, the biological factor data includes: Figure 2 The control group and PM 2.5 Lung function indicators (such as PIF, PEF, F, mv and EF50, etc.) of male and female individuals in the exposure group, Figure 3 The expression levels of lung injury-related genes (such as Spa, Spb and Spc) and Figure 4 The expression levels of X-chromosome-linked genes shown (such as Kdm6a, Kdm5c, Ace2, and Eif2s3x);

[0050] S2. Obtaining the data of pollutant elements in the body; in this embodiment, the data of pollutant elements in the body include Figure 5 The control group and PM2.5 The contents of metal elements such as calcium, sodium, potassium, iron, and magnesium in the lung tissues of male and female individuals in the exposed group;

[0051] S3. Regression model screening: Screen out the regression model with the best performance in screening key elements of pollutants;

[0052] S3-1. Regression model screening: Use several machine learning models, combine 80% of the training data and 20% of the test data, and construct multiple models to characterize different biological indicators; the several machine learning models include: linear regression, ridge regression, lasso regression, elasticnet regression, support vector regression, k-nearest neighbor regression, decision tree regression, gradient boosting regression, XGBoost regression, LightGBM regression, or multilayer perceptron regression.

[0053] S3-2. Using mean square error and coefficient of determination R 2 As an indicator for evaluating the performance of regression models;

[0054] The mean square error MSE is:

[0055]

[0056] In formula (1), y i is the true value, is the predicted value, n is the number of samples;

[0057] Coefficient of determination R 2 for:

[0058]

[0059] In formula (2), ss res is the residual sum of squares, ss tot is the total sum of squares;

[0060] Residual sum of squares ss res The calculation formula is:

[0061]

[0062] In formula (3), y i is the true value, is the predicted value, n is the number of samples;

[0063] Total sum of squares ss tot The calculation formula is:

[0064]

[0065] In formula (4), y i is the true value, is the mean of the observed values, and n is the number of samples;

[0066] S3-3. Calculate the average of each model

[0067]

[0068] S3-4. Average of each model Sort and select the best performing regression model for model building; Figure 6 As shown, in this embodiment, according to The sorting shows that XGBoost regression is closest to 1 among all models, so it can be considered that the model with the best performance is XGBoost regression;

[0069] S4. Key element screening: Use the optimal regression model XGBoost regression selected in step S3 to perform multivariate correlation analysis on important indicator parameters in the key element screening of mixed pollutants to identify elements that are highly correlated with the target variable, and further screen out the key elements with the highest weights to complete the key element screening of mixed pollutants based on gender differences. Figure 7 As shown, the conclusion shows that Cr has the highest correlation with changes in biological factors, indicating that Cr is likely to be PM 2.5 To further explore the key elements that drive male-female differences in PM 2.5 This study provides a reference for the source analysis of key elements that cause male-female differences in lung injury.

[0070] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for screening key elements of mixed pollutants based on gender differences, characterized in that: The following steps are involved: S1. Obtain biological factor data; S2. Obtaining data on pollutant elements in the body; S3. Regression model screening: Screen out the regression model with the best performance in screening key elements of pollutants; S3-1. Regression model screening: Use several machine learning models, combine 80% of the training data and 20% of the test data, and construct multiple models to characterize different biological indicators; the several machine learning models include: linear regression, ridge regression, lasso regression, elasticnet regression, support vector regression, k-nearest neighbor regression, decision tree regression, gradient boosting regression, XGBoost regression, LightGBM regression, or multilayer perceptron regression. S3-2. Using mean square error (MSE) and coefficient of determination (R) 2 As an indicator for evaluating the performance of regression models; The mean square error MSE is: In formula (1), y i is the true value, is the predicted value, n is the number of samples; Coefficient of determination R 2 for: In formula (2), ss res is the residual sum of squares, ss tot is the total sum of squares; Residual sum of squares ss res The calculation formula is: In formula (3), y i is the true value, is the predicted value, n is the number of samples; Total sum of squares ss tot The calculation formula is: In formula (4), y i is the true value, is the mean of the observed values, and n is the number of samples; S3-3. Calculate the average of each model S3-4. Average of each model Sorting and selecting the regression model with the best performance for model building; S4. Key element screening: Use the optimal regression model selected in step S3 to conduct multivariate correlation analysis on important indicator parameters in the key element screening of mixed pollutants to identify elements that are highly correlated with the target variables, and further screen out the key elements with the highest weights to complete the key element screening of mixed pollutants based on gender differences.

2. The method for screening key elements of mixed pollutants based on gender differences according to claim 1, characterized in that: In step S1, the biological factor data includes at least one of an epigenetic index, a functional index, or a physiological response index.

3. The method for screening key elements of mixed pollutants based on gender differences according to claim 2, characterized in that: Epigenetic indicators include X chromosome DNA methylation level and histone modification level; functional indicators include the mRNA abundance and protein activity of sex-differentially expressed genes and function-related biomarkers; physiological response indicators include lung function and behavioral physiology indicators.

4. The method for screening key elements of mixed pollutants based on gender differences according to claim 1, characterized in that: The correlation data of the pollutant elements in the biological sample in step S1 and step S2 are determined by biological experiments and mass spectrometry.