A benzene-induced blood abnormality marker, an evaluation model and a construction method and application thereof

By using principal component analysis, gene enrichment, and machine learning algorithms, key genes NFKB1 and PTX3 were screened out, and a benzene-induced blood abnormality assessment model was constructed. This solved the problem of assessing low-concentration benzene exposure and achieved sensitivity and reliability in early identification and risk assessment.

CN119229973BActive Publication Date: 2025-11-21JIANGSU PROVINCIAL CENTER FOR DISEASE CONTROL AND PREVENTION (PUBLIC HEALTH RESEARCH INSTITUTE OF JIANGSU PROVINCE)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411201638.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-11-21
Estimated Expiration
2044-08-29

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify and assess blood abnormalities caused by low-concentration benzene exposure, and lack sensitive and reliable biomarkers and assessment models, resulting in inadequate early intervention and risk assessment.

Method used

Using a combination of principal component analysis, gene enrichment analysis, and machine learning algorithms, differentially expressed genes were screened from the GEO database. Key genes NFKB1 and PTX3 were then identified using machine learning algorithms. A benzene-induced blood abnormality assessment model was constructed, and the expression levels of these genes were detected for evaluation.

Benefits of technology

It provides a more sensitive and reliable model for assessing benzene-induced blood abnormalities, enabling early identification of benzene poisoning risks, improving analytical efficiency and the model's versatility and reproducibility, and validating the key roles of NFKB1 and PTX3 in benzene exposure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229973B_ABST
    Figure CN119229973B_ABST
Patent Text Reader

Abstract

The present application relates to the field of biotechnology, in particular to a benzene-induced blood abnormality marker, an evaluation model and a construction method and application thereof, and the following scheme is proposed, which comprises S1: firstly, principal component analysis is performed on a benzene injury expression profile dataset, then a preliminary screening condition is set, and differential expression genes between the benzene injury expression profile datasets are screened out; S2: GO, KEGG and GSEA enrichment analysis are performed on the differential expression genes, and an analysis result is obtained; S3: different machine learning algorithms are used to screen out feature genes from the analysis result and take an intersection, key genes are obtained, and verification analysis is performed on the key genes, and benzene injury marker genes are obtained; S4: a variety of learning models are trained and tested by using the benzene injury marker genes, and a suitable learning model is selected as a benzene-induced blood abnormality evaluation model according to the test result of the learning model. The present application can provide a theoretical basis for early intervention of benzene toxicity, and can provide a more comprehensive and sensitive model for risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, in particular to a benzene-induced blood abnormality marker, an evaluation model and a construction method and application thereof. BACKGROUND

[0002] As a common organic hydrocarbon industrial solvent, benzene is widely used in pharmaceutical, pesticide, medicine, shoemaking and other industries, causing widespread occupational exposure. It can cause various blood-related diseases and is listed as a class I carcinogen by the International Agency for Research on Cancer (IARC 1982). Occupational benzene poisoning is divided into acute and chronic poisoning. The most common is occupational chronic benzene poisoning. High concentration benzene exposure can cause serious blood system diseases such as pancytopenia, aplastic anemia and leukemia; in addition, studies have shown that long-term low concentration benzene exposure can also induce blood toxicity. There is no safe exposure limit for benzene, and even exposure to low concentrations may cause blood malignancies. Therefore, it is particularly important to actively explore biomarkers of low concentration benzene exposure to provide a reference basis for screening and early intervention treatment of high-risk benzene exposure groups; machine learning is a computer program that can learn from experience in certain tasks and performance metrics. It is widely used in bioinformatics due to its easy adaptation and self-regulation, and is used by researchers to explore the potential mechanisms, potential biomarkers and treatment targets of various diseases.

[0003] Therefore, the present application provides a benzene-induced blood abnormality marker, an evaluation model and a construction method and application thereof. SUMMARY

[0004] In order to effectively perform risk assessment of benzene poisoning and identify early biomarkers of benzene poisoning abnormalities, the present application provides a benzene-induced blood abnormality marker, an evaluation model and a construction method and application thereof.

[0005] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0006] The present application provides a method for constructing a benzene-induced blood abnormality evaluation model, comprising the following steps:

[0007] S1: first performing principal component analysis on the benzene damage expression profile dataset, then setting a preliminary screening condition, and screening out differentially expressed genes between the benzene damage expression profile datasets;

[0008] Preferably, the benzene damage expression profile dataset is selected from the benzene damage expression profile datasets GSE9569 and GSE21862 in the GEO database (Gene Expression Omnibus database);

[0009] S2: performing GO, KEGG and GSEA enrichment analysis on the differentially expressed genes to obtain analysis results;

[0010] S3: screening feature genes from the analysis results by using different machine learning algorithms and taking intersection to obtain key genes, and then verifying the key genes to obtain benzene injury marker genes;

[0011] S4: training and testing various learning models by using the benzene injury marker genes, and selecting a suitable learning model as a benzene-induced blood abnormality evaluation model according to the test results of the learning model.

[0012] In some embodiments, the setting of the preliminary screening condition in S1 includes using the R software "limma" package to set the significance value P < 0.05 and the absolute value of logfoldchange ≥ 0.5 as the preliminary screening condition; and after screening out the differential expression genes, using the "pheatmap" package to draw a heat map and using the "ggplot2" package to draw a volcano plot.

[0013] In some embodiments, the step S2 includes using the "ClusterProfiler" package in the R software to perform GO, KEGG, and GSEA enrichment analysis on the differential expression genes, taking the significance value p < 0.05 as the standard, and using a bubble chart to display the enrichment analysis results.

[0014] In some embodiments, the machine learning algorithm in S3 includes a lasso regression algorithm, a SVM-RFE support vector machine recursive feature elimination algorithm, and a random forest algorithm.

[0015] In some embodiments, the step S3 includes:

[0016] S3.1: screening feature genes by using the lasso regression algorithm, the SVM-RFE support vector machine recursive feature elimination algorithm, and the random forest algorithm through the R software, and then taking intersection of the feature genes screened by the three machine learning algorithms to obtain key genes by using the venn package;

[0017] S3.2: drawing a receiver operating characteristic curve for the obtained key genes by using the R software, calculating the area under the curve (AUC) value respectively, the larger the AUC value, the better the diagnostic performance; and drawing a violin plot; for observation and analysis;

[0018] S3.3: selecting research objects, dividing the research objects into a benzene exposure group and a benzene non-exposure group, performing population verification and statistical analysis on the corresponding relationship between the key genes and benzene poisoning to obtain benzene injury marker genes.

[0019] Preferably, data analysis is performed using SPSS 27.0 software, for data conforming to normal distribution, is adopted, and independent sample t test is adopted for comparison of differences between groups; for data not conforming to normal distribution, is adopted, and non-parametric test is adopted for comparison of differences between groups, and p<0.05 is significantly different.

[0020] Preferably, the benzene injury marker genes include NFKB1 and PTX3.

[0021] In some embodiments, the plurality of learning models in S4 include a support vector machine model, a Bp neural network model, a Bayesian model and a C5.0 decision tree; the manner of selecting a suitable learning model according to the test result of the learning model includes: calculating the test result accuracy of each learning model and drawing an ROC curve, and comparing, and using GraphPad Prism 8.3 software to draw the comparison result to obtain a suitable learning model; preferably, the trained C5.0 decision tree is used as a benzene-induced blood abnormality evaluation model.

[0022] The second aspect of the present application provides a benzene-induced blood abnormality evaluation model obtained by the above construction method.

[0023] The third aspect of the present application provides an application of the above benzene-induced blood abnormality evaluation model in a blood abnormality detection model in a benzene exposure environment.

[0024] The fourth aspect of the present application provides a benzene-induced blood abnormality marker screened by using the above model, characterized in that: it includes NFKB1 and PTX3.

[0025] The present application is based on benzene exposure population (benzene exposure group) and non-benzene exposure population (control group), through detection, it is found that NFKB1 and PTX3 in the benzene exposure group are higher than those in the control group. Subsequently, according to the "diagnostic criteria for occupational benzene poisoning" GBZ68-2022, the benzene exposure group is divided into blood abnormality group and blood normality group (control group), and it is found through detection that NFKB1, PHACTR1 and PTX3 in the blood abnormality group are higher than those in the control group. By combining the two methods, it can be considered that NFKB1 and PTX3 are key risk genes of benzene exposure damage. At the same time, the present application also finds that the indirect oxidative damage indicators malondialdehyde (MDA), the DNA damage marker 8-hydroxydeoxyguanosine (8-OhdG) and the benzene exposure indicator S-PMA are obviously changed.

[0026] The fifth aspect of the present application provides an application of a primer for detecting NFKB1 and / or PTX3 in the preparation of a reagent for detecting benzene-induced blood abnormalities.

[0027] The present application has the following beneficial effects:

[0028] 1. The application first screens the key genes from the benzene injury expression profile data sets GSE9569 and GSE21862 in the GEO database (Gene Expression Omnibus) and using learning algorithms, establishes the research object, analyzes and verifies the key genes, obtains the benzene injury marker genes, uses the benzene injury marker genes to construct the training set and test set to train and test various machine learning models, and selects the optimal model as the benzene-induced blood abnormality evaluation model, thereby combining bioinformatics and machine learning and verifying from the mechanism. The application can provide a theoretical basis for early intervention of benzene toxicity and can provide a more comprehensive and sensitive model for risk assessment.

[0029] 2. The verification results show that the increase of benzene content in the benzene exposure population causes the increase of the body inflammation and oxidative stress level and thus may promote the blood toxicity of benzene;

[0030] 3. In the previous studies, most of them are focused on the relationship between NFKB and benzene exposure, and the research on the relationship between PTX3 and benzene exposure is less. The model constructed by combining machine learning and benzene injury marker genes in the present study screens out PTX3 in addition to NFKB1 in the screened benzene injury key genes, and has high screening sensitivity;

[0031] 4. The present study firstly combines principal component analysis, gene enrichment analysis and machine learning algorithm, can more comprehensively evaluate benzene injury, and has high reliability; then the differential expression genes and key genes are quickly and effectively screened out through the preliminary screening condition and machine learning algorithm, greatly improving the analysis efficiency; then the data used by us is from the public database, has strong reproducibility and universality, and can be verified and applied between different laboratories and research teams. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 FIG. 1 is an analysis result graph of the benzene injury expression profile data sets GSE9569 and GSE21862 of the present application, wherein a and b are two-dimensional graphs of the first two components in the principal component analysis results, a is before adjustment, and b is after adjustment; c is a volcano plot of the differential expression gene results, the red dots are up-regulated genes, the green dots are down-regulated genes, and the gray dots are expression non-differential genes; d is a clustering heat map of the differential expression genes;

[0033] Figure 2Figure 1 is a functional enrichment result diagram of differential genes of the present application, wherein a is a GO enrichment analysis result, the diagram is divided into three parts, from top to bottom, BP biological process, CC cell component and MF molecular function, and the color corresponds to the enrichment degree; b is a KEGG enrichment analysis result, the abscissa represents the number of genes enriched in the pathway, the ordinate represents the pathway name, and the color corresponds to the enrichment degree; c is a GSEA diagram of DEGs in the high expression group; d is a GSEA diagram of DEGs in the low expression group;

[0034] Figure 3 Figure 2 is a key gene screening result diagram of the present application, wherein a and b are diagrams of benzene injury related genes identified by LASSO and cox regression analysis; c and d are analysis result diagrams of SVM-RFE algorithm; e is a random forest algorithm diagram, the abscissa is the number of trees, the ordinate is the cross-validation error, the red line represents the error of the benzene exposure group, the green line represents the error of the control group, and the black line represents the error of all samples; f is a random forest gene importance ranking diagram, which shows the importance ranking of the first 30 genes; g is a Wayne diagram combining LASSO algorithm, SVM-RFE algorithm and random forest algorithm, the red part represents LASSO algorithm, the green part represents RF algorithm, the purple part represents SVM algorithm, the intersection part represents the identified key genes, and the numbers in the diagram represent the number of genes;

[0035] Figure 4 Figure 3 is a ROC curve diagram of the key genes of the present application, wherein a is a ROC curve diagram of NFKB1; b is a ROC curve diagram of PHACTR1; c is a ROC curve diagram of PTGS2; d is a ROC curve diagram of PTX3;

[0036] Figure 5 Figure 4 is a violin diagram of the relative expression amount of different genes of the benzene exposure group and the benzene non-exposure group of the present application, Control represents the benzene non-exposure group, and Treat represents the benzene exposure group, wherein a is a violin diagram of the expression amount of NFKB1; b is a violin diagram of the expression amount of PHACTR1; c is a violin diagram of the expression amount of PTGS2; d is a violin diagram of the expression amount of PTX3;

[0037] Figure 6For the statistical analysis comparison graph of the benzene exposure group and the benzene non-exposure group of the present application, "*" indicates that the independent sample t test is used, wherein a is the comparison result graph of 8-OHdG in the benzene exposure group and the benzene non-exposure group; b is the comparison result graph of BMI in the benzene exposure group and the benzene non-exposure group; c is the comparison result graph of MDA in the benzene exposure group and the benzene non-exposure group; d is the comparison result graph of Neutrophi in the benzene exposure group and the benzene non-exposure group; e is the comparison result graph of NFKB1 in the benzene exposure group and the benzene non-exposure group; f is the comparison result graph of PHACTR1 in the benzene exposure group and the benzene non-exposure group; g is the comparison result graph of PTGS2 in the benzene exposure group and the benzene non-exposure group; h is the comparison result graph of PTX3 in the benzene exposure group and the benzene non-exposure group; i is the comparison result graph of PLT in the benzene exposure group and the benzene non-exposure group; j is the comparison result graph of S-PMA in the benzene exposure group and the benzene non-exposure group; k is the comparison result graph of WBC in the benzene exposure group and the benzene non-exposure group;

[0038] Figure 7 For the comparison result graph of the statistical analysis of the blood abnormal group and the blood normal group of the benzene exposure group of the present application, "*" indicates that the independent sample t test is used, wherein a is the comparison result graph of 8-OHdG in the blood abnormal group and the blood normal group; b is the comparison result graph of BMI in the blood abnormal group and the blood normal group; c is the comparison result graph of MDA in the blood abnormal group and the blood normal group; d is the comparison result graph of Neutrophi in the blood abnormal group and the blood normal group; e is the comparison result graph of NFKB1 in the blood abnormal group and the blood normal group; f is the comparison result graph of PHACTR1 in the blood abnormal group and the blood normal group; g is the comparison result graph of PTGS2 in the blood abnormal group and the blood normal group; h is the comparison result graph of PTX3 in the blood abnormal group and the blood normal group; i is the comparison result graph of PLT in the blood abnormal group and the blood normal group; j is the comparison result graph of S-PMA in the blood abnormal group and the blood normal group; k is the comparison result graph of WBC in the blood abnormal group and the blood normal group;

[0039] Figure 8Figure of comparison of prediction of benzene exposure group and non-benzene exposure group by support vector machine model (SVM), Bp neural network model (BP), Bayesian model (Naive Bayes) and C5.0 decision tree (C5.0DT) of the present application, wherein (a-d) are the comparison result figures of prediction of test group and training group of benzene exposure group by four models. (a, c) are the comparison result figures of AUC and accuracy of test group by four models respectively, (b, d) are the comparison result figures of AUC and accuracy of training group by four models respectively. (e-h) are the comparison result figures of prediction of test group and training group of non-benzene exposure group by four models. (e, g) are the comparison result figures of AUC and accuracy of test group by four models respectively, (f, h) are the comparison result figures of AUC and accuracy of training group by four models respectively.

[0040] Figure 9 Figure of ROC curve of four machine learning models, i.e. support vector machine model (SVM), Bp neural network model (BP), Bayesian model (Naive Bayes) and C5.0 decision tree (C5.0DT) model of the present application and weight figure of four key genes, i.e. NFKB1, PHACTR1, PTGS2 and PTX3 genes in C5.0 decision tree model, wherein a is the ROC curve figure of four machine learning models of benzene exposure group; b is the weight figure of four key genes in C5.0 decision tree model of benzene exposure group; c is the ROC curve figure of four machine learning models in non-benzene exposure group; d is the weight figure of four key genes in C5.0 decision tree model of non-benzene exposure group.

[0041] Figure 10 Figure of volcano plot and heat map of PTX3 of the present application, wherein a is the volcano plot of PTX3, b is the heat map of PTX3, c is the GSVA analysis result figure of PTX3. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical scheme and advantages of the embodiments of the present application more clear, the technical scheme of the embodiments of the present application will be described clearly and completely below in combination with the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0043] The test materials or reagents used in the following embodiments, unless otherwise specified, are all existing and can be obtained from commercial channels. If the specific technology or condition is not specified in the embodiments, it can be carried out according to the conventional technology or condition disclosed in the art.

[0044] The abbreviations of various enrichment analyses in the present application are explained as follows:

[0045] GO: Gene Ontology, gene ontology, is a widely used term library for describing and statistical correlation of gene and protein function;

[0046] KEGG: Kyoto Encyclopedia of Genes and Genomes, is a comprehensive database, which is roughly divided into three categories of system information, genome information and chemical information;

[0047] GSEA: Gene Set Enrichment Analysis, is a computational biology method for analyzing gene expression data, aiming to reveal the gene expression pattern related to specific biological processes, pathways or functions;

[0048] Embodiment

[0049] A method for constructing a benzene-induced blood abnormality evaluation model, comprising the following steps:

[0050] S1.1: Preparation of benzene injury expression profile dataset: download benzene injury expression profile datasets GSE9569 and GSE21862 from GEO database; the GSE21862 dataset includes 83 benzene-exposed workers exposed to different benzene concentrations and 42 unexposed controls; the GSE9569 dataset includes 8 benzene-exposed workers and 8 unexposed controls, and the data obtained by comparing the data on two microarray platforms.

[0051] S1.2: Eliminate external differences of the two datasets by principal component analysis (PCA); use R software "limma" package to screen differentially expressed genes (DEGs) between GSE9569 and GSE21862 with significance value P<0.05 and ∣logfoldchange (logarithm of difference multiple) ∣≥0.5 as the preliminary screening condition, and then use "pheatmap" package to draw DEGs heat map and "ggplot2" package to draw DEGs volcano plot.

[0052] S2: using the "ClusterProfiler" package in R software to perform GO, KEGG and GSEA enrichment analysis on DEGs; GO analysis includes three parts: cell components (CC, refers to each part of the cell and the extracellular environment), molecular function (MF, refers to the activity at the molecular level, such as catalytic or binding activity) and biological process (BP, refers to a series of events produced by the ordered combination of one or more molecular functions); p<0.05 is the standard, and bubble chart is used to show the results of enrichment analysis;

[0053] S3.1: According to the results of enrichment analysis, three machine learning algorithms, lasso regression algorithm, SVM-RFE support vector machine recursive feature elimination algorithm and random forest algorithm, are used to screen feature genes by R software, and then the venn package is used to take the intersection of the feature genes screened by the three machine learning algorithms, so as to obtain key genes;

[0054] S3.2: Draw the receiver operating characteristic curve (ROC) of the obtained key genes by R software, and calculate the area under the curve (AUC) value respectively, the larger the AUC value, the better the diagnostic performance; and draw a violin plot; for observation and analysis;

[0055] S3.3: Select research subjects: establish a five-year prospective benzene-exposed worker cohort, select 445 research subjects in the first year, and divide the research subjects into benzene-exposed group and benzene-non-exposed group, of which 214 are in the benzene-exposed group and 231 are in the benzene-non-exposed group; obtain the data of key genes and subjects diagnosed as benzene poisoning by physical examination. The physical examination method is: 4mL of peripheral blood is extracted from the benzene-exposed group and the benzene-non-exposed group, of which 2mL is for biochemical detection and 2mL is for routine blood index analysis.

[0056] The method for analyzing and determining chronic mild benzene poisoning is as follows: The diagnosis of chronic mild benzene poisoning refers to the "Occupational Benzene Poisoning Diagnosis Standard" GBZ 68-2022: having a 3-month and above occupational history of close contact with benzene, which may be accompanied by dizziness, headache, fatigue, insomnia, memory loss, repeated infection and other clinical manifestations. Peripheral blood cell analysis is rechecked every 2 weeks within 3 months, and one of the following conditions is met: a) white blood cell count is less than 3.5x10 9 / L (normal range 3.5x10 9 -9.5x10 9 / L) for 4 times and above; b) neutrophil count is less than 1.8x10 9 / L (normal range 1.8x10 9 -6.3x10 9; c) platelet count < 80 x 109 / L (normal range 125 x 109 / L) 4 times or more 9 -350 x 109 / L. 9 / L).

[0057] S3.4: Data analysis was performed using SPSS 27.0 software. For data conforming to normal distribution, independent sample t test was used for comparison of differences between groups; for data not conforming to normal distribution, median and quartile were used for comparison of differences between groups, and non-parametric test was used; p < 0.05 was significantly different;

[0058] The "GSVA" software package was used to evaluate the genome-related path. GSVA: gene set variation analysis is a non-parametric unsupervised analysis method, mainly used to evaluate the sequencing gene set enrichment results.

[0059] After population verification and statistical analysis of the corresponding relationship between key genes and benzene poisoning, the benzene damage marker genes were obtained.

[0060] S4.1: According to the obtained benzene damage marker genes, a variety of learning models were trained and tested to build machine learning models:

[0061] The data set was divided into training group and test group according to the ratio of 7:3. First, the data was learned using the tabu search algorithm in R4.1.1bnlearn package to generate model network structure, and then the model structure was adjusted and optimized combined with literature materials to form the final model framework to be built. Next is to calculate four model parameters. In this study, Netica6.09 software was used to build the final model, and the parameters of the model were formed by learning the data. The following trained and tested learning models are C5.0DT, SVM, Bayes, BP, and the training, testing and construction methods of each model are as follows Tables 1-4:

[0062] Table 1 Construction method of C5.0DT model

[0063] Parameter name Parameter value Training time 0.207s Data split 0.7 Data shuffle Yes Cross-validation 10 Activation function identity Solver lbfgs Learning rate 0.1 L2 regularization term 1 Number of iterations 1000 Number of hidden first-layer neurons 100

[0064] Table 2 Construction method of SVM model

[0065] Parameter name Parameter value Training time 0.043s Data split 0.7 Data shuffle Yes Cross-validation 10 Penalty coefficient 1 Kernel function linear Kernel function coefficient scale Kernel function constant 0 Kernel function highest term number 3 Error convergence condition 0.001 Maximum number of iterations 1000 Multi-classification fusion strategy ovr

[0066] Table 3 Construction method of Bayes

[0067] Parameter name Parameter value Training time 0.021s Data split 0.7 Data shuffle Yes Cross-validation 10 Prior distribution Gaussian distribution alpha 1 Binary threshold 0

[0068] Table 4 Construction method of BP model ​

[0069] Parameter name Parameter value Training time 1.048s Data split 0.7 Data shuffle Yes Cross-validation 10 Node split evaluation measure gini Number of decision trees 100 With replacement sampling true Out-of-bag data test false Maximum feature ratio considered at split auto Minimum number of samples for internal node splits 2 Minimum number of samples for leaf nodes 1 Minimum weight of samples in leaf nodes 0 Maximum depth of tree 10 Maximum number of leaf nodes 50 Threshold for node split impurity 0

[0070] S4.2: The support vector machine model, BP neural network model, Bayesian model, and C5.0 decision tree were first trained using the test set, and then tested using the training set. The confusion matrix of different models was obtained. The accuracy of the test results of each learning model was calculated and ROC curves were plotted and compared. The comparison results were plotted using GraphPad Prism 8.3 software, and the optimal model was selected as the assessment model for benzene-induced blood abnormalities.

[0071] Results analysis:

[0072] I. Gene Analysis Results:

[0073] After eliminating the external differences between the two datasets through PCA (principal component analysis) analysis... Figure 1 As shown in the figures (a, b), there is a significant overlap between the two datasets after adjustment. Differential analysis identified 40 differentially expressed genes (DEGs), of which 38 were upregulated and 2 were downregulated. Volcano plots and heatmaps were then generated. (Volcano plot) Figure 1 c) shows the differentially regulated genes between GSE9569 and GSE21862, with red representing upregulated genes and green representing downregulated genes; cluster heatmap ( Figure 1 d) shows the specific expression of DEGs in each sample.

[0074] GO, KEGG, and GSEA enrichment analyses were performed on the screened differentially expressed genes: GO enrichment analysis revealed ( Figure 2 a) Biological processes (BP) mainly involve a series of inflammatory responses, including cytokine-mediated signaling pathways, leukocyte migration, leukocyte-cell adhesion, and responses to lipopolysaccharide; cellular components (CC) mainly involve tertiary granules, specific granules, and serine peptidase complexes; molecular functions (MF) mainly involve cytokine receptor binding, cytokine activity, and receptor-ligand activity. KEGG enrichment analysis results ( Figure 2 b) shows that the abundant pathways are mainly concentrated in the inflammation and immune response pathways, such as cytokine-cytokine receptor interactions, viral protein-cytokine and cytokine receptor interactions, chemokine signaling pathways, and NF-κB signaling pathways. Further GSEA enrichment analysis revealed that DEGs were enriched in five positive gene sets in the high-expression group, mainly involving cytokine receptor interactions, graft-versus-host disease, the JAK-STAT signaling pathway, leishmaniasis infection, and similar node receptor signaling pathways. Figure 2 c); Five positive gene sets were enriched in the low expression group, mainly involving base excision repair, Huntington's disease, oxidative phosphorylation, Parkinson's disease, and RNA degradation.Figure 2 d).

[0075] II. Verification results of key genes:

[0076] 26 feature genes were screened from DEGs using LASSO algorithm (a, b), 29 feature genes were screened using SVM-RFE algorithm (c, d), and 11 feature genes were screened using RF algorithm (e, f). The intersection of the three identified 4 genes (Fig. g), which were NFKB1, PHACTR1, PTGS2 and PTX3. Subsequently, ROC curves were drawn for the four (h), and their AUC (95% CI) were 0.822 (0.743-0.893), 0.755 (0.680-0.824), 0.802 (0.714-0.873) and 0.834 (0.756-0.901), respectively. The violin plots of the four key genes were drawn by extracting the GEO data (i), among which the four genes had large differences between the two groups. Figure 3 Figure 3 Figure 3 Figure 4 III. Population verification and statistical analysis results: Figure 5 Basic population analysis was performed on the benzene exposure group and the benzene non-exposure group. The basic characteristics of the research subjects are shown in Table 5; the statistical comparison results of blood routine and 4 genes are shown in

[0077] ; after comparison, we found that the p values of 8-OHdG, MDA, Neutrophi, PLT, S-PMA, WBC, NFKB1 and PTX3 were all less than 0.05, which had statistical significance; while the p values of PHACTR1 and PTGS2 were both greater than 0.05, which had no statistical significance; therefore, it can be considered that the changes of 8-OHdG, MDA, Neutrophi, PLT, S-PMA, WBC, NFKB1 and PTX3 in the benzene exposure group were significantly different from those in the benzene non-exposure group, which can be used as benzene damage markers, while it cannot be considered that the changes of PHACTR1 and PTGS2 in the benzene exposure group were significant, i.e., PHACTR1 and PTGS2 cannot be identified as the damage markers.

[0078] Subsequently, according to GBZ 68-2022 "Diagnosis Standard for Occupational Benzene Poisoning", the benzene exposure group was further divided into blood abnormal group and blood normal group, and statistical analysis was performed on the four key genes (see Table 6), among which the P value of PTGS2 was 0.136 > 0.05, which had no statistical significance. The statistical comparison results of various factors were plotted, as shown in Figure 6

[0079] Figure 7 ​​​​​As shown, from the results, it can be seen that the p values of PTGS2 and PLT are both greater than 0.05, and there is no statistical significance; while the p values of the rest are all less than 0.05, and there is statistical significance. It can be considered that compared with the blood normal group, the changes of 8-OHdG, MDA, Neutrophi, S-PMA, WBC, NFKB1, PHACTR1 and PTX3 in the blood abnormal group are obvious, and it cannot be considered that the changes of PTGS2 and PLT in the blood abnormal group are obvious.

[0080] Table 5, basic characteristics of 445 study subjects

[0081]

[0082]

[0083] Superscript a indicates that non-parametric rank-sum test is used

[0084] Table 6, statistical analysis of four genes of blood abnormal group and blood normal group

[0085] Blood abnormal group (n=21) Blood normal group (n=193) p-value NFKB1 a ]] 4.789(2.961-6.038) 2.215(1.659-3.725) <0.001 PHACTR1 a ]]> 2.075(0.848-2.830) 0.970(0.219-1.899) 0.004 PTGS2 b ]]> 0.997±0.501 1.237±0.714 0.136 PTX3 a ]]> 1.903(1.577-2.140) 2.272(2.003-2.847) <0.001

[0086] Superscript a indicates that non-parametric rank-sum test is used; superscript b indicates that independent sample t test is used;

[0087] IV. Determination of machine algorithm:

[0088] GraphPad Prism 8.3 software was used to draw the results obtained by different prediction models in the training set and the test set. Figure 8 It is shown that the model performance is measured by accuracy and AUC value, and combined with ROC curve figure ( Figure 9 ), four prediction models are compared, in which the area under the curve AUC value of C5.0 decision tree in benzene exposure group and benzene non-exposure group is the largest (0.823, 0.959), indicating that the diagnostic performance of C5.0 decision tree is the best. Therefore, C5.0 decision tree can better predict benzene damage, and then the weight diagram of C5.0 decision tree model is drawn ( Figure 9 b,d), according to the diagram shown, PTX3 is the most important gene. Further, the volcano plot and heat map of PTX3 are drawn ( Figure 10 a,b).

[0089] V. Results of gene set variation analysis

[0090] PTX3 is further discussed in different risk groups. Figure 10c shows that in the up-regulation group, PTX3 triggers valine and isoleucine degradation, glycine serine and threonine metabolism, and mutual transformation of pentose and glucuronic acid, etc.; in contrast, PTX3 in the down-regulation group triggers cytokine receptor interaction, nodal receptor signaling pathway, and jak stat signaling pathway, etc.

[0091] Conclusion: In summary, the present application combines bioinformatics and machine learning to explore key risk genes in benzene-induced damage, and finds that NFKB1 and PTX3 genes may be involved in benzene hematotoxicity through inflammatory response. Combined with the key risk genes verified by the present application and the machine learning algorithm, benzene poisoning risk can be accurately evaluated and patients in the early stages of symptoms can be identified.

[0092] Based on the above conclusion, the present application believes that NFKB1 and PTX3 can be used as marker genes or markers for benzene-induced blood abnormalities, that is, by judging the gene abnormalities or expression abnormalities of NFKB1 and PTX3, it can be judged that benzene-induced blood abnormalities have occurred.

[0093] At the same time, the present application proposes the use of primers for detecting NFKB1 and / or PTX3 in the preparation of reagents for detecting benzene-induced blood abnormalities.

[0094] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent substitutions or changes within the technical scope disclosed by the present application according to the technical solutions and inventive concepts of the present application, which should be covered within the protection scope of the present application.

Claims

1. A method for constructing a benzene-induced blood abnormality assessment model, characterized in that, Includes the following steps: S1: First, perform principal component analysis on the benzene damage expression profile dataset, then set initial screening conditions to screen out differentially expressed genes among the benzene damage expression profile datasets; S2: GO, KEGG, and GSEA enrichment analyses were performed on differentially expressed genes to obtain the analysis results; S3: Use different machine learning algorithms to screen out feature genes from the analysis results and take the intersection to obtain key genes. Then, perform verification analysis on the key genes to obtain benzene damage marker genes, including NFKB1 and PTX3. S4: Train and test various learning models using benzene damage marker genes, and select the appropriate learning model as the assessment model for benzene-induced blood abnormalities based on the test results of the learning models.

2. The method for constructing a benzene-induced blood abnormality assessment model according to claim 1, characterized in that, The initial screening conditions set in S1 include: using the "limma" package in R software, with a significance value of P < 0.05 and |logfoldchange| ≥ 0.5 as the initial screening conditions; after screening differentially expressed genes, using the "pheatmap" package to draw a heatmap and using the "ggplot2" package to draw a volcano plot.

3. The method for constructing a benzene-induced blood abnormality assessment model according to claim 2, characterized in that, Step S2 includes using the "ClusterProfiler" package in R software to perform GO, KEGG, and GSEA enrichment analyses on differentially expressed genes, with a significance value of p < 0.05 as the standard, and displaying the enrichment analysis results using a bubble chart.

4. The method for constructing a benzene-induced blood abnormality assessment model according to claim 3, characterized in that, The machine learning algorithms in S3 include: lasso regression algorithm, SVM-RFE support vector machine recursive feature elimination algorithm, and random forest algorithm.

5. The method for constructing a benzene-induced blood abnormality assessment model according to claim 4, characterized in that, The steps in S3 include: S3.1: Using R software, three machine learning algorithms—lasso regression, SVM-RFE support vector machine recursive feature elimination, and random forest—were used to screen feature genes. Then, the venn package was used to take the intersection of the feature genes screened by the three machine learning algorithms to obtain the key genes. S3.2: Use R software to plot receiver operating characteristic (ROC) curves for the obtained key genes, calculate the area under the curve (AUC) value, the larger the AUC value, the better the diagnostic performance; and plot a violin diagram. S3.3: Select research subjects, conduct population validation and statistical analysis on the correspondence between key genes and benzene poisoning, and obtain benzene damage marker genes.

6. The method for constructing a benzene-induced blood abnormality assessment model according to claim 5, characterized in that, The various learning models in S4 include: support vector machine model, Bp neural network model, Bayesian model, and C5.0 decision tree; the method of selecting a suitable learning model based on the test results of the learning models includes: calculating the accuracy of the test results of each learning model and plotting ROC curves, comparing them, and using GraphPad Prism 8.3 software to plot the comparison results to obtain a suitable learning model; the trained C5.0 decision tree is used as the benzene-induced blood abnormality assessment model.

7. A benzene-induced blood abnormality assessment system obtained by the construction method according to any one of claims 1-6.

8. The application of the construction method as described in any one of claims 1-6 or the benzene-induced blood abnormality assessment system as described in claim 7 in a detection model for blood abnormalities in a benzene-exposed environment.

9. A benzene-induced blood abnormality biomarker obtained by screening using the benzene-induced blood abnormality assessment system as described in claim 7, characterized in that: Including NFKB1 and PTX3.

10. Application of primers for detecting NFKB1 and / or PTX3 in the preparation of reagents for detecting benzene-induced blood abnormalities.

Citation Information

Patent Citations

  • Serum protein biomarker for esophageal squamous carcinoma and application of serum protein biomarker

    CN118330223A