Breast cancer characteristic gene screening method based on minimum classification error rate criterion

By using the minimum classification error rate criterion and bioinformatics analysis method in the selection of breast cancer characteristic genes, key genes were screened out and Bayesian network classifiers were constructed, which solved the problem of insufficient diagnostic accuracy in the existing technology, and achieved accurate diagnosis and efficient treatment of breast cancer.

CN119993267AInactive Publication Date: 2025-05-13THE FIRST AFFILIATED HOSPITAL OF SOOCHOW UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411789020.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot effectively measure the random uncertainty of the distribution in breast cancer characteristic gene screening, resulting in insufficient diagnostic accuracy and accuracy.

Method used

Using a method based on the minimum classification error rate criterion, weighted gene co-expression network analysis and protein interaction network analysis were performed by preprocessing gene expression data, key cancer genes were screened out, and Bayesian network classifiers were constructed for breast cancer diagnosis.

Benefits of technology

It improves the diagnostic accuracy of breast cancer, reduces unnecessary detection and analysis, saves time and costs, reduces the probability of misdiagnosis and missed diagnosis, and ensures that patients receive timely and correct treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993267A_ABST
    Figure CN119993267A_ABST
Patent Text Reader

Abstract

The invention discloses a breast cancer characteristic gene screening method based on a minimum classification error rate criterion, which comprises the following steps: pre-screening breast cancer characteristic genes from TCGA data by taking a minimum classification error rate as a standard, and then further screening the characteristic genes according to weighted gene co-expression network analysis and protein interaction network analysis. And finally, constructing a breast cancer Bayesian network classification model according to the screened feature genes, realizing accurate prediction of breast cancer, and evaluating the effectiveness of the verification method according to the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of gene screening, and in particular relates to a method for screening characteristic genes of breast cancer based on a minimum classification error rate criterion. Background Art

[0002] Screening of characteristic genes for breast cancer is crucial for diagnosis, prognosis and timely treatment. Inaccurate early diagnosis will lead to failure in clinical applications of advanced cancer, resulting in tumor recurrence, metastasis and death. Diagnostic methods for breast cancer include breast palpation, tumor marker detection, imaging examination and histopathological examination. Among them, tumor marker detection has the characteristics of low invasiveness, effectiveness and convenience, and has received widespread attention in the early diagnosis of breast cancer. Gene is a reliable cancer marker, and its expression level is an important indicator for determining whether a person has cancer.

[0003] With the development of high-throughput sequencing technology, data science and bioinformatics, tumor gene databases such as TCGA, GEO and GTEx have become increasingly large, providing massive data for the screening of cancer characteristic genes. The use of statistical and artificial intelligence methods to screen characteristic genes for cancer diagnosis from tumor gene databases has become an important method at present. In order to improve the accuracy and precision of diagnosis, a new screening criterion that considers distribution uncertainty should be proposed. The minimum classification error rate represents the overlapping area of ​​two distributions in a binary classification task. It can measure the uncertainty of the distribution and can therefore be used as a cancer screening criterion.

[0004] Differential expression analysis is a common and reliable method for screening cancer characteristic genes. At present, the relevant cancer gene screening methods based on differential expression analysis mainly consider the difference in the mean of gene expression in samples, and cannot measure the random uncertainty of distribution. In the process of using gene expression for cancer diagnosis, there are many uncertainties in the expression of genes in many samples. The traditional cancer characteristic gene screening method only considers the difference in the mean of gene expression in samples, and cannot measure the random uncertainty of distribution. In order to improve the diagnostic precision and accuracy, a new screening standard that considers distribution uncertainty should be proposed. The biggest obstacle to achieving accurate diagnosis of cancer patients is: in the process of cancer characteristic gene screening, whether genes with high contribution to the classification model can be accurately obtained; in the process of diagnostic model prediction, whether a machine learning model suitable for using discretized gene expression data as input features is used. Summary of the invention

[0005] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.

[0006] In view of the above problems and / or the problems existing in the prior art, the present invention is proposed.

[0007] Therefore, the purpose of the present invention is to overcome the deficiencies in the prior art and provide a method for screening breast cancer characteristic genes based on the minimum classification error rate criterion.

[0008] In order to solve the above technical problems, the present invention provides the following technical solution: a method for screening characteristic genes of breast cancer based on the minimum classification error rate criterion, characterized in that: it includes:

[0009] Preprocessing gene expression data to obtain a gene expression data set;

[0010] Pre-screening of characteristic genes according to the minimum classification error rate criterion;

[0011] Perform weighted gene co-expression network analysis on the dataset to obtain significant module gene sets;

[0012] The intersection of the pre-screened characteristic genes and the significant module gene set was taken to obtain the key cancer genes;

[0013] Construct a protein interaction network and further screen the characteristic genes for the final classification model;

[0014] The above-mentioned characteristic gene expression levels were used to construct a Bayesian network classifier for breast cancer diagnosis and prediction;

[0015] The results were presented using confusion matrices and verified using diagnostic metrics.

[0016] As a preferred embodiment of the method for screening characteristic genes of breast cancer of the present invention, the preprocessing of gene expression data obtains a gene expression data set, wherein the gene expression data is gene expression data of breast cancer samples and normal adjacent tissue samples.

[0017] As a preferred embodiment of the method for screening characteristic genes of breast cancer of the present invention, the gene expression data set is displayed in the form of a matrix.

[0018] As a preferred embodiment of the method for screening characteristic genes of breast cancer described in the present invention, the characteristic base is pre-screened according to the minimum classification error rate criterion, wherein the minimum classification error rate is obtained by calculating the overlapping area of ​​the normal distribution curve and the cancer distribution curve, and the similarity or difference of the distribution of the two groups of data is quantified by the minimum classification error rate.

[0019] As a preferred embodiment of the method for screening characteristic genes of breast cancer described in the present invention, the further screening obtains characteristic genes for the final classification model, wherein the further screening is to further screen genes by calculating the betweenness centrality of the nodes, the betweenness centrality is sorted from large to small, and the top-ranked genes are selected as the characteristic genes finally used for the classification model.

[0020] As a preferred solution of the method for screening characteristic genes of breast cancer of the present invention, the characteristic gene data set is discretized and a state variable is added, and then the data set is randomly divided into a training set and a test set.

[0021] As a preferred solution of the method for screening characteristic genes of breast cancer of the present invention, the training set retains state variables, and the test set removes state variables.

[0022] As a preferred solution of the method for screening characteristic genes of breast cancer described in the present invention, the data in the training set is calculated using a Bayesian network algorithm, a Bayesian network classifier is constructed, the greedy search algorithm is combined with the fNML scoring function, and the optimal parameters are determined using Bayesian parameter estimation.

[0023] As a preferred solution of the method for screening characteristic genes of breast cancer of the present invention, after the Bayesian network classifier is constructed, a test set is used to evaluate the performance, and a prediction function is used to predict or estimate the state variables.

[0024] As a preferred embodiment of the method for screening characteristic genes of breast cancer described in the present invention, the confusion matrix is ​​used to display the results, and the results are verified in combination with diagnostic indicators, wherein the characteristic gene screening method is verified by calculating the diagnostic indicators such as accuracy, precision, recall rate, F1 score, ROC curve and AUC area of ​​the model.

[0025] Beneficial effects of the present invention:

[0026] The present invention uses a minimum classification error rate criterion that takes into account distribution uncertainty to screen characteristic genes for cancer diagnosis, which can achieve accurate diagnosis of breast cancer. Accurate screening of characteristic genes can reduce unnecessary testing and analysis, save time and cost, and improve diagnosis and treatment efficiency. By improving classification accuracy, the probability of misdiagnosis and missed diagnosis can be reduced, ensuring that patients receive timely and correct treatment, providing doctors with new research tools and methods to help them discover new cancer markers and treatment targets in clinical research. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:

[0028] Figure 1 The present invention provides an overall flow chart of the method for screening breast cancer characteristic genes based on the minimum classification error rate criterion.

[0029] Figure 2 is the gene expression matrix.

[0030] Figure 3 Schematic diagram for calculating the minimum classification error rate criterion.

[0031] Figure 4 is the gene expression matrix of the training set.

[0032] Figure 5 is the classification error rate of the gene in the two types of samples.

[0033] Figure 6 is the correlation between modules and features.

[0034] Figure 7 To screen for significant module gene sets.

[0035] Figure 8 The Venn diagram is obtained by taking the intersection of the pre-screened genes and the significant module gene set.

[0036] Fig. 9 This is a protein interaction network diagram.

[0037] Fig.10 This is the Bayesian network structure diagram.

[0038] Fig.11 is the confusion matrix.

[0039] Fig.12 is the ROC curve and the area under the curve AUC. DETAILED DESCRIPTION

[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the embodiments of the specification.

[0041] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0042] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0043] Example 1

[0044] This embodiment provides a method for screening characteristic genes of breast cancer based on the minimum classification error rate criterion, including determining characteristic genes of breast cancer according to the minimum classification error rate criterion, WGCNA and protein interaction network analysis, and constructing a Bayesian network classifier using the expression of these genes to diagnose and predict breast cancer. The overall process is as follows Figure 1 As shown, the specific steps include:

[0045] (1) Gene expression data of breast cancer samples and normal (paracancerous tissue) samples were downloaded from the open source database and cleaned, including deletion of duplicate and low-expression genes, log2 conversion, and Z-score normalization. Finally, the gene expression was displayed in the form of a matrix. The results are shown in the figure. Figure 2 As shown, an element g in the matrix m,n is the count value of gene n in sample m.

[0046] (2) Pre-screening of characteristic genes is done by calculating the minimum classification error rate of standardized gene expression data. The calculation principle is as follows: Figure 3 As shown. The minimum classification error rate is obtained by calculating the overlapping area of ​​the normal distribution curve and the cancer distribution curve. The similarity or difference between the distributions of the two sets of data is quantified by the minimum classification error rate. The smaller the minimum classification error rate, the greater the difference between the gene in the two types of samples, and the larger the minimum classification error rate, the more similar the gene is in the two types of samples.

[0047] The intersection of the two types of samples in the distribution curve is denoted by x *, the distribution curve of the normal group is denoted as f1(x), and the distribution curve of the cancer group is denoted as f2(x). Figure 3 In the example, the area of ​​S1 represents the "false negative" in the classification task, and the area of ​​S2 represents the "false positive". The sum of the areas of the two regions is equivalent to the error rate in the classification task. The calculation process is shown in formula (1-2). The classification error rates of all genes in the two types of samples are calculated and sorted from small to large. The top 5% are taken as candidate genes for pre-screening.

[0048]

[0049] (3) Weighted gene co-expression network analysis (WGCNA) was performed on the dataset to obtain the significant module gene set.

[0050] (4) Intersect the genes pre-screened in step 2 with the significant module gene set to obtain key cancer genes.

[0051] (5) Construct a protein interaction network (PPI) and further filter genes by node degree to obtain characteristic genes for the final classification model. The PPI analysis process is as follows: Generate a network in the STRING (https: / / string-db.org) website. Then use the CytoNCA plug-in in Cytoscape to calculate the betweenness centrality of the nodes in the PPI. The betweenness centrality of a node (gene) v can be shown by formula (3).

[0052]

[0053] where σ st is the number of shortest paths from node v to node t, and σ st (v) is the number of times these paths pass through v.

[0054] (6) The characteristic genes obtained above are selected from the gene expression matrix and discretized for subsequent input into the Bayesian network. Subsequently, a state variable "state" is added to the discretized data set. "Normal" represents normal samples and "Cancer" represents cancer samples. The data set is randomly divided into a training set and a test set in a ratio of 7:3. The "state" variable is retained in the training set and removed in the test set. The gene expression matrix of the training set is as follows: Figure 4 shown.

[0055] (7) Create a Bayesian network classifier using the training set. In the process of creating the Bayesian network structure learning, a greedy search algorithm is used to gradually add or delete edges to construct the network structure, and the factorized normalized maximum likelihood (fNML) scoring function is used to evaluate the quality of each candidate structure. fNML takes into account the maximum likelihood estimation and parameter penalty to help find the optimal network structure to best fit the data and maintain the simplicity of the model. This function is suitable for learning Bayesian network structures from completely discrete data. In the parameter learning process, the Bayesian parameter estimation method is used for learning, in which the Dirichlet distribution is used to represent the prior distribution of the parameters, and the maximum a posteriori (MAP) probability criterion is used to determine the optimal parameters. At this point, the Bayesian network classifier has been built.

[0056] (8) The state variable “state” is predicted or estimated using the prediction function, which operates on the test set. The power of this function is calculated by a likelihood-weighted average using all available nodes as evidence (obviously, excluding the predicted node value). If the variable being predicted is a discrete variable, the predicted value is the expected value of its conditional distribution. In this approach, the classification model is built on a dataset with predefined labels, which excels in predicting the classification label and classifying cancer data based on the training set.

[0057] (9) Finally, the confusion matrix was used to display the results of the test set. The accuracy, precision, recall, F1 score, ROC curve and AUC area of ​​the model were calculated, and these diagnostic indicators were combined to verify the characteristic gene screening method proposed in the present invention.

[0058] Example 2

[0059] This example uses a breast cancer (BRCA) data set (including normal control samples) obtained from the TCGA (The Cancer Genome Atlas) open source database to test the method proposed in Example 1, including the following steps:

[0060] (1) Gene expression data of breast cancer samples and normal samples were downloaded from TCGA, and gene mRNA expression levels were selected for analysis. After deleting duplicate and low-expression genes, 16,476 mRNAs and 1,185 samples were obtained, including 1,086 cancer samples and 99 normal samples. The gene expression matrix was log2 transformed and Z-score normalized for subsequent analysis.

[0061] (2) By calculating the minimum classification error rate of breast cancer samples and normal samples under the gene normal distribution curve, the results are sorted from small to large, and the top 5% of genes are selected as candidate genes, a total of 800. The calculation results of the genes ranked 1st, 200th, 400th and 600th are as follows Figure 5 shown.

[0062] (3) Weighted co-expression network analysis (WGCNA) was performed on the data set to obtain two gene modules, MEbrown and MEblue, which were highly correlated with cancer. Two thresholds, gene significance (GS) and module membership (MM), were set to screen key genes, such as Figure 6 , 7 As shown in Figure 3, 518 and 779 genes were obtained, respectively, and a total of 1297 key genes were used for subsequent analysis.

[0063] (4) The genes pre-screened in step 2 are intersected with the significant module gene set to obtain cancer key genes. The overlapping genes are displayed by the Venn diagram, and 200 genes are obtained by intersecting based on the minimum classification error rate criterion and WGCNA proposed in the present invention, such as Figure 8 shown.

[0064] (5) Use the 200 genes obtained in the previous step to construct a protein-protein interaction network (PPI), such as Fig. 9 As shown. Genes were further screened by calculating the betweenness centrality of the nodes, and the betweenness centrality was sorted from large to small. The top 25 genes were selected as the final feature genes for the classification model, namely CD34, CAV1, MYH11, PPARG, CD36, GSN, LEP, SPARCL1, CLEC3B, SH3D19, CRYAB, CLDN5, LYVE1, EMCN, TGFBR2, RBP4, AP1M2, SYNM, STAT5B, LIPE, CIDEC, FABP4, ANXA1, PRNP and CNN1.

[0065] (6) After the characteristic gene set is screened out from the original gene expression matrix, it is discretized into 5 values. A state variable "state" is added to the discretized data set. "Normal" represents normal samples and "Cancer" represents cancer samples. The data set is randomly divided into a training set and a test set in a ratio of 7:3. The "state" variable is retained in the training set and removed in the test set. The training set is used to learn the structure of the Bayesian network, and the test set is used to test the performance of the diagnostic model.

[0066] (7) Create a Bayesian network classifier using the training set. The greedy search algorithm is combined with the fNML score function to learn the network structure, and the Bayesian parameter estimation is used to determine the optimal parameters. This step is implemented based on the bnlearn package in R language, and the network structure diagram is visualized with the help of Cytoscape, as shown in the figure below. Fig.10 shown.

[0067] (8) Use the test set to evaluate the performance of the Bayesian network classifier. Call the prediction function to predict or estimate the state variable "state". The function of this function is to calculate the likelihood weighted average using all available nodes as evidence.

[0068] (9)Finally, the confusion matrix is ​​used to display the results of the test set, such as Fig.11 As shown. By calculating the accuracy, precision, recall rate, F1 score, ROC curve and AUC area of ​​the model, the characteristic gene screening method proposed in the present invention was verified in combination with these diagnostic indicators. The results show that the characteristic genes screened based on the minimum classification error rate criterion proposed in the present invention show excellent results in the diagnosis of breast cancer. The diagnostic indicator results are shown in Table 1, and the area under the ROC curve is shown in Table 2. Fig.12 shown.

[0069] Table 1 Results of the model in breast cancer diagnosis

[0070]

[0071] The present invention aims to overcome the defects of the prior art and invent a method for screening characteristic genes of breast cancer based on the minimum classification error rate criterion. First, the overlapping area of ​​the expression distribution curves of genes in cancer samples and control samples is calculated, that is, the minimum classification error rate with statistical significance, and the characteristic genes of breast cancer are pre-screened from TCGA data based on the minimum overall classification error rate; secondly, characteristic genes are further screened according to WGCNA (weighted gene co-expression network analysis) and PPI (protein interaction network) analysis, and a Bayesian network for diagnosing breast cancer is constructed based on the screened genes; finally, the effectiveness of the gene screening and diagnosis method is verified using TCGA data.

[0072] The present invention proposes a method for screening characteristic genes of breast cancer based on the minimum classification error rate criterion, and on this basis, proposes a method for screening characteristic genes of breast cancer combining the minimum classification error rate criterion with two bioinformatics methods, namely, weighted gene co-expression network analysis and protein interaction network analysis. That is, the characteristic genes of breast cancer are first pre-screened from TCGA data based on the minimum classification error rate as the standard, and then the characteristic genes are further screened according to the weighted gene co-expression network analysis and the protein interaction network analysis. Finally, a Bayesian network classification model of breast cancer is constructed according to the screened characteristic genes to realize accurate diagnosis of breast cancer, and the effectiveness of the method is verified according to model evaluation.

[0073] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the present invention.

Claims

1. A method for screening characteristic genes of breast cancer based on the minimum classification error rate criterion, characterized in that: include, Preprocessing gene expression data to obtain a gene expression data set; Pre-screening of characteristic genes according to the minimum classification error rate criterion; Perform weighted gene co-expression network analysis on the dataset to obtain significant module gene sets; The intersection of the pre-screened characteristic genes and the significant module gene set was taken to obtain the key cancer genes; Construct a protein interaction network and further screen the characteristic genes for the final classification model; The above-mentioned characteristic gene expression levels were used to construct a Bayesian network classifier for breast cancer diagnosis and prediction; The results were presented using confusion matrices and verified using diagnostic metrics.

2. The method for screening breast cancer characteristic genes according to claim 1, wherein: The preprocessing of gene expression data obtains a gene expression data set, wherein the gene expression data is gene expression data of breast cancer samples and normal adjacent tissue samples.

3. The method for screening breast cancer characteristic genes according to claim 2, wherein: The gene expression dataset is displayed in the form of a matrix.

4. The method for screening breast cancer characteristic genes according to claim 1, wherein: The feature base is pre-screened according to the minimum classification error rate criterion, wherein the minimum classification error rate is obtained by calculating the overlapping area of ​​the normal distribution curve and the cancer distribution curve, and the similarity or difference between the two groups of data distributions is quantified by the minimum classification error rate.

5. The method for screening breast cancer characteristic genes according to claim 1, wherein: The further screening obtains characteristic genes for the final classification model, wherein the further screening is to further screen genes by calculating the betweenness centrality of the nodes, sorting the betweenness centrality from large to small, and selecting the top-ranked genes as the characteristic genes for the final classification model.

6. The method for screening breast cancer characteristic genes according to claim 5, characterized in that: The characteristic gene data set is discretized and then state variables are added, and then randomly divided into a training set and a test set.

7. The method for screening breast cancer characteristic genes according to claim 6, wherein: The training set retains the state variables, and the test set removes the state variables.

8. The method for screening breast cancer characteristic genes according to claim 6, wherein: The data in the training set are calculated using a Bayesian network algorithm to construct a Bayesian network classifier, the greedy search algorithm is combined with the fNML score function, and the Bayesian parameter estimation is used to determine the optimal parameters.

9. The method for screening breast cancer characteristic genes according to claim 7, wherein: After the Bayesian network classifier is constructed, the performance is evaluated using a test set, and the state variables are predicted or estimated using a prediction function.

10. The method for screening breast cancer characteristic genes according to claim 1, characterized in that: The confusion matrix was used to display the results and the results were verified by combining the diagnostic indicators, wherein the characteristic gene screening method was verified by calculating the diagnostic indicators such as accuracy, precision, recall rate, F1 score, ROC curve and AUC area of ​​the model.