A method for predicting the survival status of cancer patients based on the extraction of characteristic genes

The most valuable feature genes were screened through hierarchical clustering and differential analysis, and combined with Lasso, Ridge, RandomForestSRC models and normalization processing, the most valuable feature genes were extracted, solving the problem of low prediction accuracy in the case of small sample size and large feature variables, and achieving higher prediction accuracy.

CN115376689BActive Publication Date: 2025-06-10SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211026600.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2025-06-10
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

In the case of small sample size and many characteristic variables, existing characteristic gene extraction methods cannot effectively extract meaningful features, resulting in low prediction accuracy.

Method used

The data were classified by hierarchical clustering method, and the differential characteristic genes were determined through differential analysis, and the intersection characteristic genes were extracted based on the Lasso, Ridge, and RandomForestSRC models, combined with normalization processing, and finally the training set was constructed for prediction model training.

Benefits of technology

This improves the accuracy of survival status prediction in cancer patients, especially on datasets with small sample sizes but many characteristic variables and highly linear correlations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376689B_ABST
    Figure CN115376689B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting the survival status of cancer patients based on the extraction of characteristic genes. The method includes: classifying the data in the dataset based on the hierarchical clustering method to obtain the classified data; performing differential analysis on the classified data to determine the differential characteristic genes; training the Lasso model, Ridge model, and RandomForestSRC model based on the characteristic genes and extracting the intersection characteristic genes; performing normalization processing on the intersection characteristic genes and screening according to the normalization result to obtain the final characteristic genes; constructing a training set based on the final characteristic genes and training the prediction model to obtain the trained prediction model; predicting the survival status of cancer patients based on the trained prediction model. By using the present invention, the problem of being unable to extract meaningful features in the case of a small sample size and many feature variables can be solved, thereby improving the prediction accuracy. The present invention can be widely applied to the field of status prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of state prediction, and in particular to a method for predicting the survival state of cancer patients based on the extraction of characteristic genes. Background Art

[0002] Malignant tumors are diseases that highly affect human health and have received extensive attention in the medical field. Malignant tumors from different tissue sources have their own characteristics in morphology and biological behavior. However, there are intricate connections between different malignant tumor cells and the tumor microenvironment (TME). When normal cells transform into tumor cells, their gene expression changes to a certain extent, and a series of biomarkers are obtained. At the same time, the component cells of the TME will also undergo a series of compositional and functional changes in response to tumor cells. These changes can be reflected in gene expression and modification. State prediction methods are all based on the method of extracting characteristic genes for prediction, and the existing characteristic gene extraction methods cannot be used when the sample size is small. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for predicting the survival state of cancer patients based on the extraction of characteristic genes, which solves the problem of not being able to extract meaningful features when the sample size is small and the number of characteristic variables is large, thereby improving the prediction accuracy.

[0004] The first technical solution adopted by the present invention is: A method for predicting the survival state of cancer patients based on the extraction of characteristic genes, comprising the following steps:

[0005] Classify the data in the dataset based on the hierarchical clustering method to obtain the classified data;

[0006] Perform differential analysis on the classified data to determine the differential characteristic genes;

[0007] Train the Lasso model, Ridge model, and RandomForestSRC model based on the characteristic genes and extract the intersection characteristic genes;

[0008] Normalize the intersection characteristic genes and calculate the corresponding coefficient scores to obtain the final characteristic genes;

[0009] Construct a training set based on the final characteristic genes and the patient's conventional characteristics and train the prediction model to obtain the trained prediction model;

[0010] Predict the survival state of cancer patients based on the trained prediction model.

[0011] Further, it further includes:

[0012] Calculate the risk assessment coefficient based on the final characteristic genes;

[0013] Verify the accuracy of the risk assessment coefficient based on the KM curve and HR value.

[0014] Further, the step of classifying the data in the dataset based on the hierarchical clustering method to obtain the classified data specifically includes:

[0015] Obtain the dataset and preprocess the data in the dataset to obtain the preprocessed data;

[0016] Conduct hierarchical clustering analysis on the preprocessed data to obtain the hierarchical clustering result;

[0017] Classify the data in the dataset according to the hierarchical clustering result to obtain the classified data.

[0018] Further, the step of performing differential analysis on the classified data to determine the differential characteristic genes specifically includes:

[0019] Based on the classified data, perform differential analysis pairwise to obtain the differential analysis result;

[0020] Display the differential analysis result through a volcano plot and select the differential characteristics greater than the preset threshold between pairwise categories;

[0021] Display the differential characteristics selected for each category through a Venn diagram, obtain the intersection of the differential characteristics between all categories, and obtain the differential characteristic genes.

[0022] Further, the step of training the Lasso model, Ridge model, and RandomForestSRC model based on the differential characteristic genes and extracting the intersection characteristic genes specifically includes:

[0023] Train the Lasso model based on the characteristic genes and extract the characteristic genes with the absolute value of the non-zero characteristic coefficient before the preset number;

[0024] Train the Ridge model based on the characteristic genes and extract the characteristic genes with the absolute value of the non-zero characteristic coefficient before the preset number;

[0025] Train the RandomForestSRC model based on the characteristic genes and extract the characteristic genes with the absolute value of the non-zero characteristic coefficient before the preset number;

[0026] Take the intersection of the genes extracted by the three models to obtain the intersection characteristic genes.

[0027] Further, the calculation formula for normalization is as follows:

[0028]

[0029] In the above formula, x represents the coefficient of the characteristic gene, and x ’ represents the coefficient of the normalized characteristic gene.

[0030] Furthermore, the calculation formula of the risk assessment coefficient is as follows:

[0031]

[0032] In the above formula, represents the normalization result of the gene coefficient obtained by the lasso model, represents the normalization result of the gene coefficient obtained by the Ridge model, represents the normalization result of the gene coefficient obtained by the RandomForestSRC model, and coefficient(x i ) represents the coefficient score.

[0033] The beneficial effect of the method of the present invention is that the present invention first preliminarily screens a large number of genes in the data set through the hierarchical clustering method and the differential analysis method, and finally extracts the most valuable genes based on the three models of Lasso, Ridge, and RandomForestSRC and the normalization method. This method is particularly effective for data sets with the characteristics of a small number of samples, a large number of characteristic variables, and a high degree of linear correlation between individual variables. Description of the Drawings

[0034] Figure 1 is a step flowchart of a method for predicting the survival status of cancer patients based on characteristic gene extraction according to the present invention. Detailed Embodiments

[0035] The following further describes the present invention in detail with reference to the drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0036] As Figure 1 shown, the present invention provides a method for predicting the survival status of cancer patients based on characteristic gene extraction, and the method includes the following steps:

[0037] S1. Classify the data in the data set based on the hierarchical clustering method to obtain the classified data;

[0038] S1.1. Obtain the data set and preprocess the data in the data set to obtain the preprocessed data;

[0039] Specifically, download the gene sequencing expression dataset of cervical cancer from the TCGA official website and preprocess the data. The preprocessing steps include converting IDs and handling missing values, etc.

[0040] S1.2. Perform hierarchical clustering analysis on the preprocessed data to obtain the hierarchical clustering result;

[0041] S1.3. Classify the data in the dataset according to the hierarchical clustering result to obtain the classified data.

[0042] Specifically, first perform GSVA analysis on the data, and then use hierarchical clustering to perform clustering analysis on the GSVA analysis results. Display the results of the clustering analysis with a heatmap, and determine that the samples are divided into 3 categories according to the heatmap.

[0043] GSVA is a non-parametric and unsupervised algorithm. The full name is Gene set variation analysis. GSVA can calculate the enrichment score of a specific gene set in each sample without having to pre-group the samples in advance. It can be understood that GSVA transforms the gene expression data, converting the expression matrix with individual genes as features into an expression matrix with specific gene sets as features.

[0044] Hierarchical clustering is a type of clustering analysis. It creates a hierarchical nested clustering tree by calculating the similarity between data points. It can adopt an "agglomerative" strategy from bottom to top or a "divisive" strategy from top to bottom.

[0045] Differential analysis is a common data analysis method used to detect whether there are differences and whether the differences are significant between the experimental group and the control group in a scientific experiment, also known as a significant difference test.

[0046] S2. Perform differential analysis on the classified data to determine the differential characteristic genes;

[0047] S2.1. Based on the classified data, perform differential analysis pairwise to obtain the differential analysis result;

[0048] S2.2. Display the differential analysis result with a volcano plot and select the differential characteristics greater than the preset threshold between pairwise categories;

[0049] S2.3. Display the differential characteristics selected for each category with a Venn diagram, obtain the intersection of the differential characteristics between all categories, and obtain the differential characteristic genes.

[0050] Specifically, for samples divided into three categories, differential analysis was performed pairwise, and a volcano plot was used to display the results of the differential analysis. The threshold P_value = 0.01 was determined according to the volcano plot. Feature genes with P-value < 0.01 in the differential analysis results were selected. A Venn diagram was used to show the intersection of differential genes between every two of the three categories. Through the above data analysis and result screening, 481 genes were obtained, including both immune genes and non-immune genes. Since the research focuses on immune-related genes, only the immune genes among the 481 genes were extracted. Finally, 309 immune genes were extracted.

[0051] S3. Train the Lasso model, Ridge model, and RandomForestSRC model based on the feature genes and extract the intersection feature genes;

[0052] S3.1. Train the Lasso model based on the feature genes and extract the feature genes with the absolute value of non-zero feature coefficients before a preset number;

[0053] S3.2. Train the Ridge model based on the feature genes and extract the feature genes with the absolute value of non-zero feature coefficients before a preset number;

[0054] S3.3. Train the RandomForestSRC model based on the feature genes and extract the feature genes with the absolute value of non-zero feature coefficients before a preset number;

[0055] S3.4. Take the intersection of the genes extracted by the three models to obtain the intersection feature genes.

[0056] Specifically, extract the feature genes with the absolute value of non-zero feature coefficients before top_N, and use the dataset composed of the 309 genes extracted in step S2 as the training set. Train the Lasso, Ridge, and randomForestSRC models. Extract the top 100 genes with the absolute value of the importance coefficients in the models. Take the intersection feature genes of the top 100 feature genes of the three models.

[0057] S4. Normalize the intersection feature genes and calculate the corresponding coefficient scores to obtain the final feature genes;

[0058] Specifically, add the coefficients of the intersection feature genes of the three models after normalization to obtain the importance coefficient of the genes. According to the obtained gene coefficient situation, take the top 5 genes with coefficients greater than 1 as the genes for subsequent analysis. Five genes have been extracted from more than 60,000 genes here.

[0059] The calculation formula for normalization is as follows:

[0060]

[0061] In the above formula, x represents the coefficient of the feature gene, x’ Represents the coefficient of the normalized characteristic gene.

[0062] Calculate the coefficient score of the characteristic gene, and the formula is as follows:

[0063]

[0064] In the above formula, Represents the normalization result of the gene coefficient obtained by the lasso model, Represents the normalization result of the gene coefficient obtained by the Ridge model, Represents the normalization result of the gene coefficient obtained by the RandomForestSRC model, coefficient(x i ) represents the coefficient score.

[0065] S5. Construct a training set based on the final characteristic genes and train the prediction model to obtain a trained prediction model;

[0066] S5.1. Use the top_n feature dataset selected in step S4 to divide the data into a test set and a training set at a ratio of 3 / 7

[0067] S5.2. Use the training set to train the SVM, Randomforest, Neural Net, and AdaBoost models.

[0068] S5.3. Use the test set to test the prediction accuracy of the SVM, Randomforest, Neural Net, and AdaBoost models.

[0069] Specifically, select the dataset composed of the risk coefficient obtained in step S4 and the patient's age, patient's race, patient's cancer stage, and whether the patient smokes as features. Divide the data into a test set and a training set at a ratio of 3 / 7, and use the training set to train the RandomForest, Neural Net, and Adaboost models. The three models have good accuracy on the test set. From the prediction results of the three models on the test set, the accuracy of the prediction results on the test set reaches more than 80%. The AUC area value also reaches 90%.

[0070] S6. Predict the survival status of cancer patients based on the trained prediction model.

[0071] Further, as a preferred embodiment of this method, it further includes:

[0072] S7. Calculate the risk assessment coefficient according to the final characteristic genes;

[0073] Specifically, multiply the coefficient score of the characteristic gene by the mean value of the expression value to obtain the risk assessment coefficient of each sample, and the calculation formula is as follows:

[0074]

[0075] In the above formula, Expression(mRNA ij ) represents the expression value of the j-th gene in sample i, Coefficient(mRNA ij ) represents the coefficient score of the j-th gene in sample i, and Risk_coef i represents the risk assessment coefficient of sample i.

[0076] S8. Verify the accuracy of the risk assessment coefficient based on the KM curve and HR value.

[0077] S8.1. Divide the samples into high and low groups according to the risk assessment coefficient. Samples greater than mean(Risk_coef) are in the high group, and samples not lower than mean(Risk_coef) are in the low group;

[0078] S8.2. Draw the KM curves of the high and low groups and test the accuracy of the risk assessment coefficient through the P value;

[0079] S8.3. Calculate the HR value and test the accuracy of the risk assessment coefficient.

[0080] Specifically, use the mean of the expression values of the first 5 gene coefficients multiplied by the expression value of the gene to divide the samples into high and low groups and draw a KM curve graph. The p value is 0.00003, indicating that the High group has a higher survival rate than the low group. The HR value is 3.2, indicating that the survival rate is more than 3.2 times higher.

[0081] KM survival analysis refers to analyzing and inferring the survival time of organisms or humans based on data obtained from experiments or surveys, and studying the relationship and degree between survival time and outcome and numerous influencing factors. Kaplan-Meier curve estimation of survival rate is a non-parametric method. The Kaplan-Meier survival curve, with the vertical axis representing the survival probability and the horizontal axis representing the survival event. It presents as a descending curve. The steeper the descending slope, the lower the survival rate or the shorter the survival time, and its slope represents the death rate.

[0082] The English name of HR (hazard ratio) is Hazard Ratio. The hazard ratio is the ratio of two hazard rates. In medical and public health research, people often use the hazard ratio to represent the risk difference between the experimental group and the control group. The KM survival curve can visually represent the hazard ratio. The points on the curve represent the ratio of the number of surviving people to the total number of people in the group at this time, that is, the survival rate. The sum of the survival rate and the hazard rate is the ratio of the hazard rates of the two groups at any time point in the figure, which is the hazard ratio

[0083] This method provides a method for extracting normalized characteristic genes based on three models: Lasso, Ridge, and RandomForestSRC. First, 481 genes are screened out of more than 60,000 genes using GSVA analysis, clustering analysis, and differential analysis methods. Finally, the top_n genes are extracted using the normalization methods based on the three models of Lasso, Ridge, and RandomForestSRC. And the KM curve is used to verify whether the top_n genes are meaningful. This method is applicable to datasets with the characteristics of a small sample size, a large number of feature variables, and a high degree of linear correlation between individual variables. More than 30 cancer datasets on TCGA belong to this type of dataset, so this method can extract very meaningful genes from various cancer datasets in the TCGA database. This method solves the problem that many existing machine learning feature extraction methods cannot be used when the sample size is small but the number of feature variables is large. Because any effective machine learning must have its inductive preference, this method uses three machine learning algorithms to model and analyze the data. Finally, the learning results of the three models are normalized to extract the final target characteristic genes. This is why this method has good results on various datasets.

[0084] A cancer patient survival status prediction system based on characteristic gene extraction, comprising:

[0085] A classification module that classifies the data in the dataset based on the hierarchical clustering method to obtain the classified data;

[0086] A differential analysis module that performs differential analysis on the classified data to determine differential characteristic genes;

[0087] An extraction model training module that trains the Lasso model, Ridge model, and RandomForestSRC model based on the characteristic genes and extracts intersection characteristic genes;

[0088] A normalization module that normalizes the intersection characteristic genes and performs screening according to the normalization results to obtain the final characteristic genes;

[0089] A prediction model training module that constructs a training set based on the final characteristic genes and trains the prediction model to obtain a trained prediction model;

[0090] A prediction module that predicts the survival status of cancer patients based on the trained prediction model.

[0091] Specifically, it further includes:

[0092] A calculation module that calculates a risk assessment coefficient according to the final characteristic genes;

[0093] A verification module that verifies the accuracy of the risk assessment coefficient based on the KM curve and the HR value.

[0094] The content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented in the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0095] A device for predicting the survival status of cancer patients based on feature gene extraction:

[0096] At least one processor;

[0097] At least one memory for storing at least one program;

[0098] When the at least one program is executed by the at least one processor, the at least one processor implements the method for predicting the survival status of cancer patients based on feature gene extraction as described above.

[0099] The content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0100] A storage medium storing instructions executable by a processor, characterized in that: the instructions executable by the processor are used to implement the method for predicting the survival status of cancer patients based on feature gene extraction as described above when executed by the processor.

[0101] The content in the above method embodiments is applicable to the storage medium embodiments. The functions specifically implemented in the storage medium embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0102] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for predicting the survival status of cancer patients based on the extraction of characteristic genes, characterized in that, it includes the following steps: Classify the data in the dataset based on the hierarchical clustering method to obtain the classified data; Perform differential analysis on the classified data to determine the differential characteristic genes; Train the Lasso model, Ridge model, and RandomForestSRC model based on the differential characteristic genes and extract the intersection characteristic genes; Normalize the intersection characteristic genes and calculate the corresponding coefficient scores to obtain the final characteristic genes; Construct a training set based on the final characteristic genes and the patient's conventional characteristics and train the prediction model to obtain the trained prediction model; Predict the survival status of cancer patients based on the trained prediction model; It also includes: Calculate the risk assessment coefficient according to the final characteristic genes; Verify the accuracy of the risk assessment coefficient based on the KM curve and HR value; The step of performing differential analysis on the classified data to determine the differential characteristic genes specifically includes: Based on the classified data, perform differential analysis pairwise to obtain the differential analysis results; Display the differential analysis results through a volcano plot and select the differential characteristics greater than the preset threshold between pairwise categories; Display the differential characteristics selected for each category through a Venn diagram, obtain the intersection of the differential characteristics between all categories, and obtain the differential characteristic genes; The step of training the Lasso model, Ridge model, and RandomForestSRC model based on the differential characteristic genes and extracting the intersection characteristic genes specifically includes: Train the Lasso model based on the characteristic genes and extract the characteristic genes with the absolute value of the non-zero characteristic coefficients before the preset number; Train the Ridge model based on the characteristic genes and extract the characteristic genes with the absolute value of the non-zero characteristic coefficients before the preset number; Train the RandomForestSRC model based on the characteristic genes and extract the characteristic genes with the absolute value of the non-zero characteristic coefficients before the preset number; Take the intersection of the genes extracted by the three models to obtain the intersection characteristic genes; The calculation formula for normalization is as follows: In the above formula, x represents the coefficient of the characteristic gene, and x ’ represents the coefficient of the normalized characteristic gene; The calculation formula for the coefficient score is as follows: In the above formula, represents the normalized result of the gene coefficients obtained by the lasso model, represents the normalized result of the gene coefficients obtained by the Ridge model, represents the normalized result of the gene coefficients obtained by the RandomForestSRC model, coefficient(x i ) represents the coefficient score; The calculation formula for the risk assessment coefficient is as follows: In the above formula, Expression(mRNA ij ) represents the expression value of the j-th gene in sample i, Coefficient(mRNA ij ) represents the coefficient score of the j-th gene in sample i, and Risk_coef i represents the risk assessment coefficient of sample i.

2. The method for predicting the survival status of cancer patients based on the extraction of characteristic genes according to claim 1, characterized in that, The step of classifying the data in the dataset based on the hierarchical clustering method to obtain the classified data specifically includes: Obtain the dataset and preprocess the data in the dataset to obtain the preprocessed data; Perform hierarchical clustering analysis on the preprocessed data to obtain the hierarchical clustering results; Classify the data in the dataset according to the hierarchical clustering results to obtain the classified data.

Citation Information

Patent Citations

  • Short-term traffic flow prediction method based on gradient boosting decision tree

    CN113096388A

  • Construction method of novel 9-gene RISK (Reduced Induced Shift Keying) acute myelogenous leukemia prognosis model

    CN114220487A