Prognostic markers for gastric adenocarcinoma and clinical prognostic prediction model
By screening 8 genes as prognostic markers of gastric adenocarcinoma and constructing a random forest prognosis model, the problem of inaccurate prognostic evaluation of gastric adenocarcinoma in the prior art was solved, and more accurate prognostic evaluation and the effect of improving patient survival was achieved.
Patent Information
- Application Number
- CN202510311734.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The prior art is difficult to effectively predict the prognosis of gastric adenocarcinoma, resulting in inconsistent treatment response and prognosis, and lack of sensitive treatment methods and diagnostic indicators.
By screening the eight genes F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS and CTHRC1 as prognostic markers of gastric adenocarcinoma, a clinical prognostic prediction model based on a random forest algorithm was constructed to evaluate the prognostic risk of gastric adenocarcinoma patients.
It has achieved a more accurate prediction of the prognosis of patients with gastric adenocarcinoma, effectively evaluate the prognostic risks of different patients, and intervention early, thereby improving the 5-year survival rate of patients with gastric adenocarcinoma.
Smart Images

Figure CN119824095B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of biomedicine, and particularly relates to a prognostic marker for gastric adenocarcinoma and a clinical prognostic prediction model. Background Art
[0002] Gastric cancer is the fifth most common malignant tumor globally and the third highest in lethality among all cancers, with its incidence varying by region. Surgical operation remains the cornerstone of gastric cancer treatment. In recent years, with the continuous innovation of immune and molecular targeted drugs, the modes of perioperative treatment and advanced treatment have made great changes and progress, also increasing the survival probability of patients. Although the prognosis of current gastric cancer patients has been improved through radical surgery combined with postoperative adjuvant treatment, approximately 50% of patients still experience tumor recurrence and metastasis within 2 years after surgery and ultimately die. In addition, due to high tumor heterogeneity, the treatment responses and prognoses at different stages are inconsistent. Therefore, studying the molecular mechanisms of tumor invasion, metastasis, occurrence, and prognosis from the perspective of genomics may provide highly sensitive treatment methods, which also offers more hope and possibilities for the determination of new prognostic and diagnostic indicators and treatment targets. We aim to search for new effective prognostic risk markers for gastric adenocarcinoma and establish a new risk assessment model to provide a basis for the further diagnosis and treatment of gastric adenocarcinoma patients.
[0003] F5 gene: The coding gene of coagulation factor V, and the encoded product is an important cofactor in the blood coagulation cascade process. This factor circulates in plasma and is converted into an active form through an activation peptide released by thrombin during blood coagulation, generating a heavy chain and a light chain, which are fixed together by calcium ions. Activated protein is a cofactor that acts together with activated coagulation factor X to activate prothrombin into thrombin. Defects in this gene can lead to autosomal recessive bleeding disorders or autosomal dominant thrombotic disorders, namely activated protein C resistance.
[0004] SLC5A1 gene: The coding gene of solute carrier family 5 member 1. This gene encodes a member of the sodium-dependent glucose transporter (SGLT) family. The encoded integral membrane protein is the main mediator for the uptake of dietary glucose and galactose in the intestinal lumen. Mutations in this gene are associated with glucose-galactose malabsorption. Diseases related to SLC5A1 include glucose / galactose malabsorption and renal glycosuria. Its related pathways include digestion and absorption and transmembrane transporter disorders.
[0005] PHYHD1 gene: The encoding gene of Phytanoyl-CoA dioxygenase domain-containing 1. The expression product of this gene is a 2-oxoglutarate (2OG)-dependent dioxygenase, which has potential roles in Alzheimer's disease, certain cancers, and immune cell function.
[0006] The function of the FNDC1 (Fibronectin Type III Domain Containing 1) gene is involved in multiple biological processes, such as the cellular response to hypoxia, cardiomyocyte apoptosis, and protein phosphorylation. It has been found that overexpression or knockdown of FNDC1 can promote or inhibit tumorigenesis and metastasis, respectively. For example, in colorectal cancer and non-small cell lung cancer, upregulation of FNDC1 accelerates the survival ability of tumor cells to 5-FU or radiotherapy, while its inhibition makes tumor cells more sensitive to chemotherapy and radiotherapy.
[0007] NFE2L3 gene: Belongs to the cap 'n' collar basic region leucine zipper family and is mainly located in the endoplasmic reticulum and nuclear membrane. As an important regulator of the cellular stress response, it is expressed in various human tissues. It is involved in the regulation of multiple biological and cellular processes, such as the cell cycle, cell differentiation, or inflammation, and is significantly associated with epigenetics, cell cycle regulation, calcium signaling pathways, and steroid hormone biosynthesis.
[0008] SCUBE2 gene: The encoded protein is a secreted or membrane-bound protein initially discovered in endothelial cells (ECs). It is involved in the regulation of gene expression and post-translational control, participates in the regulation of the HH signaling pathway, promotes endochondral ossification, and promotes angiogenesis. SCUBE2 is considered a tumor suppressor and is commonly downregulated in several tumors, including gastric cancer. In patients with dyslipidemia and type 2 diabetes, high expression of SCUBE2 contributes to the integrity of the blood-brain barrier, and multiple studies have linked it to autoimmune diseases.
[0009] CTHRC1 gene: Encodes a 28 kDa secreted glycosylated protein. Many studies have confirmed the expression of CTHRC1 in various human solid tumors. It has been reported that the expression of CTHRC1 can increase the expression of the activated HIF-1α / CXCR4 signaling pathway, which may lead to the migration and invasion of gastric cancer cells.
[0010] CBS gene: Located at 21q22, it participates in key steps in the sulfur amino acid metabolic pathway in the body. It has multiple transcription start sites, can participate in multiple metabolic pathways, especially regulate the bioenergetic activity of cells by producing H2S, and is upregulated in various cancers. Its expression is closely related to tumor angiogenesis and is also related to the increased demand for sulfur amino acids by tumor cells, which is an important factor in tumor growth and metastasis. Summary of the Invention
[0011] The present invention provides a prognostic marker for gastric adenocarcinoma in view of the above technical problems existing in the prior art.
[0012] The present invention first provides a prognostic marker for gastric adenocarcinoma, which consists of 8 genes: F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS, and CTHRC1.
[0013] The present invention also provides the application of the prognostic marker for gastric adenocarcinoma as a detection marker in the preparation of a clinical prognostic prediction model for gastric adenocarcinoma. Preferably, the construction of the clinical prognostic prediction model for gastric adenocarcinoma includes the following steps:
[0014] S1: Obtain the expression data of the 8 genes in the prognostic marker for gastric adenocarcinoma in samples from gastric adenocarcinoma patients for model construction, as well as the prognostic data of gastric adenocarcinoma patients, and the prognostic data is the survival status and survival time;
[0015] S2: Use the prognostic data as the target variable and train it using the random forest algorithm to obtain the trained model, which is the clinical prognostic prediction model for gastric adenocarcinoma.
[0016] During the training process, each tree is constructed using different feature subsets and sample subsets; the splitting of the tree is based on the information gain of each gene on the prognostic target. After training each tree, the model will integrate the prediction results of these trees and obtain the final prediction result through the average value; during the model establishment process, it is specified that each tree will divide the data into 10 subsets, and the number of trees is set to 1000 to perform hyperparameter optimization, and finally a model with the best performance is obtained.
[0017] The present invention also provides the application of a reagent for detecting the gene expression level of the prognostic marker for gastric adenocarcinoma in the preparation of a gastric adenocarcinoma prognostic prediction product. Preferably, the gastric adenocarcinoma prognostic prediction includes the following steps:
[0018] (1) Detect the expression levels of the 8 genes: F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS, and CTHRC1 in a gastric adenocarcinoma tumor tissue sample from a test individual.
[0019] (2) Use the expression data of the 8 genes obtained in step (1) as input and input it into the trained gastric adenocarcinoma clinical prognosis prediction model; calculate the risk score, and the higher the risk score, the worse the prognosis.
[0020] Among them, the construction of the trained gastric adenocarcinoma clinical prognosis prediction model includes the following steps:
[0021] S1: Obtain the expression data of 8 genes among the gastric adenocarcinoma prognosis markers in the samples from gastric adenocarcinoma patients for model construction, as well as the prognosis data of gastric adenocarcinoma patients. The prognosis data is the survival status and survival time.
[0022] S2: Use the prognosis data as the target variable and train it using the random forest algorithm. The obtained trained model is the gastric adenocarcinoma clinical prognosis prediction model.
[0023] During the training process, each tree is constructed using different feature subsets and sample subsets; the splitting of the tree is based on the information gain of each gene on the prognosis target. After training each tree, the model will integrate the prediction results of these trees and obtain the final prediction result through the average value; during the model establishment process, it is specified that each tree will divide the data into 10 subsets, and the number of trees is set to 1000, so as to perform hyperparameter optimization and finally obtain the model with the best performance.
[0024] The present invention also provides a gastric adenocarcinoma prognosis prediction product, including a reagent for detecting the gene expression level of the gastric adenocarcinoma prognosis marker.
[0025] Preferably, the reagent is a reagent for detecting the expression level of each molecular marker. More preferably, the reagent is a primer for detecting the expression level of each molecular marker.
[0026] The present invention also provides a gastric adenocarcinoma prognosis evaluation model, including a detection data input module, an analysis module, and a result output module. Among them,
[0027] Detection module: Used to detect the sample from the individual to be tested and obtain the expression levels of 8 genes, namely F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS, and CTHRC1, in the sample.
[0028] Analysis module: Includes the trained gastric adenocarcinoma clinical prognosis prediction model. The analysis module is used to analyze the expression levels of each gene detected by the detection module, and the obtained expression data is used as input and input into the trained gastric adenocarcinoma clinical prognosis prediction model.
[0029] The construction of the trained gastric adenocarcinoma clinical prognosis prediction model includes the following steps:
[0030] S1: Obtain the expression data of 8 genes among the gastric adenocarcinoma prognosis markers in the samples from gastric adenocarcinoma patients for model construction, as well as the prognosis data of gastric adenocarcinoma patients, where the prognosis data is the survival status and survival time;
[0031] S2: Use the prognosis data as the target variable and train it using the random forest algorithm. The trained model is the gastric adenocarcinoma clinical prognosis prediction model.
[0032] During the training process, each tree is constructed using different feature subsets and sample subsets; the splitting of the tree is based on the information gain of each gene on the prognosis target. After training each tree, the model will integrate the prediction results of these trees and obtain the final prediction result through the average value; during the model establishment process, it is specified that each tree will divide the data into 10 subsets, and the number of trees is set to 1000 for hyperparameter optimization, and finally a model with the best performance is obtained.
[0033] Result output module: Calculate the patient risk score. The higher the patient risk score, the worse the prognosis.
[0034] The present invention uses the overall survival data of gastric adenocarcinoma patients in the TCGA and GETx2 databases, screens genes F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS, and CTHRC1 as gastric adenocarcinoma prognosis markers, constructs a risk assessment model based on F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS, and CTHRC1, verifies the accuracy of the model using an external dataset, and finds that the risk score is significantly correlated with the relatively abundant stromal cells in its tumor microenvironment. The present invention helps to better predict the prognosis of gastric adenocarcinoma patients, effectively evaluate the prognosis risks of different patients, intervene early, bring great help to guiding the clinical treatment of this disease, and thus improve the 5-year survival rate of gastric adenocarcinoma patients.
[0035] The randomly selected survival forest model of the present invention has the following advantages. First, high prediction accuracy: it has a lower prediction error rate and prediction error compared with the traditional Cox regression model, especially showing excellent performance in dealing with high-dimensional data and complex interactions. Second, strong variable identification ability: the model can effectively identify the key variables affecting prognosis and can screen variables through methods such as minimum depth, which helps to simplify the model and improve interpretability. Third, strong adaptability: it does not need to assume that the influence of variables on the risk function is linear and is applicable to various types of survival data. Fourth, good generalization ability: by adjusting parameters multiple times to find the optimal settings, the model can maintain high prediction performance on different data sets. In addition, this model also combines the VIMP method and the minimum depth method. The former has advantages such as quickly responding to changes in requirements and using UML modeling methods to guide design. The latter can improve the quality of local minima, does not require over-parameterization assumptions, and can gradually approach the global optimal solution. Therefore, the model combining the random survival forest method with the VIMP method and the minimum depth method has many advantages as mentioned above, and to a certain extent reflects the incomparable superiority of this model.
[0036] The risk assessment model of the present invention has high prediction accuracy and universality, and is significantly superior to other gastric adenocarcinoma prognosis methods in the prior art and also significantly superior to other models constructed during the research process. Description of the Drawings
[0037] Figure 1 It is a single-factor COX regression analysis diagram.
[0038] Figure 2 It is a VIMP method analysis diagram when screening gastric adenocarcinoma tumor marker genes.
[0039] Figure 3 It is the VIMP values of 8 genes screened by the random survival forest.
[0040] Figure 4 It is a diagram of the prediction error rate of models with different numbers of survival trees.
[0041] Figure 5 It is to divide the data in the training set into high-risk groups and low-risk groups through the surv_cutpoint function.
[0042] Figure 6 It is a box plot of the scores of patients with different survival situations in the training set.
[0043] Figure 7 It is to compare the differences in survival situations between the high-risk group and the low-risk group in the training set.
[0044] Figure 8 It is a risk scatter plot in the training set.
[0045] Figure 9 The difference in survival rates between high-risk and low-risk patients in the training set.
[0046] Figure 10 The reliability of the prognostic prediction of the prediction model in the training set as shown by ROC curve analysis.
[0047] Figure 11 The time-dependent ROC curve of the training set.
[0048] Figure 12 The difference in survival rates between high-risk and low-risk patients in the validation set.
[0049] Figure 13 The time-dependent ROC curve of the validation set.
[0050] Figure 14 The difference result graph of the risk score between the high group and the low group.
[0051] Figure 15 The difference result graph of the tumor purity between the high group and the low group.
[0052] Figure 16 The difference result graph of the tumor stroma between the high group and the low group, where p = 2e-08 means p = 2 -8 . Detailed implementation methods
[0053] Example 1
[0054] I. Methods
[0055] 1. Data collection.
[0056] Use the GEPIA2 online website to analyze the gene expression profile data of gastric adenocarcinoma patients, obtain the differential genes of gastric adenocarcinoma as the training set, include 81 gastric adenocarcinoma patients treated in the Zhejiang Cancer Hospital, and perform transcriptome sequencing on the tumors as the validation set.
[0057] 2. Screening of prognostic markers.
[0058] Using the data in the training set, first exclude non-coding RNAs, set the screening conditions as |log 2 FC| > 2, adjusted p value < 0.05, and 714 differential genes were screened. In the training set, univariate COX regression was used to further screen prognostic-related genes, and the threshold was set to p < 0.1. The results are as Figure 1 shown. 184 prognostic-related genes were obtained for subsequent analysis. The screening p < 0.01 is shown in the figure. Hazard ratios greater than 1 indicate that the gene is related to poor prognosis, and less than 1 indicates that the gene is related to good prognosis.
[0059] We combined two variable evaluation methods, the VIMP method and the minimum depth method, to determine the threshold. Figure 2 Figure 2 is the analysis diagram of the VIMP method for screening gastric adenocarcinoma tumor marker genes. The VIMP value of each gene was calculated using the gg_vimp function. The VIMP value represents the contribution degree of the gene to the model prediction accuracy. Genes with a VIMP value greater than 0 indicate a positive impact on the model, while variables less than 0 may reduce the prediction accuracy of the model. The gg_minimal_depth function was used to evaluate the importance of variables by the minimum depth method. The minimum depth represents the first split position of a variable in the tree. The shallower the depth, the more important the variable. Eight genes with a relatively large correlation with the prognosis of gastric adenocarcinoma were screened and used as risk markers.
[0060] 3. Model establishment.
[0061] Prepare the expression data of the screened feature genes. Use the prognosis data (survival status and survival time) of the samples as the target variable and train using the random forest algorithm. During the training process, each tree is constructed using different feature subsets (randomly selected feature genes) and sample subsets (bootstrap sampling method). The splitting of the tree is based on the information gain of each gene on the prognosis target. After training each tree, the model will integrate the prediction results of these trees and obtain the final prediction result through the average value (regression task). During the model establishment process, limit each tree to divide the data into 10 subsets and set the number of trees to 1000 for hyperparameter optimization, and finally obtain the model with the best performance.
[0062] 4. Risk classification method.
[0063] Substitute the 8 risk marker genes into the random survival forest prognosis model, draw the curve of the number of trees and the error rate, and calculate the risk score of each sample. Figure 3 Figure 3 is the VIMP value of the 8 genes screened by the random survival forest, showing the importance of these 8 genes. Figure 4 Figure 4 is the model prediction error rate diagram for different numbers of survival trees. As the number of survival trees increases, its prediction error rate decreases significantly; when the number of survival trees increases to a certain number, the error rate curve tends to be stable, and the number of decision trees is set to 1000.
[0064] Determine the high- and low-risk groups according to the surv_cutpoint function. Figure 5 Figure 5 is the division of the data in the training set into high-risk and low-risk groups by the surv_cutpoint function. Figure 6 (where **** represents p < 0.0001) is the score box plot of patients with different survival situations in the training set, and it can be seen the overall score difference between the surviving patients (0) and the dead patients (1) in the training set. Figure 7To compare the differences in survival between the high-risk group and the low-risk group in the training set. Figure 8 This is the risk scatter plot in the training set, from which the score distributions of the high-risk group and the low-risk group can be seen. Figure 9 This shows the differences in the survival rates of high-risk and low-risk patients in the training set.
[0065] 5. Evaluate the relationship between the risk model and the overall survival rate in the training set and verify the accuracy of the model.
[0066] By plotting the Kaplan-Meier curve to compare the differences in survival rates between the high- and low-risk groups, and by using the R package "survivalROC" to plot the ROC curve of the patients and calculate the area under the curve to verify the accuracy of the model. Figure 10 This is the ROC curve analysis showing the reliability of the prediction model in predicting the prognosis in the training set. Figure 11 This is the time-dependent ROC curve, demonstrating the accuracy of the model in predicting the survival rate at 2 years, 3 years, and 4 years in the training set.
[0067] 6. Verify the accuracy of the model in the validation set.
[0068] According to the boundary of high- and low-risk division confirmed in the training set, the samples in the validation set are divided into a high-risk group and a low-risk group. By plotting the Kaplan-Meier curve to compare the differences in survival rates between the high- and low-risk groups, and by using the R package "survivalROC" to plot the ROC curve of the patients and calculate the area under the curve, the accuracy of the model in the validation set is further verified. Figure 12 This shows the differences in the survival rates of high-risk and low-risk patients in the validation set. Figure 13 This is the Kaplan-Meier survival curve, showing the differences in the survival rates of the high- and low-risk score groups in the validation set.
[0069] 7. Correlation analysis of the risk score with other factors.
[0070] First, it is determined that the risk score is significantly correlated with the prognosis ( Figure 14 ), and then the relationships between the risk score and tumor purity ( Figure 15 ), and the tumor stroma ( Figure 16 ) are compared.
[0071] II. Results
[0072] 1. Apply the randomForest random forest machine learning method in the training set, combine the VIMP method with the minimal-depth method to determine the threshold, and screen out 8 prognosis-related genes in Table 1.
[0073] Table 1. Depth values and VIMP values of genes
[0074] Gene Depth value VIMP value F5 8.380 0.002 SLC5A1 8.644 0.007 PHYHD1 8.848 0.001 FNDC1 8.877 0.003 NFE2L3 8.899 0.003 SCUBE2 8.929 0.001 CBS 8.965 0.002 CTHRC1 8.988 0.003
[0075] 2. Training set risk score.
[0076] The surv_cutpoint function finds the best position to split the data by traversing all possible cutpoints of the continuous variable. Its calculation process is based on the log-rank test statistic, and the formula is: , where Z: risk score;
[0077] Q 1 : The number of observed events in the first group (such as death);
[0078] E 1 : The expected number of events in the first group (estimated based on the survival rates of all groups);
[0079] V: The variance of this test.
[0080] The surv_cutpoint function calculates that 20.63 is the threshold of the risk score. The expression levels of 8 genes, namely F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS, and CTHRC1, are measured by the kit and brought into the model to calculate the risk score of each patient. Using 20.63 as the threshold, the risk scores of gastric adenocarcinoma patients in the training set are divided into a high-risk group and a low-risk group. Kaplan-Meier curve analysis shows that the survival rate of the high-risk group is lower than that of the low-risk group ( Figure 9 , p < 0.001). The overall ROC curve shows that the AUC value is 0.86 ( Figure 10 ), and the areas under the ROC curves for predicting 2-year, 3-year, and 4-year survival rates are 0.789, 0.783, and 0.828 respectively ( Figure 11 ), further indicating that the model has good predictive ability. At the same time, we also verified the risk scores and survival information in the training set and found that the number of survivors in the high-risk group is lower.
[0081] 3. Validate the risk model in the validation set.
[0082] To validate the universality of the risk model, the validation set is divided into a high-risk group and a low-risk group using the threshold determined in the training set. Kaplan-Meier curve analysis shows that the survival rate of the high-risk group is lower than that of the low-risk group ( Figure 12 , p = 0.032), which is consistent with what was found in the training set. The areas under the ROC curves for predicting 2-year, 3-year, and 4-year survival rates are 0.757, 0.788, and 0.776 respectively ( Figure 13 ), indicating that the model also has good predictive ability in the validation set.
[0083] 4. Relationship between risk score and tumor purity, tumor stroma.
[0084] The tumor purity and tumor stroma score of each sample were calculated by the ESTIMATE algorithm in R language, and it was found that the tumor purity and tumor stroma score in the high-risk group were lower in the training set ( Figure 15 , Figure 16 ).
Claims
1. A prognostic marker for gastric adenocarcinoma, characterized in that: It consists of eight genes: F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS and CTHRC1.
2. Use of the gastric adenocarcinoma prognosis marker according to claim 1 as a detection marker in the preparation of a clinical prognosis prediction model for gastric adenocarcinoma.
3. The use according to claim 2, characterized in that: The construction of the clinical prognosis prediction model for gastric adenocarcinoma includes the following steps: S1: Obtaining the expression data of the eight genes in the gastric adenocarcinoma prognostic markers in the samples from gastric adenocarcinoma patients used for model construction, as well as the prognostic data of gastric adenocarcinoma patients, the prognostic data being the survival status and survival time; S2: The prognosis data is used as the target variable and the random forest algorithm is used for training to obtain a trained model, which is the clinical prognosis prediction model for gastric adenocarcinoma. During the training process, each tree is constructed using different feature subsets and sample subsets; the basis for tree splitting is the information gain of each gene on the prognostic target. After training each tree, the model will integrate the prediction results of these trees and obtain the final prediction result through the average value; during the model building process, each tree is limited to split the data into 10 subsets, and the number of trees is set to 1000, so as to optimize the hyperparameters and finally obtain the model with the best performance.
4. Use of a reagent for detecting the gene expression level of the gastric adenocarcinoma prognosis marker according to claim 1 in the preparation of a gastric adenocarcinoma prognosis prediction product.
5. The use according to claim 4, characterized in that: Prognosis prediction for gastric adenocarcinoma involves the following steps: (1) Detect the expression levels of eight genes, including F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS, and CTHRC1, in gastric adenocarcinoma tumor tissue samples from the tested individuals; (2) The expression data of the eight genes obtained in step (1) are used as input into the trained gastric adenocarcinoma clinical prognosis prediction model; a risk score is calculated, and the higher the risk score, the worse the prognosis; The construction of the trained gastric adenocarcinoma clinical prognosis prediction model includes the following steps: S1: Obtaining the expression data of the eight genes in the gastric adenocarcinoma prognostic markers in the samples from gastric adenocarcinoma patients used for model construction, as well as the prognostic data of gastric adenocarcinoma patients, the prognostic data being the survival status and survival time; S2: The prognosis data is used as the target variable and the random forest algorithm is used for training to obtain a trained model, which is the clinical prognosis prediction model for gastric adenocarcinoma. During the training process, each tree is constructed using different feature subsets and sample subsets; the basis for tree splitting is the information gain of each gene on the prognostic target. After training each tree, the model will integrate the prediction results of these trees and obtain the final prediction result through the average value; during the model building process, each tree is limited to split the data into 10 subsets, and the number of trees is set to 1000, so as to optimize the hyperparameters and finally obtain the model with the best performance.
6. A clinical prognosis prediction model for gastric adenocarcinoma, characterized in that: It includes a detection module, an analysis module and a result output module, wherein: Detection module: used to detect samples from individuals to be tested and obtain the expression levels of eight genes, namely, F5, SLC5A1, PHYHD1, FNDC1, NFE2L3, SCUBE2, CBS and CTHRC1, in the samples; Analysis module: including a trained gastric adenocarcinoma clinical prognosis prediction model, the analysis module is used to analyze the expression level of each gene detected by the detection module, and the obtained expression level data is used as input to the trained gastric adenocarcinoma clinical prognosis prediction model. The construction of the trained gastric adenocarcinoma clinical prognosis prediction model includes the following steps: S1: obtaining the expression data of the eight genes in the gastric adenocarcinoma prognostic markers of claim 1 from samples from gastric adenocarcinoma patients for model construction, as well as the prognostic data of gastric adenocarcinoma patients, wherein the prognostic data are survival status and survival time; S2: The prognosis data is used as the target variable and the random forest algorithm is used for training to obtain a trained model, which is the clinical prognosis prediction model for gastric adenocarcinoma. During the training process, each tree is constructed using different feature subsets and sample subsets. The basis for tree splitting is the information gain of each gene on the prognostic target. After training each tree, the model will integrate the prediction results of these trees and obtain the final prediction result through the average value. During the model building process, each tree is limited to split the data into 10 subsets, and the number of trees is set to 1000, so as to optimize the hyperparameters and finally obtain the model with the best performance. Result output module: Calculate the patient risk score. The higher the patient risk score, the worse the prognosis.
Citation Information
Patent Citations
Methods for treating small cell neuroendocrine and related cancers
US20220244263A1