A method, device, and storage medium for predicting renal clear cell carcinoma

By using genetic material and clinical information data of patients with clear cell renal tumors, A2M-related gene clusters were screened, a gene risk scoring model was constructed, and variable regression analysis was performed. This solved the problem of low prediction accuracy of clear cell renal tumors in existing technologies and achieved higher-precision survival probability prediction.

CN118711678BActive Publication Date: 2025-12-12THE FIRST AFFILIATED HOSPITAL OF JINAN UNIV
2 Cites 0 Cited by

Patent Information

Application Number
CN202411159983.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-12-12
Estimated Expiration
2044-08-22

Smart Images

  • Figure CN118711678B_ABST
    Figure CN118711678B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a kidney clear cell tumor prediction method, device, equipment and storage medium, target sample data such as genetic material test data and clinical information data of a kidney clear cell tumor patient are acquired; correlation analysis of A2M gene in kidney clear cell tumor is performed according to the target sample data, and a target gene group is determined; a gene risk score model is constructed according to the target gene group, and the risk score corresponding to each kidney clear cell tumor patient is calculated according to the gene risk score model; variable regression analysis training is performed according to the target sample data and the risk score, and a target prognosis model is obtained, wherein the independent variable data is age, TNM stage, pathological stage and risk score, and the dependent variable data is overall survival; the target prognosis model is used for prediction processing based on input information, and survival probability and survival curve are output, so that the problem of low prediction accuracy of kidney clear cell tumor can be solved, and the prediction accuracy of kidney clear cell tumor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of biomedical prediction, and in particular to a renal clear cell tumor prediction method, device, equipment and storage medium. BACKGROUND

[0002] Renal clear cell tumor is a malignant tumor derived from renal tubular epithelial cells, and is the most common type of renal cancer, accounting for 70-80% of renal cancer.

[0003] The prediction of the survival probability of renal clear cell tumor plays an important role in the entire treatment process of the patient, can provide a basis for clinical decision-making for doctors, can be used as an important indicator for evaluating the treatment effect, and can also enable the patient and family to be more clear about the treatment goal and expectation, thereby reducing anxiety and fear and improving the quality of life.

[0004] The existing prediction of the survival probability of renal clear cell tumor usually depends on clinical indicators such as tumor size and tumor grade. However, these clinical indicators have certain limitations in terms of prediction accuracy and individualized treatment plan development, resulting in relatively low prediction accuracy. SUMMARY

[0005] Embodiments of the present application provide a renal clear cell tumor prediction method, device, equipment and storage medium, which can solve the problem of low prediction accuracy of renal clear cell tumor and improve the prediction accuracy of renal clear cell tumor.

[0006] In a first aspect, embodiments of the present application provide a renal clear cell tumor prediction method, comprising:

[0007] Obtaining target sample data, the target sample data including genetic material test data and clinical information data of renal clear cell tumor patients, the clinical information including overall survival, disease-specific survival, progression-free survival, age, TNM stage, pathological stage, histological grade, survival status and gender;

[0008] Performing correlation analysis of A2M gene in renal clear cell tumor according to the target sample data, determining a screening variable, and determining a target gene group, the target gene group being a plurality of genes related to A2M;

[0009] Constructing a gene risk score model according to the target gene, and calculating a risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model;

[0010] According to the target sample data and the risk score, variable regression analysis training is performed to obtain a target prognosis model, wherein the independent variable data in the variable regression analysis training is the age, TNM stage, pathological stage and risk score in the target sample data, and the dependent variable data is the overall survival time;

[0011] Through the target prognosis model, prediction processing is performed based on input information to output a prediction result, the input information includes the age, TNM stage, pathological stage and risk score to be predicted, and the prediction result includes the survival probability and the survival curve.

[0012] Further, according to the target sample data, A2M gene correlation analysis in renal clear cell tumor is performed to screen variables and determine a target gene group, including:

[0013] Through single gene correlation analysis, A2M gene correlation analysis in renal clear cell tumor is performed based on the target sample data to obtain A2M related genes;

[0014] Through single variable COX regression analysis, A2M related gene survival prognosis analysis in renal clear cell tumor is performed based on the target sample data, and a first candidate gene meeting a preset threshold requirement is determined from the A2M related genes;

[0015] Through an L1 regularization model, a least absolute shrinkage and selection operator (LASSO) regression of overall survival time of renal clear cell carcinoma is performed based on the target sample data, and a target gene group meeting an error value condition is determined from the first candidate gene.

[0016] Further, the target gene group includes TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3;

[0017] According to the target gene group, a gene risk score model is constructed, and the risk score corresponding to each renal clear cell tumor patient in the target sample data is calculated according to the gene risk score model, including:

[0018] According to the gene expression amount of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 and the corresponding preset risk coefficient, a gene risk score model is constructed;

[0019] According to the gene risk score model, the risk score corresponding to each renal clear cell tumor patient in the target sample data is calculated.

[0020] Further, the gene risk score model is characterized by the following ways:

[0021]

[0022] Wherein, A2M GPI represents the risk score, representing the preset risk coefficient of the gene, representing the gene expression of the gene.

[0023] Further, the gene risk score model is characterized by the following ways:

[0024] A2M GPI=(-0.06057755*TIE1exp.)+(0.00416184*VWFexp.)+(-0.22620967*TCF4exp.)+(0.62309290*PTPRBexp.)+(-0.06403383*ICAM2exp.)+(-0.19163895*DOCK6exp.)+(0.08291625*RAMP3exp.)

[0025] Wherein, (-0.06057755), 0.00416184, (-0.22620967), 0.62309290, (-0.06403383), (-0.19163895) and 0.08291625 represent the preset risk coefficients of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3, respectively.

[0026] Further, according to the target sample data and the risk score, a variable regression analysis training is performed to obtain a target prognosis model, including:

[0027] Through a single variable COX regression model, a single factor analysis is performed based on the clinical information data in the target sample data to determine the target variable factor, and the target variable factor is the age, TNM stage and pathological stage in the target sample data.

[0028] Through a multivariate COX regression model, the target variable factor and the risk score are used as independent variable data, and the overall survival time in the target sample data is used as the dependent variable to perform variable regression analysis training to obtain the target prognosis model.

[0029] Further, the target prognosis model is used to perform prediction processing based on the input information, and the prediction result is output, including:

[0030] Through the target prognosis model, the age, TNM stage, pathological stage and risk score to be predicted are used to perform prediction processing to obtain a single variable prediction score, wherein the age, TNM stage, pathological stage and risk score correspond to a single variable prediction score, respectively.

[0031] The single variable prediction scores are added to obtain a model prediction total score;

[0032] According to the model, the total score is predicted, and a conversion function is converted to obtain a survival probability in a preset time range;

[0033] According to the survival probability of the preset time, a corresponding survival curve is drawn, and the survival curve takes time as the horizontal coordinate and survival probability as the vertical coordinate.

[0034] In a second aspect, the embodiments of the present application provide a renal clear cell tumor prediction device, comprising:

[0035] The sample acquisition module is configured to acquire target sample data, wherein the target sample data comprises genetic material test data and clinical information data of a renal clear cell tumor patient, and the clinical information comprises overall survival, disease-specific survival, progression-free survival, age, TNM stage, pathological stage, histological grade, survival status, and gender.

[0036] The gene correlation analysis module is configured to perform correlation analysis of A2M gene in renal clear cell tumor according to the target sample data, screen variables, and determine a target gene group, wherein the target gene group is a plurality of genes related to A2M.

[0037] The gene risk score module is configured to construct a gene risk score model according to the target gene group, and calculate a risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model.

[0038] The prognosis model training module is configured to perform variable regression analysis training according to the target sample data and the risk score to obtain a target prognosis model, wherein the independent variable data in the variable regression analysis training is the age, TNM stage, pathological stage, and risk score in the target sample data, and the dependent variable data is the overall survival.

[0039] The prediction module is configured to perform prediction processing based on input information through the target prognosis model to output a prediction result, wherein the input information comprises an age, TNM stage, pathological stage, and risk score to be predicted, and the prediction result comprises a survival probability and a survival curve.

[0040] Further, the gene correlation analysis module comprises an A2M single variable correlation analysis submodule, a first single variable COX regression submodule, and a linear fitting submodule.

[0041] The A2M single gene correlation analysis submodule is configured to perform survival status correlation analysis of A2M gene in renal clear cell tumor based on the target sample data through single gene correlation analysis to obtain A2M related genes.

[0042] a first univariate COX regression submodule configured to determine a first candidate gene from the A2M-related genes based on the target sample data by performing an A2M-related gene and renal clear cell tumor survival prognosis analysis based on a single gene correlation analysis, wherein the first candidate gene meets a preset threshold requirement;

[0043] a linear fitting submodule configured to determine a target gene group from the first candidate gene based on the target sample data by performing a least absolute shrinkage and selection operator (LASSO) regression of the overall survival of renal clear cell carcinoma based on an L1 regularization model, wherein the target gene group meets an error value condition.

[0044] Further, the target gene group includes TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6, and RAMP3.

[0045] The gene risk score module includes a score model construction submodule and a score calculation submodule.

[0046] The score model construction submodule is configured to construct a gene risk score model based on the gene expression amounts of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6, and RAMP3 and corresponding preset risk coefficients.

[0047] The score calculation submodule is configured to calculate a risk score corresponding to each renal clear cell tumor patient in the target sample data based on the gene risk score model.

[0048] Further, the gene risk score model is characterized by the following:

[0049]

[0050] wherein A2M_GPI represents the risk score, represents a preset risk coefficient of the i-th gene, represents a gene expression amount of the i-th gene.

[0051] Further, the gene risk score model is characterized by the following:

[0052] A2M_GPI=(-0.06057755*TIE1exp.)+(0.00416184*VWFexp.)+(-0.22620967*TCF4exp.)+(0.62309290*PTPRBexp.)+(-0.06403383*ICAM2exp.)+(-0.19163895*DOCK6exp.)+(0.08291625*RAMP3exp.)

[0053] ​​Wherein, (-0.06057755), 0.00416184, (-0.22620967), 0.62309290, (-0.06403383), (-0.19163895) and 0.08291625 represent the preset risk coefficients of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 respectively, (TIE1exp.), (VWFexp.), (TCF4exp.), (PTPRBexp.), (ICAM2exp.), (DOCK6exp.) and (RAMP3exp.) represent the gene expression amounts of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 respectively.

[0054] Further, the prognosis model training module comprises a second univariate COX regression submodule and a multivariate analysis submodule;

[0055] The second univariate COX regression submodule is configured to perform single-factor analysis processing based on the clinical information data in the target sample data by a univariate COX regression model, and determine target variable factors, the target variable factors being age, TNM stage and pathological stage in the target sample data;

[0056] The multivariate analysis submodule is configured to perform variable regression analysis training by taking the target variable factors and the risk score as independent variable data and taking the overall survival time in the target sample data as a dependent variable by a multivariate COX regression model, and obtain a target prognosis model.

[0057] Further, the prediction module comprises a single-item prediction submodule, a comprehensive prediction submodule, a survival probability conversion submodule and a drawing submodule;

[0058] The single-item prediction submodule is configured to perform prediction processing based on the age, TNM stage, pathological stage and risk score to be predicted by the target prognosis model, and obtain single-item variable prediction scores, wherein the age, TNM stage, pathological stage and risk score correspond to one single-item variable prediction score respectively;

[0059] The comprehensive prediction submodule is configured to add the single-item variable prediction scores to obtain a model prediction total score;

[0060] The survival probability conversion submodule is configured to perform conversion processing according to the model prediction total score and a preset conversion function, and obtain a survival probability within a preset time limit;

[0061] The drawing submodule is configured to draw a corresponding survival curve according to the survival probability within the preset time limit, the survival curve taking time as the horizontal coordinate and survival probability as the vertical coordinate.

[0062] In a third aspect, an embodiment of the present application provides a renal clear cell tumor prediction device, comprising:

[0063] a memory and one or more processors;

[0064] a memory for storing one or more programs;

[0065] When the one or more programs are executed by the one or more processors, the one or more processors implement the renal clear cell tumor prediction method as in the first aspect.

[0066] In a fourth aspect, an embodiment of the present application provides a storage medium storing computer-executable instructions for performing the renal clear cell tumor prediction method as in the first aspect when executed by a computer processor.

[0067] The embodiments of the present application can obtain target sample data when making a prognosis of a renal clear cell tumor, perform correlation analysis of A2M gene in the renal clear cell tumor according to the target sample data, screen variables, determine a target gene group significantly related to the A2M gene, construct a gene risk score model according to the target gene group, calculate a risk score corresponding to each clear cell tumor patient in the target sample data according to the gene risk score model, perform variable regression analysis training according to the target sample data and the risk score, obtain a target prognosis model, perform prediction processing based on an age to be predicted, TNM staging, pathological staging and the risk score through the target prognosis model, and output a survival probability and a survival curve. The above technical means can obtain the target prognosis model through analysis and training of the target sample data and the target gene group significantly related to the A2M gene, perform prediction of the survival probability through the target prognosis model and draw the survival curve, so as to avoid the problem of low prediction accuracy of the renal clear cell tumor. Compared with the existing survival probability prediction method based on clinical indicators, the embodiments perform prediction of the survival probability of the renal clear cell tumor based on the target prognosis model, greatly improve the reliability of the prediction, and thus improve the accuracy of the prediction of the survival probability corresponding to the preset survival time of the renal clear cell tumor patient. BRIEF DESCRIPTION OF DRAWINGS

[0068] Figure 1 is a flowchart of a renal clear cell tumor prediction method provided by an embodiment of the present application;

[0069] Figure 2 is a prediction platform page display schematic diagram provided by an embodiment of the present application;

[0070] Figure 3 is a structural schematic diagram of a renal clear cell tumor prediction device provided by an embodiment of the present application;

[0071] Figure 4Fig. 1 is a structural schematic diagram of a renal clear cell tumor prediction device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0072] In order to make the purposes, technical solutions and advantages of the present application clearer, the following further describes specific embodiments of the present application with reference to the drawings. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application. In addition, it should be noted that, for the convenience of description, only parts related to the present application are shown in the drawings, but not all. Before discussing the example embodiments in more detail, it should be mentioned that some example embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The processes can be terminated when the operations are completed, but can also have additional steps not included in the drawings. The processes can correspond to methods, functions, procedures, subroutines, etc.

[0073] The existing prediction of the survival probability of renal clear cell tumor usually depends on clinical indicators such as tumor size and tumor grade. However, these clinical indicators have certain limitations in prediction accuracy and individualized treatment plan, resulting in relatively low prediction accuracy. Based on this, the renal clear cell tumor prediction method, device, equipment and storage medium provided by the embodiments of the present application are provided. The renal clear cell tumor prediction method, device, equipment and storage medium provided by the present application are aimed at predicting the prognosis of renal clear cell tumor, obtaining target sample data, performing correlation analysis of A2M gene in renal clear cell tumor according to the target sample data, screening variables, determining the target gene group significantly related to A2M gene, constructing a gene risk score model according to the target gene group, and calculating the risk score corresponding to each clear cell tumor patient in the target sample data according to the gene risk score model. According to the target sample data and the risk score, variable regression analysis training is performed to obtain a target prognosis model. The target prognosis model is used for prediction processing based on the age, TNM stage, pathological stage and risk score to be predicted, and the survival probability and survival curve are output. By using the above technical means, the target prognosis model can be obtained by analyzing and training the target sample data and the target gene group significantly related to A2M gene. The survival probability is predicted by the target prognosis model and the survival curve is drawn. In this way, the problem of low prediction accuracy of renal clear cell tumor can be avoided. Compared with the existing survival probability prediction based on clinical indicators, the present embodiment predicts the survival probability of renal clear cell tumor based on the target prognosis model, greatly improves the reliability of the prediction, and thus improves the accuracy of predicting the survival probability of the preset survival time of the renal clear cell tumor patient.

[0074] Figure 1 A flowchart of a kidney clear cell tumor prediction method provided in an embodiment of the present application is given. The kidney clear cell tumor prediction method provided in the embodiment can be executed by a kidney clear cell tumor prediction device. The kidney clear cell tumor prediction device can be implemented in software and / or hardware. The kidney clear cell tumor prediction device can be composed of two or more physical entities, or one physical entity. Generally, the kidney clear cell tumor prediction device can be an electronic device, such as a computer device.

[0075] The following describes an example in which a computer device is taken as a main body for executing the kidney clear cell tumor prediction method. Referring to FIG. 1, the kidney clear cell tumor prediction method includes the following steps. Figure 1 The kidney clear cell tumor prediction method specifically includes the following steps.

[0076] S101, target sample data is acquired, the target sample data including genetic material test data and clinical information data of a kidney clear cell tumor patient, the clinical information including overall survival, disease-specific survival, progression-free survival, age, TNM stage, pathological stage, histological grade, survival status, and gender.

[0077] Pan-cancer data is acquired from a public database, the pan-cancer data including genetic material test data and clinical information data of tumor patients of each type. For example, the pan-cancer data includes genetic material test data and clinical information data of patients of sarcoma (SARC), skin melanoma (SKCM), kidney clear cell carcinoma (KIRC), lung squamous cell carcinoma (LUSC), and gastric adenocarcinoma (STAD), etc. The genetic material test data is RNA-seq data, and the clinical information data includes overall survival (OS), disease-specific survival (DDS), progression-free survival (PFI), age (Age), TNM stage, pathological stage (Pathologic Stage, hereinafter referred to as Stage), histological grade (Histologic Grade), survival status (Status), and gender, etc. From the pan-cancer data, genetic material test data and clinical information data of kidney clear cell tumor (i.e., kidney clear cell carcinoma) patients can be acquired as target sample data. Subsequently, prognosis analysis of kidney clear cell tumor and establishment of a target prognosis model can be performed based on the target sample data.

[0078] Overall survival (OS) is the time from randomization (or treatment initiation in single-arm trials) to death due to any cause. It focuses on the overall survival time of patients, regardless of the specific cause of death. Disease-specific survival (DDS) is the time interval from the diagnosis of a specific disease (such as renal clear cell tumor) to death due to that disease or the last follow-up. It excludes deaths caused by other non-disease-related factors and focuses on evaluating the impact of a specific disease on patient survival. Progression-free survival (PFI) is the time interval from the start of a specific treatment (such as surgery, radiotherapy, or chemotherapy) to the observation of disease progression (such as tumor volume increase, new lesion appearance, etc.) or death due to any cause. TNM staging includes T staging, N staging, and M staging. Among them, T staging represents the size and extent of the tumor, which is divided into T0, Tis, T1-T4, where T0 represents no primary tumor found, Tis represents carcinoma in situ, the tumor is limited to the skin layer and has not broken through the basement membrane, and the larger the number of T1-T4 represents the larger the range and invasion of the tumor. N staging describes the involvement of regional lymph nodes by the tumor. N staging levels include: N0: no lymph node involvement; N1-N3: the staging gradually increases with the increase of lymph node involvement and extent. For example, N1 may indicate only a few lymph nodes are involved, while N3 may indicate multiple lymph nodes or lymph node groups are involved. If the lymph node metastasis situation cannot be determined, use Nx. M staging describes whether the tumor has distant metastasis. M staging levels include: M0: no distant metastasis. M1: distant metastasis. Pathologic stage is a stage that divides the changes in the structure and morphology of cells in the development of the disease. Early pathologic stage: usually indicates that the disease is in an early stage, the tumor is small, the infiltration is shallow, and the lymph node metastasis is less or none. Patients in this stage usually have a better prognosis and more treatment options. Middle pathologic stage: indicates that the disease is moderately severe, the tumor is larger or has deeper infiltration, and may have local lymph node metastasis. Patients in this stage need more aggressive treatment and closer follow-up. Late pathologic stage: usually indicates that the disease has reached a later stage, the tumor may be large or has distant metastasis. Patients in this stage have a greater difficulty in treatment and a relatively poor prognosis. Histologic grade is the evaluation of the differentiation degree, atypia, and mitotic figures of tumor cells through histological and cytological examination, to judge the malignancy of the tumor. Histologic grade can be divided into three or four levels, but the specific grading criteria vary depending on the type of tumor and the pathological diagnosis system, so the specific number of histologic grade can be determined according to the actual situation.Illustratively, histological grading can be divided into: G1 (Grade 1) represents low-grade or well-differentiated tumors, cell morphology is close to normal, and the degree of malignancy is low; G2 (Grade 2) represents intermediate or moderately differentiated tumors, cell morphology and degree of malignancy are between G1 and G3; G3 (Grade 3) and G4 (Grade 4) represent high-grade or poorly differentiated tumors, cell morphology is abnormal, and the degree of malignancy is high.

[0079] After obtaining the genetic material test data and clinical information data of the renal clear cell tumor patient from the public database, the obtained genetic material test data and clinical information data are format-converted and log2-transformed to obtain the corresponding gene expression matrix and clinical information data. Subsequently, corresponding analysis and training processing can be performed on the format-converted genetic material test data and clinical information data to obtain the corresponding target prognosis model.

[0080] The above-mentioned genetic material test data and clinical information data obtained from the public database are subjected to corresponding analysis and training processing, which provides a large amount of sample data for the construction of the target prognosis model, thereby improving the reliability of the target prognosis model, and further improving the accuracy of the prediction result based on the target prognosis model.

[0081] S102, according to the target sample data, the correlation of A2M gene in renal clear cell tumor is analyzed, the variables are screened, and the target gene group is determined, and the target gene group is a plurality of genes related to A2M.

[0082] A2M gene is also known as α-2-macroglobulin gene, which is located on chromosome 12. The protein encoded by this gene is a protease inhibitor and a cytokine transporter; uses decoy and trap mechanisms to inhibit a wide range of proteases, including trypsin, thrombin and collagenase; can also inhibit inflammatory cytokines, thereby disrupting the inflammatory cascade. A2M gene is associated with a variety of diseases, so this example analyzes the correlation of A2M gene in renal clear cell tumor to determine the influence of genes significantly related to A2M gene on the prognosis of renal clear cell tumor, and subsequently determines the gene most related to renal clear cell tumor in the A2M-related gene group as the target gene based on the correlation between the two, that is, the variables are screened, thereby determining the target gene group. Subsequently, risk scoring and target prognosis model creation are performed based on the target gene group to increase the prognostic influence of A2M gene on renal clear cell tumor, thereby improving the consideration dimension of the prognosis of renal clear cell tumor, and further improving the prognosis accuracy of renal clear cell tumor.

[0083] After obtaining the pan-cancer data from the public database, the correlation analysis of the A2M gene is performed according to the pan-cancer data. According to the analysis result, the expression of the A2M gene is related to the overall survival prognosis of multiple cancers. The Kaplan-Meier survival analysis of the pan-cancer data based on the A2M gene is performed, and the analysis result includes: the high A2M expression patient group shows a longer overall survival in renal clear cell carcinoma, skin melanoma and sarcoma. The differential expression of A2M is related to the disease-specific survival (DDS) prognosis of renal clear cell carcinoma, sarcoma, breast invasive carcinoma, cervical squamous cell carcinoma and endocervical adenocarcinoma. The analysis result also includes: in the renal clear cell carcinoma, the disease-specific survival (DDS) of the high A2M expression patient group is significantly better than that of the low A2M expression patient group. Therefore, according to the foregoing analysis result, the A2M gene is related to the survival probability prediction of renal clear cell tumor. Therefore, the correlation analysis of the A2M gene in the renal clear cell tumor can be performed according to the target sample data, the variables are screened, and the target gene group related to the prognosis of the renal clear cell tumor is determined, wherein the target gene group is a plurality of genes significantly related to the A2M gene.

[0084] In an embodiment, after obtaining the genetic material test data and the clinical information data of the renal clear cell tumor patients from the public database, the obtained genetic material test data and the clinical information data are format-converted and log2-transformed to obtain a corresponding gene expression matrix and clinical information data. The correlation of the A2M gene in the renal clear cell tumor is analyzed based on the target sample data through single gene correlation analysis to obtain A2M-related genes. Illustratively, according to the gene expression matrix, the samples with A2M expression levels ranked in 20%-80% are filtered out, and the samples with A2M gene expression levels higher than 80% are retained as a high expression group, and the samples with A2M gene expression levels lower than 20% are retained as a low expression group. The correlation of the A2M gene in the renal clear cell tumor is analyzed according to the sample data of the high expression group and the low expression group to determine 42 characteristic genes (i.e., A2M-related genes) closely related to the A2M gene associated with the renal clear cell tumor, which are PECAM1, MMRN2, NES, CD93, PCDH12, CD34, CDH5, TIE1, ERG, CLEC14A, ESAM, ENG, SHROOM4, VWF, TCF4, JAM3, SOX17, EPAS1, EHD4, ACVRL1, S1PR1, MYCT1, GIPC3, RASIP1, GSN, GIMAP8, ARHGEF15, HSPA12B, ROBO4, SEC14L1, PLXND1, ITGA9, PTPRB, PPM1F, TBXA2R, ICAM2, ECSCR, PLVAP, IL3RA, DOCK6, FILIP1 and RAMP3. The mutation of the aforementioned 42 characteristic genes is analyzed by using the “maftools” R package, and the mutation is visualized, and the genes with higher mutation frequencies are displayed in the form of a histogram. According to the visualization result of the mutation, it can be known that about 13.4% (45 / 336) of the renal clear cell tumor patients have (A2M) gene mutations, and the top 10 mutation genes are listed, in which the mutation frequency of VWF is the highest (18%), and the mutation frequencies of the other 9 genes are between 4% and 9%, and missense mutations account for the majority of mutation types.

[0085] After the 42 feature genes are determined, the survival prognosis analysis of A2M-related genes (i.e., the 42 feature genes) and renal clear cell tumor is performed based on the target sample data by univariate COX regression to determine the first candidate gene that meets the preset threshold requirement from the (42) A2M-related genes. For example, the univariate COX regression is used to evaluate whether the 42 feature genes have an impact on the survival status of renal clear cell tumor, and the threshold is adjusted to be less than 0.05, and the genes PLXND1 and IL3RA that do not meet the preset threshold condition requirement are excluded, and the first candidate gene that meets the preset threshold requirement is determined to be PECAM1, MMRN2, NES, CD93, PCDH12, CD34, CDH5, TIE1, ERG, CLEC14A, ESAM, ENG, SHROOM4, VWF, TCF4, JAM3, SOX17, EPAS1, EHD4, ACVRL1, S1PR1, MYCT1, GIPC3, RASIP1, GSN, GIMAP8, ARHGEF15, HSPA12B, ROBO4, SEC14L1, ITGA9, PTPRB, PPM1F, TBXA2R, ICAM2, ECSCR, PLVAP, DOCK6, FILIP1, and RAMP3.

[0086] After the first candidate gene is determined, the number of genes needs to be further reduced to obtain the corresponding more closely related target genes. The correlation analysis of A2M-related genes and overall survival of renal clear cell tumor is performed based on the target sample data by an L1 regularization model to determine the target gene group that meets the error value condition from the first candidate gene. For example, the glmnet R package is used to perform LASSO regression and ten-fold cross-validation on the first candidate gene to select the "lambda min" value that is most relevant to the overall survival of renal clear cell tumor patients in the TCGA-KIRC cohort, to further reduce the candidate genes and construct the most suitable prognosis risk features (i.e., target genes). Through the above analysis and processing, 7 A2M genes are finally determined as target genes, which are TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6, and RAMP3. The 7 A2M genes form a target gene group. Subsequently, further prognosis analysis can be performed based on these target genes to construct a final target prognosis model to obtain the prediction of the survival probability of renal clear cell tumor based on A2M genes.

[0087] The above, by analyzing the correlation between the target sample data and the A2M related genes and the prognosis of the renal clear cell tumor patients, the target gene group is determined, and the subsequent prognosis model of the survival probability of the renal clear cell tumor can be constructed based on the target gene group, so that the survival probability of the renal clear cell tumor patients can be predicted based on A2M and its related genes, and the expression amount of A2M gene is closely related to renal clear cell tumor based on the analysis result, so the reliability of the target prognosis model and the accuracy of the subsequent prediction can be improved.

[0088] S103, constructing a gene risk score model according to the target gene group, and calculating the risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model.

[0089] Based on the close correlation between A2M gene and overall survival of renal clear cell tumor, after obtaining the target gene group most closely related to renal clear cell tumor in the foregoing A2M gene, a gene risk score model is constructed according to the target gene group, and the risk score corresponding to each renal clear cell tumor patient in the target sample data is calculated according to the gene risk score model. Based on the risk score and the corresponding overall survival analysis, it can be determined that the overall survival of patients with high risk score is longer, and the overall survival of patients with low risk score is shorter.

[0090] The foregoing target gene is TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3, and a gene risk score model is constructed according to the gene expression amount of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 and the corresponding preset risk coefficient; the risk score corresponding to each renal clear cell tumor patient in the target sample data is calculated according to the gene risk score model. The gene risk score model is characterized by the following way:

[0091]

[0092] Wherein, A2M_GPI represents the risk score, the preset risk coefficient of the first gene, the gene expression amount of the first gene. It should be noted that the gene expression amount of each gene can be obtained from a public database.

[0093] For example, the gene risk score model is characterized by the following way:

[0094] A2M GPI = (-0.06057755*TIE1 exp.) + (0.00416184*VWF exp.) + (-0.22620967*TCF4 exp.) + (0.62309290*PTPRB exp.) + (-0.06403383*ICAM2 exp.) + (-0.19163895*DOCK6 exp.) + (0.08291625*RAMP3 exp.)

[0095] wherein, (-0.06057755), 0.00416184, (-0.22620967), 0.62309290, (-0.06403383), (-0.19163895) and 0.08291625 represent the preset risk coefficients of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3, respectively. TIE1 exp. represents the gene expression amount of TIE1, VWF exp. represents the gene expression amount of VWF, TCF4 exp. represents the gene expression amount of TCF4, PTPRB exp. represents the gene expression amount of PTPRB, ICAM2 exp. represents the gene expression amount of ICAM2, DOCK6 exp. represents the gene expression amount of DOCK6, and RAMP3 exp. represents the gene expression amount of RAMP3.

[0096] In an embodiment, the risk score A2M GPI of the corresponding patient in the target sample data can be calculated according to the gene risk score model. The target sample data is divided into a low A2M GPI group and a high A2M GPI group according to the median value of all A2M GPI in the target sample data, and can be divided into multiple low A2M GPI groups and multiple high A2M GPI groups according to actual conditions. The target prognosis model can be verified based on grouping. According to the risk score A2M GPI calculated in the foregoing, principal component analysis (PCA) is performed using the "stats" package, and Kaplan-Meier analysis is performed on the target sample data with overall survival (OS) > 30 days using the "survival" and "survminer" packages. The analysis results show that A2M GPI is significantly correlated with various clinical characteristics, such as histological grade (G1-G4), pathological stage (I-IV), T stage (T1-T4), N stage (N0-N1), and M stage (M0-M1). With the increase of the level, A2M GPI presents significant low expression, and the A2M GPI level of the surviving patients is significantly higher than that of the dead patients.

[0097] According to the above, the gene risk score model is constructed according to the target gene group, and the risk score corresponding to each renal clear cell tumor patient in the target sample data is calculated according to the gene risk score model. Subsequently, the risk score is used as an independent variable to participate in the training of the target prognosis model, and the risk score is used as an input information during subsequent prediction, thereby improving the importance of the A2M gene in the prediction of renal clear cell tumor. Based on the significant correlation of the A2M gene with various clinical characteristics of renal clear cell tumor, the introduction of the risk score corresponding to the A2M gene can improve the reliability of the target prognosis model obtained by training, and improve the accuracy of the survival probability prediction based on the target prognosis model, thereby improving the user prediction experience.

[0098] In S104, variable regression analysis training is performed according to the target sample data and the risk score, and a target prognosis model is obtained. In the variable regression analysis training, the independent variable data is the age, TNM stage, pathological stage and risk score in the target sample data, and the dependent variable data is the overall survival time.

[0099] According to the target sample data and the risk score, variable regression analysis training is performed. The age, TNM stage (mainly M stage) and pathological stage in the target sample data and the risk score are used as independent variables, and the overall survival time is used as a dependent variable for variable regression analysis training. A target prognosis model is obtained. The output result corresponding to the target prognosis model is the survival probability and the survival curve drawn based on the survival probability. The survival probability can be understood as the ratio of the number of total survival times satisfying the preset year limit (for example, 1 year, 3 years or 5 years) to the total sample data amount.

[0100] In an embodiment, in the target sample data, all data of different stages of one patient are combined into one sample data, and each patient is analyzed and trained by using one sample data.

[0101] In an embodiment, the target prognosis model is a Nomogram model, which can be processed by a univariate COX regression model based on the target sample data for single factor analysis to determine the target variable factors, which are age, TNM stage and pathological stage in the target sample data. For example, the single factor analysis is performed on the age (Age), TNM stage, pathological stage (Pathologic Stage), histological grade (Histologic Grade), survival status (Status) and gender in the target sample data to determine the multiple factors (i.e., target variable factors) closely related to the overall survival period from these single factors. According to the analysis result, the target variable factors are determined as age, TNM stage and pathological stage. Then, the variable regression analysis training is performed by a multivariate COX regression model with the target variable factors and risk score as the independent variable data and the overall survival period in the target sample data as the dependent variable to obtain the target prognosis model.

[0102] In an embodiment, in the target prognosis model, Points represents the single score corresponding to each prediction variable in the model under different grouping / values, such as the single score of 17.35 corresponding to the Stage III group in the pathological stage (Stage). Age, M stage, Stage and A2M GPI all represent each prediction variable, i.e., the independent variables are age, TNM stage and pathological stage. The score difference under the same numerical distance of the numerical variable is the same (the HR difference in the multivariate Cox regression analysis is also the same). The score difference between different levels of the ordinal variable is different. The results (score / HR value) of the binary classification variable are the same. Total Points represents the total score obtained by adding the single scores corresponding to all variable values, which can be used to calculate the total score corresponding to the total y value (Linear Predictor) of the model, and can also be used to calculate the event occurrence probability (x-year Survival Probability) corresponding to a certain total score. Linear Predictor represents the linear prediction value, and the total y value of the multiple factor model can be compared to Total Points to determine the total score corresponding to different y values, or can be compared to the following event occurrence probability (x-year Survival Probability). x-year Survival Probability represents the prediction probability corresponding to a specified prediction time range, i.e., the survival probability in the embodiment. In the target prognosis model (Nomogram model) constructed in the foregoing, the points value corresponding to the clinical characteristic variable (Age, M stage, Stage and A2M GPI) is found and summed to obtain the predicted survival probability of 1, 2, 3 and 5 years corresponding to total points.

[0103] It should be noted that the preset years provided in the embodiment are 1, 2, 3 and 5 years, which are merely for illustration, and the specific values of the preset years can be set according to actual conditions.

[0104] In the above, the target prognosis model is obtained by taking age, TNM stage, pathological stage and risk score as independent variables and taking overall survival time as dependent variable for model training. Subsequently, survival probability prediction and survival curve drawing can be performed through the target prognosis model, so as to avoid the problem of low prediction accuracy of renal clear cell tumor. Compared with the existing survival probability prediction based on clinical indicators, the survival probability prediction of renal clear cell tumor based on the target prognosis model established by the machine learning method in the embodiment greatly improves the reliability of prediction, thereby improving the accuracy of predicting the survival probability corresponding to the preset survival years of the renal clear cell tumor patient.

[0105] In S105, prediction processing is performed based on the input information through the target prognosis model, and a prediction result is output. The input information includes age, TNM stage, pathological stage and risk score to be predicted, and the prediction result includes survival probability and survival curve.

[0106] After the target prognosis model is determined, the prognosis of the renal clear cell tumor patient can be performed through the target prognosis model. The age, TNM stage, pathological stage and risk score to be predicted are taken as input information and input into the target prognosis model. Prediction processing is performed based on the age, TNM stage, pathological stage and risk score to be predicted through the target prognosis model, and a single variable prediction score is obtained. The age, TNM stage, pathological stage and risk score correspond to a single variable prediction score respectively. The single prediction scores are added to obtain a model prediction total score or a model linear prediction value, which can correspond to the survival probability within the preset year range. According to the obtained model prediction total score or model linear prediction value, conversion processing is performed through a preset conversion function to obtain the survival probability within the preset year range. For example, the survival probability of 1-year survival time, the survival probability of 2-year survival time, the survival probability of 3-year survival time and the survival probability of 5-year survival time are obtained. According to the survival probability of the preset years, the corresponding survival curve is drawn, and the survival curve takes time as the horizontal coordinate and survival probability as the vertical coordinate. Doctors and patients can know the current prognosis by checking the corresponding survival probability and survival curve, which provides a treatment reference for doctors and a prognosis reference for patients.

[0107] In order to improve the convenience of users to know the prognosis result, a corresponding open prediction platform can be provided for doctors or patients to predict the survival probability of renal clear cell tumor. Figure 2 is a prediction platform page display schematic diagram provided by an embodiment of the present application, which is referred to Figure 2In the display page of the prediction platform, an input information entry is provided, and a user can input information in the corresponding input information entry. The input information entry includes an age information input (age), a pathological stage input (stage), a TNM stage input (M), and a score input (A2M GPI). The input information entry can be set in a selection manner or in a fill-in manner, and the specific setting manner can be set according to actual conditions. After the user inputs the corresponding age, TNM stage, pathological stage, and risk score information based on the input information entry, the prediction platform performs prediction processing through the target prognosis model, outputs the corresponding survival probability and survival curve, and displays the prediction results in the page of the prediction platform (for example, the right page of FIG. 1). Figure 2 The prediction results displayed include the survival probability of 1-year survival time, the survival probability of 2-year survival time, the survival probability of 3-year survival time, the survival probability of 5-year survival time, and the survival curve. For example, the input information is that the age is 50 years old, the pathological stage is stage 1, the TNM stage is M0, and the risk score is 0.1. Through the prediction processing of the target prognosis model, the prediction results of the patient are obtained, the probability of 1-year survival time is 92%, the survival probability of 2-year survival time is 86%, the survival probability of 3-year survival time is 79%, the survival probability of 5-year survival time is 67%, and the survival probability corresponding to the 0-10-year survival time of the patient can be obtained according to the survival curve.

[0108] In an embodiment, after obtaining the target prognosis model, the target prognosis model can be tested and calibrated through a test set to obtain the corresponding calibrate curve, DCA curve, and ROC curve. The performance of the target prognosis model is evaluated through the calibrate curve, DCA curve, and ROC curve to improve the prediction reliability and accuracy of the target prognosis model.

[0109] The horizontal coordinate of the calibration curve is the survival probability predicted by the model, and the vertical coordinate is the survival probability actually observed. There are four lines in the calibration curve, each of which represents the comparison between the model-predicted 1-year, 2-year, 3-year, and 5-year survival and the actual situation, as well as the most ideal line (diagonal line). The closer the model-predicted line to the diagonal line, the better the fitting. The points on the calibration curve represent the model-predicted survival probability and the actually observed survival probability (similar to the probability corresponding to different scores in the lowermost part of the Nomogram model). The vertical line corresponding to the point on the calibration curve represents the confidence interval at that position. The calibration curve shows the results corrected by stratified Kaplan-Meier. The vertical line at the top of the calibration curve represents the survival probability of the specific sample (the distribution of the survival rate), and the denser it is, the more samples have a survival probability in this probability interval. The calibration curve of the target prognosis model provided in this embodiment shows that the fitting between the model-predicted 1-year, 2-year, 3-year, and 5-year overall survival probability of renal clear cell tumor patients and the actual survival rate is good, and is close to the ideal diagonal line. The vertical line at the top shows that most of the renal clear cell tumor patients in the original training data obtained from the public database have a survival probability between 0.6 and 0.8.

[0110] The DCA curve considers the clinical utility or patient benefit of the target prognosis model, finds a reasonable range / interval value as the threshold, reduces false positives and false negatives (population impact), increases true positives and true negatives (population impact), and observes and compares the clinical utility (net benefit) of the model. The x-axis of the DCA curve represents the probability threshold or threshold probability, and the y-axis represents the net benefit. The threshold probability refers to the risk of a certain value of the evaluation method (target prognosis model) acting on the cohort at a certain time point, which will result in the corresponding mortality of the entire cohort. By setting a threshold for mortality, the cohort is divided, and this threshold is the threshold probability at a certain x (threshold probability). The DCA curve corresponding to the target prognosis model provided in this embodiment shows that in the initial region of x, this line is almost coincident with all positive. After a certain x (threshold probability), this line begins to be higher than all positive. The difference between this line and the all negative line at a certain threshold probability interval is the net benefit. When comparing multiple models, observe which model is consistently higher than the other, and the higher part (y value difference) is the net benefit obtained at a certain x value. The DCA curve obtained by the present embodiment shows that in the decision curve analysis (DCA) of the third and fifth years of the patients, the DCA curve is always higher than other single clinical characteristic variables and A2M_GPI scores in the 0.2-0.5 interval, and the clinical patient benefit, specificity, and sensitivity are stronger than other variables.

[0111] The horizontal coordinate of the ROC curve represents 1-Specificity (FPR), i.e., "1-specificity", representing the false positive rate. The greater the FPR, the higher the predicted false positive rate. The vertical coordinate represents Sensitivity (TRP), i.e., sensitivity, representing the true positive rate. The ROC curve is a comprehensive indicator reflecting the continuous variables of sensitivity and specificity, and reflects the mutual relationship between sensitivity and specificity by mapping method. When the value of a variable (risk factor) is the trend of promoting the occurrence of an event, the AUC of the molecule is >0.5, and the larger the area (the closer the AUC value to 1), the better the prognostic performance. When the value of a variable (protective factor) is opposite to the trend of the occurrence of an event, then the AUC of this molecule is <0.5, and the smaller the area (the closer the AUC value to 0), the better the prognostic performance. The present embodiment analyzes the diagnostic values of 1, 2, 3 and 5 years of 7 closely related genes (i.e., target genes), A2M-GPI model (i.e., gene risk score model) and Nomogram model (i.e., target prognosis model) through ROC curve analysis. The area under the ROC curve (AUC) shows that the Nomogram model has higher diagnostic value and accuracy in predicting the prognosis of renal clear cell tumor patients (1-year AUC=0.130, 3-year AUC=0.186, 5-year AUC=0.199); the C-index of the model is 0.786, and the 95% CI is (0.768-0.804).

[0112] As described above, renal clear cell tumor is a common type of renal cancer in clinical practice, and its prognosis prediction is of great significance for treatment and patient survival rate evaluation. Compared with the traditional prognosis method relying only on clinical indicators (such as tumor size and grade), the present embodiment greatly improves the accuracy of prognosis prediction and individual treatment plan determination by means of correlation analysis and machine learning, combined with feature genes and clinical variables.

[0113] As described above, the potential prognostic related genes (i.e., target genes) are found by correlation analysis of the significant prognostic factor A2M gene, and the non-linear relationship between genes and clinical variables is mined by Lasso machine learning algorithm, so that the target prognosis model has higher accuracy and prediction ability. In addition, the performance of the target prognosis model is evaluated by using ROC curve, DCA curve and Calibration curve, and the target prognosis model based on machine learning is obviously superior to the traditional prediction method based on clinical indicators in terms of accuracy. The prediction result of the target prognosis model provided in the present embodiment is closer to the actual prognosis of the patient.

[0114] The traditional prognosis prediction method based on a single clinical variable is susceptible to individual differences, environment, susceptibility and other factors, and has poor stability in different data sets or sample sets. The target prognosis model constructed by the multivariate method based on machine learning in the embodiment can better capture the differences between individuals by considering multiple target genes and clinical variables. The multivariate model construction method provided in the embodiment can improve the stability of the target prognosis model, so that the prediction results have better consistency in different sample sets. In addition, the data verification results of the test set also show that the target prognosis model provided in the embodiment has good stability.

[0115] The traditional prognosis prediction method often only considers some basic characteristics of patients, while the machine learning model (i.e., the target prognosis model) provided in the embodiment combines genes (A2M gene) and clinical variables to more accurately develop individualized treatment plans. Through the target prognosis model provided in the embodiment, specific prognosis evaluation and corresponding treatment recommendations can be provided for each renal clear cell tumor patient, thereby improving the survival rate and quality of life of the patient.

[0116] The feature selection algorithm and multivariate Cox regression analysis are used to select target genes and clinical variables related to prognosis in the embodiment, so that the prediction results of the model are more interpretable. Doctors can understand the impact of important feature coefficients on prognosis according to the model output, so as to better understand and explain the prediction results and provide reasonable treatment recommendations for patients.

[0117] Based on the renal clear cell tumor prediction method provided in the embodiment, an open platform (website) for users to use is established, and the target prognosis model can be applied in real time through an online platform or a mobile application to provide guidance and decision support for clinicians. Doctors can input the gene and clinical variable data of patients into the target prognosis model for prediction, thereby quickly obtaining prognosis evaluation results and making corresponding treatment decisions based on the same. This real-time application method greatly improves the efficiency and accuracy of clinical work.

[0118] The above, by acquiring target sample data when making a prognosis of renal clear cell tumor, performing correlation analysis of A2M gene in renal clear cell tumor according to the target sample data, screening variables, determining a target gene group significantly related to the A2M gene, constructing a gene risk score model according to the target gene group, and calculating the risk score corresponding to each clear cell tumor patient in the target sample data according to the gene risk score model, performing variable regression analysis training according to the target sample data and the risk score, and obtaining a target prognosis model, the survival probability and the survival curve are output by performing prediction processing on the target prognosis model based on the age, TNM stage, pathological stage and risk score to be predicted. By using the above technical means, the target prognosis model can be obtained by analyzing and training the target sample data and the target gene group significantly related to the A2M gene, and the survival probability is predicted and the survival curve is drawn by using the target prognosis model, so as to avoid the problem that the prediction accuracy of renal clear cell tumor is low. Compared with the existing survival probability prediction method based on clinical indicators, the survival probability of renal clear cell tumor is predicted based on the target prognosis model, which greatly improves the reliability of the prediction, thereby improving the accuracy of predicting the survival probability corresponding to the preset survival time of the renal clear cell tumor patient.

[0119] On the basis of the above embodiment, Figure 3 The structure diagram of a renal clear cell tumor prediction device provided by the embodiment of the present application is shown in the figure. Figure 3 The renal clear cell tumor prediction device provided by the embodiment of the present application specifically comprises: a sample acquisition module 21, a gene correlation analysis module 22, a gene risk score module 23, a prognosis model training module 24 and a prediction module 25.

[0120] The sample acquisition module 21 is used to acquire target sample data, and the target sample data includes genetic material test data and clinical information data of renal clear cell tumor patients. The clinical information includes overall survival, disease-specific survival, progression-free survival, age, TNM stage, pathological stage, histological grade, survival status and gender.

[0121] The gene correlation analysis module 22 is used to perform correlation analysis of A2M gene in renal clear cell tumor according to the target sample data, screen variables, determine a target gene group, and the target gene group is a plurality of genes related to the A2M gene.

[0122] The gene risk score module 23 is used to construct a gene risk score model according to the target gene group, and calculate the risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model.

[0123] The prognosis model training module 24 is configured to perform variable regression analysis training according to the target sample data and the risk score, to obtain a target prognosis model, wherein the independent variable data in the variable regression analysis training is the age, TNM stage, pathological stage and risk score in the target sample data, and the dependent variable data is the overall survival time;

[0124] The prediction module 25 is configured to perform prediction processing based on input information by using the target prognosis model, and output a prediction result, wherein the input information includes the age, TNM stage, pathological stage and risk score to be predicted, and the prediction result includes the survival probability and the survival curve.

[0125] Further, the gene correlation analysis module 22 includes an A2M single variable correlation analysis submodule, a first single variable COX regression submodule and a linear fitting submodule.

[0126] The A2M single gene correlation analysis submodule is configured to perform survival state correlation analysis of the A2M gene in renal clear cell tumor based on the target sample data by single gene correlation analysis, to obtain A2M related genes.

[0127] The single variable COX regression submodule is configured to perform survival prognosis analysis of the A2M related genes in renal clear cell tumor based on the target sample data by single gene correlation analysis, to determine a first candidate gene meeting a preset threshold requirement from the A2M related genes.

[0128] The linear fitting submodule is configured to perform least absolute shrinkage and selection operator (LASSO) regression of the overall survival time of renal clear cell carcinoma based on the target sample data by an L1 regularization model, to determine a target gene group meeting an error value condition from the first candidate gene.

[0129] Further, the target gene group includes TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3.

[0130] The gene risk score module 23 includes a score model construction submodule and a score calculation submodule.

[0131] The score model construction submodule is configured to construct a gene risk score model according to the gene expression amounts of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 and corresponding preset risk coefficients.

[0132] The score calculation submodule is configured to calculate the corresponding risk score of each renal clear cell tumor patient in the target sample data according to the gene risk score model.

[0133] Further, the gene risk score model is characterized by the following manner:

[0134]

[0135] wherein A2M GPI represents a risk score, represents a preset risk coefficient of a gene, represents a gene expression of a gene.

[0136] Further, the gene risk score model is characterized by the following:

[0137] A2M GPI=(-0.06057755*TIE1exp.)+(0.00416184*VWFexp.)+(-0.22620967*TCF4exp.)+(0.62309290*PTPRBexp.)+(-0.06403383*ICAM2exp.)+(-0.19163895*DOCK6exp.)+(0.08291625*RAMP3exp.)

[0138] wherein (-0.06057755), 0.00416184, (-0.22620967), 0.62309290, (-0.06403383), (-0.19163895) and 0.08291625 represent preset risk coefficients of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 respectively, (TIE1exp.), (VWFexp.), (TCF4exp.), (PTPRBexp.), (ICAM2exp.), (DOCK6exp.) and (RAMP3exp.) represent gene expressions of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 respectively.

[0139] Further, the prognosis model training module 24 comprises a second univariate COX regression submodule and a multivariate analysis submodule;

[0140] The second univariate COX regression submodule is configured to perform single factor analysis processing based on the target sample data by a univariate COX regression model, to determine target variable factors, the target variable factors being age, TNM stage and pathological stage in the target sample data;

[0141] The multivariate analysis submodule is configured to perform variable regression analysis training by a multivariate COX regression model, with the target variable factors and the risk score as independent variable data, and with the overall survival time in the target sample data as a dependent variable, to obtain the target prognosis model.

[0142] Further, the prediction module 25 comprises a single-item prediction sub-module, a comprehensive prediction sub-module, a survival probability conversion sub-module and a drawing sub-module;

[0143] The single-item prediction sub-module is configured to perform prediction processing based on the age, TNM stage, pathological stage and risk score to be predicted by using the target prognosis model to obtain a single-item variable prediction score, wherein the age, TNM stage, pathological stage and risk score correspond to one single-item variable prediction score respectively.

[0144] The comprehensive prediction sub-module is configured to add the single-item variable prediction scores to obtain a model prediction total score.

[0145] The survival probability conversion sub-module is configured to perform conversion processing on the model prediction total score and a preset conversion function to obtain a survival probability within a preset time limit.

[0146] The drawing sub-module is configured to draw a corresponding survival curve according to the survival probability within the preset time limit, wherein the survival curve takes time as the horizontal coordinate and survival probability as the vertical coordinate.

[0147] In the above, when making a prognosis of the renal clear cell tumor, target sample data is obtained, correlation analysis of the A2M gene in the renal clear cell tumor is performed according to the target sample data, variables are screened, a target gene group significantly related to the A2M gene is determined, a gene risk score model is constructed according to the target gene group, and a risk score corresponding to each clear cell tumor patient in the target sample data is calculated according to the gene risk score model. The target prognosis model is obtained by performing variable regression analysis training according to the target sample data and the risk score. The survival probability and the survival curve are output by performing prediction processing based on the age, TNM stage, pathological stage and risk score to be predicted by using the target prognosis model. By using the above technical means, the target prognosis model can be obtained by analyzing and training the target sample data and the target gene group significantly related to the A2M gene. The survival probability is predicted by using the target prognosis model, and the survival curve is drawn. In this way, the problem of low prediction accuracy of the renal clear cell tumor can be avoided. Compared with the existing method of predicting the survival probability based on clinical indicators, the survival probability of the renal clear cell tumor is predicted based on the target prognosis model in the embodiment, which greatly improves the reliability of the prediction, thereby improving the accuracy of predicting the survival probability corresponding to the preset survival time of the renal clear cell tumor patient.

[0148] The renal clear cell tumor prediction device provided by the embodiment of the present application can be used to execute the renal clear cell tumor prediction method provided by the above-mentioned embodiment, and has corresponding functions and beneficial effects.

[0149] The embodiment of the present application provides a renal clear cell tumor prediction device, which is referred to Figure 4The renal clear cell tumor prediction device includes a processor 31, a memory 32, a communication module 33, an input device 34, and an output device 35. The number of processors in the renal clear cell tumor prediction device can be one or more, and the number of memories in the renal clear cell tumor prediction device can be one or more. The processor, memory, communication module, input device, and output device of the renal clear cell tumor prediction device can be connected by a bus or other means.

[0150] The memory 32, as a computer readable storage medium, can be used to store software programs, computer executable programs, and modules, such as program instructions / modules corresponding to the renal clear cell tumor prediction method described in any embodiment of the present application (for example, the sample acquisition module, the gene correlation analysis module, the gene risk score module, the prognosis model training module, and the prediction module in the renal clear cell tumor prediction device). The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required by a function; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0151] The communication module 33 is used for data transmission.

[0152] The processor 31 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory, that is, implements the renal clear cell tumor prediction method described above.

[0153] The input device 34 can be used to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device 35 can include a display device such as a display screen.

[0154] The renal clear cell tumor prediction device provided above can be used to execute the renal clear cell tumor prediction method provided in the above embodiments, and has corresponding functions and beneficial effects.

[0155] The embodiment of the present application further provides a storage medium storing computer executable instructions, which, when executed by a computer processor, are used to perform a renal clear cell tumor prediction method, the renal clear cell tumor prediction method comprising: obtaining target sample data, the target sample data comprising genetic material test data and clinical information data of a renal clear cell tumor patient, the clinical information comprising overall survival, disease-specific survival, progression-free survival, age, TNM stage, pathological stage, histological grade, survival status and gender; performing correlation analysis of A2M gene in the renal clear cell tumor according to the target sample data, determining screening variables, and determining a target gene group, the target gene group being a plurality of genes related to A2M; constructing a gene risk score model according to the target gene, and calculating a risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model; performing variable regression analysis training according to the target sample data and the risk score, to obtain a target prognosis model, wherein the independent variable data in the variable regression analysis training is the age, TNM stage, pathological stage and risk score in the target sample data, and the dependent variable data is the overall survival; performing prediction processing based on input information through the target prognosis model, and outputting a prediction result, the input information comprising an age, TNM stage, pathological stage and risk score to be predicted, and the prediction result comprising a survival probability and a survival curve.

[0156] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include an installation medium, e.g., a CD-ROM, floppy disks, or tape device; a computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; a non-volatile memory such as a magnetic medium (e.g., a hard drive or optical storage); registers or other similar types of memory elements upon which a computer processes data. The memory can also include other types of storage medium and combinations thereof. Moreover, the memory can be located in a first computer system while the processing operations are implemented on a second computer system, across a network such as the Internet. The second computer system can provide the program instructions to the first computer system for execution. The term "storage medium" can include two or more memory devices, which reside in different locations, e.g., in different computer systems that are connected over a network such as the Internet. The memory can store program instructions (e.g., a computer program) that can be executed by one or more processors.

[0157] Of course, the storage medium storing computer executable instructions provided by the embodiment of the present application is not limited to the renal clear cell tumor prediction method as described above, and can also perform the related operations in the renal clear cell tumor prediction method provided by any embodiment of the present application.

[0158] The renal clear cell tumor prediction apparatus, the storage medium and the renal clear cell tumor prediction device provided in the above embodiments can perform the renal clear cell tumor prediction method provided in any of the embodiments of the present application. Technical details not described in detail in the above embodiments can be referred to the renal clear cell tumor prediction method provided in any of the embodiments of the present application.

[0159] The above are only preferred embodiments of the present application and the technical principles applied. The present application is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments and replacements that can be made by those skilled in the art will not deviate from the protection scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments. More other equivalent embodiments can be included without deviating from the concept of the present application, and the scope of the present application is determined by the scope of the claims.

Claims

1. A method of predicting renal clear cell tumor, characterized by, The method comprises the following steps: obtaining target sample data, wherein the target sample data comprises genetic material test data and clinical information data of a renal clear cell tumor patient, and the clinical information comprises overall survival, disease-specific survival, progression-free survival, age, TNM stage, pathological stage, histological grade, survival status and gender; performing correlation analysis of A2M gene in renal clear cell tumor according to the target sample data, screening variables, and determining a target gene group, wherein the target gene group is a plurality of genes related to A2M; constructing a gene risk score model according to the target gene group, and calculating a risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model; performing variable regression analysis training according to the target sample data and the risk score to obtain a target prognosis model, wherein the independent variable data in the variable regression analysis training is age, TNM stage, pathological stage and risk score in the target sample data, and the dependent variable data is overall survival; performing prediction processing based on input information by using the target prognosis model to output a prediction result, wherein the input information comprises age, TNM stage, pathological stage and risk score to be predicted, and the prediction result comprises survival probability and survival curve; wherein the correlation analysis of A2M gene in renal clear cell tumor according to the target sample data, the screening of variables and the determination of the target gene group comprise: performing correlation analysis of A2M gene in renal clear cell tumor based on the target sample data by single gene correlation analysis to obtain A2M related genes; performing A2M related gene and renal clear cell tumor survival prognosis analysis based on the target sample data by single variable COX regression analysis to determine first candidate genes meeting a preset threshold requirement from the A2M related genes; performing least absolute shrinkage and selection operator (LASSO) regression of overall survival of renal clear cell tumor based on the target sample data by L1 regularization model to determine a target gene group meeting an error value condition from the first candidate genes.

2. The method of claim 1, wherein, The target gene group comprises TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3. The construction of the gene risk score model according to the target gene group and the calculation of the risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model comprise: constructing a gene risk score model according to the gene expression amount of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 and a corresponding preset risk coefficient; calculating the risk score corresponding to each renal clear cell tumor patient in the target sample data according to the gene risk score model.

3. The method of claim 2, wherein, The gene risk score model is characterized in the following manner: wherein A2M GPI represents the risk score of the gene A2M, represents a preset risk coefficient of the gene, represents the gene expression amount of the gene.

4. The method of claim 3, wherein, The gene risk score model is characterized in the following manner: A2M GPI = (-0.06057755 * TIE1 exp.) + (0.00416184 * VWF exp.) + (-0.22620967 * TCF4 exp.) + (0.62309290 * PTPRB exp.) + (-0.06403383 * ICAM2 exp.) + (-0.19163895 * DOCK6 exp.) + (0.08291625 * RAMP3 exp.) Wherein, -0.06057755, 0.00416184, -0.22620967, 0.62309290, -0.06403383, -0.19163895 and 0.08291625 represent the preset risk coefficients of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 respectively, and TIE1 exp., VWF exp., TCF4 exp., PTPRB exp., ICAM2 exp., DOCK6 exp. and RAMP3 exp. represent the gene expression amounts of TIE1, VWF, TCF4, PTPRB, ICAM2, DOCK6 and RAMP3 respectively.

5. The method of claim 1, wherein, The variable regression analysis training according to the target sample data and the risk score obtains a target prognosis model, comprising: Single factor analysis processing is performed on the clinical information data in the target sample data based on a single variable COX regression model to determine a target variable factor, and the target variable factor is age, TNM stage and pathological stage in the target sample data; Variable regression analysis training is performed on the total survival time in the target sample data by taking the target variable factor and the risk score as independent variable data based on a multivariate COX regression model to obtain a target prognosis model.

6. The method of claim 1, wherein, The target prognosis model is used for prediction processing based on input information to output a prediction result, comprising: The target prognosis model is used for prediction processing based on the age, TNM stage, pathological stage and risk score to be predicted to obtain a single variable prediction score, wherein the age, TNM stage, pathological stage and risk score correspond to a single variable prediction score respectively; The single variable prediction scores are added to obtain a model prediction total score; The model prediction total score is converted according to a preset conversion function to obtain a survival probability within a preset time limit; A survival curve corresponding to the survival probability within the preset time limit is drawn, and the survival curve takes time as the horizontal coordinate and the survival probability as the vertical coordinate.

7. A renal clear cell tumor prediction device, characterized by, Comprising: A sample acquisition module is configured to acquire target sample data, wherein the target sample data comprises genetic material test data and clinical information data of a renal clear cell tumor patient, and the clinical information comprises total survival time, disease-specific survival time, progression-free survival time, age, TNM stage, pathological stage, histological grade, survival status and gender; a gene correlation analysis module configured to perform correlation analysis of A2M gene in renal clear cell carcinoma based on the target sample data, screen variables, and determine a target gene group, which is a plurality of genes related to A2M; a gene risk score module configured to construct a gene risk score model based on the target gene group, and calculate a risk score corresponding to each renal clear cell carcinoma patient in the target sample data based on the gene risk score model; a prognosis model training module configured to perform variable regression analysis training based on the target sample data and the risk score, and obtain a target prognosis model, wherein the independent variable data in the variable regression analysis training is age, TNM stage, pathological stage, and risk score in the target sample data, and the dependent variable data is overall survival time; a prediction module configured to perform prediction processing based on input information by using the target prognosis model, and output a prediction result, wherein the input information includes age, TNM stage, pathological stage, and risk score to be predicted, and the prediction result includes survival probability and survival curve; the gene correlation analysis module includes an A2M univariate correlation analysis submodule, a first univariate COX regression submodule, and a linear fitting submodule; the A2M univariate correlation analysis submodule is configured to perform survival state correlation analysis of A2M gene in renal clear cell carcinoma based on target sample data by single-gene correlation analysis, and obtain A2M-related genes; the univariate COX regression submodule is configured to perform survival prognosis analysis of A2M-related genes in renal clear cell carcinoma based on target sample data by single-gene correlation analysis, and determine first candidate genes meeting a preset threshold requirement from the A2M-related genes; the linear fitting submodule is configured to perform least absolute shrinkage and selection operator (LASSO) regression of overall survival time of renal clear cell carcinoma based on target sample data by an L1 regularization model, and determine a target gene group meeting an error value condition from the first candidate genes.

8. A renal clear cell tumor prediction device, comprising: comprise: a memory and one or more processors; the memory is configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method of any one of claims 1-6.

9. A storage medium storing computer-executable instructions, wherein: the computer executable instructions, when executed by the processor, are configured to perform the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Immune gene prognosis model for predicting hepatocellular carcinoma tumor immune infiltration and postoperative survival time

    CN112011616A

  • Evaluation model for predicting prognosis and adjuvant chemotherapy benefit of colon cancer and application

    CN115011689A