A method for evaluating the malignancy degree of a tumor, an electronic device and a storage medium

By constructing a comprehensive hazard ratio index model and using transcriptome data to screen differentially expressed genes, the complexity and cost issues of tumor malignancy assessment in existing technologies have been resolved, enabling quantitative analysis of tumor malignancy and accurate assessment of early-stage patient survival.

CN116312789BActive Publication Date: 2026-05-08OMIXSCIENCE (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
OMIXSCIENCE (SHENZHEN) CO LTD
Filing Date
2023-03-31
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing tumor malignancy assessment techniques are complex, expensive, and time-consuming, and cannot quantitatively analyze tumor malignancy or assess the survival of patients with early-stage tumors.

Method used

By constructing a comprehensive hazard ratio index model, using transcriptomic data of the tumor under test, differentially expressed genes are screened, survival analysis is performed, prognosis-related genes are screened, and a comprehensive hazard ratio index model is constructed to assess the malignancy of the tumor.

Benefits of technology

It enables simple, low-cost, and rapid assessment of tumor malignancy, quantitatively analyzes tumor malignancy, and accurately assesses the survival of cancer patients at various stages, exhibiting greater robustness and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312789B_ABST
    Figure CN116312789B_ABST
Patent Text Reader

Abstract

The application discloses a tumor malignancy evaluation method, an electronic device and a storage medium. The tumor malignancy evaluation method comprises the following steps: inputting a tumor sample to be measured into a comprehensive hazard ratio index model; outputting a comprehensive hazard ratio index from the comprehensive hazard ratio index model according to the tumor sample to be measured; and judging the malignancy degree of the tumor sample to be measured according to the comprehensive hazard ratio index. The malignancy degree of the tumor to be measured can be evaluated only by obtaining the transcriptome data of the tumor of a patient, and the method is simple, low in cost, short in time consumption, and capable of realizing quantitative analysis of the tumor malignancy degree and accurately evaluating the survival condition of tumor patients at different periods according to the comprehensive hazard ratio index obtained from the molecular level information of the tumor to be measured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of tumor malignancy assessment technology, and in particular to a method, electronic device and storage medium for assessing tumor malignancy. Background Technology

[0002] For cancer patients, precision medicine is currently recognized as the best treatment approach. In precision medicine, the malignancy of a tumor is highly valuable in guiding the selection of treatment plans and the assessment of patient prognosis. Currently, tumor malignancy is primarily assessed through tumor staging based on the size of the primary tumor and its extent of spread within the body. Tumor staging helps clinicians develop appropriate treatment plans, predict patient prognosis, and evaluate the effectiveness of treatments. Tumor staging is currently derived from various examinations, including physical examinations, imaging studies (X-rays, CT scans, etc.), and laboratory tests (such as complete blood counts, urinalysis, etc.). Patient staging is primarily based on the TNM system, which evaluates tumors, lymph nodes, and metastasis across three dimensions. T-staging describes the progression of the primary tumor. Based on tumor size, depth of growth, and whether it has spread to adjacent tissues, it is divided into T0, T1, T2, T3, and T4. A higher stage value indicates a larger tumor and deeper growth location. T0 represents a lack of evidence of a primary tumor. N-staging assesses the extent of lymph node spread, and is divided into N0, N1, N2, N3, and N4. A higher stage value indicates more lymph nodes affected by the tumor. N0 indicates no lymph node spread. M-staging is divided into M0 and M1. M0 indicates no metastasis, while M1 indicates metastasis. Combining TNM staging values ​​yields an overall stage (Stage I, II, III, IV). Generally, a lower overall stage value indicates an earlier stage of the tumor and a better prognosis; a higher overall stage value indicates a later stage of the tumor, requiring more complex treatment and resulting in a worse prognosis.

[0003] Assessing the malignancy of a tumor based on its stage can provide some guidance for treatment plans and prognosis in cancer patients. However, this method has three main problems: First, tumor staging requires multiple examinations and the acquisition of tumor tissue for examination, which is complex, expensive, and time-consuming. Second, current tumor staging methods rely on clinicians' experience to determine TNM staging based on examination results, without considering any molecular-level tumor information, making it more qualitative than quantitative and limiting its ability to guide precision medicine. Third, for early-stage cancer patients, it is difficult to determine their survival based solely on tumor stage. Summary of the Invention

[0004] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a method, electronic device, and storage medium for assessing tumor malignancy, which can solve the problems of existing tumor malignancy assessment technologies, such as complex examinations, high costs, long processing times, inability to quantitatively analyze tumor malignancy, and difficulty in assessing the survival of patients with early-stage tumors.

[0005] A method for assessing tumor malignancy according to a first aspect of this application includes: inputting a tumor sample to be tested into a comprehensive hazard ratio index model, wherein the tumor sample to be tested includes transcriptomic data of the tumor to be tested; the comprehensive hazard ratio index model outputs a comprehensive hazard ratio index based on the tumor sample to be tested; and determining the malignancy of the tumor sample to be tested based on the comprehensive hazard ratio index; the comprehensive hazard ratio index model is constructed by: acquiring multiple tumor sample data and adjacent normal sample data corresponding to the tumor sample data, wherein the tumor sample data includes first transcriptomic data and patient prognostic information, and the first transcriptomic data is... Transcriptome data of tumor samples, wherein the adjacent normal sample data includes second transcriptome data, which is transcriptome data of adjacent normal samples; differentially expressed genes are screened based on the first and second transcriptome data; the tumor sample data is divided into a training set and a test set; survival analysis is performed on the training set based on the differentially expressed genes, and prognosis-related genes are screened from the training set, wherein the prognosis-related genes are the differentially expressed genes related to the patient's prognosis; a comprehensive hazard ratio index model is constructed based on the prognosis-related genes; the comprehensive hazard ratio index model is validated using the test set.

[0006] A tumor malignancy assessment method according to an embodiment of the first aspect of this application has at least the following beneficial effects:

[0007] By acquiring multiple tumor sample data and corresponding adjacent normal sample data, the tumor sample data includes first transcriptome data and patient prognostic information, while the adjacent normal sample data includes second transcriptome data. Differentially expressed genes are screened based on the first and second transcriptome data. The tumor sample data is divided into training and test sets. Survival analysis is performed on the training set based on the differentially expressed genes. Prognostic-related genes are screened from the training set, and a comprehensive hazard ratio index model is constructed based on these genes. The comprehensive hazard ratio index model is validated using the test set. The tumor sample to be tested is input into the comprehensive hazard ratio index model, which outputs a comprehensive hazard ratio index based on the tumor sample. The malignancy of the tumor sample is determined based on the comprehensive hazard ratio index. Compared to traditional tumor malignancy assessment techniques, this method for assessing tumor malignancy according to the first aspect of this application only requires obtaining the transcriptome data of the patient's tumor to assess its malignancy. It is simple, inexpensive, and time-efficient. The comprehensive hazard ratio index model can obtain a comprehensive hazard ratio index based on the molecular-level information of the tumor, enabling quantitative analysis of tumor malignancy and accurate assessment of the survival status of cancer patients at various stages.

[0008] According to some embodiments of this application, the step of performing survival analysis on the training set based on the differentially expressed genes and screening out prognosis-related genes from the training set includes: obtaining the expression level of the differentially expressed genes in the training set; dividing the training set into a high expression level group and a low expression level group based on the expression level of the differentially expressed genes in the training set; performing survival analysis on the high expression level group and the low expression level group using the Kaplan-Meier method to obtain the P-value and HR value corresponding to the differentially expressed genes; and screening out the prognosis-related genes from the training set based on the P-value.

[0009] According to some embodiments of this application, the types of prognosis-related genes include suspected proto-oncogenes, suspected upregulated protective genes, suspected tumor suppressor genes, and suspected downregulated protective genes. The suspected proto-oncogenes are differentially expressed genes that are upregulated in the tumor sample data and lead to a worse prognosis. The suspected upregulated protective genes are differentially expressed genes that are upregulated in the tumor sample data and lead to a better prognosis. The suspected tumor suppressor genes are differentially expressed genes that are downregulated in the tumor sample data and lead to a worse prognosis. The suspected downregulated protective genes are differentially expressed genes that are downregulated in the tumor sample data and lead to a better prognosis.

[0010] According to some embodiments of this application, the step of screening the prognosis-related genes from the training set based on the P-value includes: screening candidate genes from the training set based on the P-value, wherein the candidate genes are differentially expressed genes related to the patient's prognosis; sorting the candidate genes of each type from smallest to largest according to the P-value to obtain significant sequences; and selecting the top forty candidate genes from the significant sequences of each type as the prognosis-related genes.

[0011] According to some embodiments of this application, the formula for calculating the comprehensive hazard ratio index is as follows:

[0012]

[0013] in, To form a comprehensive risk ratio index, For the first The HR values ​​corresponding to the prognosis-related genes. For the first The calculation coefficients corresponding to the prognosis-related genes are obtained by: obtaining the expression level of each prognosis-related gene in all the tumor sample data; determining a threshold based on the expression level of each prognosis-related gene in all the tumor sample data; obtaining the expression level of the prognosis-related gene in the tumor sample to be tested; determining the calculation coefficient based on the threshold, the type of the prognosis-related gene, and the expression level of the prognosis-related gene in the tumor sample to be tested; if the prognosis-related gene is the suspected proto-oncogene or the suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is higher than the threshold. If the prognosis-related gene is the suspected proto-oncogene or the suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is lower than the threshold, then the calculation coefficient is -1; if the prognosis-related gene is the suspected upregulated protective gene or the suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is higher than the threshold, then the calculation coefficient is -1; if the prognosis-related gene is the suspected upregulated protective gene or the suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is lower than the threshold, then the calculation coefficient is 1.

[0014] According to some embodiments of this application, obtaining multiple tumor sample data and adjacent normal sample data corresponding to the tumor sample data includes: obtaining multiple tumor sample data including multiple cancer types and adjacent normal sample data corresponding to the tumor sample data.

[0015] According to some embodiments of this application, the step of screening differentially expressed genes based on the first transcriptome data and the second transcriptome data includes: screening the differentially expressed genes from the first transcriptome data using DESeq2 software based on the first transcriptome data and the second transcriptome data.

[0016] An electronic device according to a second aspect of this application includes: at least one processor; at least one memory for storing at least one program; and when at least one of the programs is executed by at least one of the processors, it implements a tumor malignancy assessment method as described above.

[0017] An electronic device according to a second aspect embodiment of this application has at least the following beneficial effects:

[0018] By acquiring multiple tumor sample data and corresponding adjacent normal sample data, the tumor sample data includes first transcriptome data and patient prognostic information, while the adjacent normal sample data includes second transcriptome data. Differentially expressed genes are screened based on the first and second transcriptome data. The tumor sample data is divided into a training set and a test set. Survival analysis is performed on the training set based on the differentially expressed genes. Prognostic-related genes are screened from the training set, and a comprehensive hazard ratio index model is constructed based on these genes. The comprehensive hazard ratio index model is validated using the test set. The tumor sample to be tested is input into the comprehensive hazard ratio index model, which outputs a comprehensive hazard ratio index based on the tumor sample. The malignancy of the tumor sample is determined based on the comprehensive hazard ratio index. According to an embodiment of the second aspect of this application, an electronic device only needs to obtain the transcriptome data of the patient's tumor to be tested to assess the malignancy of the tumor. This method is simple, inexpensive, and time-saving. The comprehensive hazard ratio index model can obtain a comprehensive hazard ratio index based on the molecular-level information of the tumor, achieving quantitative analysis of tumor malignancy and accurately assessing the survival of cancer patients at various stages.

[0019] A computer-readable storage medium according to a third aspect of this application stores a processor-executable program, which, when executed by a processor, is used to implement a tumor malignancy assessment method as described above.

[0020] A computer-readable storage medium according to an embodiment of a third aspect of this application has at least the following advantages:

[0021] By acquiring multiple tumor sample data and corresponding adjacent normal sample data, the tumor sample data includes first transcriptome data and patient prognostic information, while the adjacent normal sample data includes second transcriptome data. Differentially expressed genes are screened based on the first and second transcriptome data. The tumor sample data is divided into training and test sets. Survival analysis is performed on the training set based on the differentially expressed genes. Prognostic-related genes are screened from the training set, and a comprehensive hazard ratio index model is constructed based on these genes. The comprehensive hazard ratio index model is validated using the test set. The tumor sample to be tested is input into the comprehensive hazard ratio index model, which outputs a comprehensive hazard ratio index based on the tumor sample. The malignancy of the tumor sample is determined based on the comprehensive hazard ratio index. According to a computer-readable storage medium according to an embodiment of the third aspect of this application, only the transcriptome data of the patient's tumor to be tested is needed to assess the malignancy of the tumor. This method is simple, inexpensive, and time-saving. The comprehensive hazard ratio index model can obtain a comprehensive hazard ratio index based on the molecular-level information of the tumor, achieving quantitative analysis of tumor malignancy and accurately assessing the survival of cancer patients at various stages.

[0022] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0023] The present application will be further described below with reference to the accompanying drawings and embodiments, wherein:

[0024] Figure 1 This is a flowchart of a tumor malignancy assessment method according to one embodiment of this application;

[0025] Figure 2 A flowchart of a method for constructing a comprehensive hazard ratio index model in one embodiment of this application;

[0026] Figure 3 This is a flowchart illustrating the screening of prognosis-related genes in one embodiment of this application;

[0027] Figure 4 This is a flowchart illustrating the selection of prognosis-related genes from candidate genes in one embodiment of this application;

[0028] Figure 5 This is a flowchart of a method for obtaining calculation coefficients in one embodiment of this application;

[0029] Figure 6 This is a schematic diagram of an electronic device according to an embodiment of this application;

[0030] Figure 7 This is a graph showing the correlation between SAHR values ​​and survival rates for colon cancer in the training set in one embodiment of this application.

[0031] Figure 8 This is a graph showing the correlation between SAHR values ​​and survival rates for colon cancer in a test set according to one embodiment of this application.

[0032] Figure 9 This is a graph showing the correlation between SAHR values ​​and survival rates in the training set for clear cell renal cell carcinoma in one embodiment of this application.

[0033] Figure 10 This is a graph showing the correlation between SAHR values ​​and survival rates in a test set for clear cell renal cell carcinoma in one embodiment of this application.

[0034] Figure 11 This is a graph showing the correlation between SAHR values ​​and survival rates in the training set for renal papillary cell carcinoma in one embodiment of this application.

[0035] Figure 12 This is a graph showing the correlation between SAHR values ​​and survival rates in a test set for renal papillary cell carcinoma in one embodiment of this application.

[0036] Figure 13 This is a graph showing the correlation between SAHR values ​​and survival rates of lung adenocarcinoma in the training set in one embodiment of this application.

[0037] Figure 14 This is a graph showing the correlation between SAHR values ​​and survival rates for lung adenocarcinoma in a test set according to one embodiment of this application.

[0038] Figure 15 This is a correlation graph showing the SAHR value and survival rate of squamous cell carcinoma of the lung in the training set in one embodiment of this application.

[0039] Figure 16 This is a graph showing the correlation between SAHR values ​​and survival rates in a test set for squamous cell carcinoma of the lung in one embodiment of this application.

[0040] Figure 17 This is a correlation diagram of SAHR value and survival rate of rectal adenocarcinoma in the training set in one embodiment of this application;

[0041] Figure 18 This is a graph showing the correlation between SAHR values ​​and survival rates for rectal adenocarcinoma in a test set according to one embodiment of this application.

[0042] Figure 19 This is a correlation diagram of SAHR value and survival rate of gastric cancer in the training set in one embodiment of this application;

[0043] Figure 20 This is a graph showing the correlation between SAHR values ​​and survival rates for gastric cancer in a test set according to one embodiment of this application.

[0044] Figure 21This is a graph showing the correlation between SAHR values ​​and survival rates in the training set for invasive ductal carcinoma of the breast in one embodiment of this application.

[0045] Figure 22 This is a graph showing the correlation between SAHR values ​​and survival rates in the test set for invasive ductal carcinoma of the breast in one embodiment of this application.

[0046] Figure 23 This is a correlation diagram of SAHR value and survival rate for laryngeal cancer in the training set in one embodiment of this application;

[0047] Figure 24 This is a graph showing the correlation between SAHR values ​​and survival rates for laryngeal cancer in a test set, according to one embodiment of this application.

[0048] Figure 25 This is a correlation diagram of SAHR value and survival rate of tongue cancer in the training set in one embodiment of this application;

[0049] Figure 26 This is a graph showing the correlation between SAHR values ​​and survival rates for tongue cancer in a test set according to one embodiment of this application.

[0050] Figure 27 This is a graph showing the correlation between SAHR values ​​and survival rates in the training set for bladder urothelial carcinoma in one embodiment of this application.

[0051] Figure 28 This is a graph showing the correlation between SAHR values ​​and survival rates in a test set for bladder urothelial carcinoma in one embodiment of this application.

[0052] Figure 29 This is a graph showing the correlation between SAHR values ​​and survival rates for endometrioid adenocarcinoma in the training set in one embodiment of this application.

[0053] Figure 30 This is a graph showing the correlation between SAHR values ​​and survival rates for endometrioid adenocarcinoma in a test set according to one embodiment of this application.

[0054] Figure 31 This is a graph showing the correlation between SAHR values ​​and survival rates for thyroid cancer in the training set in one embodiment of this application.

[0055] Figure 32 This is a graph showing the correlation between SAHR values ​​and survival rates in a test set for thyroid cancer in one embodiment of this application.

[0056] Figure 33 This is a graph showing the correlation between SAHR values ​​and survival rates in the training set for multifocal gliomas in one embodiment of this application.

[0057] Figure 34This is a graph showing the correlation between SAHR values ​​and survival rates in a test set for multifocal gliomas in one embodiment of this application.

[0058] Figure 35 Figure a shows a prediction of early-stage colon cancer survival in one embodiment of this application;

[0059] Figure 36 Figure b shows a prediction of early-stage colon cancer survival in one embodiment of this application.

[0060] Figure 37 Figure a shows a prediction of clinically advanced colon cancer survival in one embodiment of this application;

[0061] Figure 38 Figure b shows a prediction of survival rate for clinically advanced colorectal cancer in one embodiment of this application.

[0062] Figure label:

[0063] Electronic device 100, processor 110, memory 120. Detailed Implementation

[0064] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0065] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0066] In the description of this application, "multiple" refers to two or more. The use of "first" and "second" is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance, or implicitly indicating the number of technical features indicated, or the order in which the technical features are indicated.

[0067] In the description of this application, unless otherwise expressly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.

[0068] like Figure 1 As shown, Figure 1This is a flowchart of a tumor malignancy assessment method according to an embodiment of this application. The tumor malignancy assessment method according to an embodiment of this application includes:

[0069] Step S100: Input the tumor sample to be tested into the Comprehensive Hazard Ratio (SAHR) model. The tumor sample to be tested includes the transcriptome data of the tumor to be tested.

[0070] Step S200: The SAHR model outputs the SAHR value based on the tumor sample to be tested;

[0071] Step S300: Determine the malignancy of the tumor sample to be tested based on the SAHR value;

[0072] like Figure 2 As shown, Figure 2 This is a flowchart of a method for constructing the Comprehensive Risk Ratio Index (SAHR) model according to an embodiment of this application. The SAHR model is constructed using the following method:

[0073] Step S400: Obtain multiple tumor sample data and adjacent normal sample data corresponding to the tumor sample data. The tumor sample data includes first transcriptome data and patient prognostic information. The first transcriptome data is the transcriptome data of the tumor sample. The adjacent normal sample data includes second transcriptome data, which is the transcriptome data of the adjacent normal sample.

[0074] Step S500: Screen differentially expressed genes based on the first transcriptome data and the second transcriptome data;

[0075] Step S600: Divide the tumor sample data into a training set and a test set;

[0076] Step S700: Perform survival analysis on the training set based on differentially expressed genes, and screen out prognosis-related genes from the training set. Prognosis-related genes are differentially expressed genes that are related to the patient's prognosis.

[0077] Step S800: Construct a SAHR model based on prognosis-related genes;

[0078] Step S900: Validate the SAHR model using a test set.

[0079] In this step, multiple tumor sample data and corresponding adjacent normal sample data are acquired. The tumor sample data includes first transcriptome data and patient prognostic information, while the adjacent normal sample data includes second transcriptome data. Differentially expressed genes are screened based on the first and second transcriptome data. The tumor sample data is divided into a training set and a test set. Survival analysis is performed on the training set based on the differentially expressed genes. Prognostic-related genes are screened from the training set. A SAHR model is constructed based on the prognostic-related genes. The SAHR model is validated using the test set. The tumor sample to be tested is input into the SAHR model, and the SAHR model outputs the SAHR value based on the tumor sample to be tested. The malignancy of the tumor sample to be tested is determined based on the SAHR value. The tumor malignancy assessment method according to the first aspect of this application, compared with traditional tumor malignancy assessment techniques, only requires obtaining the transcriptome data of the patient's tumor to assess its malignancy. This method is simple, inexpensive, and time-efficient. The SAHR model can obtain SAHR values ​​based on the molecular-level information of the tumor, enabling quantitative analysis of tumor malignancy and accurately assessing the survival of cancer patients at various stages. The SAHR model is constructed based on the expression information of multiple genes, rather than a single biomarker gene, thus exhibiting stronger robustness. The SAHR model can be used alone or in conjunction with clinical staging to provide more accurate patient prognosis prediction.

[0080] Understandably, patient prognostic information can include overall survival, or it can use data such as disease-free survival and disease-free recurrence survival. Tumor sample data and adjacent normal sample data were obtained by extracting from the TCGA (The Cancer Genome Atlas) database.

[0081] like Figure 3 As shown, Figure 3 This is a flowchart of screening prognosis-related genes in one embodiment of this application. In one embodiment of this application, step S700, "performing survival analysis on the training set based on differentially expressed genes and screening prognosis-related genes from the training set", is further explained. Step S700 includes, but is not limited to, steps S710, S720, S730 and S740.

[0082] Step S710: Obtain the expression levels of differentially expressed genes in the training set;

[0083] Step S720: Divide the training set into a high expression level group and a low expression level group according to the expression level of differentially expressed genes in the training set;

[0084] Step S730: Perform survival analysis on the high expression group and the low expression group using the Kaplan-Meier method to obtain the hazard ratio (HR) and P value corresponding to the differentially expressed gene. The P value is the probability of making a mistake.

[0085] Step S740: Select prognosis-related genes from the training set based on the P-value.

[0086] In this step, the expression levels of differentially expressed genes in the training set are obtained. Based on the expression levels of differentially expressed genes in the training set, the training set is divided into a high expression level group and a low expression level group. Then, the Kaplan-Meier method is used to perform survival analysis on the high expression level group and the low expression level group to obtain the P-value corresponding to the differentially expressed genes. Based on the P-value, prognosis-related genes are screened from the training set, which can accurately screen prognosis-related genes from differentially expressed genes.

[0087] In one embodiment, using the Kaplan-Meier method based on the Log-rank test, each differentially expressed gene can obtain a P-value and an HR value of the high expression group relative to the low expression group. Then, prognosis-related genes are screened based on the P-value, for example, using a P-value less than 0.05 as the screening condition to screen prognosis-related genes from differentially expressed genes.

[0088] Understandably, other p-value ranges can be set as filtering criteria, such as p-values ​​less than 0.04.

[0089] It should be noted that survival analysis can also be performed on the high expression group and the low expression group through regression analysis to obtain the P value corresponding to the differentially expressed gene.

[0090] In one embodiment of this application, the types of prognosis-related genes include suspected proto-oncogenes, suspected upregulated protective genes, suspected tumor suppressor genes, and suspected downregulated protective genes. Suspected proto-oncogenes are differentially expressed genes that are upregulated in tumor sample data and lead to a worse prognosis. Suspected upregulated protective genes are differentially expressed genes that are upregulated in tumor sample data and lead to a better prognosis. Suspected tumor suppressor genes are differentially expressed genes that are downregulated in tumor sample data and lead to a worse prognosis. Suspected downregulated protective genes are differentially expressed genes that are downregulated in tumor sample data and lead to a better prognosis.

[0091] It should be noted that the types of prognosis-related genes may include one or more of the following: suspected proto-oncogenes, suspected upregulated protective genes, suspected tumor suppressor genes, and suspected downregulated protective genes.

[0092] like Figure 4 As shown, Figure 4This is a flowchart of selecting prognosis-related genes from candidate genes in one embodiment of this application. In one embodiment of this application, step S740 "screening prognosis-related genes from the training set according to the P value" is further explained. Step S740 includes, but is not limited to, steps S741, S742 and S743.

[0093] Step S741: Select candidate genes from the training set based on the P-value. Candidate genes are differentially expressed genes that are related to the patient's prognosis.

[0094] Step S742: Sort the candidate genes of each type according to the P-value from smallest to largest to obtain the significant sequence;

[0095] Step S743: Select the top forty candidate genes from the significant sequences of each type as prognostic-related genes.

[0096] In this step, candidate genes are selected from the training set based on the P-value. These candidate genes are differentially expressed genes that are related to the patient's prognosis. The candidate genes of each type are sorted from smallest to largest according to the P-value to obtain the significance sequence. The top forty candidate genes in each type of significance sequence are selected as prognosis-related genes. This allows the SAHR model to accurately output the SAHR value based on the tumor sample to be tested according to the prognosis-related genes.

[0097] It should be noted that there is no limit to the number of prognostic genes selected from each type of significant sequence. For example, the top sixty prognostic genes in the significant sequence can also be selected as prognostic genes.

[0098] Understandably, significant sequences can also be obtained by sorting each type of candidate gene based on the HR value.

[0099] In one embodiment of this application, the formula for calculating the SAHR value in step 200 is as follows:

[0100]

[0101] in, To form a comprehensive risk ratio index, For the first HR values ​​for each prognosis-related gene For the first The calculated coefficients for each prognosis-related gene, such as Figure 5 As shown, Figure 5 This is a flowchart of a method for obtaining calculation coefficients in one embodiment of this application. The method for obtaining calculation coefficients is as follows:

[0102] Step S210: Obtain the expression level of each prognosis-related gene in all tumor sample data;

[0103] Step S220: Determine the threshold based on the expression level of each prognosis-related gene in all tumor sample data;

[0104] Step S230: Obtain the expression levels of prognosis-related genes in the tumor sample to be tested;

[0105] Step S240: Determine the calculation coefficients based on the threshold, the type of prognosis-related genes, and the expression level of prognosis-related genes in the tumor sample to be tested;

[0106] If the prognosis-related gene is a suspected proto-oncogene or a suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample being tested is higher than the threshold, the calculation coefficient is 1; if the prognosis-related gene is a suspected proto-oncogene or a suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample being tested is lower than the threshold, the calculation coefficient is -1; if the prognosis-related gene is a suspected upregulated protective gene or a suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample being tested is higher than the threshold, the calculation coefficient is -1; if the prognosis-related gene is a suspected upregulated protective gene or a suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample being tested is lower than the threshold, the calculation coefficient is 1.

[0107] In this step, the expression level of each prognosis-related gene is obtained in all tumor sample data. A threshold is determined based on the expression level of each prognosis-related gene in all tumor sample data. The expression level of the prognosis-related gene in the tumor sample to be tested is then obtained. A calculation coefficient is determined based on the threshold, the type of prognosis-related gene, and the expression level of the prognosis-related gene in the tumor sample to be tested. If the prognosis-related gene is a suspected proto-oncogene or a suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is higher than the threshold, then the calculation coefficient is 1. If the prognosis-related gene is a suspected proto-oncogene or a suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is lower than the threshold, then the calculation coefficient is 1. The calculation coefficient is -1. If the prognosis-related gene is a suspected upregulated protective gene or a suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample is higher than the threshold, the calculation coefficient is -1. If the prognosis-related gene is a suspected upregulated protective gene or a suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample is lower than the threshold, the calculation coefficient is 1. The corresponding calculation coefficient can be determined according to the type of prognosis-related gene and the level of its expression in the tumor sample. The corresponding calculation coefficient is multiplied by the corresponding HR value and then summed to accurately obtain the SAHR value. The larger the SAHR value, the more severe the tumor malignancy. It can effectively quantitatively assess the degree of tumor malignancy.

[0108] In one embodiment, the median expression level of all prognosis-related genes across all tumor sample data was selected as the threshold.

[0109] One embodiment of this application further explains step S400, "acquiring multiple tumor sample data and adjacent normal sample data corresponding to the tumor sample data". Step S400 includes: acquiring multiple tumor sample data including multiple cancer types and adjacent normal sample data corresponding to the tumor sample data.

[0110] In this step, multiple tumor sample data and adjacent normal sample data corresponding to various cancer types are obtained. This allows for the screening of multiple differentially expressed genes from multiple tumor sample data of various cancer types based on the adjacent normal sample data. Furthermore, prognosis-related genes are screened from these differentially expressed genes, making the constructed SAHR model more universal and suitable for assessing the malignancy of tumors of different cancer types.

[0111] It should be noted that the selected cancer types include fourteen mainstream cancers: colon cancer, papillary renal cell carcinoma, clear cell renal cell carcinoma, lung adenocarcinoma, squamous cell lung carcinoma, rectal adenocarcinoma, gastric cancer, invasive ductal carcinoma of the breast, laryngeal cancer, tongue cancer, urothelial carcinoma of the bladder, endometrioid adenocarcinoma, thyroid cancer, and multiple glioma. Understandably, there is no limitation on the number or types of cancer types selected; other types and numbers can be chosen.

[0112] In one embodiment of this application, step S500, "screening differentially expressed genes based on first transcriptome data and second transcriptome data," is further described. Step S500 includes: screening differentially expressed genes from the first transcriptome data using DESeq2 software based on the first transcriptome data and the second transcriptome data.

[0113] In this step, differentially expressed genes are screened based on the first transcriptome data and the second transcriptome data. The DESeq2 software can accurately screen differentially expressed genes from the first transcriptome data.

[0114] It should be noted that differentially expressed genes can also be screened based on the first and second transcriptome data using other software or algorithms, such as edgeR, limma, cuffdiff, ballgown, and sleuth.

[0115] Figures 7 to 34 The correlation plots between SAHR values ​​and survival rates for colon cancer, clear cell renal cell carcinoma, papillary renal cell carcinoma, lung adenocarcinoma, squamous cell lung carcinoma, rectal adenocarcinoma, gastric cancer, invasive ductal carcinoma of the breast, laryngeal cancer, tongue cancer, urothelial carcinoma of the bladder, endometrioid adenocarcinoma, thyroid cancer, and multiform glioma are shown in the training and test sets. Figures 7 to 34The results showed that the SAHR value can predict patient survival rate well. The SAHR value can also be combined with clinical staging to more accurately predict patient prognosis. For example... Figures 35 to 38 This study demonstrates the prognostic prediction of colorectal cancer patients based on SAHR values ​​combined with clinical staging information. The results show that for patients in the early clinical stages (Stage I and Stage II), patients with positive SAHR values ​​have a significantly worse prognosis compared to patients with negative SAHR values.

[0116] In addition, such as Figure 6 As shown, one embodiment of this application also discloses an electronic device 100, including: at least one processor 110 and at least one memory 120 for storing at least one program, which, when executed by at least one processor 110, implements a tumor malignancy assessment method as described in any of the preceding embodiments.

[0117] By acquiring multiple tumor sample data and corresponding adjacent normal sample data, the tumor sample data includes first transcriptome data and patient prognostic information, while the adjacent normal sample data includes second transcriptome data. Differentially expressed genes are screened based on the first and second transcriptome data. The tumor sample data is divided into training and test sets. Survival analysis is performed on the training set based on the differentially expressed genes. Prognostic-related genes are screened from the training set, and a SAHR model is constructed based on these genes. The SAHR model is validated using a test set. The tumor sample to be tested is input into the SAHR model, and the SAHR model outputs an SAHR value based on the tumor sample. The malignancy of the tumor sample is determined based on the SAHR value. According to an embodiment of this application, an electronic device only needs to obtain the transcriptome data of the patient's tumor to be tested to assess the malignancy of the tumor. This method is simple, inexpensive, and time-saving. The SAHR model can obtain an SAHR value based on the molecular-level information of the tumor, achieving quantitative analysis of tumor malignancy and accurately assessing the survival of cancer patients at various stages.

[0118] In addition, one embodiment of this application discloses a computer-readable storage medium storing computer-executable instructions for performing a tumor malignancy assessment method as described in any of the preceding embodiments.

[0119] By acquiring multiple tumor sample data and corresponding adjacent normal sample data, the tumor sample data includes first transcriptome data and patient prognostic information, while the adjacent normal sample data includes second transcriptome data. Differentially expressed genes are screened based on the first and second transcriptome data. The tumor sample data is divided into training and test sets. Survival analysis is performed on the training set based on the differentially expressed genes. Prognostic-related genes are screened from the training set, and a SAHR model is constructed based on these genes. The SAHR model is validated using a test set. The tumor sample to be tested is input into the SAHR model, and the SAHR model outputs an SAHR value based on the tumor sample. The malignancy of the tumor sample is determined based on the SAHR value. According to an embodiment of this application, a computer-readable storage medium only requires obtaining the transcriptome data of the patient's tumor to be tested to assess the malignancy of the tumor. This method is simple, inexpensive, and time-saving. The SAHR model can obtain SAHR values ​​based on the molecular-level information of the tumor, achieving quantitative analysis of tumor malignancy and accurately assessing the survival of cancer patients at various stages.

[0120] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0121] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.

Claims

1. A processor-executable tumor malignancy assessment program, characterized in that, The tumor malignancy assessment methods implemented when it is executed by the processor include: The tumor sample to be tested is input into the comprehensive hazard ratio index model, wherein the tumor sample to be tested includes the transcriptome data of the tumor to be tested; The comprehensive hazard ratio index model outputs a comprehensive hazard ratio index based on the tumor sample to be tested. The malignancy of the tumor sample to be tested is determined based on the comprehensive risk ratio index. The comprehensive hazard ratio index model is constructed using the following method: Acquire multiple tumor sample data and adjacent normal sample data corresponding to the tumor sample data. The tumor sample data includes first transcriptome data and patient prognostic information. The first transcriptome data is the transcriptome data of the tumor sample. The adjacent normal sample data includes second transcriptome data, which is the transcriptome data of the adjacent normal sample. Differentially expressed genes were screened based on the first transcriptome data and the second transcriptome data; The tumor sample data is divided into a training set and a test set; Obtain the expression levels of the differentially expressed genes in the training set; The training set is divided into a high expression level group and a low expression level group based on the expression level of the differentially expressed genes in the training set. Survival analysis was performed on the high expression group and the low expression group to obtain the P-value and HR value corresponding to the differentially expressed gene, where the P-value is the probability of making a mistake. Candidate genes are selected from the training set based on the P-value, and the candidate genes are differentially expressed genes that are related to the patient's prognosis. The candidate genes of each type are sorted from smallest to largest according to the P-value to obtain a significant sequence; The top 40 candidate genes from each type of significant sequence are selected as prognostic-related genes. These prognostic-related genes are differentially expressed genes associated with patient prognosis. The types of prognostic-related genes include suspected proto-oncogenes, suspected upregulated protective genes, suspected tumor suppressor genes, and suspected downregulated protective genes. Suspected proto-oncogenes are differentially expressed genes that are upregulated in the tumor sample data and lead to a worse prognosis. Suspected upregulated protective genes are differentially expressed genes that are upregulated in the tumor sample data and lead to a better prognosis. Suspected tumor suppressor genes are differentially expressed genes that are downregulated in the tumor sample data and lead to a worse prognosis. Suspected downregulated protective genes are differentially expressed genes that are downregulated in the tumor sample data and lead to a better prognosis. A comprehensive hazard ratio index model was constructed based on the prognosis-related genes; the formula for calculating the comprehensive hazard ratio index is as follows: in, To form a comprehensive risk ratio index, For the first The HR values ​​corresponding to the prognosis-related genes. For the first The calculation coefficients corresponding to the prognosis-related genes are obtained by: obtaining the expression level of each prognosis-related gene in all the tumor sample data; determining a threshold based on the expression level of each prognosis-related gene in all the tumor sample data, wherein the threshold is the median of the expression level; obtaining the expression level of the prognosis-related gene in the tumor sample to be tested; determining the calculation coefficient based on the threshold, the type of the prognosis-related gene, and the expression level of the prognosis-related gene in the tumor sample to be tested; if the prognosis-related gene is the suspected proto-oncogene or the suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is... If the expression level of the prognosis-related gene is higher than the threshold, the calculation coefficient is 1; if the prognosis-related gene is the suspected proto-oncogene or the suspected downregulated protective gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is lower than the threshold, the calculation coefficient is -1; if the prognosis-related gene is the suspected upregulated protective gene or the suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is higher than the threshold, the calculation coefficient is -1; if the prognosis-related gene is the suspected upregulated protective gene or the suspected tumor suppressor gene, and the expression level of the prognosis-related gene in the tumor sample to be tested is lower than the threshold, the calculation coefficient is 1. The comprehensive hazard ratio index model was validated using the test set.

2. The tumor malignancy assessment procedure according to claim 1, characterized in that, The step of acquiring multiple tumor sample data and adjacent normal sample data corresponding to the tumor sample data includes: acquiring multiple tumor sample data including multiple cancer types and adjacent normal sample data corresponding to the tumor sample data.

3. The tumor malignancy assessment procedure according to claim 1, characterized in that, The step of screening differentially expressed genes based on the first transcriptome data and the second transcriptome data includes: screening the differentially expressed genes from the first transcriptome data using DESeq2 software based on the first transcriptome data and the second transcriptome data.

4. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When at least one of the programs is executed by at least one of the processors, the tumor malignancy assessment program as described in any one of claims 1 to 3 is implemented.

5. A computer-readable storage medium, characterized in that, It contains a processor-executable program, which, when executed by the processor, is used to implement the tumor malignancy assessment program as described in any one of claims 1 to 3.