Colorectal cancer prognosis evaluation method and device based on multi-modal data, and medium
By integrating multimodal data and using feature detection models, combined with image verification, the limitations of traditional colorectal cancer prognostic assessment methods have been addressed, enabling more accurate prognostic assessment and support for personalized treatment plans.
Patent Information
- Application Number
- CN202411701757.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Traditional prognostic assessment methods for colorectal cancer rely on single clinicopathological indicators, which are difficult to comprehensively and accurately predict individualized prognosis and cannot provide sufficient information to support personalized treatment plans and predict the risk of disease recurrence and metastasis.
A multimodal data integration approach was adopted, including genomics, proteomics, microbiome and metabolomics data, combined with imaging data. Key features were extracted through a pre-set colorectal cancer feature detection model and verified by images, forming a two-way verification closed loop to ensure the accuracy and reliability of feature extraction.
It enables comprehensive capture of complex molecular biological changes in colorectal cancer patients after surgery, improves the accuracy and reliability of prognostic assessment, better adapts to individual differences, and provides prognostic assessments that are more in line with actual situations.
Smart Images

Figure CN119889640B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of colorectal cancer prognosis, in particular to a colorectal cancer prognosis evaluation method based on multi-modal data, a device and a medium. BACKGROUND
[0002] Colorectal cancer is one of the malignant tumors with high morbidity and mortality worldwide, which poses a serious threat to human health. According to the relevant data of the World Health Organization, its morbidity ranks among the top in various cancers, and in recent years it has shown a rising trend, especially in some developed countries and regions, the prevention and control situation of colorectal cancer is more severe.
[0003] In the clinical diagnosis and treatment process of colorectal cancer, prognosis evaluation has always been a very key link. Accurate prognosis evaluation can provide an important basis for developing personalized treatment plans, predicting disease recurrence and metastasis risk, and reasonably planning follow-up strategies, thereby significantly improving the survival rate and quality of life of patients. However, the traditional colorectal cancer prognosis evaluation method mainly relies on clinical pathological indicators such as tumor TNM stage, histological grade, tumor marker level, and patient age, gender and other factors. Although these indicators can reflect the severity of the patient's condition and prognosis to some extent, they can only provide relatively limited information and are difficult to comprehensively and accurately predict the individual prognosis of colorectal cancer patients. SUMMARY
[0004] In order to improve the problem that the traditional colorectal cancer prognosis evaluation method is difficult to comprehensively and accurately predict the individual prognosis of colorectal cancer patients, the present application provides a colorectal cancer prognosis evaluation method based on multi-modal data, a device and a medium.
[0005] In a first aspect, the present application provides a colorectal cancer prognosis evaluation method based on multi-modal data, which adopts the following technical solution:
[0006] A colorectal cancer prognosis evaluation method based on multi-modal data, the multi-modal data including genomics data, proteomics data, microbiomics data, metabolomics data and imaging data, the imaging data including one of endoscopy imaging, PET-CT molecular imaging and magnetic resonance imaging; the colorectal cancer prognosis evaluation method comprising:
[0007] acquire a multi-modal to-be-detected data set of a target patient, the multi-modal to-be-detected data set comprising examination data and imaging data, the examination data comprising at least two of genomic data, proteomic data, microbiomic data and metabolomic data of the patient collected after colorectal cancer surgery, and the imaging data being image imaging data for reflecting a current surgical site and a tissue or organ around the surgical site of the target patient;
[0008] input the examination data into a preset colorectal cancer feature detection model to obtain a plurality of first key features and a plurality of second key features, the first key features comprising a plurality of sub-features corresponding to one or more types of omics data in the examination data that are abnormal, and the second key features being image manifestations of the first key features;
[0009] verify the second key features according to the image imaging data to obtain a verification result, the verification result comprising a first key feature being correct and the first key feature being incorrect;
[0010] when the verification result is that the first key feature is correct, evaluate the prognosis of colorectal cancer of the target patient based on the first key feature to obtain an evaluation result.
[0011] By adopting the above technical solution, the advantages of multi-modal data are fully integrated, the examination data covers information from the genetic level to the microbial community and the metabolic pathway, and can more comprehensively capture the complex molecular biological change characteristics in the body of a colorectal cancer patient after surgery. The preset colorectal cancer feature detection model is used to analyze and mine these data, and the first key features and the corresponding image manifestations of the second key features are accurately extracted, thereby effectively avoiding the limitations and one-sidedness of a single data source;
[0012] In addition, in the verification link, the second key features obtained from other omics data are verified by means of the image imaging data in the multi-modal data, thereby further ensuring the accuracy and reliability of the extraction of key features. When the verification result shows that the first key feature is correct, the prognosis of colorectal cancer of the patient is evaluated based on these first key features that have been strictly screened and verified, thereby significantly improving the accuracy of the prognosis evaluation.
[0013] In a specific implementable embodiment, before the examination data is input into the preset colorectal cancer feature detection model, the method further comprises:
[0014] acquiring a plurality of multi-modal reference data sets of a plurality of specific patients, the specific patients refer to all patients with stage II or III colorectal cancer after radical surgery, and each of the specific patients is labeled with one of a first label indicating a good postoperative prognosis, a second label indicating the occurrence of colorectal cancer recurrence, and a third label indicating the occurrence of colon cancer metastasis, the multi-modal reference data set includes three sets of genomic data, proteomic data, microbiomic data and metabolomic data corresponding to each label respectively;
[0015] establishing a colorectal cancer feature detection model, and training the colorectal cancer feature detection model through the test data;
[0016] The training of the colorectal cancer feature detection model through the test data includes:
[0017] inputting all the test data into the established colorectal cancer feature detection model;
[0018] According to all the test data, a plurality of standard key features corresponding to the second label and the third label are obtained to complete the training.
[0019] In a specific embodiment, the standard key features refer to key features that exhibit consistent or typical performance in the specific patient population corresponding to one of the second label or third label, and each of the labels corresponds to one or more standard key features.
[0020] In a specific embodiment, the genomic data includes a plurality of genetic mutation indicators of single nucleotide polymorphism, copy number variation and insertion / deletion mutation, the proteomic data includes a plurality of protein expression level indicators of tumor marker protein, cell cycle regulation protein and apoptosis related protein; the microbiomic data includes a plurality of microbial community composition indicators classified by genus and door level, and microbial diversity indicators; the metabolomic data includes concentration indicators of small molecule metabolites.
[0021] In a specific embodiment, the obtaining of a plurality of first key features and a plurality of second key features includes:
[0022] calculating the matching degree of each data in the test data according to all standard key features corresponding to the second label and the third label respectively, obtaining the recurrence matching degree corresponding to the first label and the metastasis matching degree corresponding to the second label, the recurrence matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient corresponding to the second label, and the metastasis matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient corresponding to the third label;
[0023] obtaining a first key feature according to the recurrence matching degree and the metastasis matching degree;
[0024] obtaining a second key feature according to the first key feature;
[0025] the calculating the matching degree of each data in the test data according to all standard key features corresponding to the second label and the third label respectively, obtaining the recurrence matching degree corresponding to the first label and the metastasis matching degree corresponding to the second label includes:
[0026] determining target sub-features corresponding to each label in the test data according to the standard key features corresponding to the second label and the third label respectively, the number of target sub-features corresponding to each label is consistent with and one-to-one corresponding to the number of standard key features corresponding to the corresponding label;
[0027] obtaining the recurrence matching degree according to all target sub-features corresponding to the second label, and obtaining the metastasis matching degree according to all target sub-features corresponding to the third label.
[0028] In a specific implementation, the obtaining a first key feature according to the recurrence matching degree and the metastasis matching degree includes:
[0029] determining a recurrence matching threshold and a metastasis matching threshold, the recurrence matching threshold = W1* the number of target sub-features corresponding to the second label in the test data, and the metastasis matching threshold = W2* the number of target sub-features corresponding to the third label in the test data; wherein W1 and W2 are adjustment coefficients, used to change the size of the recurrence / metastasis matching threshold according to actual conditions;
[0030] when the recurrence matching degree is greater than or equal to the recurrence matching threshold, all target sub-features corresponding to the second label and meeting the corresponding standard key features are all taken as first key features;
[0031] When the transfer matching degree is greater than or equal to the transfer matching threshold, all target sub-features corresponding to the third label that meet the corresponding standard key feature are determined as first key features.
[0032] In a specific implementation, the verification result includes a first key feature correct and a first key feature incorrect; and the verification of the second key feature according to the imaging data of the image yields a verification result, which includes:
[0033] If the actual image performance in the imaging data of the image matches the second key feature, the verification result is determined to be a first key feature correct;
[0034] Otherwise, the verification result is determined to be a first key feature incorrect, and the steps of determining the recurrence matching threshold and the transfer matching threshold are returned.
[0035] In a specific implementation, when the verification result is a first key feature correct, the prognosis of colorectal cancer of the target patient is evaluated according to a recurrence matching degree and a transfer matching degree, and an evaluation result is obtained, which includes:
[0036] A first weight value of the recurrence matching degree and a second weight value of the transfer matching degree are respectively determined;
[0037] A comprehensive risk score of the target patient is calculated according to the weights of the recurrence matching degree and the transfer matching degree, the comprehensive risk score = first weight value * (recurrence matching degree / number of all target sub-features corresponding to the second label in the test data) + second weight value * (transfer matching degree / number of all target sub-features corresponding to the third label in the test data), and the first weight value + the second weight value = 1;
[0038] The prognosis level is divided according to the comprehensive risk score, and an evaluation result is obtained, which includes a prognosis condition relatively optimistic, an evaluation result of a prognosis condition at a medium level, and a prognosis condition not optimistic;
[0039] The evaluation result obtained by dividing the prognosis level according to the comprehensive risk score includes:
[0040] A first threshold and a second threshold are determined;
[0041] When the comprehensive risk score is less than or equal to the first threshold, the evaluation result is determined to be a prognosis condition relatively optimistic;
[0042] When the comprehensive risk score is greater than the first threshold and less than the second threshold, the evaluation result is determined to be a prognosis condition at a medium level;
[0043] When the comprehensive risk score is greater than the second threshold value, the evaluation result is determined as a prognosis condition not optimistic.
[0044] In a second aspect, the present application provides an intelligent terminal, which adopts the technical solution as follows:
[0045] An intelligent terminal, characterized in that it comprises a memory and a processor, the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to realize the multi-modal data-based colorectal cancer prognosis evaluation method according to the first aspect.
[0046] In a third aspect, the present application provides a computer-readable storage medium, which adopts the technical solution as follows:
[0047] A computer-readable storage medium, the readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to realize the multi-modal data-based colorectal cancer prognosis evaluation method according to the first aspect.
[0048] In summary, the present application includes at least one of the following beneficial technical effects:
[0049] 1. Multi-dimensional data integration advantage: By integrating genomics, proteomics, microbiomics and metabolomics data, the complex molecular biology change network in the body of colorectal cancer patients after surgery can be comprehensively reflected, not limited to the one-sided perspective of a single data type, but considered from multiple aspects such as gene variation, protein expression regulation, microbial community ecology and metabolic pathway remodeling, so as to more accurately capture key features related to colorectal cancer recurrence and metastasis; these multi-modal data are correlated and complementary to each other, and together build a more complete and accurate colorectal cancer prognosis judgment system, greatly improving the accuracy of predicting the development of the patient's condition.
[0050] 2. The multi-modal reference data set of a large number of historical specific patients (stage II / III colorectal cancer patients after radical surgery) is used to train the colorectal cancer feature detection model, so that the model can learn the common features and difference patterns of the patient population under different prognosis labels (good postoperative prognosis, colorectal cancer recurrence and colorectal cancer metastasis). For each patient's unique test data, this personalized model trained based on big data can better adapt to the heterogeneity between different patient individuals, avoiding the limitations of traditional general models when facing complex individual differences, thereby providing more accurate prognosis evaluation for each patient, further improving the accuracy and reliability of prognosis judgment.
[0051] 3. After the model extracts the first key feature related to recurrence and metastasis, it further generates the corresponding image performance second key feature according to the first key feature, and then verifies the first key feature through the second key feature and the image feature. This process forms a closed loop of bidirectional verification, which confirms the key features from different angles. In this way, it can effectively exclude false feature extraction caused by data noise, model error or other uncertain factors, greatly improving the accuracy and reliability of the extracted first key feature. For example, if only the model extracts the first key feature unilaterally, it may misjudge some accidental data features that are not closely related to the actual disease condition as key features. However, after the generation of the image performance associated with the second key feature and the subsequent reverse verification, it can more accurately filter out the key features that are closely related to colorectal cancer recurrence and metastasis.
[0052] 4. The first key feature is an abnormal feature related to disease development extracted from multi-modal data, while the second key feature is the direct representation of these abnormal features on the image. By combining and verifying the second key feature with the actual image imaging data, it realizes the cross-verification from the perspective of molecular biology (multi-modal data) to the perspective of anatomical imaging (image imaging data). Different perspectives of data confirm each other, making the determined key features more convincing. For example, from the perspective of genomics, it is found that a certain gene mutation occurs (first key feature), and the corresponding image performance may be a specific change in shape or density of the surgical site intestine (second key feature). When the image imaging data indeed presents such changes, it further confirms the effectiveness of the association between the gene mutation and colorectal cancer recurrence and metastasis, ensuring that the key features can truly reflect the development trend of the disease. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 is a flowchart of a colorectal cancer prognosis evaluation method based on multi-modal data according to an embodiment of the present application. DETAILED DESCRIPTION
[0054] To make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0055] The embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0056] An embodiment of the present application discloses a colorectal cancer prognosis evaluation method based on multi-modal data. The multi-modal data in the present application includes genomics data, proteomics data, microbiomics data, metabolomics data, and imaging data. The imaging data includes one of endoscopy imaging, PET-CT molecular imaging, and magnetic resonance imaging.
[0057] Referring to Figure 1 A colorectal cancer prognosis evaluation method based on multi-modal data includes the following steps.
[0058] S100, acquiring a multi-modal data set to be detected of a target patient;
[0059] The target patient refers to a patient who needs to perform colorectal cancer prognosis evaluation. The test data includes test data and imaging data. The test data includes at least two of the genomics data, proteomics data, microbiomics data, and metabolomics data collected from the patient after colorectal cancer surgery. The imaging data is imaging data for reflecting the current surgical site and the surrounding tissue or organ of the target patient.
[0060] The genomics data includes multiple genetic mutation indicators such as single nucleotide polymorphism (SNP), copy number variation (CNV), and insertion / deletion mutation (Indel). These genetic mutation indicators can be obtained by using existing gene sequencing technology or gene chip technology (i.e., corresponding genetic detection report).
[0061] The proteomics data includes multiple protein expression level indicators of common tumor marker proteins (embryonic antigen CEA, carbohydrate antigen 19-9, etc.), cell cycle regulatory proteins (cell cycle proteins Cyclins, cell cycle protein-dependent kinases CDKs, and their inhibitors CKIs, etc.), and apoptosis-related proteins (apoptosis proteins Bax, Bad, anti-apoptotic proteins (Bcl-2, Bcl-xL, etc.). These indicators can be obtained by using existing mass spectrometry technology, enzyme-linked immunosorbent assay technology, etc. (i.e., blood test results).
[0062] The microbiomics data includes multiple microbial community composition indicators classified by genus (Fusobacterium nucleatum, Bacteroides fragilis, Bifidobacterium, Lactobacillus, etc.) / classified by phylum (Firmicutes, Bacteroidetes, Proteobacteria, Actinobacteria, etc.) and microbial diversity indicators (species richness, Shannon diversity index, and Simpson diversity index). These indicators can be obtained by using metagenomic sequencing, analysis software and algorithms based on sequencing data (such as QIIME and Mothur), etc. (i.e., fecal / intestinal mucosa tissue / blood sample test results).
[0063] The metabolomics data includes concentration indicators of small molecule metabolites (amino acid metabolites, sugar metabolites, and fatty acid metabolites), which can be obtained by existing liquid chromatography-mass spectrometry, metabolic flow analysis, etc. (i.e. blood / urine / tissue / fecal sample detection results).
[0064] S200, input the test data into a preset colorectal cancer feature detection model to obtain a plurality of first key features and a plurality of second key features;
[0065] In this embodiment, the image imaging data is taken as an example of a plurality of abdominal CT scan images of a target patient; the first key features include a plurality of abnormal sub-features corresponding to one or more omics data in the test data, wherein the genomics data corresponds to a plurality of gene sub-features, the proteomics data corresponds to a plurality of protein sub-features, the microbiomics data corresponds to a plurality of microorganism sub-features, and the metabolomics data corresponds to a plurality of metabolite sub-features; and the second key features are image manifestations of the first key features.
[0066] It should be noted that the preset colorectal cancer feature detection model in S200 is mainly used to obtain a plurality of first key features corresponding to the target patient according to the test data of the target patient, and the first key features are certain abnormal features related to the development of colorectal cancer in the corresponding omics data. The meaning of the first key features is briefly described as follows:
[0067] For genomics data, the corresponding gene sub-feature is the mutation of certain specific genes related to colorectal cancer, which can be expressed as XX gene mutation; for example, KRAS gene mutation is relatively common in colorectal cancer. The protein encoded by this gene participates in the intracellular signal transduction pathway and is used to regulate cell proliferation, differentiation, etc. However, the mutated KRAS protein is in a continuous activation state, which makes the growth-promoting signal in the cell continuously transmitted, thereby enhancing the proliferation ability, invasiveness, and migration ability of tumor cells, and further causing the tumor to be more likely to relapse and metastasize.
[0068] For proteomics data, the corresponding protein sub-feature is the abnormal expression level of certain key proteins related to colorectal cancer, which can be expressed as XX protein expression abnormality; too high or too low expression of some proteins can disturb normal cell metabolism and signal transduction pathways, thereby promoting the progression of colorectal cancer. For example, the classic colorectal cancer-related protein marker carcinoembryonic antigen (CEA) often causes an increase in serum concentration in the blood of patients when the tumor recurs and metastasizes, because when tumor cells recur or metastasize to other parts, tumor cells will continue to secrete CEA in the current part, causing the CEA level in the blood to continuously increase, so the change of CEA can be dynamically monitored to assist in judging whether the tumor has relapsed or metastasized.
[0069] For microbiome data, the corresponding microbial sub-feature is the imbalance state of certain specific microbial community related to colorectal cancer, which can be expressed as XX microbial community state imbalance; the specific microbial community includes harmful microbial community such as Fusobacterium nucleatum and Bacteroides fragilis, and beneficial microbial community such as Bifidobacterium and Lactobacillus, wherein the overgrowth of harmful microorganisms or the continuous decrease of beneficial microorganisms can affect the intestinal microenvironment, immune regulation and the like, thereby creating conditions for the growth and deterioration of colorectal cancer; for example, the overgrowth of Fusobacterium nucleatum can promote the recurrence and metastasis of colorectal cancer through various mechanisms, it can adhere to the surface of tumor cells, activate related signaling pathways in tumor cells, enhance the invasiveness of tumor cells, and also regulate the immune response in the tumor microenvironment, inhibit the body's anti-tumor immunity, so that tumor cells are more likely to escape immune surveillance, and then spread in the body, increasing the risk of recurrence and metastasis.
[0070] For metabolomics data, the corresponding metabolic sub-feature is the abnormality of the change of the body's metabolic products related to colorectal cancer, which can be expressed as XX metabolic product change abnormality; in this embodiment, the abnormality of the change of the body's metabolic products includes abnormality of the metabolite level of energy / amino acid and change of the metabolic pathway, and when the metabolomics data of the target patient meets one of the above two (there is an abnormality in the metabolite level of energy / amino acid), the metabolic sub-feature is XX metabolic product change abnormality; for example, the uptake and metabolism of amino acids in the colorectal tumor cells will change, for example, the uptake of glutamine increases: glutamine can participate in various biological synthesis pathways in cells, providing nitrogen source, carbon source and other material basis for tumor cells, supporting the rapid proliferation of tumor cells and the metabolic demand in the metastasis process, therefore, the increase of glutamine uptake is likely to help the proliferation / metastasis of tumors, leading to an increased risk of recurrence and metastasis, therefore, the accumulation of such metabolic intermediates as above can be regarded as one of the manifestations in the development of colorectal cancer.
[0071] In order to better illustrate the prognosis method of colorectal cancer provided in this embodiment, before specifically explaining the steps of S200, the construction and training process of the preset colorectal cancer feature detection model in S200 will be explained first through steps S110-S120:
[0072] S110, obtaining a plurality of specific patient's multi-modal reference data sets in history;
[0073] Wherein, the specific patients refer to all the patients with stage II / III colorectal cancer after radical surgery, and each specific patient has a label, which can be one of good postoperative prognosis (first label), occurrence of colorectal cancer recurrence (second label) and occurrence of colon cancer metastasis (third label), and the number of specific patients under each label in this embodiment is 100.
[0074] S120, establishing a colorectal cancer feature detection model, and training the colorectal cancer feature detection model through the multi-modal reference data set;
[0075] Specifically, the step of training the colorectal cancer feature detection model in S120 includes:
[0076] S121, inputting all the multi-modal reference data sets into the established colorectal cancer feature detection model;
[0077] Wherein, the colorectal cancer feature detection model will classify all the multi-modal reference data sets according to the labels of the specific patients corresponding to the multi-modal reference data sets.
[0078] S122, obtaining a plurality of standard key features corresponding to the second label and the third label respectively according to all the multi-modal reference data sets to complete the training;
[0079] Wherein, the standard key feature refers to a key feature that presents consistent or typical performance in the specific patient population corresponding to a certain label (second label or third label), and each label can correspond to one or more standard key features. Specifically, the standard key feature is mainly determined according to the statistical analysis results of each index in the corresponding omics data. In this embodiment, the multi-modal reference data set of the specific patient population with the label of good postoperative prognosis (first label) is mainly used as a reference group, and then the multi-modal reference data sets corresponding to the other two labels are compared with the reference group to obtain the corresponding standard key features. The following is an example of how to obtain the standard key features of the specific patient population with the label of colorectal cancer recurrence (second label) (the principle of obtaining the standard key features corresponding to the third label is the same as the following steps, which will not be repeated in this application):
[0080] For genomics data, after collecting the relevant data such as genome sequencing of specific patients under a plurality of second labels, the indicators of each gene (such as SNP site frequency, CNV condition, etc.) are statistically summarized; for example, for a SNP site of a specific gene, the frequency of the occurrence of a specific variation at the site in all specific patients is counted, and the average value is calculated. If this average value is significantly higher than that of the specific patient population with good postoperative prognosis, and this increase is consistent in the recurrent patient population, then the corresponding frequency value is determined as the relevant value in the standard key feature. Similarly, for the copy number variation of oncogenes and tumor suppressor genes, the numerical range of their respective copy number variations in the recurrent patient population is also counted, and the representative average value (such as the CNV of the oncogene reaching a certain value and the CNV of the tumor suppressor gene reaching a certain value) is taken to define it. This reflects the common genetic mutation characteristics of the specific patient population with recurrence of colorectal cancer at the genome level;
[0081] For proteomics data, the expression levels of each tumor marker protein, cell cycle regulatory protein, and apoptosis-related protein are statistically analyzed; for example, for a specific tumor marker protein, the expression level of the protein in all specific patients is collected, and the average expression level is calculated. If this level is significantly higher than that of the patient population with good postoperative prognosis, and this high expression is generally present in the recurrent patient population, then this high expression value is set as the expression level value of the protein in the standard key feature (such as the expression level of XX protein reaching a certain value). For cell cycle protein-dependent kinase inhibitors, which are key proteins for regulating cell cycle, their expression levels in the recurrent patient population are also counted, and if they are generally low, the average value of their low expression (such as the expression level of cell cycle protein-dependent kinase inhibitor being low to a certain value) is taken as the standard key feature;
[0082] For microbiomics data, the number of specific microorganisms (such as harmful bacteria such as Fusobacterium nucleatum or beneficial bacteria such as Bifidobacterium) in the intestines of each specific patient is counted, and the average number is calculated. If it is found that the number of these bacteria increases (such as the number of XX bacteria increases to a certain value) or decreases in the recurrent patient population, and this change is closely related to tumor recurrence, then this number change value is determined as one of the standard key features. At the same time, for the relative proportion between different bacterial phyla (such as Firmicutes and Bacteroidetes), the relative proportion of all corresponding specific patients is counted. If this relative proportion is significantly and consistently increased or decreased compared to the specific patient population with good postoperative prognosis (such as the relative proportion of XX bacterial phylum to XX bacterial phylum increases / decreases to a certain value), it is included in the standard key feature category;
[0083] For metabolomics data, the synthesis amount or concentration of small molecule metabolites such as glutamine, glucose, lactic acid, etc. in the corresponding specific patient population is counted. If the synthesis amount increases or decreases universally (e.g., the synthesis amount of XX metabolite increases / decreases to a certain value), and the change conforms to the metabolic common characteristics of the population, the corresponding change value is taken as part of the standard key feature.
[0084] The following is the standard key feature corresponding to the specific patient population with the label of occurrence of colorectal cancer recurrence (second label):
[0085] The standard key feature corresponding to the genomic data is that the SNP site frequency of XX gene increases to a certain value, the CNV of oncogene is as high as a certain value, and the CNV of tumor suppressor gene is as low as a certain value; the standard key feature corresponding to the proteomics data is that the expression level of XX protein is as high as a certain value, and the expression level of cell cycle protein-dependent kinase inhibitor is as low as a certain value; the standard key feature corresponding to the microbiome data is that the number of XX bacteria increases to a certain value, and the relative proportion of XX bacterial phylum to XX bacterial phylum increases to a certain value; the standard key feature corresponding to the metabolomics data is that the synthesis amount of XX metabolite decreases to a certain value; wherein the "certain value" in the above is the average value of all corresponding parameters (one of the above SNP site frequency, CNV of oncogene, expression level of XX protein, etc.) in the corresponding specific patient population in this embodiment.
[0086] Specifically, the step of obtaining the first key feature and the second key feature in S200 includes:
[0087] S210, respectively, according to all the standard key features corresponding to the second label and the third label, the matching degree of each data in the test data is calculated, and the recurrence matching degree corresponding to the first label and the metastasis matching degree corresponding to the second label are obtained;
[0088] Wherein, the recurrence matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient with the label of occurrence of colorectal cancer recurrence, and the metastasis matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient with the label of occurrence of colorectal cancer recurrence. Specifically, the calculation of the recurrence matching degree and the metastasis matching degree in this embodiment adopts the way of quantitative scoring:
[0089] S211, respectively, according to the standard key features corresponding to the second label and the third label, the target sub-feature corresponding to each label in the test data is determined;
[0090] The target sub-feature corresponding to each label is consistent with and one-to-one corresponds to the standard key feature corresponding to the label, such as the standard key feature of "SNP site frequency of gene a increases to X1" corresponding to the target sub-feature of "SNP site frequency of gene a is X2", and the standard key feature of "relative proportion of b phylum and c phylum decreases to Y1" corresponding to the target sub-feature of "relative proportion of b phylum and c phylum is Y2".
[0091] S212, obtaining a recurrence matching degree according to all target sub-features corresponding to the second label, and obtaining a metastasis matching degree according to all target sub-features corresponding to the third label;
[0092] Specifically, the calculation method of the recurrence matching degree is as follows (the calculation of the metastasis matching degree is the same, and is not described here):
[0093] S2121, if the target sub-feature meets the corresponding standard key feature, the recurrence matching degree score is increased by one; otherwise, the recurrence matching degree score remains unchanged;
[0094] The target sub-feature meeting the corresponding standard key feature means that the difference between the feature value in the target sub-feature and the feature value in the standard key feature is within a preset error range. In this embodiment, the error range is taken as an example of 2% of the feature value in the standard key feature.
[0095] It should be particularly noted that the error range in step S2121 is actually determined according to the specificity of each standard key feature reflecting the development of colorectal cancer and the like. At the same time, for the same group of characteristic data, if there are multiple corresponding standard key features, the matching value corresponding to each standard key feature needs to be weighted and summed according to a certain weight. The weight can be determined in advance according to factors such as the importance of each standard key feature in the disease recurrence mechanism, for example, the weight of gene mutation is set to 0.6, and the weight of gene expression regulation is set to 0.4. In addition, in actual situations, the detection completed by the patient may not be complete due to physical / economic / time reasons and the like, so in this embodiment, if the target sub-feature corresponding to a certain standard key feature is missing in the test data, the standard key feature is directly ignored and considered not to meet the error range condition (i.e. the original recurrence / metastasis matching degree score remains unchanged). Next, the calculation of the recurrence matching degree score is described through an example in combination with the above weight distribution method:
[0096] Assuming that in the current colorectal cancer feature detection model, the standard key features corresponding to the second label are: ① the SNP site frequency of gene a increases to X1; ② the CNV of the oncogene is as high as Z1; ③ the relative proportion of b and c bacterial flora decreases to Y1; ④ the expression level of d protein is as high as S1; ⑤ the synthesis amount of e metabolite is reduced to T1; the target sub-features corresponding to the second label in the test data of the target patient are: (1) the SNP site frequency of gene a is X2; (2) the CNV of the oncogene is Z2; (3) the relative proportion of b and c bacterial flora is Y2; (5) the synthesis amount of e metabolite is T2; if only ① and (1), ⑤ and (5) meet the error range condition, then the recurrence matching degree of the target patient = 0.6*1+0.4*0+0+0+1 = 1.6.
[0097] S220, obtaining the first key feature according to the recurrence matching degree and the metastasis matching degree;
[0098] Specifically, S220 includes:
[0099] S221, determining the recurrence matching threshold and the metastasis matching threshold;
[0100] Wherein, the recurrence matching threshold = W1*the number of target sub-features corresponding to the second label in the test data, the metastasis matching threshold = W2*the number of target sub-features corresponding to the third label in the test data; wherein, W1, W2 are adjustment coefficients, used to change the size of the recurrence / metastasis matching threshold according to the actual situation, and W1 and W2 of the embodiment are taken as 0.5.
[0101] S222, when the recurrence matching degree is greater than or equal to the recurrence matching threshold, all target sub-features corresponding to the second label that meet the corresponding standard key features are regarded as the first key feature;
[0102] Wherein, the recurrence matching degree greater than or equal to the recurrence matching threshold represents that the target patient has a greater probability of recurrence of colorectal cancer, and all target sub-features that meet the corresponding standard key features represent that the target sub-feature is abnormal and similar to the features exhibited by the recurrence of colorectal cancer. The first key feature can be understood as an abnormal conclusion related to the corresponding key sub-feature.
[0103] S223, when the metastasis matching degree is greater than or equal to the metastasis matching threshold, all target sub-features corresponding to the third label that meet the corresponding standard key features are regarded as the first key feature;
[0104] Wherein, the metastasis matching degree greater than or equal to the metastasis matching threshold represents that the target patient has a greater probability of colorectal cancer metastasis, and all target sub-features that meet the corresponding standard key features represent that the target sub-feature is abnormal and similar to the features exhibited by the metastasis of colorectal cancer.
[0105] S230, obtaining a second key feature according to the first key feature;
[0106] The second key feature is a feature in which the first key feature is manifested in the image of the surgical site of the patient in a certain form, including morphological features, density features of the intestinal tract at the surgical site in the image data, and relative position relationship features of the surrounding tissue organs and the surgical site, and the like. For example, when the first key feature is a mutation of the KRAS gene, the corresponding second key feature is unclear image boundary, lobular change, surrounding tissue infiltration, and uneven density, and the like. When the first key feature is an imbalance of the state of the Fusobacterium nucleatum colony, the corresponding second key feature is intestinal mucosal hyperemia, edema, rough surface, erosion, ulcer formation, and the like. Specifically, the association relationship between different target sub-features and the second key feature is stored in the intelligent terminal for implementing the method of the present application, so as to be associated and called at any time. The association relationship is common knowledge in the art, and can be known through medical research papers and reviews, image differential diagnosis manuals, and the like, and will not be described here in detail.
[0107] S300, verifying the second key feature according to the image imaging data to obtain a verification result;
[0108] In the step S300 of the embodiment, the second key feature is mainly verified by a professional doctor against the image data of the target patient. Specifically, the step S300 includes:
[0109] S310, if the actual image in the image imaging data is consistent with the second key feature, it is determined that the verification result is that the first key feature is correct.
[0110] S320, otherwise, it is determined that the verification result is that the first key feature is incorrect, and the step S221 is returned to re-determine a new recurrence matching threshold and a metastasis matching threshold.
[0111] The new recurrence matching threshold or the metastasis matching threshold is mainly determined by modifying the adjustment coefficients W1 or W2.
[0112] S400, when the verification result is that the first key feature is correct, the prognosis of the colorectal cancer of the target patient is evaluated according to the recurrence matching degree and the metastasis matching degree to obtain an evaluation result.
[0113] Specifically, the step S400 includes:
[0114] S410, respectively determining a first weight value of the recurrence matching degree and a second weight value of the metastasis matching degree.
[0115] The weights of the recurrence matching degree and the metastasis matching degree are mainly determined according to clinical research and past experience. Since the recurrence and metastasis of colorectal cancer have different influences on the final condition of patients in the prognosis of colorectal cancer, the allocation of the weights of the two is also different. In this embodiment, the weight of the influence of the recurrence of colorectal cancer on the long-term survival and quality of life of patients is 0.6, and the weight of the influence of metastasis is 0.4 (the weight values here are only examples, and in actual application, they need to be adjusted according to in-depth and accurate analysis and research on a large number of cases. Generally, multiple factors need to be considered, such as the difficulty of treatment after the recurrence and metastasis of colorectal cancer, the damage to the functions of various organs, and the relative importance of the two in increasing the risk of death of patients).
[0116] S420, calculating the comprehensive risk score of the target patient according to the weights of the recurrence matching degree and the metastasis matching degree;
[0117] The comprehensive risk score = first weight value * (recurrence matching degree / number of all target sub-features corresponding to the second label in the test data) + second weight value * (metastasis matching degree / number of all target sub-features corresponding to the third label in the test data), and the first weight value + the second weight value = 1.
[0118] S430, dividing the prognosis level according to the comprehensive risk score to obtain the evaluation result;
[0119] The evaluation result includes a relatively optimistic prognosis, an evaluation result of a prognosis at a medium level, and a prognosis that is not optimistic. Specifically, S330 includes:
[0120] S431, determining the first threshold and the second threshold;
[0121] The first threshold is less than the second threshold. In this embodiment, the first threshold is 0.3, and the second threshold is 0.7.
[0122] S432, when the comprehensive risk score is less than or equal to the first threshold, determining that the evaluation result is a relatively optimistic prognosis;
[0123] The comprehensive risk score less than or equal to the first threshold indicates that the possibility of the recurrence and metastasis of colorectal cancer of the target patient is relatively low, and the prognosis is relatively optimistic. In subsequent treatment, it may be mainly regular consolidation treatment and regular review. It is expected that the patient's condition will remain stable for a long period of time, and the quality of life will be relatively less affected by the disease. The 5-year survival rate (only a reference index, which can be adjusted according to the actual situation) may reach 80% or more.
[0124] S433, when the comprehensive risk score is greater than the first threshold value and less than the second threshold value, determining that the evaluation result is that the prognosis is at a medium level;
[0125] Wherein, the comprehensive risk score greater than the first threshold value and less than the second threshold value means that the target patient has a certain degree of risk of recurrence and metastasis of colorectal cancer, and the prognosis is at a medium level; for such patients, the disease needs to be closely monitored, and the adjuvant treatment measures may need to be strengthened, such as appropriately adjusting the chemotherapy regimen, increasing targeted therapy, etc., and the review interval time needs to be shortened, so as to timely discover the possible recurrence or metastasis, and the 5-year survival rate may be between 30%-80%.
[0126] S434, when the comprehensive risk score is greater than the second threshold value, determining that the evaluation result is that the prognosis is not optimistic;
[0127] Wherein, the comprehensive risk score greater than the second threshold value indicates that the target patient has a high risk of recurrence and metastasis of colorectal cancer, and the prognosis is not optimistic; the clinician needs to actively adjust the treatment strategy, consider using more aggressive comprehensive treatment measures, such as combining multiple chemotherapy drugs, cooperating with immunotherapy, etc., and at the same time, psychological counseling and life care and other support work for the patient need to be done well, and the 5-year survival rate of such patients is often less than 30%.
[0128] Based on the same inventive concept as above, the embodiments of the present application also disclose an intelligent terminal, which comprises a memory and a processor, the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to realize a colorectal cancer prognosis evaluation method based on multi-modal data provided by the above method embodiments.
[0129] Based on the same inventive concept as above, the embodiments of the present application also disclose a computer readable storage medium, which stores at least one instruction, at least one program, a code set or an instruction set, and at least one instruction, at least one program, a code set or an instruction set can be loaded and executed by a processor to realize a colorectal cancer prognosis evaluation method based on multi-modal data provided by the above method embodiments.
[0130] It should be understood that "multiple" referred to herein means two or more. "And / or", which describes the association relationship of the associated objects, means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The character " / " generally represents that the front and rear associated objects are in an "or" relationship.
[0131] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or can be instructed by a program to complete the related hardware, and the program can be stored in a computer readable storage medium, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various storage medium capable of storing program codes.
[0132] The above only describes optional embodiments of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for colorectal cancer prognosis assessment based on multi-modal data, characterized in that, The multi-modal data includes genomics data, proteomics data, microbiomics data, metabolomics data, and image imaging data, the image imaging data including one of endoscopic examination image, PET-CT molecular image, and magnetic resonance image; the colorectal cancer prognosis evaluation method includes: Obtaining a multi-modal to-be-detected data set of a target patient, the multi-modal to-be-detected data set including test data and image examination data, the test data including at least two of genomics data, proteomics data, microbiomics data, and metabolomics data collected from the patient after colorectal cancer surgery, and the image examination data being image imaging data reflecting the current surgical site and the tissue or organ around the surgical site of the target patient; Inputting the test data into a preset colorectal cancer feature detection model to obtain a plurality of first key features and a plurality of second key features, the first key features including a plurality of sub-features corresponding to one or more genomics data in the test data that are abnormal, and the second key features being image manifestations of the first key features; Verifying the second key features according to the image imaging data to obtain a verification result, the verification result including first key feature correct and first key feature incorrect; When the verification result is first key feature correct, evaluating the colorectal cancer prognosis of the target patient based on the first key features to obtain an evaluation result.
2. The colorectal cancer prognosis method according to claim 1, characterized in that, Before the test data is input into the preset colorectal cancer feature detection model, the method further includes: Obtaining a multi-modal reference data set of a plurality of specific patients in history, the specific patients referring to all stage II or III colorectal cancer patients who have undergone radical surgery, and each of the specific patients having a label, the label being one of a first label indicating good postoperative prognosis, a second label indicating recurrence of colorectal cancer, and a third label indicating metastasis of colon cancer, the multi-modal reference data set including three groups of genomics data, proteomics data, microbiomics data, and metabolomics data corresponding to each type of label; Establishing a colorectal cancer feature detection model and training the colorectal cancer feature detection model through the test data; The training of the colorectal cancer feature detection model through the test data includes: Inputting all the test data into the established colorectal cancer feature detection model; Obtaining a plurality of standard key features corresponding to the second label and the third label respectively according to all the test data to complete the training.
3. The colorectal cancer prognosis method according to claim 2, characterized in that, The standard key features refer to key features that exhibit consistent or typical performance in the specific patient population corresponding to one of the second label or the third label, and each type of label corresponds to one or more standard key features.
4. The colorectal cancer prognosis method according to claim 3, characterized in that, The genomic data includes multiple genetic mutation indicators of single nucleotide polymorphism, copy number variation, and insertion / deletion mutation, the proteomic data includes multiple protein expression level indicators of tumor marker protein, cell cycle regulation protein, and apoptosis related protein, the microbiomic data includes multiple microbial community composition indicators classified by genus and door level, and microbial diversity indicators, and the metabolomic data includes concentration indicators of small molecule metabolites.
5. The colorectal cancer prognosis method according to claim 2, characterized in that, The obtaining of the first key features and the second key features comprises: respectively according to all standard key features corresponding to the second label and the third label, calculating the matching degree of each data in the test data, obtaining the recurrence matching degree corresponding to the first label and the metastasis matching degree corresponding to the second label, the recurrence matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient corresponding to the second label, and the metastasis matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient corresponding to the third label; obtaining the first key features according to the recurrence matching degree and the metastasis matching degree; obtaining the second key features according to the first key features; respectively according to all standard key features corresponding to the second label and the third label, calculating the matching degree of each data in the test data, obtaining the recurrence matching degree corresponding to the first label and the metastasis matching degree corresponding to the second label, the recurrence matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient corresponding to the second label, and the metastasis matching degree is used to reflect the similarity between the test data of the target patient and the multi-modal reference data set of the specific patient corresponding to the third label; respectively according to the standard key features corresponding to the second label and the third label, determining the target sub-features corresponding to each label in the test data, the number of target sub-features corresponding to the label is consistent with and one-to-one corresponding to the number of standard key features corresponding to the label; obtaining the recurrence matching degree according to all target sub-features corresponding to the second label, and obtaining the metastasis matching degree according to all target sub-features corresponding to the third label.
6. The colorectal cancer prognosis method according to claim 5, characterized in that, The obtaining of the first key features according to the recurrence matching degree and the metastasis matching degree comprises: determining the recurrence matching threshold and the metastasis matching threshold, the recurrence matching threshold is W1*the number of target sub-features corresponding to the second label in the test data, and the metastasis matching threshold is W2*the number of target sub-features corresponding to the third label in the test data; wherein W1 and W2 are adjustment coefficients, used to change the size of the recurrence / metastasis matching threshold according to the actual situation; when the recurrence matching degree is greater than or equal to the recurrence matching threshold, all target sub-features corresponding to the second label and meeting the corresponding standard key features are all taken as the first key features; when the metastasis matching degree is greater than or equal to the metastasis matching threshold, all target sub-features corresponding to the third label and meeting the corresponding standard key features are all taken as the first key features.
7. The colorectal cancer prognosis method according to claim 6, characterized in that, The verification result includes correct first key feature and incorrect first key feature; the verification of the second key feature according to the imaging data of the image includes: If the actual image in the imaging data of the image is consistent with the second key feature, it is determined that the verification result is correct first key feature; Otherwise, it is determined that the verification result is incorrect first key feature and the steps of determining the recurrence matching threshold and the metastasis matching threshold are returned.
8. The colorectal cancer prognosis method according to claim 7, characterized in that, When the verification result is correct first key feature, the prognosis of the target patient with colorectal cancer is evaluated according to the recurrence matching degree and the metastasis matching degree, and the evaluation result includes: The first weight value of the recurrence matching degree and the second weight value of the metastasis matching degree are determined respectively; The comprehensive risk score of the target patient is calculated according to the weights of the recurrence matching degree and the metastasis matching degree, the comprehensive risk score = first weight value * (recurrence matching degree / number of all target sub-features corresponding to the second label in the test data) + second weight value * (metastasis matching degree / number of all target sub-features corresponding to the third label in the test data), first weight value + second weight value = 1; The prognosis level is divided according to the comprehensive risk score, and the evaluation result includes that the prognosis is relatively optimistic, the evaluation result is that the prognosis is at a medium level, and the prognosis is not optimistic; The evaluation result includes: The first threshold and the second threshold are determined; When the comprehensive risk score is less than or equal to the first threshold, it is determined that the evaluation result is that the prognosis is relatively optimistic; When the comprehensive risk score is greater than the first threshold and less than the second threshold, it is determined that the evaluation result is that the prognosis is at a medium level; When the comprehensive risk score is greater than the second threshold, it is determined that the evaluation result is that the prognosis is not optimistic.
9. A smart terminal, characterized by The memory and the processor, the memory stores at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to realize the prognosis evaluation method of colorectal cancer based on multi-modal data in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, at least one program, a code set or an instruction set is loaded and executed by the processor to realize the prognosis evaluation method of colorectal cancer based on multi-modal data in any one of claims 1-8.
Citation Information
Patent Citations
Combined marker related to early and middle stage colon cancer, detection test kit and detection system
CN111690747A
Smart microarray cancer detection system
WO2008088322A2