A colorectal cancer multi-modal image data processing method and system based on uncertainty constraint evidence fusion

CN122617801APending Publication Date: 2026-08-21QIDONG FUDAN INSTITUTE OF MEDICAL INNOVATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610759198.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

当不同模态信息方向不一致时,模型难以表达模态间证据冲突,也难以提示输出结果的可信程度

Benefits of technology

第一,本发明将CT图像和全切片病理图像由普通特征输入转化为包含类别信念质量和认知不确定性的证据输出,实现了从确定性多模态特征融合到证据驱动图像数据处理的转变。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122617801A_ABST
    Figure CN122617801A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on uncertainty constraint evidence fusion colorectal cancer multi-modal image data processing method and system, belong to medical artificial intelligence and medical image processing field.The present application includes: obtaining the CT image and whole section pathological image of colorectal cancer patient;Respectively extract macroscopic image representation and microscopic pathological representation;Each modal representation is mapped to include belief quality and cognitive uncertainty evidence output;Based on clinical pathology prior, cross-scale dynamic calibration is carried out;Using Dempster-Shafer evidence theory, intra-modal and inter-modal evidence fusion is carried out;When any mode is missing, based on uncertainty constraint mechanism, missing mode compensation is carried out;Finally, the output of explainable image processing result.The present application can cross-scale fusion macroscopic image and microscopic pathological evidence, improve the reliability of image evidence fusion, while avoiding overconfident output in the scene of missing mode, suitable for colorectal cancer multi-modal image auxiliary analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of medical artificial intelligence and medical image processing, and particularly relates to a method and system for processing multimodal image data of colorectal cancer based on uncertainty-constrained evidence fusion. Background Technology

[0002] Colorectal cancer is one of the most common malignant tumors worldwide. In the diagnosis, treatment, and medical research of colorectal cancer, preoperative enhanced CT images, postoperative hematoxylin-eosin stained whole-section pathological images, and corresponding clinicopathological information constitute important multi-source data. Enhanced CT images can reflect the overall morphology of the tumor, tumor boundaries, peritumoral tissue, and imaging enhancement characteristics, while whole-section pathological images can present microscopic pathological information such as tumor glandular structure, invasion front, tumor budding, stromal reaction, and immune cell distribution. These two types of images represent the macroscopic imaging scale and the microscopic tissue scale, respectively, and are naturally complementary at the data analysis level.

[0003] Existing multimodal medical image processing models typically employ feature stitching, attention weighting, or deterministic fusion to directly combine features from different modalities and output a single score or category. While these methods can utilize multi-source information, they generally lack explicit modeling of cognitive uncertainty. When the information from different modalities is inconsistent in direction, the model struggles to express intermodal evidence conflicts and indicates the credibility of the output results.

[0004] Furthermore, in real-world data processing, situations frequently arise where a particular modality is temporarily missing, image quality is poor, or feature extraction results are incomplete. Traditional multimodal models often rely on complete paired data; when CT images or whole-slice pathological images are missing, the model may fail to function or give overconfident outputs under insufficient information.

[0005] Existing methods typically limit interpretability to heatmap visualization, lacking a mechanism for quantitatively aligning the model's region of interest with known image phenotypes and pathological morphological features of colorectal cancer. Therefore, a multimodal image data processing method and system for colorectal cancer is needed. This method can fuse macroscopic CT imaging information with microscopic pathological information from whole-slice sections across scales, explicitly output cognitive uncertainty, maintain stable data processing capabilities in modality-deficient scenarios, and provide interpretable image processing results. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a method for processing multimodal image data of colorectal cancer based on uncertainty-constrained evidence fusion, comprising: Obtain CT images and whole-section pathological images of colorectal cancer patients; The CT images and the whole-slice pathological images are preprocessed and feature-encoded respectively to obtain CT modality characterization and pathological modality characterization; The CT modal representation and the pathological modal representation are respectively mapped into modal evidence outputs that include belief quality and cognitive uncertainty; The modal evidence output is dynamically calibrated across scales based on colorectal cancer-related clinicopathological priors. Based on the Dempster-Shafer evidence theory, evidence fusion is performed on the calibrated modal evidence output to obtain fused category beliefs, fused cognitive uncertainty and evidence conflict information; When any mode is missing or determined to be invalid, missing mode compensation is performed based on the uncertainty constraint mechanism; Interpretable image processing results are generated based on the fusion category beliefs and the fusion cognitive uncertainty.

[0007] Optionally, acquiring CT images and whole-section pathological images of colorectal cancer patients further includes acquiring structured clinicopathological information that corresponds to the CT images and the whole-section pathological images.

[0008] Optionally, the preprocessing includes: The CT images are subjected to tumor region delineation, spatial resampling, grayscale window width and level adjustment, intensity normalization, and tumor center region cropping. The whole-slice pathological images are subjected to tissue region detection, quality control, removal of non-tumor or artifact regions, image block segmentation, color normalization, and image block-level feature preparation.

[0009] Optionally, the feature encoding includes: A CT image feature encoder is used to extract three-dimensional features from the preprocessed CT image data to obtain patient-level CT modal characterization. A pathological image feature encoder is used to extract features from the preprocessed pathological image block data, and patient-level pathological modality representations are obtained through multi-instance learning or attention aggregation.

[0010] Optionally, mapping the CT modal representation and the pathological modal representation to modal evidence outputs that include belief quality and cognitive uncertainty includes: Input any modality representation into at least two basic belief assignment heads to generate normalized categorical belief quality and non-negative evidence strength, respectively. Dirichlet evidence parameters are constructed based on the normalized category belief quality and the non-negative evidence strength, and the belief quality of the modality for the preset category and the corresponding cognitive uncertainty are calculated.

[0011] Optionally, the cross-scale dynamic calibration of the modal evidence output based on colorectal cancer-related clinicopathological priors includes: Based on the whole-section pathological images, tumor budding grade, invasion front characterization, stromal reaction characterization or immune infiltration characterization are obtained; based on the CT images, tumor enhancement characteristics, tumor boundary characteristics or peritumoral enhancement characteristics are obtained. The obtained pathological and CT characterizations are input into a dynamic calibration network to obtain the prior calibration coefficients for colorectal cancer. The prior calibration coefficients for colorectal cancer are used to adjust the evidence parameters, belief quality, or modality credibility weights for at least one modality.

[0012] Optionally, the evidence fusion based on Dempster-Shafer evidence theory for the calibrated modal evidence output includes: Intramodal evidence fusion is performed, which combines the evidence outputs generated by different basic belief assigners within the same modality. Intermodal evidence fusion is performed, which combines CT modal evidence and pathological modal evidence, and the overconfidence output caused by modal inconsistency is suppressed by evidence conflict normalization term; When the CT modal evidence and pathological modal evidence do not support the preset category in the same direction, the evidence conflict information increases, and the uncertainty of the fusion cognition increases accordingly.

[0013] Optionally, the missing mode compensation based on the uncertainty constraint mechanism includes: Construct a modality existence mask to determine whether a CT modality or pathological modality exists and is valid; Modal belief quality, cognitive uncertainty, and the modality existence mask are input into a dynamic evidence weighting network to obtain sample-level modality importance weights. When a certain modality is missing, false evidence of the missing modality is generated using the evidence output of the existing modalities and the cross-modal correlation matrix, and an uncertainty enhancement term is applied to the false evidence. Calculate the degree of conflict between real evidence and false evidence. If the degree of conflict is higher than a preset threshold, discard the false evidence. If the degree of conflict is not higher than the preset threshold, perform conflict perception fusion on the real evidence and the false evidence.

[0014] Optionally, the interpretable image processing result includes heatmaps or saliency maps generated for the CT modality coding branch and the pathology modality coding branch, respectively; the method further includes: calculating the tumor boundary concentration index based on the CT heatmap, calculating the high-evidence pathology image block load or invasion front enrichment degree based on the pathology image block heatmap, and jointly outputting the tumor boundary concentration index and / or the high-evidence pathology image block load with the interpretable image processing result.

[0015] This invention also provides a multimodal image data processing system for colorectal cancer based on uncertainty-constrained evidence fusion, used to implement the method described above, the system comprising: The data acquisition module is used to acquire CT images, whole-slice pathological images, and accompanying structured clinicopathological information of colorectal cancer patients. The preprocessing module is used to standardize the CT images and the whole-slice pathological images; The feature encoding module is used to generate CT modal representations and pathological modal representations, respectively; The evidence generation module is used to map the representations of each modality to category belief quality and cognitive uncertainty; A priori calibration module is used to modulate cross-scale modal evidence based on priors of colorectal cancer-related images and pathology. The evidence fusion module is used to fuse evidence from various modalities based on the Dempster-Shafer evidence theory. The missing modality compensation module is used to generate and filter false evidence when any modality is missing or invalid. The output module is used to output interpretable image processing results.

[0016] Compared with the prior art, the present invention has the following advantages and technical effects: First, this invention transforms CT images and whole-slice pathological images from ordinary feature inputs into evidence outputs that include category belief quality and cognitive uncertainty, realizing a shift from deterministic multimodal feature fusion to evidence-driven image data processing.

[0017] Second, this invention dynamically calibrates cross-scale evidence through colorectal cancer-related images and pathological priors, enabling microscopic morphological features such as the pathological invasion front and tumor budding to form mutual constraints with macroscopic imaging features such as CT tumor boundaries and peritumoral enhancement during the fusion process.

[0018] Third, this invention explicitly calculates intermodal evidence conflicts during the evidence fusion process and increases fusion cognitive uncertainty when conflicts increase, thereby reducing overconfident output caused by modality inconsistency.

[0019] Fourth, this invention treats missing modalities as a data processing state under uncertainty constraints. Through modality existence verification, dynamic evidence weight network, false evidence generation, and conflict perception fusion, relatively stable fusion processing results can still be obtained when CT or pathological modalities are missing.

[0020] Fifth, this invention can output interpretable image processing indicators such as saliency of CT tumor boundaries, high-evidence pathological image blocks, tumor boundary concentration index, and high-evidence pathological image block load, so that the fusion score and category tendency results correspond to the image phenotypes and pathological morphological features related to colorectal cancer, which facilitates multimodal image data analysis and model result display. Attached Figure Description

[0021] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of a method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the model architecture in an embodiment of the present invention. Detailed Implementation

[0022] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0024] Example 1 like Figure 1-2 As shown, this embodiment provides a method for processing multimodal image data of colorectal cancer based on uncertainty-constrained evidence fusion, characterized by including: Obtain CT images and whole-section pathological images of colorectal cancer patients; The CT images and the whole-slice pathological images are preprocessed and feature-encoded respectively to obtain CT modality characterization and pathological modality characterization; The CT modal representation and the pathological modal representation are respectively mapped into modal evidence outputs that include belief quality and cognitive uncertainty; The modal evidence output is dynamically calibrated across scales based on colorectal cancer-related clinicopathological priors. Based on the Dempster-Shafer evidence theory, evidence fusion is performed on the calibrated modal evidence output to obtain fused category beliefs, fused cognitive uncertainty and evidence conflict information; When any mode is missing or determined to be invalid, missing mode compensation is performed based on the uncertainty constraint mechanism; Interpretable image processing results are generated based on the fusion category beliefs and the fusion cognitive uncertainty.

[0025] Specifically, the following steps are included: S1. Obtain multimodal data of colorectal cancer patients; Obtain preoperative enhanced CT images, postoperative H&E stained whole-section pathological images, and clinicopathological information from patients with pathologically confirmed colorectal cancer. Clinicopathological information may include age, sex, tumor location, histological grade, AJCC / TNM stage, pT stage, N stage, M stage, lymphovascular invasion, nerve invasion, circumferential resection margin status, number of lymph nodes, tumor deposition, recurrence status, overall survival, and disease-free survival.

[0026] In this embodiment, the CT images are taken in the venous phase or the enhanced phase, and the primary tumor area is delineated by a radiologist or an automatic / semi-automatic segmentation algorithm; the pathological images are quality controlled by a pathologist to remove folded, contaminated, out-of-focus, blank background and non-tumor areas.

[0027] S2. Perform multimodal data preprocessing; Spatial resampling, intensity cropping, window width and level adjustment, Z-score normalization, and tumor center cropping are performed on CT images to obtain three-dimensional tumor region data with uniform resolution and input size. A peritumoral ring region around the tumor boundary can be further constructed for subsequent interpretive analysis.

[0028] Tissue region detection, color normalization, and image block segmentation are performed on whole-slice pathological images to divide the slices into several fixed-size pathological image blocks. A multi-instance learning approach can be used to aggregate the image block-level representations into patient-level pathological representations.

[0029] S3. Extract CT modal characterization and pathological modal characterization; The preprocessed CT tumor region is input into a CT image feature encoder to obtain a patient-level CT modal representation. The CT image feature encoder can employ a three-dimensional medical image basic model, a three-dimensional convolutional neural network, or a three-dimensional vision Transformer.

[0030] Pathological image blocks are input into a pathological image feature encoder to obtain image block-level pathological representations, and patient-level pathological modal representations are obtained through attention aggregation or multi-instance learning methods. The pathological image feature encoder can employ a basic pathological model, a visual-language pathological pre-trained model, a convolutional neural network, or a visual Transformer.

[0031] S4. Generate modal evidence output; For each modality m∈{CT,Path}, its modality representation is input into the evidence generation network to obtain the degree of evidence support for the low-risk and high-risk categories of that modality. Preferably, the evidence generation network includes two basic belief assignment heads: the first basic belief assignment head generates normalized belief quality through a softmax function; the second basic belief assignment head generates non-negative evidence strength through a ReLU function.

[0032] In this implementation, the number of risk categories K is set to 2, representing low risk and high risk respectively; Dirichlet parameters are constructed from non-negative evidence.

[0033] ; Where m represents the modality type, including CT modality and pathological image modality; k represents the category index, and K represents the total number of categories. In this embodiment, K=2, corresponding to low-risk category and high-risk category, respectively; x (m) This represents the feature representation vector of the m-th mode; , as well as These represent the weight parameter and bias parameter corresponding to the k-th class in the two basic belief assignment heads, respectively; Indicates the quality of normalized beliefs. Indicates the strength of non-negative evidence; Indicates the parameters of the Dirichlet distribution; This represents the total amount of evidence for that modality; This indicates the cognitive uncertainty of the modality; the higher the value, the less valid evidence the modality provides.

[0034] The above statement means that the model not only outputs the risk propensity, but also the sufficiency of evidence for its judgment of that risk.

[0035] S5. Perform prior dynamic calibration for colorectal cancer; To ensure consistent constraints between macroscopic CT imaging phenotypes and microscopic pathological morphological phenotypes in the evidence space, this embodiment introduces a priori dynamic calibration mechanism for colorectal cancer. Specifically, tumor budding grade, invasion front characteristics, stromal reaction, or immune infiltration-related characteristics are obtained from pathological images or reports; tumor enhancement descriptors, tumor boundary features, peritumoral enhancement features, or peritumoral extension features are obtained from CT images.

[0036] The pathological and CT-related characterizations are input into a dynamic calibration network to obtain calibration coefficients. These calibration coefficients are used to adjust the pathological modality evidence parameters, CT modality evidence parameters, or modality confidence weights. When the pathological and CT-related characterizations are consistent, the fusion contribution of the corresponding modality evidence is increased; when they are inconsistent, the possibility of over-reliance on a single modality is reduced.

[0037] S6. Perform Dempster-Shafer two-stage evidence fusion; This embodiment employs a two-stage Dempster-Shafer evidence fusion strategy. The first stage is intramodal fusion, which integrates evidence outputs generated by multiple basic belief assignment heads within the same modality to enhance intramodal consistency. The second stage is intermodal fusion, which integrates CT modality evidence and pathological modality evidence to obtain the final fused risk beliefs, fused cognitive uncertainty, and evidence conflict information.

[0038] For the belief quality and uncertainty from two sources of evidence, A and B, the conflict term C can be calculated based on the cross-product of the belief quality for different risk categories. C increases when A and B provide opposing support for different risk categories. The fused category k belief quality can be jointly determined by the category consistency term, the A belief and B uncertainty term, and the A uncertainty and B belief term, and normalized using 1-C; the fused uncertainty can be obtained by normalizing the product of the uncertainties from the two sources of evidence through conflict.

[0039] ; ; in, This represents the quality of belief in the k-th type of risk after fusion; and These represent the degree of evidence support for the k-th risk in the pathological modality and the CT modality, respectively. and These represent the cognitive uncertainties corresponding to the pathological modality and the CT modality, respectively. represents the overall cognitive uncertainty after fusion; c represents the conflict coefficient between different modal evidences, and the larger the value, the more obvious the risk judgment disagreement between the two modalities.

[0040] Through the above mechanism, when the CT modality and the pathological modality provide consistent high-risk evidence, the model outputs a higher high-risk belief; when the two evidence directions are inconsistent, the model does not force a definite risk judgment, but outputs a higher cognitive uncertainty at the same time, thus indicating that the patient belongs to a low-confidence case or a case that needs further evaluation.

[0041] S7. Perform missing mode compensation; In real-world clinical workflows, some patients may lack high-quality CT images or digitized whole-slice pathology images. This embodiment addresses this situation by implementing a modality existence verification mechanism. This mechanism includes document-level validity checks, feature integrity checks, and clinical plausibility checks, resulting in a modality existence mask.

[0042] When a modality is missing, the belief quality, cognitive uncertainty, and existence mask of the existing modalities are input into a dynamic evidence weighting network to obtain sample-level modality importance weights. Subsequently, cross-modal association matrices are used to map existing modality evidence to pseudo-evidence of the missing modality, and an uncertainty enhancement term is applied to the pseudo-evidence to prevent it from being regarded as high-confidence evidence equivalent to the true modality during the fusion process.

[0043] Furthermore, the degree of conflict between real evidence and false evidence is calculated. When the degree of conflict exceeds a preset threshold, false evidence is discarded and risk inference is made solely based on real evidence; when the degree of conflict does not exceed the preset threshold, real evidence and false evidence are fused using conflict perception. Thus, this embodiment can provide robust and non-overconfident risk assessment under conditions of missing modalities.

[0044] S8. Output multimodal evidence fusion score, category bias result, and interpretation result; A multimodal evidence fusion score is generated based on the quality of fused category beliefs and the fusion cognitive uncertainty. This score characterizes the overall support level of the CT and pathology modalities in the evidence space. Category bias results are output according to preset thresholds or rules, along with a cognitive uncertainty value or confidence level.

[0045] In this embodiment, heatmaps or saliency maps are generated for the CT modality branch and the pathology modality branch, respectively. The CT heatmap can be superimposed on the tumor area and the peritumoral area to observe the tumor boundary or peritumoral enhancement area of ​​interest in the model; the pathology heatmap can be superimposed on the whole slice image to locate high evidence image blocks, invasion fronts, tumor budding or stromal reaction areas.

[0046] Furthermore, a tumor boundary concentration index is calculated based on CT heatmaps, representing the degree of concentration of salience signals at the tumor boundary or peritumoral region; high-evidence pathological image block load is calculated based on pathological image block clustering or salience distribution, representing the proportion of high-evidence morphological regions in pathological sections. The above interpretable image processing results can be output together with multimodal evidence fusion scoring, category tendency results, confidence level hints, and structured clinicopathological information for multimodal image data analysis and model result display.

[0047] like Figure 3As shown, this embodiment also provides a multimodal image data processing system for colorectal cancer based on uncertainty-constrained evidence fusion, used to implement the method described above. The system includes: The data acquisition module is used to acquire CT images, whole-slice pathological images, and accompanying structured clinicopathological information of colorectal cancer patients. The preprocessing module is used to standardize the CT images and the whole-slice pathological images; The feature encoding module is used to generate CT modal representations and pathological modal representations, respectively; The evidence generation module is used to map the representations of each modality to category belief quality and cognitive uncertainty; A priori calibration module is used to modulate cross-scale modal evidence based on priors of colorectal cancer-related images and pathology. The evidence fusion module is used to fuse evidence from various modalities based on the Dempster-Shafer evidence theory. The missing modality compensation module is used to generate and filter false evidence when any modality is missing or invalid. The output module is used to output interpretable image processing results.

[0048] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for processing multimodal image data of colorectal cancer based on uncertainty-constrained evidence fusion, characterized in that, include: Obtain CT images and whole-section pathological images of colorectal cancer patients; The CT images and the whole-slice pathological images are preprocessed and feature-encoded respectively to obtain CT modality characterization and pathological modality characterization; The CT modal representation and the pathological modal representation are respectively mapped into modal evidence outputs that include belief quality and cognitive uncertainty; The modal evidence output is dynamically calibrated across scales based on colorectal cancer-related clinicopathological priors. Based on the Dempster-Shafer evidence theory, evidence fusion is performed on the calibrated modal evidence output to obtain fused category beliefs, fused cognitive uncertainty and evidence conflict information; When any mode is missing or determined to be invalid, missing mode compensation is performed based on the uncertainty constraint mechanism; Interpretable image processing results are generated based on the fusion category beliefs and the fusion cognitive uncertainty.

2. The method according to claim 1, characterized in that, The acquisition of CT images and whole-section pathological images of colorectal cancer patients also includes: acquiring structured clinicopathological information that corresponds to the CT images and the whole-section pathological images.

3. The method according to claim 1, characterized in that, The preprocessing includes: The CT images are subjected to tumor region delineation, spatial resampling, grayscale window width and level adjustment, intensity normalization, and tumor center region cropping. The whole-slice pathological images are subjected to tissue region detection, quality control, removal of non-tumor or artifact regions, image block segmentation, color normalization, and image block-level feature preparation.

4. The method according to claim 1, characterized in that, The feature encoding includes: A CT image feature encoder is used to extract three-dimensional features from the preprocessed CT image data to obtain patient-level CT modal characterization. A pathological image feature encoder is used to extract features from the preprocessed pathological image block data, and patient-level pathological modality representations are obtained through multi-instance learning or attention aggregation.

5. The method according to claim 1, characterized in that, Mapping CT modal representations and pathological modal representations into modal evidence outputs that include belief quality and cognitive uncertainty includes: Input any modality representation into at least two basic belief assignment heads to generate normalized categorical belief quality and non-negative evidence strength, respectively. Dirichlet evidence parameters are constructed based on the normalized category belief quality and the non-negative evidence strength, and the belief quality of the modality for the preset category and the corresponding cognitive uncertainty are calculated.

6. The method according to claim 1, characterized in that, The cross-scale dynamic calibration of the modal evidence output based on colorectal cancer-related clinicopathological priors includes: Based on the whole-section pathological images, tumor budding grade, invasion front characterization, stromal reaction characterization or immune infiltration characterization are obtained; based on the CT images, tumor enhancement characteristics, tumor boundary characteristics or peritumoral enhancement characteristics are obtained. The obtained pathological and CT characterizations are input into a dynamic calibration network to obtain the prior calibration coefficients for colorectal cancer. The prior calibration coefficients for colorectal cancer are used to adjust the evidence parameters, belief quality, or modality credibility weights for at least one modality.

7. The method according to claim 1, characterized in that, The evidence fusion based on Dempster-Shafer evidence theory for the calibrated modal evidence output includes: Intramodal evidence fusion is performed, which combines the evidence outputs generated by different basic belief assigners within the same modality. Intermodal evidence fusion is performed, which combines CT modal evidence and pathological modal evidence, and the overconfidence output caused by modal inconsistency is suppressed by evidence conflict normalization term; When the CT modal evidence and pathological modal evidence do not support the preset category in the same direction, the evidence conflict information increases, and the uncertainty of the fusion cognition increases accordingly.

8. The method according to claim 1, characterized in that, The missing mode compensation based on the uncertainty constraint mechanism includes: Construct a modality existence mask to determine whether a CT modality or pathological modality exists and is valid; Modal belief quality, cognitive uncertainty, and the modality existence mask are input into a dynamic evidence weighting network to obtain sample-level modality importance weights. When a certain modality is missing, false evidence of the missing modality is generated using the evidence output of the existing modalities and the cross-modal correlation matrix, and an uncertainty enhancement term is applied to the false evidence. Calculate the degree of conflict between real evidence and false evidence. If the degree of conflict is higher than a preset threshold, discard the false evidence. If the degree of conflict is not higher than the preset threshold, perform conflict perception fusion on the real evidence and the false evidence.

9. The method according to claim 1, characterized in that, The interpretable image processing results include heatmaps or saliency maps generated for the CT modality coding branch and the pathology modality coding branch, respectively; the method further includes: calculating the tumor boundary concentration index based on the CT heatmap, calculating the high-evidence pathology image block load or invasion front enrichment degree based on the pathology image block heatmap, and jointly outputting the tumor boundary concentration index and / or the high-evidence pathology image block load with the interpretable image processing results.

10. A multimodal image data processing system for colorectal cancer based on uncertainty-constrained evidence fusion, characterized in that, The system for implementing the method of any one of claims 1-9 comprises: The data acquisition module is used to acquire CT images, whole-slice pathological images, and accompanying structured clinicopathological information of colorectal cancer patients. The preprocessing module is used to standardize the CT images and the whole-slice pathological images; The feature encoding module is used to generate CT modal representations and pathological modal representations, respectively; The evidence generation module is used to map the representations of each modality to category belief quality and cognitive uncertainty; A priori calibration module is used to modulate cross-scale modal evidence based on priors of colorectal cancer-related images and pathology. The evidence fusion module is used to fuse evidence from various modalities based on the Dempster-Shafer evidence theory. The missing modality compensation module is used to generate and filter false evidence when any modality is missing or invalid. The output module is used to output interpretable image processing results.