Multimodal multi-task probability prediction method and system for prognosis of neoadjuvant therapy of breast cancer
By integrating various data from breast cancer patients using a multimodal, multi-task probabilistic prediction method, the potential outcomes of neoadjuvant therapy can be predicted. This addresses the issue of incomplete multimodal data integration in existing systems, improves the accuracy and interpretability of prognostic assessment, and provides objective support for individualized treatment plans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MACAO POLYTECHNIC INST
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-29
Smart Images

Figure CN122117368A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer deep learning technology, specifically relating to a multimodal, multi-task probability prediction method and system for the prognosis of neoadjuvant therapy for breast cancer. Background Technology
[0002] Neoadjuvant therapy (NAT) is a standard treatment for locally advanced breast cancer, and in recent years it has been increasingly extended to patients with smaller, more biologically aggressive tumors. NAT reduces tumor burden, improving surgical feasibility and facilitating breast-conserving and axillary-conserving surgeries. Achieving pathological complete response (pCR) after treatment is generally associated with good long-term prognosis and can provide important guidance for subsequent treatment plans.
[0003] According to clinical guidelines, current NAT (Non-Adjuvant Therapy) regimens are fixed for different molecular subtypes. For example, multi-drug combination chemotherapy with anthracyclines and taxanes as the backbone is used for triple-negative patients; chemotherapy combined with HER2 antibodies is used for HER2+ tumors; and neoadjuvant chemotherapy or endocrine therapy can be used for ER / PR+ breast cancer. However, individual responses to NAT vary considerably. Some HER2+ tumors can achieve pCR within 14 days of starting treatment (the entire treatment process may require up to 6 cycles), but the toxic side effects increase with each additional treatment cycle. Some tumors do not respond to treatment, requiring timely adjustment of the NAT regimen to maximize treatment benefits. Therefore, individualized prediction of NAT response is crucial to determine whether a patient can achieve pCR or prolong survival, thereby optimizing the choice of NAT regimen and avoiding ineffective or overtreatment.
[0004] Current breast cancer prognostic systems are primarily based on molecular subtypes, TNM staging, genomic risk assessment using gene expression profiling, and evolving prognostic biomarkers such as tumor-infiltrating lymphocytes, but do not consider the heterogeneity observed in breast cancer imaging and histopathological sections. Imaging biomarkers obtained through MRI, such as functional tumor volume from dynamic contrast-enhanced MRI (DCE-MRI) and apparent diffusion coefficients from diffusion-weighted imaging, have been found to be associated with response to neoadjuvant therapy (NAT). High-quality predictive features can also be extracted from unstructured medical reports, including radiology and pathology reports.
[0005] In recent years, deep learning has been widely used to evaluate NAT responses, enabling the automatic extraction of predictive features across modalities, thereby improving pCR prediction and risk stratification. However, previous studies have been limited to single NAT protocols, failing to consider counterfactual results of alternatives and lacking model interpretability. Selection bias in non-randomized controlled trials further complicates the comparison of treatment effects.
[0006] Previous research, such as balanced representation learning, aimed to mitigate selection bias by mapping patient features to a latent space to achieve a better balance of covariates between the treatment and control groups. While effective in handling confounding factors, these methods are typically applied to tabular data and do not explicitly model the uncertainty of counterfactual predictions. Other methods utilize generative models, such as generative adversarial networks and variational autoencoders, to infer hidden unobserved variables and estimate the distribution of counterfactual outcomes.
[0007] However, these studies have primarily focused on single-modal data, neglecting the multimodal integration of clinical variables, imaging, and textual reports, which is crucial for oncology decision-making. Combining multimodal fusion with these representation learning techniques promises to enable a more comprehensive and interpretable approach to estimating treatment outcomes. Summary of the Invention
[0008] To overcome the problems existing in the prior art, the present invention provides a multimodal, multi-task probability prediction method and system for neoadjuvant therapy prognosis of breast cancer. This method can integrate various AI-driven pre-neoadjuvant therapy information to perform multiple prognostic predictions for neoadjuvant therapy of breast cancer. At the same time, it can classify patients according to the prediction scores and recommend them to suitable treatment plans to achieve individualized treatment decisions.
[0009] To solve the above-mentioned technical problems and achieve the above-mentioned technical effects, the present invention is implemented through the following technical solution: A multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy in breast cancer includes the following steps: Step 1) Collect breast DCE-MRI images and their paired segmented images from multiple breast cancer patients who have undergone various neoadjuvant therapies (NAT) in the past, as well as clinical information, radiological imaging reports, pathological reports of diagnoses and case information, to form multiple training data samples and construct a training dataset; Step 2) Preprocess the breast DCE-MRI images and their paired segmented images, as well as the radiological imaging reports and diagnostic pathology reports in the training dataset; Step 3) Use the preprocessed training dataset to train the constructed multimodal multitask probabilistic prediction model, so that the multimodal multitask probabilistic prediction model can predict the pathological complete response (pCR), disease-free survival (DFS) and overall survival (OS) of breast cancer patients after various neoadjuvant therapies (NAT), assess the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and identify the types of benign and malignant tumors and the level of Breast Imaging Reporting and Data System (BI-RADS); Step 4) Input the clinical information of the breast cancer patient to be tested, along with preprocessed breast DCE-MRI images or preprocessed breast DCE-MRI images + radiological imaging reports or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images + case information, into the trained multimodal multi-task probability prediction model for multimodal multi-task probability prediction. Finally, generate prediction results for pathological complete remission, disease-free survival, and overall survival for various neoadjuvant therapy regimens, assessment results for the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results for benign and malignant tumor types and breast imaging reports and data system levels. Step 5) Generate a comprehensive report by combining the predicted results of pathological complete remission, disease-free survival, and overall survival of breast cancer patients after neoadjuvant therapy, the assessment results of the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and the identification results of benign and malignant tumor types and breast imaging reports and data systems. The report will be visualized and output for doctors to select the best treatment plan.
[0010] Furthermore, the breast DCE-MRI images are fat-suppressed breast DCE-MRI images, including pre-constrast and post-constrast breast DCE-MRI images. Case information includes the patient's age, ethnicity, and presence of extranodal metastases. Segmented images are breast ROI segmentation images manually annotated by the physician. Clinical information is structured clinical tabular data, including molecular subtype, stage, and genomic risk information.
[0011] Furthermore, for breast DCE-MRI images, the preprocessing procedure includes two steps: acquisition protocol standardization and image normalization.
[0012] First, based on the fat-suppressed pre-constrast and post-constrast breast DCE-MRI sequences in the acquired breast DCE-MRI images, the Seq2Seq method is used to generate domain-unified synthesized pre-constrast and synthesized post-constrast breast DCE-MRI sequences, thereby achieving consistent expression of cross-protocol data and reducing domain shifts introduced by differences in different scanning protocols and enrollment times; For cases where fat suppression is absent in the breast DCE-MRI sequence before and after contrast, plain MRI and b=0, 800, 1000 s / mm were first used. 2 The pre- and post-contrast breast DWI-MRI sequences were used to synthesize corresponding breast DCE-MRI images. Based on this, the TSF-Seq2Seq method was used to generate fat-suppressed pre- and post-contrast breast DCE-MRI sequences. Then, based on the generated fat-suppressed pre- and post-contrast breast DCE-MRI sequences, the Seq2Seq method was used to generate domain-unified synthesized pre- and post-contrast breast DCE-MRI images. Then, the generated domain-unified pre-synthesis contrast and post-synthesis contrast breast DCE-MRI sequences are fused into a dual-channel 3D volume data, and the dual-channel 3D volume data is subjected to a maximum intensity projection to obtain a comprehensive MIP (2D MRI) image. Finally, the generated Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI sequences and their corresponding segmented images were sampled to a voxel resolution of 2×2×2 mm and unified to a voxel resolution of 88×176×176 mm by center cropping or zero padding. Then, the images were normalized to the [0, 1] interval by 99.5% intensity truncation to achieve image normalization.
[0013] Furthermore, the training and testing methods for the multimodal, multi-task probabilistic prediction model are as follows: Step 3.1) The multimodal multitask probabilistic prediction model performs unified feature modeling on multi-source text information from case information, radiological imaging reports and pathology reports through its contained text encoder.
[0014] Step 3.2) The multimodal multitask probabilistic prediction model performs unified feature modeling on breast DCE-MRI images and breast MIP images before and after synthetic contrast through its contained image encoder.
[0015] Step 3.3) The multimodal multi-task probabilistic prediction model extracts features from the structured clinical table data of clinical information through its table encoder, thereby generating corresponding table features as the output feature vector of the table encoder.
[0016] Step 3.4) The multimodal multitask probabilistic prediction model performs image-text semantic alignment on multi-source integrated text features, MRI image features and table features to achieve consistent modeling of multimodal information in the shared semantic space and improve the collaborative expression ability of each modality in subsequent fusion and prediction tasks.
[0017] Step 3.5) The multimodal multitask probabilistic prediction model uses its multimodal fusion unit to splice, fuse, and weight the aligned multi-source integrated text features, MRI image features, and table features to obtain multimodal fused features.
[0018] Step 3.6) The multimodal multitask probabilistic prediction model processes the multimodal fusion features simultaneously through its six feature mapping networks to obtain multiple corresponding task prediction results, which are then used as the task prediction output of the multimodal multitask probabilistic prediction model.
[0019] Step 3.7) Using counterfactual estimation methods, under non-randomized controlled data conditions, the impact of different neoadjuvant therapy regimens on the potential outcomes (pCR and risk assessment) of breast cancer patients is predicted. By comparing the predicted outcomes of breast cancer patients under multiple neoadjuvant therapy regimens, the potential outcomes of patients receiving other treatment regimens are simulated, thereby assessing the individualized efficacy differences of various neoadjuvant therapy regimens for the same breast cancer patient. This testing method provides more objective and robust support for clinical treatment decisions and effectively reduces selection bias caused by non-randomized treatment allocation.
[0020] Step 3.8) Optimize the prediction task by assessing the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and by identifying benign and malignant tumor types and breast imaging reporting and data system levels in the order of predicting pathological complete remission, disease-free survival and overall survival of breast cancer patients after neoadjuvant therapy.
[0021] Furthermore, the loss functions used when training the multimodal multi-task probabilistic prediction model include: a loss function for prediction error and feature distribution constraints, a loss function for predicting pathological complete remission, a loss function for predicting overall survival (OS) and disease-free survival (DFS), a loss function for predicting the probability of breast cancer metastasis to various organs, a loss function for predicting the probability of breast cancer recurrence, a loss function for the classification of benign and malignant tumor types, a loss function for BI-RADS classification, and an alignment loss function for calculating 10 sets of alignment features.
[0022] A multimodal, multi-task probabilistic prediction system for the prognosis of neoadjuvant therapy for breast cancer, comprising: The receiving unit is responsible for receiving the patient's breast DCE-MRI images and their paired segmented images, as well as clinical information (including molecular subtype, stage and genomic risk information), radiological imaging reports, diagnostic pathology reports and case information; The preprocessing unit is responsible for preprocessing the collected breast DCE-MRI images and their paired segmented images, as well as the radiological imaging reports and diagnostic pathology reports; The task prediction unit is responsible for performing multimodal, multi-task probability predictions on clinical information of breast cancer patients and preprocessed breast DCE-MRI images, or preprocessed breast DCE-MRI images + radiological imaging reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images + case information using a multimodal probability prediction model. This results in generating predictions for pathological complete response (pCR), disease-free survival (DFS), and overall survival (OS) for various neoadjuvant therapy regimens, assessments of the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results of benign and malignant tumor types and Breast Imaging Reporting and Data System (BI-RADS) levels. The output unit is responsible for generating a comprehensive report from the pathological complete response (pCR), disease-free survival (DFS), and overall survival (OS) predictions of various neoadjuvant therapy (NAT) regimens generated by the multimodal probabilistic prediction model, the risk probability assessment results of breast cancer metastasis to different organs and recurrence probability assessment results, and the identification results of benign and malignant tumor types and breast imaging reporting and data system (BI-RADS) level, and then visualizing the output.
[0023] Furthermore, the multimodal multitask probabilistic prediction model consists of a text encoder, an image encoder, a table encoder, a multimodal fusion unit, and six independent feature mapping networks. The text encoder is used to perform unified feature modeling on multi-source text information from radiological imaging reports, pathology reports, and case information to generate multi-source integrated text features; An image encoder is used to extract imaging features from synthesized pre-constrast and synthesized post-constrast breast DCE-MRI images and breast MIP (2D MRI) images to generate MRI image features that match the dimensions of multi-source integrated text features. A table encoder is used to extract features from structured clinical table data to generate corresponding table features. A multimodal fusion unit is used to concatenate, fuse, and weight interactively combine aligned multi-source integrated text features, MRI image features, and tabular features to obtain multimodal fused features. The six feature mapping networks are used to process the weighted interactive multimodal fusion features to obtain prediction results of pathological complete remission, disease-free survival and overall survival of breast cancer patients after neoadjuvant therapy, assessment results of the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results of benign and malignant tumor types and breast imaging reporting and data system level. All of these are used as task prediction output results of the multimodal multi-task probability prediction model.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes an interpretable multimodal, multi-task probabilistic prediction model for neoadjuvant therapy prognosis. This model employs multimodal data fusion technology, utilizing not only traditional clinical information such as molecular subtypes, stages, and genomic risk information, but also combining multidimensional data such as breast cancer MRI images, histopathological information, and unstructured radiological and pathological reports. This enables a multidimensional and accurate characterization of breast tumor heterogeneity, significantly improving the accuracy and comprehensiveness of neoadjuvant therapy prognostic assessment.
[0025] 2. This invention uses a multimodal probabilistic modeling approach to predict the response and survival of different treatment methods, explicitly describing the uncertainty of the prediction. This overcomes the shortcomings of existing methods that only output deterministic results and are difficult to interpret the reliability of counterfactual predictions, thus improving the credibility and interpretability of the prediction results.
[0026] 3. This invention can perform counterfactual estimation of the potential outcomes of different neoadjuvant therapy regimens under non-randomized controlled data conditions, thereby providing a more objective and robust basis for comparing the efficacy of different neoadjuvant therapy regimens and making treatment decisions, and effectively reducing the impact of selection bias.
[0027] 4. This invention uses a deep learning model to perform end-to-end feature extraction from clinical data, medical images, and text reports, avoiding the reliance on manually designed features in traditional methods and improving the model's adaptability and generalization ability in complex clinical scenarios.
[0028] 5. Based on the predicted response probability and risk score, this invention can classify patients into different levels of treatment benefit and recommend low-toxicity regimens, high-toxicity regimens or clinical trial registration pathways accordingly, providing actionable individualized decision support for neoadjuvant therapy of breast cancer.
[0029] 6. This invention can integrate various AI-driven pre-neoadjuvant therapy information to predict the prognosis of various neoadjuvant therapies for breast cancer, such as predicting pathological complete response (pCR), disease-free survival (DFS), and overall survival (OS) of breast cancer patients after different neoadjuvant therapies, assessing the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and realizing the classification of benign and malignant tumors and the Breast Imaging Reporting and Data System (BI-RADS) grading, which helps physicians to make individualized treatment decisions.
[0030] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the invention and to implement it according to the contents of the specification, the preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings. Specific embodiments of the present invention are given in detail below with reference to the accompanying drawings. Attached Figure Description
[0031] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating the steps of the multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy for breast cancer according to the present invention.
[0032] Figure 2 This is a schematic diagram of a set of training data, including fat-suppressed DCE-MRI, radiological imaging reports, pathology reports, case information, and clinical information, collected for the multimodal, multi-task probability prediction method for neoadjuvant therapy prognosis of breast cancer of the present invention.
[0033] Figure 3 This is a schematic diagram of the process of synthesizing pre- and post-contrast breast DCE-MRI sequences using a unified generation domain of fat-suppressed pre- and post-contrast breast DCE-MRI sequences in the preprocessing step of the multimodal multi-task probability prediction method for neoadjuvant breast cancer treatment prognosis of the present invention.
[0034] Figure 4 In the preprocessing step of the multimodal, multi-task probability prediction method for neoadjuvant breast cancer treatment prognosis of this invention, plain MRI and b=0, 800, 1000 s / mm are used. 2 A schematic diagram of the process of generating fat-suppressed breast DWI-MRI sequences before and after the comparison of pre-contrast and post-contrast breast DCE-MRI sequences.
[0035] Figure 5 This is a schematic diagram of the text encoder structure in the multimodal, multi-task probabilistic prediction model used in this invention.
[0036] Figure 6This is a schematic diagram of the image encoder structure in the multimodal, multi-task probabilistic prediction model used in this invention.
[0037] Figure 7 This is a schematic diagram of the structure of the self-trained encoder and its Transformer Block and Fusion Block in the image encoder of the multimodal multi-task probabilistic prediction model used in this invention.
[0038] Figure 8 This is a diagram showing the overall architecture of the multimodal, multi-task probabilistic prediction model used in this invention.
[0039] Figure 9 This is a schematic diagram of the similarity matrix generated in the feature alignment step of the multimodal multitask probability prediction method for neoadjuvant breast cancer treatment prognosis according to the present invention. Detailed Implementation
[0040] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings to provide a clearer understanding of the invention's purpose, features, and advantages. It should be understood that the embodiments shown in the drawings are not intended to limit the scope of the invention, but are merely illustrative of the essential spirit of the invention's technical solution.
[0041] See Figure 1 As shown, this invention provides a multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy for breast cancer, comprising the following steps: Step 1) Collection of training data: We collected breast DCE-MRI images and their paired segmented images from multiple breast cancer patients who had undergone various neoadjuvant therapies (NAT) in the past, as well as radiological imaging reports, diagnostic pathology reports, case information, and clinical information (including molecular subtype, stage, and genomic risk information) to form multiple training data samples and construct a training dataset. Furthermore, the breast DCE-MRI images are fat-suppressed breast DCE-MRI images, including pre-constrast and post-constrast breast DCE-MRI images with fat suppression.
[0042] Furthermore, case information includes the patient's age, race, and whether there is a history of extranodal metastasis.
[0043] Furthermore, clinical information consists of structured clinical tabular data, such as classification information from the TNM staging system for tumors.
[0044] Furthermore, the segmented image is a breast ROI (Region of Interest) segmented image obtained by manual annotation by the doctor or other segmentation methods.
[0045] See Figure 2 As shown, a training data sample contains fat-suppressed pre-constrast and post-constrast breast DCE-MRI images, case information, radiological imaging reports, pathology reports, and clinical information for the same patient. The case information included "age 60 years, white race, no extranodal metastasis"; the pathology report stated "ER-, PR+, Her2+, a molecular subtype assessed as Her2 positive"; the imaging report stated "patient has a solitary tumor, not involving the contralateral breast"; the clinical information included the TNM staging system classification information for tumors, which included "T1, N2, M0 and T1, N1, M1". T (Tumor) represents the size and depth of invasion of the primary tumor site, divided into T1-T4, with larger numbers indicating higher tumor volume or degree of invasion; N (Node) refers to the regional lymph node involvement, divided into N0-N3, where N0 indicates no regional lymph node metastasis, and N1-N3 indicates an increased number or extent of metastatic lymph nodes; M (Metastasis) indicates whether the tumor has metastasized to distant sites, divided into M0 (no distant metastasis) and M1 (distant metastasis).
[0046] Step 2) Preprocessing of training data: Preprocessing is performed on breast DCE-MRI images and their paired segmented images, as well as radiological imaging reports and diagnostic pathology reports from the training dataset.
[0047] Furthermore, for breast DCE-MRI images, the preprocessing workflow mainly includes two steps: acquisition protocol standardization and image normalization. Specifically: First, see Figure 3 As shown, based on the fat-suppressed pre-constrast and post-constrast breast DCE-MRI sequences in the acquired breast DCE-MRI images, the Seq2Seq method was used to generate domain-uniform synthesized pre-constrast and synthesized post-constrast breast DCE-MRI sequences, thereby achieving consistent expression of cross-protocol data and reducing domain shifts introduced by differences in different scanning protocols and enrollment times; Then, the generated domain-unified Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI sequences are fused into a dual-channel 3D volume data, and the dual-channel 3D volume data is subjected to a maximum intensity projection to obtain a comprehensive MIP (2D MRI) image. Finally, the generated Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI sequences and their corresponding segmented images were sampled to a voxel resolution of 2×2×2 mm and unified to a voxel resolution of 88×176×176 mm by center cropping or zero padding. Then, the images were normalized to the [0, 1] interval by 99.5% intensity truncation to achieve image normalization.
[0048] Furthermore, in cases where fat-suppressed pre-constrast and post-constrast breast DCE-MRI sequences are missing, see [reference needed]. Figure 4 As shown, plain MRI and b=0, 800, 1000 s / mm were first used. 2 Pre-contrast and post-contrast breast DWI-MRI sequences were used to synthesize corresponding breast DCE-MRI images. Based on these, the TSF-Seq2Seq method was used to generate fat-suppressed pre-contrast and post-contrast breast DCE-MRI sequences. Then, based on the generated fat-suppressed pre-contrast and post-contrast breast DCE-MRI sequences, the Seq2Seq method was used to generate domain-unified synthesized pre-contrast and synthesized post-contrast breast DCE-MRI images.
[0049] Furthermore, for radiological imaging reports and diagnostic pathology reports, the preprocessing workflow mainly includes two steps: word segmentation and padding, to ensure a uniform length of 512 tokens.
[0050] For missing or unavailable radiological imaging reports, diagnostic pathology reports, case information, and clinical information, blank text is used as a substitute.
[0051] Step 3) Training and testing of the multimodal probabilistic prediction model: Using a preprocessed training dataset, the constructed multimodal multitask probabilistic prediction model is trained, enabling it to predict pathological complete response (pCR), disease-free survival (DFS), and overall survival (OS) in breast cancer patients after various neoadjuvant therapies (NAT), assess the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and identify benign and malignant tumor types and Breast Imaging Reporting and Data System (BI-RADS) level.
[0052] The neoadjuvant therapy (NAT) regimens include: anthracycline-based regimens, anthracycline plus taxane-based regimens, single HER2 antibody regimens, multiple HER2 antibody regimens, and hormone therapy-based regimens.
[0053] Furthermore, the training and testing methods for the multimodal, multi-task probabilistic prediction model are as follows: Step 3.1) The multimodal multitask probabilistic prediction model performs unified feature modeling on multi-source text information from case information, radiological imaging reports and pathology reports through its contained text encoder.
[0054] Specifically, firstly, the text encoder performs context modeling and semantic feature extraction on radiological image reports, pathology reports, and case information respectively through its three independent text feature extraction pathways, thereby obtaining the corresponding intermediate text features of radiological image reports, pathology reports, and case information. Subsequently, the text encoder concatenates the obtained intermediate text features of radiological image reports, pathology reports, and case information in terms of feature dimensions through its connection layer, thereby forming a joint text feature. Finally, the text encoder performs cross-source information interaction and feature fusion on the formed joint text feature through a text feature fusion module, ultimately obtaining a multi-source comprehensive text feature of size [B, 1024]. f txt This is used as the output feature vector of the text encoder, which also retains the radiological image report features. f t-Rad and pathology report characteristics f t-Path .
[0055] Step 3.2) The multimodal multitask probabilistic prediction model performs unified feature modeling on Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI images and breast MIP images through its contained image encoder, and completes image-text semantic alignment.
[0056] Specifically, firstly, the image encoder simultaneously extracts features from synthesized pre-constrast and synthesized post-constrast breast DCE-MRI images and breast MIP images using three independent image sub-encoders: a self-trained encoder, a general medical large model (MedSAM), and an MRI visual large model (MRI-CORE), thus obtaining three intermediate image features. Then, the image encoder concatenates the intermediate image features from the three image sub-encoders through a connection layer to form joint image features. Finally, the image encoder performs feature mapping on the formed joint image features through its contained TransformerBlock, matching the dimension of the joint image features with the dimension of the text encoder, ultimately obtaining MRI image features of size [B, 1024]. f img As the output feature vector of the image encoder, this output feature vector retains the imaging features of breast DCE-MRI images. f i-MRI and breast MIP imaging features f i-MIP .
[0057] The method for extracting image features using the self-trained encoder is as follows: 1) When the input data is 3D data from Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI images, the Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI images are first stitched together along the channel dimension to obtain a stitched breast DCE-MRI image. This stitched breast DCE-MRI image, along with its paired segmented image, is then used as the input image for the self-trained encoder. After the input image enters the self-trained encoder, the first convolutional layer performs a convolution operation on the stitched breast DCE-MRI image to achieve block processing of the original 3D image, thereby dividing the stitched breast DCE-MRI image into 1331 16×16×16 3D image patches. Subsequently, each 3D image patch is flattened, and the flattened 3D image patches are mapped to a unified feature space through the first linear mapping layer, so that the feature dimension is mapped from the original dimension to [B, 4096, ...].
[1024] , thereby obtaining first-level image features; simultaneously, feature extraction is performed on the segmented image through the second convolutional layer and the second linear mapping layer to obtain first-level local features; then, the first-level image features and the first-level local features are fused through the first Fusion Block to obtain first-level fused features, and the first-level fused features are updated and enhanced through the first Transformer Block to obtain second-level image features; on this basis, the first-level local features are extracted through the first local feature extraction path to obtain second-level local features; then, the second-level image features and the second-level local features are fused through the second Fusion Block to obtain second-level fused features, and the second-level fused features are updated and enhanced through the second Transformer Block to obtain third-level image features; and so on, through 3 layers of local feature extraction paths and 4 corresponding Fusion Blocks and Transformer Blocks, the "local feature extraction-fusion feature generation-image feature extraction" process is repeated 4 times to achieve deep contextual modeling and feature extraction of features, so as to enhance the expressive ability of global dependencies and improve the discriminative ability and robustness of features layer by layer, finally obtaining fifth-level image features;Finally, the feature dimensions of the five-level image features are compressed through the fifth and sixth Transformer Blocks, combined with the average pooling operation of an average pooling module, resulting in a DCE-MRI image feature of size [B, 1024]. This feature serves as the output feature vector of the self-trained encoder, representing the intermediate features between the synthesized pre-constrast and synthesized post-constrast breast DCE-MRI images.
[0058] 2) When the input data is 2D data from a breast MIP image, the breast MIP image is directly used as the input image of the self-trained encoder. After the input image enters the self-trained encoder, the breast MIP image is first processed by a convolutional layer to divide the original 2D image into 121 16×16 2D image blocks. Then, each 2D image block is flattened and mapped to a unified feature space by a first linear mapping layer, so that the feature dimension is mapped from the original dimension to [B, 4096, 1024], thereby obtaining the first-level image features. Subsequently, the first-level image features are updated and enhanced four times by the first, second, third, and fourth Transformer Blocks to improve the discriminative ability and robustness of the features layer by layer, thereby obtaining the fifth-level image features. Finally, the features are processed by the fifth and sixth Transformer Blocks, combined with an average pooling module. The pooling operation compresses the feature dimensions of the five-level image features, thereby obtaining MIP image features of size [B, 1024], which serve as the output feature vector of the self-trained encoder, i.e., the intermediate image features of the breast MIP image.
[0059] Step 3.3) The multimodal, multi-task probabilistic prediction model extracts features from the structured clinical tabular data of clinical information through its included tabular encoder, thereby generating corresponding tabular features. f tab , which serves as the output feature vector of the table encoder.
[0060] Step 3.4) The multimodal, multi-task probabilistic prediction model performs image-text semantic alignment on multi-source integrated text features, MRI image features, and table features; Specifically, the extracted radiological image report features are first... f t-Rad Characteristics of pathology reports f t-Path Imaging features of breast DCE-MRI imagesf i-MRI Breast MIP imaging features f i-MIP and table features f tab Alignment is performed by combining pairs of features, resulting in 10 sets of alignment features, including... fi-MRI & fi-MIP , fi-MRI & ft-Rad , fi-MRI & ft-Path , fi- MRI & ftab , fi-MIP & ft-Rad , fi-MIP & ft-Path , fi-MIP & ftab , ft-Rad & ft-Path , ft-Rad & ftab , ft-Path&ftab .
[0061] Then, the 10 sets of alignment features were normalized to eliminate the influence of feature amplitude differences on similarity calculation. The specific formula for normalization is as follows: (1); In equation (1), fk and fv These represent two features from each group of alignment features before normalization. The 2-norm of the eigenvectors, and These represent two features in each group of aligned features after normalization.
[0062] After normalization, the cosine similarity between any two features in each aligned feature set is calculated within the shared feature space to characterize their semantic matching degree. The formula for calculating cosine similarity is as follows: (2); In equation (2), Skv This represents the semantic similarity between two features in each group of aligned features after normalization. and ...
[0063] See Figure 9As shown, the similarity matrix for feature alignment is constructed through the above cosine similarity calculation, so that semantically related image features and text features have high similarity values in the feature space, while semantically unrelated feature pairs have low similarity values.
[0064] Finally, the InfoNCE loss function is used as the alignment loss, serving as an alignment constraint across modalities or feature spaces. During the training phase, the semantic alignment relationship between two features in each pair of aligned features is constrained and optimized to enhance the consistency and discriminability of different modal features in the common representation space.
[0065] Specifically, within the same training batch, feature representations from different modalities but corresponding to the same object or case are grouped into positive sample pairs, while feature representations from different objects or cases are grouped into negative sample pairs. By calculating the similarity between the two feature representations in each aligned feature pair, and exponentially amplifying the similarity of positive samples while normalizing and suppressing the similarity of negative samples, the model maximizes the similarity of positive sample pairs during training while minimizing the similarity between them and negative sample pairs. This can be expressed as: (14); Based on this, the InfoNCE loss function is defined as: (15); In equation (15), τ is the smoothing coefficient, used to adjust the smoothness of the similarity distribution, s() represents the feature similarity calculation function, N represents the length of the feature vector, and exp() represents the exponential function. UI , vi These represent two different feature vectors.
[0066] For the overall alignment loss function, this invention uses the following process to calculate the loss.
[0067] For the overall alignment loss function under 10 different alignment features, this invention performs weighted summation based on their direct correlation. Specifically, it identifies the DCE-MRI image with the strongest correlation to the radiological image report, and assigns [the appropriate feature]. fi- MRI & ft-Rad The highest weight is 1, followed by DCE-MRI images and pathology reports, which are assigned... fi-MRI & ft-Path Weighted by α, followed by DCE-MRI images and clinical information, assigned... fi-MRI & ftab The weight is β, and finally, for other losses, they are all assigned... fi-MRI & fi-MIP , fi-MIP & ft-Rad , fi-MIP & ft-Path , fi-MIP & ftab , ft-Rad & ft-Path , ft-Rad & ftab , ft-Path&ftab The weights γ exist such that 1 > α > β > γ > 0. Here, these three values are set to 0.75, 0.55, and 0.35 in this invention. The overall alignment loss function is: (16).
[0068] After 50 rounds of training, the change in the overall alignment loss of all 10 alignment features is evaluated. If the reduction in the overall alignment loss is less than 5% after 5 rounds of training, the weights are adjusted to enhance the alignment effect. Specifically, the weight coefficients are strengthened to obtain 1≥α>β>γ>0.25, thereby strengthening the alignment constraints on different features by enhancing the weight coefficients.
[0069] For the above 10 sets of alignment features, the present invention uses the above method to perform similarity calculation and semantic alignment, thereby realizing consistent modeling of multimodal information in the shared semantic space and improving the collaborative expression ability of each modality in subsequent fusion and prediction tasks.
[0070] Step 3.5) The multimodal multitask probabilistic prediction model uses its multimodal fusion unit to splice, fuse, and weight the aligned multi-source integrated text features, MRI image features, and table features to obtain multimodal fused features.
[0071] Specifically, when the aligned multi-source integrated text features, MRI image features, and table features are input into the multimodal fusion unit, a concatenation layer first concatenates the aligned multi-source integrated text features, MRI image features, and table features along the feature dimension to form a unified multimodal feature. Then, a Transformer submodule models the global dependencies between different modal features to form a multimodal fused feature. Finally, a linear layer and a Softmax layer are used to calculate the weights of the multimodal fused feature to obtain the importance coefficients of each modality feature, and a weighted interaction is performed with the multimodal fused feature to strengthen important feature information.
[0072] Step 3.6) The multimodal multitask probabilistic prediction model processes the weighted interactive multimodal fusion features simultaneously through its six feature mapping networks to obtain multiple corresponding task prediction results, which are then used as the task prediction output of the multimodal multitask probabilistic prediction model.
[0073] Specifically, the weighted and interactive multimodal fusion features are input into the first, second, third, fourth, fifth, and sixth feature mapping networks, respectively. The first, second, third, fourth, fifth, and sixth feature mapping networks are optimized by backpropagation based on the true labels of their respective tasks and the corresponding task loss functions, extracting information related to their respective tasks and completing the final mapping. Ultimately, this enables the prediction of pCR, DFS, and OS of breast cancer patients after NAT, the assessment of the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and the identification of benign and malignant tumor types and BI-RADS levels.
[0074] Each prediction task has a corresponding task loss function. After backpropagation, the predicted value of each prediction task will be aligned with the true label based on its true label and its own task loss function, thereby achieving the prediction effect.
[0075] 1) Loss function for prediction error and feature distribution constraints: This invention considers both prediction error and feature distribution constraints during model training. To improve prediction accuracy while maintaining the stability and consistency of feature representation, the Evidence Lower Bound (ELBO) of the response model is used as a loss function to optimize the model, and its definition is as follows: (3); In equation (3), Represents distribution and Kullback–Leibler (KL) divergence between them; λ KL The balance coefficient in ELBO is set to 1 in this invention.
[0076] Therefore, the loss function for the prediction error and the feature distribution constraint can be defined as: (4); In equation (4), λKL This refers to the balance coefficient in ELBO. Ztab , Ztxt and Zimg The latent variable is a Gaussian distribution obtained through the features of its respective encoder. Specifically, the mean and average difference are calculated using the feature map vectors encoded by the respective encoders, and then the Gaussian distribution is obtained from these two parameters, as shown in the following formula: (5); In equation (5), xm Indicates the first Modal input features, μm and σ2mThese represent the mean vector and variance parameter, respectively, calculated from the output features of the corresponding modal encoder, used to parameterize latent variables. Zm The posterior distribution.
[0077] 2) Loss function for predicting pathological complete remission (pCR): For the task of predicting pathological complete remission (pCR), this invention employs a cross-entropy loss function, the form of which is as follows: (6); In equation (6), y pcr The true label is the pCR state, and the label is not the pCR state. y pcr When pCR is 0 y pcr =1, This represents the pCR score predicted by the response model.
[0078] 3) Loss functions for predicting overall survival (OS) and disease-free survival (DFS): For the overall survival (OS) prediction task, the Cox partial likelihood function is used as the loss function, which is defined as follows: (7); In equation (7), xi Represents the features of the i-th sample. In response to the risk score predicted by the model, E i This indicates the follow-up status of patient i (1 indicates death, 0 indicates survival). T i This indicates the survival time of patient i.
[0079] 4) Loss function for predicting the probability of breast cancer metastasizing to various organs: To predict the risk of breast cancer metastasis to various organs, this invention uses a binary cross-entropy loss function and a Focal Loss loss function for loss calculation.
[0080] The binary cross-entropy loss function has the following formula: (8); In equation (8), k represents the number of organs. The dataset contains 5 body parts. The predicted probability value. This is a real label.
[0081] The loss function for Focal Loss is given by the following formula: (9); In equation (9), αk γ is a balancing factor, and γ > 0 is a hyperparameter that adjusts the weights of easy / difficult samples.
[0082] By optimizing the loss function, the model can accurately predict the probability of breast cancer metastasizing to various organs and effectively address the problem of sample imbalance, thereby improving the reliability of clinical risk assessment.
[0083] Therefore, the loss function for the probability of breast cancer metastasizing to various organs can be obtained as the sum of the two, defined as: (10).
[0084] 5) Loss function for predicting the probability of breast cancer recurrence: For predicting the risk of breast cancer recurrence, this invention employs a binary cross-entropy loss function, assuming the recurrence event label is... y rec ∈{0,1}, the model predicts the recurrence probability as p rec The loss function for predicting the probability of breast cancer recurrence is defined as follows: (11); By incorporating this loss function, the model can effectively learn the probability distribution of breast cancer recurrence risk, achieve accurate prediction of recurrence events, and provide reliable assistance for clinical follow-up and treatment decisions.
[0085] 6) Loss function for grading benign and malignant tumor types: For the task of identifying benign / malignant tumors in breast imaging, this invention uses a binary cross-entropy loss function to optimize the model. Let the true tumor label be... y b / m ∈{0,1}, where 0 and 1 represent benign and malignant tumors, respectively. The loss function for predicting the type of benign or malignant tumor classification is defined as follows: (12); By combining this loss function, the model can effectively distinguish between benign and malignant tumors, achieve accurate identification of tumor characteristics, and provide a reference for clinical diagnosis and treatment.
[0086] 7) Loss function for the BI-RADS grading task: For the BI-RADS classification task in breast imaging, this invention uses multi-class cross-entropy loss for network optimization. Given a dataset with 7 levels, the loss function for the BI-RADS classification task is defined as follows: (13); In equation (13), 1{y =c} is an indicator function that takes the value 1 when the actual level is c, and 0 otherwise. p c Predict the probability of the model belonging to level c.
[0087] Step 3.7) Using counterfactual estimation methods, under non-randomized controlled data conditions, the impact of different neoadjuvant therapy regimens on the potential outcomes (pCR and risk assessment) of breast cancer patients is predicted. By comparing the predicted outcomes of breast cancer patients under multiple neoadjuvant therapy regimens, the potential outcomes of patients receiving other treatment regimens are simulated, thereby assessing the individualized efficacy differences of various neoadjuvant therapy regimens for the same breast cancer patient. This testing method provides more objective and robust support for clinical treatment decisions and effectively reduces selection bias caused by non-randomized treatment allocation.
[0088] Step 3.8) Optimize the prediction task by assessing the risk probability of breast cancer metastasis to different organs and the probability of recurrence according to the prediction of pCR, DFS and OS of various post-NAT breast cancer patients, and the order of identifying benign and malignant tumor types and BI-RADS levels.
[0089] Step 4) Input the clinical information of the breast cancer patient to be tested, along with preprocessed breast DCE-MRI images, or preprocessed breast DCE-MRI images + radiological imaging reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images + case information, into the trained multimodal multi-task probability prediction model for multimodal multi-task probability prediction. Finally, generate pCR, DFS, and OS prediction results for various NAT schemes, assessment results of the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results of benign and malignant tumor types and BI-RADS levels.
[0090] Step 5) Generate a comprehensive report by combining the predicted pCR, DFS, and OS of various breast cancer patients after NAT, the assessment results of the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and the identification results of benign and malignant tumor types and BI-RADS levels. The report will be visualized and output to help physicians select the best treatment plan.
[0091] See Figure 8As shown in one embodiment of the present invention, the predicted pCR and risk of metastasis to different organs for five different types of post-NAT breast cancer patients—Ab (anthracycline-based treatment), ApTb (anthracycline combined with taxanes), S-HER2-a (single HER2-targeting antibody treatment), M-HER2-a (multiple HER2-targeting antibody treatment), and Ht (endocrine therapy-based treatment)—are as follows: Ab's pCR prediction result is 28%, and its risk prediction result is 54%; ApTb's pCR prediction result is 35%, and its risk prediction result is 56%; S-HER2-a's pCR prediction result is 53%, and its risk prediction result is 22%; M-HER2-a's pCR prediction result is 68%, and its risk prediction result is 13%; and Ht's pCR prediction result is 12%, and its risk prediction result is 77%.
[0092] Through the above technical solutions, this invention maintains the stability of semantic alignment of multi-feature information while introducing multi-task collaboration to achieve deep fusion of multimodal and multi-task functions, effectively improving the accuracy and clinical applicability of pCR prediction, survival prediction, risk metastasis prediction, tumor type identification, and BI-RADS level identification after neoadjuvant therapy.
[0093] This invention provides a multimodal, multi-task probabilistic prediction system for the prognosis of neoadjuvant therapy for breast cancer, the structure of which includes: The receiving unit is responsible for receiving the patient's breast DCE-MRI images and their paired segmented images, as well as clinical information (including molecular subtype, stage, and genomic risk information), radiological imaging reports, diagnostic pathology reports, and case information.
[0094] The preprocessing unit is responsible for preprocessing the collected breast DCE-MRI images and their paired segmented images, as well as the radiological imaging reports and diagnostic pathology reports. For preprocessing methods, please refer to the data preprocessing steps in the multimodal probabilistic prediction model training and testing procedures.
[0095] The task prediction unit is responsible for performing multimodal, multi-task probability predictions on clinical information of breast cancer patients and preprocessed breast DCE-MRI images, or preprocessed breast DCE-MRI images + radiological imaging reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images + case information using a multimodal probability prediction model. This results in generating predictions for pathological complete response (pCR), disease-free survival (DFS), and overall survival (OS) for various neoadjuvant therapy regimens, assessments of the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results for benign and malignant tumor types and Breast Imaging Reporting and Data System (BI-RADS) levels.
[0096] The output unit is responsible for generating a comprehensive report from the pCR, DFS, and OS prediction results of various NAT schemes generated by the multimodal probabilistic prediction model, the risk probability assessment results of breast cancer metastasis to different organs and recurrence probability assessment results, and the identification results of benign and malignant tumor types and BI-RADS levels, and then visualizing the output.
[0097] Further, see Figure 8 As shown, this invention specifically developed a multimodal multitask probabilistic prediction model, the structure of which includes a text encoder, an image encoder, a table encoder, a multimodal fusion unit, and six independent feature mapping networks.
[0098] The text encoder is used to perform unified feature modeling on multi-source text information from radiological imaging reports, diagnostic pathology reports, and case information to generate multi-source integrated text features.
[0099] The image encoder is used to extract imaging features from Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI images and breast MIP images to generate MRI image features that match the dimensions of multi-source integrated text features.
[0100] The table encoder is used to extract features from structured clinical table data of clinical information to generate corresponding table features.
[0101] The multimodal fusion unit is used to concatenate, fuse, and weight the aligned multi-source integrated text features, MRI image features, and table features to obtain multimodal fused features.
[0102] The six feature mapping networks are used to process the weighted interactive multimodal fusion features to obtain the prediction results of pCR, DFS and OS of breast cancer patients after NAT, the assessment results of the risk probability of breast cancer metastasis to different organs and the recurrence probability, and the identification results of benign and malignant tumor types and BI-RADS levels. All of these are used as the task prediction output results of the multimodal multi-task probability prediction model.
[0103] See Figure 5 As shown, the text encoder consists of three independent text feature extraction paths, a splicing layer, and a text feature fusion module.
[0104] The three text feature extraction pathways are each composed of two Transformer sub-modules and corresponding Layer Normalization and ReLU modules, which are used to perform context modeling and semantic feature extraction on radiological image reports, pathological reports, and case information, respectively, to obtain the corresponding intermediate features of radiological image report text, intermediate features of pathological report text, and intermediate features of case information text.
[0105] The splicing layer is used to splice the intermediate features of radiological image report text, pathology report text, and case information text from the three text feature extraction pathways in terms of feature dimensions to form a joint text feature.
[0106] The text feature fusion module consists of a Transformer submodule and its corresponding LayerNormalization and ReLU. It is responsible for performing cross-text source information interaction and feature fusion on the formed joint text features, thereby obtaining multi-source comprehensive text features of size [B, 1024]. f txt This serves as the output feature vector of the text encoder, while simultaneously preserving radiological image report features. f t-Rad and pathology report characteristics f t-Path .
[0107] See Figure 6 As shown, the image encoder consists of three independent image sub-encoders: a self-trained encoder, a general medical large model (MedSAM), and an MRI visual large model (MRI-CORE), as well as a connection layer and a Transformer Block.
[0108] In the image encoder, the self-trained encoder, MedSAM, and MRI-CORE are all responsible for receiving 3D data from synthesized pre-constrast and synthesized post-constrast breast DCE-MRI images and 2D data from breast MIP images, and extracting radiographic features to obtain corresponding intermediate image features. Three image sub-encoders are used to better extract radiographic features from synthesized pre-constrast and synthesized post-constrast breast DCE-MRI images and breast MIP images. The connection layer is responsible for concatenating the three intermediate image features from the self-trained encoder, MedSAM, and MRI-CORE to obtain joint image features. The Transformer Block is responsible for feature mapping the joint image features output from the connection layer, matching its dimensions with the dimensions of the text encoder to obtain MRI image features. f img As the output feature vector of the image encoder, which preserves the imaging features of breast DCE-MRI images. f i-MRI and breast MIP imaging features f i-MIP .
[0109] See Figure 7 As shown, the self-trained encoder includes two convolutional layers, two linear mapping layers, four feature blocks, six Transformer blocks, one average pooling module, and three local feature extraction pathways. The self-trained encoder employs a Transformer-based network architecture, which can reduce the feature gap between different modalities or different feature spaces.
[0110] The first convolutional layer is responsible for performing convolution operations on the breast DCE-MRI stitched image and the breast MIP image obtained by stitching the Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI images in the channel dimension. This achieves block processing of the original image, thereby dividing the breast DCE-MRI stitched image and the breast MIP image into 1331 three-dimensional image patches of size 16×16×16 and 121 two-dimensional image patches of size 16×16.
[0111] The first linear mapping layer maps the flattened 3D and 2D image patches to a unified feature space, mapping their feature dimensions from the original dimensions to [B, 4096, 1024] to obtain first-level image features. The second convolutional layer and the second linear mapping layer extract features from the segmented image to obtain first-level local features. The first fusion block fuses the first-level image features with the first-level local features to obtain first-level fused features.
[0112] The first Transformer Block is responsible for updating and enhancing the primary fused features or primary image features to obtain secondary image features. The first local feature extraction path is responsible for extracting features from the primary local features to obtain secondary local features. The second Fusion Block is responsible for fusing the secondary image features with the secondary local features to obtain secondary fused features.
[0113] The second Transformer Block is responsible for updating and enhancing the secondary fused features or secondary image features to obtain tertiary image features. The second local feature extraction path is responsible for extracting features from the secondary local features to obtain tertiary local features. The third Fusion Block is responsible for fusing the tertiary image features with the tertiary local features to obtain tertiary fused features.
[0114] The third Transformer Block is responsible for updating and enhancing the three-level fused features or three-level image features to obtain four-level image features. The third local feature extraction path is responsible for extracting features from the three-level local features to obtain four-level local features.
[0115] The fourth Fusion Block is responsible for fusing fourth-level image features with fourth-level local features to obtain fourth-level fused features. The fourth Transformer Block is responsible for updating and enhancing the fourth-level fused features or fourth-level image features to obtain fifth-level image features.
[0116] The fifth and sixth Transformer Blocks and the average pooling module are responsible for updating and enhancing the five-level image features again, and compressing their dimensions through the average pooling operation to obtain DCE-MRI image features and MIP (2D MRI) image features of size [B, 1024].
[0117] See Figure 7As shown, in the self-trained encoder, each of the six Transformer Blocks consists of four Transformer sub-modules, Layer Normalization, and ReLU. During feature update and enhancement, the fused features or image features from the previous level are first input to the first Transformer sub-module for preliminary feature modeling to obtain the first intermediate features. Then, the first intermediate features are sequentially input to the second and third Transformer sub-modules for processing, followed by normalization by the normalization module and activation by the nonlinear activation module to obtain the second intermediate features. Subsequently, the second intermediate features are fused with the first intermediate features output from the first Transformer sub-module using a residual connection to obtain residual features. Finally, the residual features are input to the fourth Transformer sub-module to obtain the next level of image features, thus completing the feature calculation and output of the corresponding Transformer Block. In the self-trained encoder, each of the four Fusion Blocks consists of four linear mapping layers and a Softmax normalization module. During feature fusion, the previous level image features of the DCE-MRI image are first mapped through the first, second, and third linear mapping layers, generating query vectors Q. MRI Key vector K MRI and value vector V MRI Simultaneously, the fourth linear mapping layer is used to perform feature mapping on the local features extracted from the previous level based on the ROI, resulting in the mask feature vector A. mask ; then the query vector Q MRI With key vector K MRI Perform matrix multiplication to obtain the feature correlation matrix, and then combine this correlation matrix with the adaptively adjusted mask features (λS×A). mask The weighted results of the summation are then superimposed; the superimposed weighted results are then normalized using the Softmax normalization module to generate an attention weight map; finally, the attention weight map is compared with the value vector V. MRI Element-wise dot product or matrix multiplication operations are performed to obtain the next level of image features, thereby achieving effective fusion and output of the DCE-MRI image features (image features) of the corresponding Fusion Block and the prior information of ROI segmentation (local features). Each of the three local feature extraction pathways consists of a Transformer submodule and its corresponding Layer Normalization and ReLU, responsible for extracting features from the previous level's local features to output the next level's local features.
[0118] The table encoder consists of two Transformer sub-modules and their corresponding Layer Normalization and ReLU, and is responsible for extracting features from the input structured clinical table data to generate the corresponding table feature sequence.
[0119] See Figure 8 As shown, the multimodal fusion unit consists of a stitching layer, a Transformer submodule, a linear layer, and a Softmax layer.
[0120] In the multimodal fusion processor, the stitching layer is responsible for stitching together the aligned multi-source integrated text features, MRI image features, and tabular features along the feature dimension to form a unified multimodal feature. The Transformer submodule is responsible for modeling the global dependencies between different modal features to form multimodal fusion features. The linear layer and Softmax layer are responsible for calculating the weights of the multimodal fusion features to obtain the importance coefficients of each modality feature, and for weighted interaction with the multimodal fusion features to strengthen important feature information.
[0121] To adapt to the needs of different tasks, this invention configures separate feature mapping networks for each task. Although all feature mapping networks receive the same multimodal fusion features as input, they can be optimized through backpropagation based on the true labels of their respective tasks and the corresponding loss functions. This allows the six feature mapping networks to perform nonlinear transformations on the multimodal fusion features through their respective Transformer sub-modules and their corresponding Layer Normalization and ReLU, thereby extracting information related to their respective tasks and completing the final mapping. This results in predictions of pCR, DFS, and OS for breast cancer patients after NAT, assessments of the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and identification results of benign and malignant tumor types and BI-RADS levels, thus enabling prediction and output for different tasks.
[0122] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy for breast cancer, characterized in that, Includes the following steps: Step 1) Collect breast DCE-MRI images and their paired segmented images from multiple breast cancer patients who have undergone various neoadjuvant therapies in the past, as well as clinical information, radiological imaging reports, pathological reports of diagnoses and case information, to form multiple training data samples and construct a training dataset; Step 2) Preprocess the breast DCE-MRI images and their paired segmented images, as well as the radiological imaging reports and diagnostic pathology reports in the training dataset; Step 3) Use the preprocessed training dataset to train and test the constructed multimodal multitask probabilistic prediction model, so that the multimodal multitask probabilistic prediction model can predict the pathological complete remission, disease-free survival and overall survival of breast cancer patients after various neoadjuvant therapies, assess the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and identify the types of benign and malignant tumors and the level of breast imaging reporting and data systems. Step 4) Input the clinical information of the breast cancer patient to be tested, along with preprocessed breast DCE-MRI images or preprocessed breast DCE-MRI images + radiological imaging reports or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images + case information, into the trained multimodal multi-task probability prediction model for multimodal multi-task probability prediction. Finally, generate prediction results for pathological complete remission, disease-free survival, and overall survival for various neoadjuvant therapy regimens, assessment results for the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results for benign and malignant tumor types and breast imaging reports and data system levels. Step 5) Generate a comprehensive report from all the predictions, evaluations and identification results generated by the multimodal multi-task probabilistic prediction model and output it in a visual format.
2. The multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy for breast cancer according to claim 1, characterized in that: The breast DCE-MRI images are fat-suppressed breast DCE-MRI images, including pre-contrast and post-contrast fat-suppressed breast DCE-MRI images; the segmented images are breast ROI segmented images manually annotated by the physician; the case information includes the patient's age, ethnicity, and whether there is a record of extranodal metastasis; the clinical information is structured clinical tabular data, including molecular subtype, stage, and genomic risk information.
3. The multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy for breast cancer according to claim 1, characterized in that, For breast DCE-MRI images, the preprocessing procedure includes two steps: acquisition protocol standardization and image normalization, as detailed below: First, based on the pre- and post-contrast breast DCE-MRI sequences with fat suppression in the acquired breast DCE-MRI images, the Seq2Seq method is used to generate domain-unified synthetic pre- and post-contrast breast DCE-MRI sequences, thereby achieving consistent expression of cross-protocol data and reducing domain shifts introduced by differences in different scanning protocols and enrollment times. For cases where fat suppression is absent in the breast DCE-MRI sequence before and after contrast, plain MRI and b=0, 800, 1000 s / mm were first used. 2 The pre- and post-contrast breast DWI-MRI sequences were used to synthesize corresponding breast DCE-MRI images. Based on this, the TSF-Seq2Seq method was used to generate fat-suppressed pre- and post-contrast breast DCE-MRI sequences. Then, based on the generated fat-suppressed pre- and post-contrast breast DCE-MRI sequences, the Seq2Seq method was used to generate domain-unified synthesized pre- and post-contrast breast DCE-MRI images. Then, the generated domain-unified pre-synthesis contrast and post-synthesis contrast breast DCE-MRI sequences are fused into a dual-channel 3D volume data, and the dual-channel 3D volume data is subjected to a maximum intensity projection to obtain a comprehensive MIP image. Finally, the generated Synthesized Pre-constrast and Synthesized Post-constrast breast DCE-MRI sequences and their corresponding segmented images were sampled to a voxel resolution of 2×2×2 mm and unified to a voxel resolution of 88×176×176 mm by center cropping or zero padding. Then, the images were normalized to the [0, 1] interval by 99.5% intensity truncation to achieve image normalization.
4. The multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy for breast cancer according to claim 1, characterized in that, The training and testing methods for multimodal, multi-task probabilistic prediction models are as follows: Step 3.1) The multimodal multitask probabilistic prediction model performs unified feature modeling on multi-source text information from case information, radiological imaging reports and pathology reports through its contained text encoder; Specifically, firstly, the text encoder performs context modeling and semantic feature extraction on radiological imaging reports, pathology reports, and case information through three independent text feature extraction pathways, respectively, to obtain the corresponding intermediate text features of radiological imaging reports, pathology reports, and case information. Subsequently, the text encoder concatenates these intermediate text features along the feature dimension through a connection layer, forming a joint text feature. Finally, the text encoder uses a text feature fusion module to perform cross-source information interaction and feature fusion on the formed joint text feature, ultimately obtaining a multi-source comprehensive text feature of size [B, 1024], which serves as the output feature vector of the text encoder. This output feature vector also includes the radiological imaging report features. f t-Rad and pathology report characteristics f t-Path ; Step 3.2) The multimodal multitask probabilistic prediction model performs unified feature modeling on breast DCE-MRI images and breast MIP images before and after synthetic contrast through its contained image encoder; Specifically, firstly, the image encoder simultaneously extracts features from both pre- and post-synthetic contrast-enhanced breast DCE-MRI and breast MIP images using three independent image sub-encoders: a self-trained encoder, a general medical large model, and an MRI visual large model. This yields three corresponding intermediate image features. Then, the image encoder concatenates these intermediate features from the three sub-encoders using a connection layer, forming a joint image feature. Finally, the image encoder performs feature mapping on the formed joint image feature using its Transformer Block, matching the dimension of the joint image feature with that of the text encoder. This results in an MRI image feature of size [B, 1024], which serves as the output feature vector of the image encoder. , The output feature vector contains imaging features from breast DCE-MRI images. f i-MRI and breast MIP imaging features f i-MIP ; Step 3.3) The multimodal, multi-task probabilistic prediction model extracts features from the structured clinical tabular data of clinical information through its included tabular encoder, thereby generating corresponding tabular features. f tab , which serves as the output feature vector of the table encoder; Step 3.4) The multimodal multitask probabilistic prediction model performs image-text semantic alignment on multi-source integrated text features, MRI image features and table features to achieve consistent modeling of multimodal information in the shared semantic space and improve the collaborative expression ability of each modality in subsequent fusion and prediction tasks. Specifically, the extracted radiological image report features are first... f t-Rad Characteristics of pathology reports f t-Path Imaging features of breast DCE-MRI images f i-MRI Breast MIP imaging features f i-MIP and table features f tab Alignment is performed by combining pairs of features, resulting in 10 sets of alignment features, including... f i-MRI & f i-MIP , f i-MRI & f t-Rad , f i-MRI & f t-Path , f i-MRI & f tab , f i-MIP & f t-Rad , f i-MIP & f t-Path , f i-MIP & f tab , f t-Rad &f t-Path , f t-Rad & f tab , f t-Path &f tab Then, the 10 sets of alignment features were normalized to eliminate the influence of feature amplitude differences on similarity calculation. Next, the cosine similarity between the two features in each set of alignment features was calculated in the shared feature space to characterize the degree of matching between them at the semantic level. Finally, the InfoNCE loss function was used as the alignment loss as an alignment constraint across modalities or across feature spaces. During the training phase, the semantic alignment relationship between the two features in each set of alignment features was constrained and optimized to enhance the consistency and discriminability of features of different modalities in the common representation space. Step 3.5) The multimodal multitask probabilistic prediction model uses its multimodal fusion unit to concatenate, fuse, and weight the aligned multi-source integrated text features, MRI image features, and table features to obtain multimodal fused features; Specifically, when the aligned multi-source integrated text features, MRI image features, and table features are input into the multimodal fusion unit, a concatenation layer first concatenates the aligned multi-source integrated text features, MRI image features, and table features along the feature dimension to form a unified multimodal feature. Then, a Transformer submodule models the global dependencies between different modal features to form a multimodal fusion feature. Finally, a linear layer and a Softmax layer are used to calculate the weights of the multimodal fusion feature to obtain the importance coefficients of each modality feature, and a weighted interaction is performed with the multimodal fusion feature to strengthen important feature information. Step 3.6) The multimodal multitask probabilistic prediction model processes the multimodal fusion features simultaneously through its six feature mapping networks to obtain multiple corresponding task prediction results, which are then used as the task prediction output of the multimodal multitask probabilistic prediction model. Specifically, the weighted and interactive multimodal fusion features are input into the first, second, third, fourth, fifth, and sixth feature mapping networks, respectively. The first, second, third, fourth, fifth, and sixth feature mapping networks are optimized based on the true labels of their respective tasks and the backpropagation of the corresponding task loss functions, extracting information related to their respective tasks and completing the final mapping. Ultimately, this enables the prediction of pathological complete remission, disease-free survival, and overall survival in breast cancer patients after neoadjuvant therapy, the assessment of the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and the identification of benign and malignant tumor types and breast imaging reporting and data system level. Step 3.7) Using the counterfactual estimation method, under non-randomized controlled data conditions, the multimodal multi-task probabilistic prediction model is tested to predict each task. By comparing the prediction results of each task for breast cancer patients under various neoadjuvant therapy regimens, the potential risks of breast cancer patients receiving various neoadjuvant therapy regimens are simulated, thereby evaluating the individualized efficacy differences of each neoadjuvant therapy regimen for the same breast cancer patient. Step 3.8) Optimize the prediction task by assessing the risk probability of breast cancer metastasis to different organs and the probability of recurrence, and by identifying benign and malignant tumor types and breast imaging reporting and data system levels in the order of predicting pathological complete remission, disease-free survival and overall survival of breast cancer patients after neoadjuvant therapy.
5. The multimodal, multi-task probability prediction method for the prognosis of neoadjuvant therapy for breast cancer according to claim 4, characterized in that, The loss function used when training the multimodal, multi-task probabilistic prediction model is as follows: 1) The loss function for prediction error and feature distribution constraints is defined as: (4); In equation (4), λ KL This refers to the balance coefficient in ELBO. Z tab , Z txt and Z img The latent variable is a Gaussian distribution, obtained through the features of their respective encoders. Specifically, the mean and average difference are calculated using the feature map vectors encoded by their respective encoders, and then the Gaussian distribution is obtained from the mean and average difference. 2) The loss function for predicting complete pathological remission is defined as: (6); In equation (6), y pcr The true label is the pCR state, and the label is not the pCR state. y pcr When pCR is 0 y pcr =1, This represents the pCR score predicted by the response model; 3) The loss function for predicting overall survival and disease-free survival is defined as follows: (7); In equation (7), x i Represents the features of the i-th sample. In response to the risk score predicted by the model, E i This indicates the follow-up status of patient i. E i for 1 indicates death. E i for 0 indicates survival. T i Indicates the survival time of patient i; 4) The loss function for predicting the probability of breast cancer metastasizing to various organs is defined as: (10); In equation (10), L meta-BCE The binary cross-entropy loss function is... L meta-Focal Focal Loss is the loss function. 5) The loss function for predicting the probability of breast cancer recurrence is defined as: (11); In equation (11), y rec ∈{0,1} represents the recurrence event label. p rec The recurrence probability predicted by the model; 6) The loss function for grading benign and malignant tumor types is defined as: (12); In equation (12), y b / m ∈{0,1} represents the true tumor label, where 0 and 1 represent benign / malignant tumors, respectively. Indicates the predicted tumor type; 7) The loss function for the BI-RADS grading task is defined as: (13); In equation (13), 1 {y=c} This is an indicator function that takes the value 1 when the actual grade is c, and 0 otherwise. p c Predict the probability of belonging to class c for the model; 8) The method for calculating the overall alignment loss of the 10 alignment features is as follows: First, the InfoNCE loss function is used to calculate the alignment loss for each of the 10 groups of aligned features. Specifically, in the same training batch, feature representations from different modalities but corresponding to the same object or case are formed into positive sample pairs, while feature representations from different objects or cases are formed into negative sample pairs. By calculating the similarity between the two feature representations in each group of aligned features, and exponentially amplifying the similarity of positive samples and normalizing and suppressing the similarity of negative samples, the model maximizes the similarity of positive sample pairs during training while minimizing the similarity between them and negative sample pairs. The formula for calculating the InfoNCE loss function is as follows: (15); In equation (15), τ is the smoothing coefficient, used to adjust the smoothness of the similarity distribution, s() represents the feature similarity calculation function, N represents the length of the feature vector, and exp() represents the exponential function. u i , v i These represent two different feature vectors; Then, weights are summed based on the correlation between each pair of alignment features, specifically assigning... fi-MRI & ft- Rad With a weight of 1, assign fi-MRI & ft-Path The weight is α, and it is assigned fi-MRI & ftab The weight is β, assigned fi-MRI & fi- MIP , fi-MIP & ft-Rad , fi-MIP & ft-Path , fi-MIP & ftab , ft-Rad & ft-Path , ft-Rad & ftab , ft- Path&ftab All weights are γ, and 1 > α > β > γ > 0; the formula for calculating the overall alignment loss function is as follows: (16); After 50 rounds of training, the change in the overall alignment loss of all 10 alignment features is evaluated. If the reduction in the overall alignment loss is less than 5% after 5 rounds of training, the weights are adjusted to enhance the alignment effect. Specifically, the weight coefficients are strengthened to obtain 1≥α>β>γ>0.25, thereby strengthening the alignment constraints on different features by enhancing the weight coefficients.
6. A system for implementing the multimodal, multi-task probabilistic prediction method for prognosis of neoadjuvant therapy for breast cancer as described in any one of claims 1-5, characterized in that, include: The receiving unit is responsible for receiving the patient's breast DCE-MRI images and their paired segmented images, as well as clinical information, radiological imaging reports, diagnostic pathology reports, and case information; The preprocessing unit is responsible for preprocessing the collected breast DCE-MRI images and their paired segmented images, as well as the radiological imaging reports and diagnostic pathology reports; The task prediction unit is responsible for performing multimodal, multi-task probability predictions on clinical information of breast cancer patients and preprocessed breast DCE-MRI images, or preprocessed breast DCE-MRI images + radiological imaging reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images, or preprocessed breast DCE-MRI images + radiological imaging reports + pathology reports + segmented images + case information using a multimodal probability prediction model. This results in generating predictions for pathological complete remission, disease-free survival, and overall survival for various neoadjuvant therapy regimens, assessments of the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results of benign and malignant tumor types and breast imaging reports and data system levels. The output unit is responsible for generating a comprehensive report from the pathological complete remission, disease-free survival, and overall survival prediction results of various neoadjuvant treatment regimens generated by the multimodal probabilistic prediction model, the risk probability assessment results of breast cancer metastasis to different organs and recurrence probability assessment results, the identification results of benign and malignant tumor types and breast imaging reports and data system level identification results, and then visualizing the report.
7. The system according to claim 6, characterized in that: The multimodal multitask probabilistic prediction model consists of a text encoder, an image encoder, a table encoder, a multimodal fusion unit, and six independent feature mapping networks. The text encoder is used to perform unified feature modeling on multi-source text information from radiological imaging reports, diagnostic pathology reports, and case information to generate multi-source integrated text features; The image encoder is used to extract imaging features from breast DCE-MRI images and breast MIP images before and after synthetic comparison, so as to generate MRI image features that match the dimensions of multi-source integrated text features. The table encoder is used to extract features from structured clinical table data of clinical information to generate corresponding table features; The multimodal fusion unit is used to concatenate, fuse, and weight the aligned multi-source integrated text features, MRI image features, and table features to obtain multimodal fused features. The six feature mapping networks are used to process the weighted interactive multimodal fusion features to obtain prediction results of pathological complete remission, disease-free survival and overall survival of breast cancer patients after neoadjuvant therapy, assessment results of the risk probability of breast cancer metastasis to different organs and recurrence probability, and identification results of benign and malignant tumor types and breast imaging reporting and data system level. All of these are used as task prediction output results of the multimodal multi-task probability prediction model.
8. The system according to claim 7, characterized in that: The text encoder consists of three independent text feature extraction paths, a splicing layer, and a text feature fusion module. The three text feature extraction pathways are each composed of two Transformer sub-modules and corresponding layer normalization and nonlinear activation modules, which are used to perform context modeling and semantic feature extraction on radiological image reports, pathological reports, and case information, respectively, to obtain the corresponding intermediate features of radiological image report text, intermediate features of pathological report text, and intermediate features of case information text. The splicing layer is used to splice the intermediate features of radiological image report text, pathology report text, and case information text from the three text feature extraction pathways in terms of feature dimensions to form a joint text feature; The text feature fusion module consists of a Transformer sub-module and its corresponding layer normalization module and nonlinear activation combination module. It is responsible for performing cross-text source information interaction and feature fusion on the formed joint text features, thereby obtaining multi-source comprehensive text features with a size of [B, 1024]. ftxt This serves as the output feature vector of the text encoder, while simultaneously preserving radiological image report features. ft-Rad and pathology report characteristics ft-Path .
9. The system according to claim 7, characterized in that: The image encoder consists of three independent image sub-encoders: a self-trained encoder, a general medical large model, and an MRI visual large model, as well as a connection layer and a TransformerBlock. In the image encoder, the self-trained encoder, the general medical large model, and the MRI visual large model are all responsible for receiving 3D data from the pre- and post-synthetic contrast-enhanced breast DCE-MRI images and 2D data from the breast MIP images, and extracting imaging features to obtain the corresponding intermediate image features. The use of three image sub-encoders is to better extract imaging-related features from the pre- and post-synthetic contrast-enhanced breast DCE-MRI images and breast MIP images. In the image encoder, the connection layer is responsible for stitching together the intermediate features of the three images output by the self-trained encoder, the general medical big model, and the MRI visual big model to obtain joint image features; In the image encoder, the Transformer Block is responsible for feature mapping the joint image features output from the connection layer, matching their dimensions with those of the text encoder to obtain MRI image features. fimg As the output feature vector of the image encoder, which preserves the imaging features of breast DCE-MRI images. fi-MRI and breast MIP imaging features fi-MIP .
10. The system according to claim 9, characterized in that: The self-trained encoder includes 2 convolutional layers, 2 linear mapping layers, 4 usion blocks, 6 transformer blocks, 1 average pooling module, and 3 local feature extraction paths; The first convolutional layer is responsible for performing convolution operations on the breast DCE-MRI stitched image or breast MIP image obtained by stitching the breast DCE-MRI images before and after the synthesis and contrast in the channel dimension. This enables the block processing of the original image, thereby dividing the breast DCE-MRI stitched image or breast MIP image into 1331 three-dimensional image blocks of size 16×16×16 or 121 two-dimensional image blocks of size 16×16. The first linear mapping layer is responsible for mapping the flattened 3D or 2D image blocks to a unified feature space, so that its feature dimensions are mapped from the original dimensions to [B, 4096, 1024] to obtain first-level image features; The second convolutional layer and the second linear mapping layer are responsible for feature extraction from the segmented image to obtain first-level local features; The first Fusion Block is responsible for fusing the primary image features with the primary local features to obtain the primary fused features; The first Transformer Block is responsible for updating and enhancing the first-level fused features or first-level image features to obtain second-level image features; The first local feature extraction pathway is responsible for extracting features from the first-level local features to obtain the second-level local features; The second Fusion Block is responsible for fusing secondary image features with secondary local features to obtain secondary fused features; The second Transformer Block is responsible for updating and enhancing the secondary fused features or secondary image features to obtain tertiary image features; The second local feature extraction pathway is responsible for extracting features from the secondary local features to obtain the tertiary local features; The third Fusion Block is responsible for fusing the third-level image features with the third-level local features to obtain the third-level fused features; The third Transformer Block is responsible for updating and enhancing the three-level fused features or three-level image features to obtain four-level image features. The third local feature extraction pathway is responsible for extracting features from the third-level local features to obtain the fourth-level local features; The fourth Fusion Block is responsible for fusing the fourth-level image features with the fourth-level local features to obtain the fourth-level fused features; The fourth Transformer Block is responsible for updating and enhancing the fourth-level fused features or fourth-level image features to obtain fifth-level image features. The fifth and sixth Transformer Blocks and the average pooling module are responsible for updating and enhancing the five-level image features again, and compressing their dimensions through the average pooling operation to obtain DCE-MRI image features or MIP image features of size [B, 1024].