A medical data processing system and method for generating breast cancer prognosis reference information

CN122822366APending Publication Date: 2026-09-25JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611316444.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

不足一:多模态融合形同虚设

Benefits of technology

第一,在多模态语义对齐与可信融合方面,通过跨模态交叉注意力机制实现模态间语义交互而非简单物理拼接,通过VAE概率补全解决了临床数据缺失问题,通过质量感知双重加权区分真实数据与重建数据,使融合结果更加可靠。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822366A_ABST
    Figure CN122822366A_ABST
Patent Text Reader

Abstract

The application discloses a medical data processing system and method for generating breast cancer prognosis reference information, and belongs to the technical field of medical image analysis and deep learning. Four main deficiencies in the comprehensive analysis of breast examination results are solved. Cross-modal cross-attention is used to realize modal semantic interaction, probability completion is used to process clinical missing data, quality perception double weighting is used to distinguish real data from reconstructed data, and fusion reliability is improved. The model explicitly encodes irregular time intervals, uses an incremental feature extractor to take the adjacent follow-up change as a prognosis feature, and quantifies the change rate. The global features of the tumor are decoupled, and a learnable sub-region node is modeled without manual ROI annotation. Relying on four interpretable mechanisms, the prognosis information is traceable. Shared feature representation is used to provide inductive bias, and uncertainty weighting is used to balance multi-task loss, so that multi-task collaborative optimization is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of medical image analysis and deep learning, specifically relating to a medical data processing system and method for generating prognostic reference information for breast cancer. Background Technology

[0002] Breast cancer is the most common malignant tumor among women worldwide. In clinical practice, a breast cancer patient typically needs to undergo the following examinations in sequence from diagnosis to the development of a treatment plan.

[0003] Mammography is used to observe calcifications and structural distortions, capturing the morphological features of the tumor. Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is used to assess tumor blood supply and perfusion dynamics, capturing the hemodynamic characteristics of the tumor. Ultrasound is used to observe tumor morphology and elasticity, capturing the tissue elasticity characteristics of the tumor. Whole-slide pathological imaging (H&E staining) after biopsy is used to confirm histological type and grade, capturing the cellular morphological characteristics of the tumor. Simultaneously, clinical indicators such as age, hormone receptor status, and molecular subtyping are collected to reflect the image expression characteristics of the tumor.

[0004] The five tests mentioned above each capture the characteristics of tumors in different biological dimensions, and relying on any one of them alone cannot comprehensively assess the invasiveness and prognostic risk of a tumor. However, in actual clinical practice, the challenge faced by doctors is to integrate these test results from different "semantic systems" and to complete cross-modal semantic alignment and empirical synthesis in the brain.

[0005] In actual medical practice, there are still the following shortcomings in the comprehensive analysis of breast examination results: Shortcoming 1: Multimodal fusion is practically useless. Existing multimodal methods mostly employ direct concatenation of feature vectors or mean pooling before feeding them into the classifier. Features of different modalities are independently encoded in their respective autonomous semantic spaces; the concatenation operation merely physically connects vectors without semantic interaction. Mammography X-ray features are unaware of which information in MRI features is relevant to their own calcification description, and MRI features are also unaware of which information in mammography X-ray features is complementary to their perfusion description. Furthermore, in clinical practice, patients commonly have missing parts of their examinations; existing methods mostly employ zero-filling or directly discard samples containing missing modalities. The former introduces artifact interference, while the latter causes selection bias.

[0006] Second shortcoming: Time-related information is systematically ignored. Existing models almost entirely rely on single imaging data from the patient's baseline (at initial diagnosis), completely discarding longitudinal changes accumulated during follow-up. Breast cancer patients typically require long-term follow-up after surgery, and the changes in the tumor between two follow-up examinations—whether it is shrinking (treatment effective) or enlarging (progression)—are themselves extremely important prognostic signals. Furthermore, the irregularity of follow-up intervals (intensive follow-up in the early postoperative period, gradually lengthening after stabilization) is also a technical obstacle. Standard time-series models assume equal data collection intervals and cannot handle real-world clinical follow-up data.

[0007] Shortcoming 3: Tumor spatial heterogeneity is smoothed out Existing models typically perform global pooling on tumor features, compressing different functional subregions within the tumor (core necrosis zone, invasion front zone, stromal reaction zone, etc.) into a unified vector representation. However, these subregions exhibit vastly different biological behaviors and prognostic contributions, and global pooling leads to the complete loss of spatial structural information.

[0008] Shortcoming 4: The prediction process is not traceable. Deep learning models are often black-box structures, outputting predictions but failing to explain which input features drive those predictions. If a model outputs "a 65% probability of recurrence in 5 years," doctors need to know the basis for this conclusion—is it primarily due to MRI perfusion characteristics, a high pathological stratification index, or some combination of both? This lack of interpretability prevents clinicians from validating the model's output, making it difficult to incorporate the model results into clinical decision-making processes. Summary of the Invention

[0009] To address the four clinical challenges mentioned above, this invention provides a medical data processing system and method for generating prognostic reference information for breast cancer.

[0010] The system comprises a four-level architecture: The first level is a multimodal feature extraction module, which uses a dedicated feature extractor for each of the five modalities: mammogram images, breast DCE-MRI images, breast ultrasound images, breast whole slide pathology images, and clinical feature data, to map the raw data of each modality to a feature vector space of the same dimension. The second level is the cross-modal Transformer fusion module, which achieves semantic alignment and fusion of heterogeneous modalities through four sub-processes: missing modality probability completion, cross-modal semantic interaction, dual credibility weighting, and high-order association modeling between modalities, to obtain a global fusion vector. ; The third level is the spatiotemporal graph network module, which receives the global fusion vector. A separate architecture is adopted, which first models the time dimension, then the spatial dimension, and finally performs gated fusion, to obtain spatiotemporal fusion features. ; The fourth level is the multi-task prediction head module, which simultaneously outputs disease-free survival, overall survival, risk stratification, metastasis pattern, and treatment response probability based on shared features, and automatically balances the losses of each task through uncertainty weighting; it also integrates four interpretability mechanisms.

[0011] Furthermore, after receiving the mammogram image feature extractor, the local contrast is first enhanced by contrast-limited adaptive histogram equalization. The enhanced image then extracts features through two parallel paths: the upper path outputs a feature map capturing global contextual information, and the lower path outputs a feature map extracting local detail features. The outputs of the two paths are concatenated and compressed into a feature vector by an attention pooling layer. After receiving the image, the DCE-MRI feature extractor uses a three-way parallel branch to extract complementary features. Branch 1 processes the original image spatial features and outputs spatiotemporal attention features; Branch 2 processes deep convolutional spatial features and outputs hierarchical spatial features; Branch 3 processes quantitative physiological features and outputs pharmacokinetic features. The outputs of the three branches are weighted and summed by a multi-scale attention fusion module to obtain the feature vector. After receiving an image, the ultrasonic feature extractor first extracts spatial features through a convolutional neighborhood network. Then, it models the dynamic changes between frames by sliding a 3D convolution along the time dimension. Next, it uses a temporal attention mechanism to weight and aggregate the features from different frames, outputting spatiotemporal features. Finally, it compresses the features using global average pooling to obtain the feature vector. The WSI feature extractor for whole-section pathological images receives the image, performs global localization at low magnification to extract the region of interest, and performs patch extraction at high magnification. Then, attention-based multi-instance learning is employed, using learnable attention weights to weighted aggregate features from the tumor area, stromal area, and lymphocyte area respectively. The aggregated results from the three areas are then concatenated and compressed into a feature vector through a linear projection layer. The clinical feature encoder receives structured tabular data. Categorical variables are encoded into dense vectors through an embedding layer; continuous variables are standardized to zero mean and unit variance through a normalization layer. All variables are concatenated and input into a fully connected network, and finally mapped into feature vectors through a linear projection layer. .

[0012] Furthermore, the ultrasound feature extractor also includes an elastography auxiliary branch, which is an optional path. The input data of the elastography auxiliary branch is the elastography image of the tissue to be detected. The elastography image is first subjected to preliminary feature extraction and downsampling by the first convolutional layer. After batch normalization and ReLU activation, the output is sent to the second convolutional layer for further extraction of high-level mechanical features and downsampling again. Subsequently, the spatial dimension is compressed by adaptive global average pooling. After flattening, it is mapped to an elastic feature vector by a fully connected layer. The obtained elastic feature vector is expanded by padding zeros on the right side and fused element-wise with the feature vector obtained by temporal average pooling of the main branch.

[0013] Furthermore, missing modality probability completion is performed through a cross-modal variational autoencoder submodule. The operation involves: establishing independent variational autoencoders for the five modal features extracted from the multimodal feature extraction module; and obtaining independent latent variables for each modality from the posterior distribution through reparameterization. In the joint encoding stage, a joint encoder is set up to concatenate the independent latent variables of all available modalities and map them into low-dimensional joint latent variables. During the conditional decoding phase, for any missing mode Conditional decoder to combine latent variables and modality Independent latent variables Together as conditional inputs, we reconstruct the feature estimate of the missing mode. ; Cross-modal semantic interaction is performed through a cross-modal multi-head attention interaction submodule, where the following operations are performed: for each modality to its characteristics The query is mapped through three independent linear layers. ,key Sum At the same time, other modalities Mapping to key Sum Then, parallel computation Scaled dot product cross attention with all other modalities; for each pair of modalities ,Will and After concatenation, the weights are fed into a two-layer fully connected network and then normalized using Softmax to generate adaptive weights. All attention results are multiplied by learnable adaptive weights. After aggregation by summation nodes, the result is then joined with the original input via residual connections. Adding them together, we get five fused features, which are the attention fusion features.

[0014] Furthermore, the dual credibility weighting is performed through an adaptive weighted fusion submodule, which involves: performing a dual-weighted summation on the attention fusion features to obtain the dual-weighted modal features. The dual weighting mechanism is as follows: After the content adaptive weight generator concatenates the attention fusion features of each modality into a joint vector along the feature dimension, it generates a weight vector reflecting the importance of content discrimination of each modality through a fully connected layer and a Softmax function. The quality-aware weight generator stacks the modality availability mask into a 5-dimensional vector, and generates a quality score vector reflecting the credibility of the data source of each modality through a multilayer perceptron and a Sigmoid function mapping. The two weights are multiplied element by element and normalized, and then a weighted sum is performed on the attention fusion features of each modality. High-order intermodal correlation modeling is performed through a graph convolutional fusion enhancement submodule. The operations are as follows: First, a modal relationship graph is constructed, with the five modalities as graph nodes. A learnable adjacency matrix is ​​initialized and normalized using Softmax. End-to-end training automatically mines modal collaboration and redundancy relationships. Then, two layers of graph convolution are used, with the attention fusion features of each modality as initial node features. Through multiple rounds of neighborhood aggregation and nonlinear transformation, high-order topological dependencies of the modalities are captured. Finally, the output features are aggregated using the global mean to obtain the graph convolutional enhanced features. Combine it with the double-weighted modal features Element-by-element fusion, followed by a projection layer to obtain the final fused feature. .

[0015] Furthermore, the input to the spatiotemporal graph network module is the patient's longitudinal medical visit sequence. }, each of which For the first The global multimodal fusion features output by the cross-modal Transformer fusion module for each patient visit include the actual time intervals between adjacent visits. Temporal dimension modeling is performed through a temporal branch: first, time location encoding is performed, based on the actual time intervals of each patient's visits, using a learnable linear projection mapping to convert them into time interval embedding vectors. These time embedding vectors are then element-wise added to the corresponding global fusion vectors to obtain the time-encoded feature sequence. This time-encoded feature sequence is fed into a temporal Transformer encoder, which outputs context-enhanced temporal dependency features for each time point. An incremental feature extractor calculates the feature differences between adjacent visit time points. ,based on Tumor growth rate modeling and treatment response modeling were performed separately, and their outputs were spliced ​​and linearly projected before being fused to obtain the temporal dynamic features. .

[0016] Furthermore, the spatiotemporal graph network module models the spatial structure through spatial branches. At the corresponding time point of medical visit, the tumor imaging sub-region and regional lymph nodes are used as graph nodes. The graph's connection edges are constructed based on spatial proximity relationships and anatomical drainage pathways. The resulting graph is represented by a node feature matrix and an adjacency matrix. By using a learnable adjacency matrix and normalizing it with Softmax, and then combining it with a two-layer graph convolutional network for message propagation and node feature aggregation, the spatial structure features are obtained. ; Time dynamic characteristics Spatial structural features Joint encoding is performed using a gated addition unit:

[0017] in, ∈(0,1) is the gating coefficient, which controls the contribution ratio of the temporal branch and the spatial branch to the final fused feature.

[0018] Furthermore, spatiotemporal fusion features The global fusion vector corresponding to the time point of visit output by the cross-modal Transformer fusion module. f fused Cross-module fusion is performed using gated addition units to obtain a shared spatiotemporal fusion feature representation. :

[0019]

[0020] in, For learnable weight vectors, For bias terms; For learnable scalar gating coefficients, by and After splicing, it is adaptively generated through linear mapping and the Sigmoid function; Multi-task prediction head module in shared spatiotemporal fusion feature representation Based on this, four parallel task-specific prediction heads are set up, and four clinical prediction results are output simultaneously: Cox proportional hazards model branch: The neural Cox proportional hazards model is adopted, and the hazard ratio is output through a linear layer and exponential activation. This branch has two parallel sub-output heads that share the basic hazard function and predict the 1-year / 3-year / 5-year survival probability of disease-free survival and overall survival, respectively. Risk stratification tri-classifier branch: Patients are divided into three risk groups: low risk, medium risk, and high risk through linear layer and Softmax activation; The binary classifier branch for the transfer pattern outputs the predicted probabilities of local transfer and distant transfer after a linear layer and sigmoid activation. pCR prediction classifier branch: outputs the predicted probability of complete pathological remission after linear layer and Sigmoid activation.

[0021] Furthermore, the four-fold interpretability mechanism is as follows: The first level is Grad-CAM attention heatmap visualization: generating a multimodal attention heatmap, highlighting the lesion areas that have the greatest impact on prediction on the original image; extracting the adaptive attention weight matrix in cross-modal fusion to generate an intermodal attention heatmap, showing the intensity and direction of information interaction between each modality; The second level is SHAP attribution analysis: using the SHAP method, the contribution of each modality and each clinical feature to the individual prediction is quantified by bar chart, the Shapley marginal contribution value of each modality feature and each clinical feature under all possible modality combinations is calculated, and the average incremental contribution of each input feature to the final prediction result is quantified. The third level is prototype network case reasoning: by displaying representative cases most similar to the current patient through thumbnails, case-based reasoning is realized, providing clinicians with intuitive analogy references; The fourth level, counterfactual reasoning, presents the impact of changes in key variables on prediction results in the form of text bubbles. By intervening in specific modal feature values ​​and observing the direction and magnitude of changes in prediction results, it generates reasoning explanations and provides quantitative basis for adjusting treatment plans.

[0022] The method of the present invention includes: Multimodal feature extraction: For five modalities, namely mammogram images, breast DCE-MRI images, breast ultrasound images, breast whole slide pathology images and clinical feature data, a corresponding dedicated feature extractor is used to map the original data of each modality to a feature vector space of the same dimension. Cross-modal Transformer fusion: Through a four-level sub-process, including missing modality probability completion, cross-modal semantic interaction, dual credibility weighting, and high-order intermodal association modeling, semantic alignment and fusion of heterogeneous modalities are achieved, resulting in a global fusion vector. ; Spatiotemporal graph network fusion: receiving global fusion vector A separate architecture is adopted, which first models the time dimension, then the spatial dimension, and finally performs gated fusion, to obtain spatiotemporal fusion features. ; Multi-task prediction: Based on shared features, it simultaneously outputs disease-free survival, overall survival, risk stratification, metastasis pattern, and treatment response probability, and automatically balances the loss of each task through uncertainty weighting; it also integrates four interpretability mechanisms.

[0023] The beneficial effects of the medical data processing system for generating breast cancer prognostic reference information described in this invention are as follows: First, in terms of multimodal semantic alignment and reliable fusion, the cross-modal cross-attention mechanism is used to achieve semantic interaction between modalities rather than simple physical splicing. The VAE probabilistic completion solves the problem of missing clinical data. The quality-perception dual weighting distinguishes between real data and reconstructed data, making the fusion results more reliable.

[0024] Second, in terms of fully utilizing longitudinal change information, it is the first time that irregular time intervals are explicitly encoded as model inputs, and the first time that changes in adjacent follow-ups are explicitly encoded as independent prognostic features (DeltaEncoder), making "rate of change" a quantifiable prognostic biomarker.

[0025] Third, in terms of preserving tumor spatial heterogeneity, it is the first to decompose the global features of the tumor into learnable sub-region nodes for decoupling modeling, which can automatically differentiate into different functional regions (including necrotic areas, invasion front areas, matrix reaction areas, etc.) without the need for manual ROI annotation.

[0026] Fourth, regarding the traceability of the prediction process, a four-fold interpretability mechanism, including Grad-weighted Class Activation Mapping (Grad-CAM) visualization, SHAP (SHapley Additive exPlanations) attribution analysis, prototype network case reasoning, and counterfactual reasoning, makes the generated prognostic reference information traceable.

[0027] Fifth, in terms of multi-task collaborative optimization, the prediction tasks provide inductive biases to each other by sharing feature representations, and the losses of each task are automatically balanced by uncertainty weighting. Attached Figure Description

[0028] Figure 1 This is a diagram of the medical data processing system architecture for generating breast cancer prognostic reference information in an embodiment of the present invention. Figure 2 This is an example of mammography imaging in an embodiment of the present invention, showing typical images of bilateral breast CC and MLO views, including different manifestations of benign and malignant lesions; Figure 3 This is an example of a DCE-MRI dynamic enhancement sequence image in an embodiment of the present invention, demonstrating the three-dimensional voxel data features at different time phases; Figure 4This is an example of B-mode ultrasound imaging in an embodiment of the present invention, demonstrating typical ultrasound image features (morphology, margins, echo type, etc.) of benign and malignant lesions. Figure 5 This is an example of an H&E-stained whole slide image in an embodiment of the present invention, demonstrating pixel-level high-resolution image features of benign and malignant lesions; Figure 6 This is a diagram illustrating the architecture of a mammography feature extractor in an embodiment of the present invention. Figure 7 This is a diagram of the DCE-MRI feature extractor architecture in an embodiment of the present invention; Figure 8 This is a diagram of the ultrasonic feature extractor architecture in an embodiment of the present invention; Figure 9 This is a diagram of the WSI feature extractor architecture for whole-slice pathological images in an embodiment of the present invention; Figure 10 This is a diagram of the clinical feature encoder architecture in an embodiment of the present invention; Figure 11 This is an architecture diagram of the cross-modal Transformer fusion module in an embodiment of the present invention; Figure 12 This is a structural diagram of the spatiotemporal graph network module in an embodiment of the present invention; Figure 13 This is a structural diagram of the multi-task prediction head module in an embodiment of the present invention; Figure 14 This is a structural diagram of the interpretability module in an embodiment of the present invention; Figure 15 This is a schematic diagram comparing the consistency indices of various models in an embodiment of the present invention; Figure 16 This is a schematic diagram of the time-dependent AUC curves of various models in embodiments of the present invention; Figure 17 This is a schematic diagram illustrating Brier scores (measurement of calibration performance) for various models in embodiments of the present invention; Figure 18 This is a schematic diagram illustrating the consistency (assessment of universality) among different patient subgroups in an embodiment of the present invention; Figure 19 This is a schematic diagram illustrating the quantification of the contributions of different modes through ablation studies in an embodiment of the present invention. Detailed Implementation

[0029] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0030] Example 1 This embodiment provides a medical data processing system for generating prognostic reference information for breast cancer. The system architecture diagram is shown below. Figure 1 As shown.

[0031] The system includes a four-stage pipeline architecture: The first level is a multimodal feature extraction module, targeting five modalities: mammogram images, DCE-MRI (dynamic contrast-enhanced MRI) images, ultrasound images, whole-slide pathology images, and clinical feature data (see examples for each modality). Figures 2 to 5 Each modality's original data is mapped to a feature vector space of a unified dimension using its own dedicated feature extractor.

[0032] The second level is the cross-modal Transformer fusion module (see...). Figure 11 Through a four-level subprocess, including missing modality probability completion, cross-modal semantic interaction, dual credibility weighting, and graph convolutional collaborative enhancement, semantic alignment and reliable fusion of heterogeneous modalities are achieved.

[0033] The third level is the spatiotemporal graph network module (see...). Figure 12 The reception process involves first modeling the time dimension, then modeling the spatial dimension, and finally gating and fusing.

[0034] The fourth level is the multi-task prediction head module (see...). Figure 13 Based on shared features, it simultaneously outputs disease-free survival, overall survival, risk stratification, metastasis pattern, and treatment response probability, and automatically balances the losses of each task through uncertainty weighting; it also integrates four interpretability mechanisms (see...). Figure 14 ).

[0035] All prognostic reference information generated by this invention are probabilistic or risk-based values. The output format includes disease-free survival prediction (probability of 1 year / 3 years / 5 years), overall survival prediction (probability of 1 year / 3 years / 5 years), low / medium / high risk level probability, independent probability of each metastatic site, and pathological complete remission probability. This information is only used as reference data to assist doctors in making clinical decisions and is not directly equivalent to a diagnostic conclusion.

[0036] Example 2 This embodiment further defines Embodiment 1, and provides further explanation of the source, composition and annotation of the dataset.

[0037] 1. Data Source This embodiment is a retrospective, multicenter study covering patients diagnosed with primary breast cancer between 2013 and 2023. The training set consisted of 1300 patients from a single institution, while the external validation set comprised 456 patients. The publicly available dataset used in this embodiment exempted the need for informed consent for retrospective data analysis.

[0038] Inclusion criteria: (1) Histologically confirmed invasive breast cancer; (2) At least two imaging modalities; (3) Complete clinical and follow-up data; (4) No history of cancer.

[0039] 2. Data Composition Multimodal data includes two-dimensional mammography (see [link]). Figure 2 ), 3D / 4D magnetic resonance imaging sequences (see) Figure 3 ), ultrasound video (see Figure 4 ), pixel-level pathological images (see Figure 5 ), and structured clinical record forms.

[0040] (a) Mammogram: Images of the bilateral breasts were acquired in four views: cephalothorax (CC) and endoscopic oblique (MLO). Each view was 768×768 pixels in resolution and was a grayscale image.

[0041] (b) DCE-MRI images: Three-dimensional voxel data containing 5 to 7 time phases were acquired, with a spatial resolution of 128×128×64 (height×width×depth) for each time phase, and were dynamically enhanced sequences.

[0042] (c) Pathological images of whole slides: H&E stained whole slides were acquired at 40x magnification, which are pixel-level high-resolution images. Multi-resolution pathological blocks were obtained by sampling at four magnification levels: 5x, 10x, 20x and 40x.

[0043] (d) Ultrasound imaging: B-mode ultrasound video sequences were acquired, consisting of 16 consecutive images, each with a resolution of 224×224 pixels. Elastography images were also acquired as an optional branch in some cases.

[0044] (e) Clinical characteristic data: Collect structured clinical record forms, including continuous variables (age, tumor size, number of days from diagnosis to surgery, etc.) and categorical variables (molecular markers ER (estrogen receptor), PR (progesterone receptor), HER2 (human epidermal growth factor receptor 2), histological grade, stage, etc.).

[0045] 3. Data labeling This embodiment uses a benign / malignant binary label as the basic annotation. A dataset of breast lesions confirmed by surgical pathology was constructed, with malignant samples covering various histological pathological types such as invasive ductal carcinoma, special types of invasive carcinoma, and carcinoma in situ (Hist.type). The dataset is equipped with standardized structured annotation fields, covering patient baseline information, ultrasound image files, tumor ultrasound morphology (Shape), breast tissue classification, menopausal status; ER / PR / HR (Hormone Receptor) / HER2 molecular markers and molecular classification (Mol.Subtype); T / N / M tumor staging, Nottingham pathological grade, lymph node metastasis status; surgical plan, pathological complete response PCR, imaging response assessment (Imaging_Response); and three survival prognostic time endpoints: DMFS (Distant Metastasis-Free Survival), OS (Overall Survival), and RFS (Recurrence-Free Survival), comprehensively covering imaging features, clinicopathology, molecular classification, and long-term prognostic multidimensional information.

[0046] Example 3 This embodiment further defines Embodiment 1 and provides a detailed description of the technical solution for a medical data processing system used to generate prognostic reference information for breast cancer.

[0047] 1. Multimodal feature extraction module This module is configured to extract features for each of the five modalities using their respective dedicated feature extractors, mapping the original data of each modality to a feature vector space of a unified dimension (the unified embedding dimension is set to 512, and this value can be adjusted in the range of 256 to 1024 through cross-validation).

[0048] (a) A dedicated feature extractor for mammogram images (see Mammogram Image Feature Extractor, see [link]). Figure 6 ) The encoder receives four 768×768 mammogram images (bilateral CC and MLO views), and first performs contrast-limited adaptive histogram equalization (CLAHE) to enhance local contrast.

[0049] The enhanced image extracts features through two parallel paths. The upper path uses a Shifted Window Transformer (Swin Transformer), which captures global contextual information through alternating stacks of window-based multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA) modules, outputting a feature map of size 4×48×48×192. The lower path uses an EfficientNet (Efficient Convolutional Network), which extracts local detail features through stacked convolutions of inverted residual bottleneck convolution modules (MBConv), outputting a feature map of size 4×48×48×128.

[0050] The outputs of the two paths are concatenated into a 4×48×48×320 multi-scale feature vector across the channel dimension. Subsequently, a symmetry perception module performs cross-attention calculation on the left and right breast views, using the features of the affected side view as the query and the corresponding features of the healthy side view as the reference to explicitly model bilateral asymmetry—a known important prognostic indicator—outputting a 2×512 differential feature vector. Finally, this is compressed into a 512-dimensional feature vector by an attention pooling layer. It is used for cross-modal fusion.

[0051] The medical principle behind this is that the healthy tissue can automatically serve as an anatomical reference for the diseased side, and the detection sensitivity of the lesion area is enhanced by comparing the asymmetry between the two sides.

[0052] (b) Dedicated feature extractor for DCE-MRI dynamic enhancement sequences (DCE-MRI feature extractor, see [link]) Figure 7 ) The encoder receives five phases of three-dimensional voxel data, with dimensions of 5×128×128×64 (time×height×width×depth), and uses three parallel branches to extract complementary features.

[0053] Branch 1 processes the original image spatial features, using a 3D VisionTransformer for block embedding (block size 8×8×8×1). A stacked 3D Transformer encoder captures global spatiotemporal dependencies, followed by global average pooling to output 192-dimensional features. Branch 2 processes deep convolutional spatial features, using a 3D Residual Grouped Convolutional Network (3D ResNeXt) with stacked residual bottleneck blocks to extract deep spatiotemporal features. 3D global average pooling outputs 256-dimensional features. Branch 3 processes quantitative physiological features, using the dynamic enhancement time-intensity curves corresponding to each voxel of the lesion as temporal input. A 3-layer fully connected network implicitly encodes the curve morphology, automatically learning the DCE-MRI pharmacokinetic pattern features corresponding to the curves, outputting 64-dimensional temporal features, implicitly representing the volume transfer constant (VLC). K trans ), rate constant ( k ep ), extracellular space volume ratio ( V e This reflects information about vascular permeability (a metabolic kinetic indicator of contrast agents in tissues).

[0054] The outputs of the three branches are weighted and summed by a multi-scale attention fusion module, automatically focusing on the feature scale most valuable for prognosis prediction, and finally outputting a 512-dimensional feature vector. .

[0055] (c) Dedicated feature extractor for ultrasound images (ultrasound feature extractor, see Figure 8 ) The encoder receives 16 frames of ultrasound video with a size of 224×224. First, spatial features are extracted using a second-generation convolutional neighborhood network (ConvNeXt V2), with the feature maps progressively downsampled to 16×7×7×768. Then, a 3D convolution (3×3×3 kernel) is used to slide along the temporal dimension to model the dynamic changes between frames. Next, a temporal attention mechanism is used to weighted aggregate the features from different frames, outputting a 7×7×512 spatiotemporal feature vector. Finally, global average pooling is used to compress the vector into a 512-dimensional feature vector. .

[0056] Ultrasound elastography is a medical ultrasound imaging technique whose core principle is that tissues in different pathological states have different mechanical stiffness (elastic modulus). Malignant tumors, due to the high proliferation of tumor cells, stromal fibrosis, and extracellular matrix remodeling, are usually harder than benign lesions and normal tissues (i.e., less elastic and more rigid).

[0057] In clinical practice, an ultrasound probe applies slight pressure (or utilizes acoustic radiation force pulses) to the tissue and tracks minute displacements within it. By calculating strain or shear wave propagation velocity, a color or grayscale image reflecting the relative stiffness distribution of the tissue is reconstructed—an elastography image. Using this image as a supplement to conventional B-mode ultrasound can significantly improve the specificity of differentiating between benign and malignant breast cancer: areas with higher stiffness have a greater risk of malignancy.

[0058] In this system, elastography images are used as optional auxiliary inputs to inject prior knowledge of tissue mechanical properties into the model, thereby enhancing the accuracy of prognostic prediction.

[0059] Simultaneously, an elastography auxiliary branch can be added as an optional path to integrate mechanical features. The specific data processing flow is as follows: The elastography image (single channel, reflecting the spatial distribution of tissue stiffness) is first processed by the first convolutional layer (32 3×3 convolutional kernels, stride 2) for preliminary feature extraction and downsampling. After batch normalization and ReLU activation, the output is fed into the second convolutional layer (64 3×3 convolutional kernels, stride 2) for further extraction of high-level mechanical features and downsampling again. Subsequently, adaptive global average pooling compresses the spatial dimension to 1×1, flattens it, and maps it to a 128-dimensional elastic feature vector through a fully connected layer. The obtained elastic feature vector is then extended to 512 dimensions by right-side zero padding, and fused element-wise with the 512-dimensional feature vector obtained from the main branch's temporal average pooling. Finally, it is processed through a projection layer (Linear→LayerNorm→GELU) to generate a 512-dimensional ultrasound feature vector incorporating mechanical properties. .

[0060] (d) Dedicated feature extractor for whole-slide pathological images (WSI) (see Whole-Slide Pathological Image Feature Extractor, see...) Figure 9 ) The encoder receives a pyramid structure from a whole-slice pathological image, performs global localization at a low magnification of 5x to extract the region of interest, and extracts 256×256 patches at a high magnification of 40x. Each patch is encoded into a 768-dimensional feature vector by a pre-trained ViT (VisionTransformer) encoder.

[0061] Subsequently, attention-based multi-instance learning was employed. Learnable attention weights (including Softmax normalization) were used to weight and aggregate the features of the three regions: the tumor region, the stromal region, and the lymphocyte region. The aggregated results of the three regions were then concatenated and compressed into a 512-dimensional feature vector through a linear projection layer. To capture the comprehensive spatial characteristics of the tumor microenvironment.

[0062] (e) A dedicated feature extractor for clinical feature data (clinical feature encoder, see below) Figure 10 ) The encoder receives structured tabular data. Categorical variables (including ER, PR, HER2, histological grade, stage, etc.) are encoded into dense vectors through an embedding layer; continuous variables (including age, tumor size, etc.) are standardized to zero mean and unit variance through a normalization layer.

[0063] After concatenating all variables, the result is fed into a 3-layer fully connected network (dimensionality changes: input → 128 → 64 → 256), with each layer followed by a ReLU activation function. Finally, a linear projection layer maps the 256 dimensions to a 512-dimensional feature vector. This aligns its dimensions with the image modal features, facilitating subsequent cross-modal fusion.

[0064] Each modal feature is mapped to a uniform 512-dimensional feature vector space via a linear projection layer, and the mapping relationship is as follows:

[0065] in and These are trainable parameters.

[0066] 2. Cross-modal Transformer fusion module The architecture diagram of the cross-modal Transformer fusion module is as follows: Figure 11 As shown, this module receives feature vectors from five encoders. Each of them is 512-dimensional. The complete processing flow is a four-stage pipeline of "missing modality probability completion → cross-modal semantic interaction → dual credibility weighting → graph convolutional collaborative enhancement".

[0067] (a) Cross-modal variational autoencoder submodule (missing mode probability completion) To address the common problem in clinical practice that breast cancer patients cannot have all five modalities of examination data collected, a cross-modal variational autoencoder is constructed to probabilistically complete the missing modalities.

[0068] During the independent encoding phase, for each modality m Construct an independent variational autoencoder. The encoder will convert modal features... Mapped to the mean of the underlying normal distribution Sum of logarithmic variance Independent latent variables for each modality are obtained by sampling from the posterior distribution using reparameterization techniques. Its mathematical representation is:

[0069]

[0070] in, Indicates that in the given first Features of each modality Under the condition of encoder (Parameters are) The latent variables inferred The approximate posterior distribution (i.e., variational posterior). express Follow the mean The covariance matrix is The multivariate normal distribution (i.e., a multidimensional Gaussian distribution in which each dimension is independent and has the same variance). This represents the variance (scalar value) of the normal distribution in each dimension. I Represents the identity matrix, used to construct the diagonal covariance matrix. This ensures that the dimensions of the latent variables are independent of each other and have consistent variance. Describes a random noise vector sampled from a standard normal prior distribution (and...). (same dimension), used to implement reparameterization techniques to make the sampling process differentiable; express The value of follows a multivariate standard normal distribution with a mean of zero and a covariance matrix of identity. As a noise source for reparameterization, its randomness is injected during training and fixed to zero during inference to ensure deterministic output.

[0071] In the joint encoding stage, a joint encoder is set up to concatenate the independent latent variables of all available modalities and map them into low-dimensional joint latent variables. To encode cross-modal shared semantic information.

[0072] During the conditional decoding phase, for any missing mode Conditional decoder to combine latent variables and modality Independent latent variables Together as conditional inputs, we reconstruct the feature estimate of the missing mode. Specifically, the training phase The encoder corresponding to this mode extracts the true features Posterior distribution Sampling is used to learn conditions for generating mappings using complete supervised signals; during the inference phase, due to the actual absence of this modality, From the standard normal prior distribution Mid-sampling, combined latent variables This provides cross-modal semantic constraints from other available modalities. The concatenated constraints are then fed into a conditional decoder composed of fully connected layers or multi-layer perceptrons. ( | , The function outputs the estimated features of the missing modes, thereby obtaining the complete five-mode feature set after imputation.

[0073] The medical principle behind this is that the manifestations of the same tumor in different modalities are intrinsically related, and the combined information of available modalities can probabilistically constrain the possible range of values ​​for the missing modality.

[0074] (b) Cross-modal multi-head attention interaction submodule (cross-modal semantic interaction) For each mode The module will feature The query is mapped through three independent linear layers. ,key Sum At the same time, other modalities Mapping to key Sum Then, parallel computation Scaled dot product cross attention with all other modalities.

[0075] For each pair of modes ,Will and After concatenation, the weights are fed into a two-layer fully connected network and then normalized using Softmax to generate adaptive weights. This is used to measure the importance of interactions between different modalities. All attention results are multiplied by learnable adaptive weights. After aggregation by summation nodes, the result is then joined with the original input via residual connections. Adding them together, we get five fused features, namely attention fusion features, which are mathematically represented as follows:

[0076] in, Attention dimension per head (in this embodiment) =64, which means the 512-dimensional total features are split into 8 dimensions per head. This represents the scaled dot product cross attention function, whose input is the query vector. Key matrix Sum matrix The output is an attention-weighted context vector.

[0077] Here Modality in the above text Self-mapping key These are two different concepts: It is modal From its own characteristics The key vector obtained by linear projection mapping is used when other modes In terms of modality When performing cross-attention calculations for a target, it acts as a key (i.e., a modality). (providing information about itself to other modalities); and This indicates that the modality will be excluded. Key vectors of all other modalities ( The key matrix is ​​formed by concatenating elements along the sequence dimension. , dimension ( (total number of modes), used for modes Retrieve relevant information from all other modalities. Similarly, This indicates that the modality will be excluded. The value vectors of all other modalities. ( The value matrix formed by concatenating along the sequence dimension, i.e. .

[0078] The fundamental difference between this computation and general self-attention lies in the fact that queries and key-value pairs come from different modalities, and the computation results reflect the semantic associations between modalities rather than the self-associations within a modality. This allows each modality to actively retrieve complementary information from other modalities that is semantically related to itself. Multiple cross-modal association patterns are captured in parallel through a multi-head attention mechanism (8 heads).

[0079] (c) Adaptive weighted fusion submodule (dual credibility weighting) The attention fusion features of each modality are summed using a double weighting method to obtain the double-weighted modal features. .

[0080] A dual weighting mechanism is designed. The content-adaptive weight generator concatenates the attention fusion features from each modality along the feature dimension into a joint vector. Afterwards, a weight vector reflecting the importance of each modality's content is generated through a fully connected layer (2560→5) and a Softmax function. The quality-aware weight generator stacks modal availability masks (1 for actual data acquisition, 0 for missing data or data reconstructed via VAE) into a 5-dimensional vector. This vector is then mapped using a multilayer perceptron and a sigmoid function to generate a quality score vector reflecting the reliability of the data source for each modality. , After element-wise multiplication of the two weights and normalization (ensuring the sum of the weights is 1), a weighted summation is performed on the attention fusion features of each modality. This ensures that the fusion weights simultaneously reflect the two orthogonal dimensions of content importance and data credibility, resulting in double-weighted modal features. .

[0081] (d) Graph convolutional fusion enhancement submodule (modeling of higher-order intermodal associations) Double weighting Although cross-modal semantic information has been integrated, the weighted summation aggregation alone fails to explicitly model high-order topological relationships between modalities. Therefore, a graph convolutional fusion enhancement submodule is introduced to capture the structured dependencies between modalities using graph neural networks.

[0082] (1) Modal relation graph construction: Treat the five modes as graph nodes and initialize the learnable adjacency matrix. After being normalized along the row direction by Softmax, Representing modes Towards mode The normalization strength of the transmitted information. This adjacency matrix is ​​optimized end-to-end during training, automatically discovering potential cooperative and redundant relationships between modalities without the need for manually predefined graph structures.

[0083] (2) Two-layer graph convolutional propagation: feature fusion with attention from each modality ( ) as the initial representation of the nodes, stacked to form a feature matrix , The first layer of graph convolution will convert the adjacency matrix... and Matrix multiplication is performed to achieve cross-modal information propagation. After linear transformation (512→512), GELU activation, and layer normalization, an intermediate representation is obtained: H ( ¹ ) = LayerNorm( GELU( A·H (0) ·W1) ) in This is the first layer learnable weight matrix. This represents the node features after one round of neighborhood aggregation and nonlinear transformation. The second layer of graph convolution uses the same mechanism... Based on this, further propagation and aggregation are performed to capture indirect modal associations within the two-hop neighborhood: H ( ² ) = LayerNorm( GELU( A·H ( ¹ ) ·W2) ) in This is the second-layer learnable weight matrix. The node features, after two rounds of information propagation, have fully encoded the higher-order topological dependencies between modalities.

[0084] (3) Global graph aggregation: The average value of the five node features output by the second-layer graph convolution is taken along the modality dimension to obtain the graph-level representation. It encodes the global semantics after high-order topological associations between modalities.

[0085] (4) Dual-path fusion and final projection: Enhancing features through graph convolution. With dual-weighted features Element-wise addition achieves complementary fusion of modal independent weighting and inter-modal graph structure modeling. The final output of the cross-modal Transformer fusion module is then obtained through a final projection layer (fully connected 512→512 + LayerNorm + GELU). = Projection( + )

[0086] This design makes It simultaneously encodes three levels of information: cross-modal semantic interaction (attention), data credibility awareness (double weighting), and intermodal topological association (graph convolution), providing a complete fusion representation for downstream multi-task prediction.

[0087] In this invention, the use of three feature fusion methods is designed based on the characteristics of multimodal feature vectors. Multimodal feature vectors have the following characteristics: (1) Strong heterogeneity of modal sources - the five modalities are respectively from mammograms, magnetic resonance imaging, ultrasound, pathological images and clinical structured data. The 512-dimensional features output by each encoder are in different semantic subspaces. Direct splicing or simple weighting makes it difficult to align semantic differences; (2) High clinical missing rate - in actual diagnosis and treatment, patients rarely complete all five examinations. If missing modalities are simply set to zero or ignored, it will lead to information loss and prediction bias; (3) Uneven modal quality and credibility - there is an essential difference in information fidelity between the modalities collected in reality and the modalities reconstructed by VAE. The model needs to maintain an appropriate awareness of the uncertainty of the reconstructed features; (4) There are both complementary correlations and redundant interferences between modalities - the representation of the same tumor in different modalities is driven by common biological mechanisms, but different modalities contribute differently to specific prediction tasks. Some modalities may contain noise or low-correlation information.

[0088] Therefore, in this invention: the problem of missing clinical data is solved by using VAE probabilistic completion. The probabilistic constraints of cross-modal joint latent variables are used to generate conditions for missing modalities, so that downstream modules always run on the complete modal set, avoiding inference interruption and performance degradation caused by missing modalities. Attention fusion features are obtained through a cross-modal multi-head attention interaction submodule. These features reflect fine-grained semantic complementarity relationships between modalities—each modality actively retrieves relevant information from other modalities as its own query, and adaptively weights each modality pairwise. The differential importance of interactions between different modal pairs was further quantified, enabling complementary information to be selectively enhanced and redundant information to be naturally suppressed. The weighted modal features are obtained through an adaptive weighted fusion submodule. These weighted modal features can overcome the problem of uneven reliability of modal data sources—content weight. The contribution of each modality is evaluated from a semantic discrimination perspective, with quality weights. The credibility of each modality is evaluated from the perspective of data acquisition source (real acquisition tends to 1, VAE reconstruction tends to a lower value). The two-dimensional joint constraint ensures that the model neither ignores the supplementary information of the reconstructed modality nor over-relys on it. By fusing and enhancing the above three features through graph convolution and then jointly outputting them, it is possible to simultaneously achieve the synergistic optimization of three orthogonal objectives: robust missing feature completion, cross-modal semantic interaction, and quality-aware credibility assessment. VAE completion ensures modal integrity, attention interaction mines cross-modal complementarity, and dual weighting guarantees fusion credibility. The three form a progressively advancing fusion pipeline, resulting in a final fused feature. It possesses robust and complete representation capabilities even under real-world clinical conditions characterized by incomplete modalities and uneven quality, providing a reliable multimodal fusion foundation for downstream prognostic prediction.

[0089] After two layers of graph convolution, global average pooling is performed on all node features to obtain graph-level aggregated features. This feature encodes intermodal interaction information and is fused with the output of the subsequent spatiotemporal graph network at the final merging node via a gated addition unit.

[0090] 3. Spatiotemporal Graph Network Module The structure diagram of the spatiotemporal graph network module is as follows: Figure 12 As shown, this module receives the patient's longitudinal medical visit sequence { }, each of which For the first Global multimodal fusion features output by the cross-modal Transformer fusion module during the second medical visit. The sequence also includes the actual time interval between each adjacent visit. This is used for temporal location encoding of time-series branches. Within each time point, the global fusion features are decomposed into multiple spatial nodes (corresponding to tumor image sub-regions and regional lymph nodes) by a node encoder, and a tumor spatial topology map is constructed through a learnable spatial adjacency matrix to characterize the spatial heterogeneity distribution of the tumor at that time point. Therefore, the longitudinal medical visit sequence { The system fully records the multimodal semantic evolution trajectory of patients from baseline to each follow-up visit and the dynamic changes in tumor spatial structure, providing an information foundation for spatiotemporal joint modeling. The module adopts a separate architecture of "first modeling the time dimension, then modeling the spatial dimension, and finally gated fusion" to obtain spatiotemporal fusion feature vectors.

[0091] Spatiotemporal graph construction Vertical medical visit sequence It is transformed into a multi-time-point feature sequence. Each of these sequences... That is, the global fusion vector output by the cross-modal Transformer fusion module for this medical visit. , Essentially, it is a 512-dimensional real vector that encodes the integrated semantics of the patient's five modalities after fusion at the t-th visit; the sequence fully records the evolution trajectory of the patient's multimodal semantics from baseline to each follow-up visit.

[0092] (1) Temporal branching - time dynamic modeling First, time-location coding is performed: based on the actual time intervals between each patient's visit and the baseline. ( ,in Corresponding to baseline medical visits), it is mapped to a temporal embedding vector through a learnable linear projection:

[0093] in, Represents the time projection weight matrix. This represents the time bias vector.

[0094] Temporal embedding vector Global fusion vector at the corresponding time point Add elements one by one to obtain the time-coded feature sequence.

[0095] ,in, This indicates the corresponding point in time.

[0096] in This represents the actual number of days (or months) between the i-th and i+1-th visits, reflecting the non-uniform sparsity of follow-up—the frequency and interval of follow-up visits vary significantly among different patients. This encoding enables the model to perceive the true length of the time scale rather than treating each visit as equally spaced. The temporally encoded feature sequence is fed into a temporal Transformer encoder, which relies on a multi-head self-attention mechanism to capture the evolutionary pattern of disease progression over time, outputting context-enhanced temporal dependent features at each time point.

[0097] Subsequently, the feature differences between adjacent visit time points are calculated using an incremental feature extractor:

[0098] in, This represents the time-dependent feature vector (dimension 512) output by the Transformer encoder at the time of the t-th visit. Let || denote the time-dependent feature vector of the (t+1)th visit, and || denote the vector concatenation operation. and A 1024-dimensional joint vector is formed by connecting the first and last parts along the feature dimension. DeltaEncoder is an incremental feature encoding network composed of fully connected layers (1024→512), which compresses the joint representation of two adjacent visits into a compact encoding of a single change. The output vector encodes the semantic changes in lesions between the t-th and t+1-th visits. Positive dimensions indicate the direction of feature enhancement, while negative dimensions indicate the trend of feature weakening.

[0099] based on Perform tumor growth rate modeling and treatment response modeling separately: the former will... The data is fed into a velocity prediction network composed of fully connected layers to quantitatively analyze the dynamic rate of change of lesion volume or extent; the latter will... The data is fed into a trend analysis network composed of fully connected layers to analyze the morphological and functional changes of lesions under therapeutic interventions (such as before and after neoadjuvant chemotherapy). The outputs of both are spliced ​​and linearly projected before being fused to obtain the temporal dynamic features. This feature encodes the speed and direction of disease evolution along the time axis, and is a compact summary of longitudinal follow-up information.

[0100] (2) Spatial branching - spatial structure modeling The latest time point (the Tth visit) after Transformer encoding is used to construct an anatomical spatial graph structure. The graph nodes are the tumor image sub-region (obtained through feature clustering, set to 8 nodes in this embodiment) and regional lymph nodes. Connection edges are constructed based on spatial proximity and anatomical drainage pathways. The spatial proximity is determined when the Euclidean distance d(i,j) between two nodes in the anatomical coordinate system is less than the spatial proximity threshold δ. δ is set to 5mm based on medical priors, which roughly corresponds to the typical range of tumor microinfiltration in pathology. Anatomical drainage pathways refer to the anatomical connection paths between adjacent drainage sites in the lymphatic system, such as the chain drainage direction from the sentinel lymph node to the axillary first-order lymph node to the axillary second-order lymph node. The resulting graph is represented by a node feature matrix and an adjacency matrix, where each row of the node feature matrix corresponds to a 512-dimensional initial representation of the tumor sub-region or lymph node after linear projection.

[0101] The graph structure is similar to the modal relationship graph in the graph convolutional fusion enhancement submodule (d). Both use a learnable adjacency matrix, normalized by Softmax rows, in conjunction with a two-layer graph convolutional network for message propagation and node feature aggregation. The learnable adjacency matrix is ​​initialized with random values ​​and optimized end-to-end through backpropagation during training, enabling the model to automatically discover which spatial locations have stronger information interaction needs, without relying on manually predefined fixed connection weights. The core difference between this submodule and the graph convolutional fusion enhancement submodule (d) lies in the semantic objects of the graph nodes: here it is a tumor spatial subregion (spatial decomposition within the same patient), while in the graph convolutional fusion enhancement submodule (d) it is a modality type (semantic association across modalities).

[0102] After two layers of graph convolution, the pre-contribution weights of each spatial node are learned through a node attention mechanism. The features of all nodes are then weighted, summed, and aggregated to obtain the spatial structure features. This feature compresses local information from different spatial locations into a single global vector, encoding the degree of spatial heterogeneity of the tumor at the current time point—the greater the difference in features between sub-regions, the stronger the heterogeneity—as well as the potential activation patterns of lymph node metastasis pathways.

[0103] (3) Spatiotemporal feature fusion Time dynamic characteristics Spatial structural features Joint encoding is performed using a gated addition unit:

[0104] in, ∈(0,1) represents the gating coefficient, which controls the proportion of contribution of the temporal branch and the spatial branch to the final fused feature. When the model approaches 1, it focuses on longitudinal evolution information. When approaching 0, the focus is on lateral spatial structure information; The temporal dynamic features of the time-series branch output encode the evolution of the disease along the time axis; The spatial structural features output by the spatial branch encode the spatial heterogeneity distribution of tumors; It is a spatiotemporal fusion feature vector that simultaneously encodes two complementary types of information: how the tumor changes over time and how the tumor is distributed spatially.

[0105] Gating coefficient Adaptive computation via gated generative networks:

[0106] in, The learnable weight matrix of the gated generative network has dimensions (1024, 1), and the concatenated 1024-dimensional joint features are mapped to scalar logit values. For the learnable bias term of the gated generative network, a scalar; This means concatenating two 512-dimensional features along their respective feature dimensions to form a 1024-dimensional joint vector; Sigmoid(·) is the Sigmoid activation function, defined as... Map any real number to the interval (0,1) such that It has a probabilistic interpretation—which can be understood as the relative reliability of temporal information relative to spatial information in the current sample.

[0107] This adaptive gating mechanism enables the model to dynamically adjust the fusion ratio based on sample characteristics: for patients with multiple follow-up visits and long time spans... The trend is towards increasing to fully utilize longitudinal evolution information; for patients with only baseline data and no follow-up records, As the number of followers decreases, the model degenerates into one that primarily relies on spatial structure features, thus maintaining compatibility with patients of varying follow-up completeness within a unified framework. Final output. ∈ 5 ¹² serves as a spatiotemporal fusion feature vector for use by downstream multi-task prediction heads.

[0108] Finally, the spatiotemporal graph network module outputs... As a spatiotemporal fusion feature vector. Subsequently, Final output with cross-modal Transformer fusion module f fused (At the latest point in time) Cross-module fusion is performed via gated addition units:

[0109]

[0110] in, For learnable weight vectors, For bias terms; For learnable scalar gating coefficients, by and After splicing, it is adaptively generated through linear mapping and the Sigmoid function; It encodes joint information of longitudinal temporal dynamics and tumor spatial structure; The final output of the cross-modal Transformer fusion module already contains complete information on cross-modal semantic interaction, dual credibility weighting, and inter-modal graph convolution enhancement. After gated weighting fusion, It simultaneously encodes information at three levels: cross-modal semantic interaction, modal topological association, and spatiotemporal evolution patterns. It provides the evolution trajectory and spatial distribution of cross-modal semantics along the time axis. It provides a cross-modal static snapshot of the current time point, and the complementary fusion of the two enables the final representation to possess both vertical dynamic perception and horizontal multimodal semantic depth. After further transformation by a multi-layer network, it serves as a shared feature representation, which is then used in parallel by the five downstream prediction heads.

[0111] 4. Multi-task prediction head module The structure diagram of the multi-task prediction head module is as follows: Figure 13 As shown, in the shared spatiotemporal fusion feature representation Based on this, four parallel task-specific prediction heads are set up, which simultaneously output four clinical prediction results.

[0112] (a) Cox Proportional Hazard Model Branch: Employing a neural Cox proportional hazards model, this branch outputs hazard ratios via a linear layer and exponential activation, supplemented by survival curve visualization. Internally, this branch has two parallel sub-output heads sharing a base hazard function, predicting 1-year / 3-year / 5-year survival probabilities for disease-free survival (DFS) and overall survival (OS), respectively. The activation function is linear, and the loss function is the Cox negative log-partial likelihood loss, specifically designed to address the right censoring problem prevalent in medical survival data, enabling unbiased estimation even with incomplete follow-up data.

[0113] (b) Risk Stratification Tri-Classifier Branch: Patients are divided into three risk groups—low-risk, medium-risk, and high-risk—through a linear layer and Softmax activation, and the results are displayed using a bar chart. The output is a 3-dimensional probability vector, and the loss function is cross-entropy loss.

[0114] (c) Transfer Pattern Binary Classifier Branch: After a linear layer and sigmoid activation, the output is the predicted probability of local transfer and distant transfer. The output is a probability value, and the loss function is the binary cross-entropy loss.

[0115] (d) pCR prediction classifier branch: Outputs the predicted probability of pathological complete remission after linear layer and sigmoid activation. The output is a probability value, and the loss function is Focal loss, specifically designed to address the class imbalance problem caused by low response rate to neoadjuvant chemotherapy. This is achieved by adjusting the focus parameter. This makes the model pay more attention to positive samples that are difficult to classify.

[0116] Uncertainty-weighted multi-task loss module: Assigns a learnable noise parameter to each prediction task. The weight of each task's loss in the total loss is automatically adjusted using this noise parameter, and the total loss function is:

[0117] Where K=4 represents the total number of tasks, corresponding to four tasks: Cox survival analysis, risk stratification, migration pattern prediction, and pCR prediction. This is the sub-loss function for the k-th task; The learnable noise parameters for the k-th task are automatically optimized during training—high-noise tasks automatically receive lower weights, and low-noise tasks automatically receive higher weights.

[0118] This design addresses the issues of varying convergence speeds, inconsistent loss dimensions, and different label noise levels across different clinical endpoints in medical prognostic tasks, avoiding the high cost of manually adjusting the weights of each task.

[0119] 5. Interpretability Analysis Module The structure diagram of the interpretability analysis module is as follows: Figure 14 As shown, the interpretability module receives intermediate feature maps, attention weights, and gradient information from the fusion and prediction layers. Internally, the module provides transparent interpretation for clinical decision-making through four complementary methods.

[0120] The first level—Grad-CAM attention heatmap visualization—generates a multimodal attention heatmap, highlighting the lesion areas with the greatest impact on prediction on the original image. It extracts the adaptive attention weight matrix from cross-modal fusion to generate an inter-modal attention heatmap, demonstrating the intensity and direction of information interaction between different modalities. During forward propagation, cross-modal attention weights, view attention weights, and spatial node attention weights are all stored internally and can be directly converted into heatmaps.

[0121] The second level—SHAP attribution analysis—uses the SHAP method to quantify the contribution of each modality and each clinical feature to individual predictions using bar charts. It calculates the Shapley marginal contribution of each modality feature and each clinical feature under all possible modality combinations, quantifying the average incremental contribution of each input feature to the final prediction result. Its mathematical representation is as follows:

[0122] in: This represents the prediction result of the j-th modality (or clinical feature) for the i-th patient. i The Shapley value (SHAP value) is the average marginal contribution of the feature across all possible modal combinations. A positive value indicates that the feature increases the prediction risk, while a negative value indicates that the feature decreases the prediction risk. The larger the absolute value, the stronger the influence of the feature on individual predictions. m represents the total number of input features involved in the interpretation, which is the sum of the number of all modal features and clinical features (in this embodiment, m=6, corresponding to five modal features plus clinical structured features). S represents a subset of features selected from m features. S represents any subset (including the empty set) of all remaining features after excluding feature j. ), total Possible subset combinations; The cardinality of subset S represents the number of features contained in subset S (i.e., the cardinality of the subset), and its value ranges from 0 to m. 1; express factorial, i.e. , used to calculate the number of permutations of S; express The factorial of , which is the number of permutations of the features that remain unparticipated after the subset S is formed; The normalized weighting factor is equal to the probability that, when all m features are randomly arranged, S is the set of features that comes before feature j. This factor ensures the equivalence of all features being considered in all possible permutations, and makes the Shapley value satisfy the four axioms of efficiency, symmetry, dummy nature, and additivity. This represents the expected prediction value of the model when only the features corresponding to the modality subset S are used as input. The calculation method is as follows: fix the actual values ​​of the features in the subset S, randomly sample the values ​​of the features (including j) that are not in S on all samples, and take the mean of the model output. This represents the expected prediction value of the model when additional feature j is added to subset S; This represents the marginal contribution of feature j given that subset S already exists—that is, the change in the model's predicted value after adding feature j. This difference can be positive or negative, reflecting the incremental impact of feature j under a specific feature combination.

[0123] The third level – Prototype Network Case Reasoning: By displaying representative cases most similar to the current patient through thumbnails, case-based reasoning is realized, providing clinicians with intuitive analogy references.

[0124] The fourth level – counterfactual reasoning: This presents the impact of changes in key variables on prediction results in the form of text bubbles. By intervening in specific modal feature values ​​and observing the direction and magnitude of changes in prediction results, it generates reasoning explanations such as "how the prognosis will change if a certain modal feature changes," providing quantitative evidence for adjusting treatment plans.

[0125] Example 4 This embodiment further defines Embodiment 1 and specifically implements a medical data processing system for generating prognostic reference information for breast cancer.

[0126] 1. Training Data and Preprocessing Five modalities of data and corresponding follow-up labels were collected from breast cancer patients. The follow-up labels included disease-free survival (days), overall survival (days), risk level (low / intermediate / high), metastatic site, and pathological complete remission status.

[0127] The data preprocessing steps are as follows: CLAHE contrast enhancement and breast region cropping are performed on mammogram images. Motion correction and temporal intensity curve registration are performed on DCE-MRI sequences. Noise reduction and region of interest extraction are performed on ultrasound images. Multi-resolution pyramid sampling and color normalization are performed on whole-slide images. Missing value imputation and numerical normalization are performed on clinical features.

[0128] The clinical characteristics dataset includes the following variables: Continuous variables: age (years), tumor size (maximum diameter, mm), number of days from diagnosis to surgery (days).

[0129] Categorical variables: Menopausal status; Molecular markers: Hormone receptor (HR) status, estrogen receptor (ER), progesterone receptor (PR), human epidermal growth factor receptor 2 (HER2); Molecular subtype; TNM staging: Tumor size T stage, lymph node N stage, distant metastasis M stage; Histological grading: Tubule formation, tumor grading based on nuclear features, tumor grading based on fission count, malignancy grade (Nottingham); Histological type; Lesion characteristics (bilateral breast cancer, multifocal, contralateral breast involvement); Invasion characteristics: Lymphadenopathy.

[0130] 2. Detailed parameters of the model architecture (a) Breast X-ray Image Feature Extractor The input size is 4×768×768. After CLAHE enhancement, it operates in dual-path parallelism. The upper-path Swin Transformer uses swinv2_base_window12_192_22k pre-trained weights, with a window size of 12×12, an input resolution of 192×192, and 192 output channels. The lower-path EfficientNet uses EfficientNet-V2 pre-trained weights, with 128 output channels. After concatenation, the number of channels is 320. After symmetric perception cross-attention (8 heads), the output is 2×512, and after attention pooling, the output is 512-dimensional.

[0131] (ii) DCE-MRI Feature Extractor The input dimensions are 5×128×128×64. Branch 1: 3D ViT: Block size 8×8×8×1, embedding dimension 192, 6-layer Transformer encoder, 8 heads, output 192 dimensions. Branch 2: 3D ResNeXt: Residual block stacking, output 256 dimensions. Branch 3: Pharmacokinetic encoder: 3 fully connected layers (64→128→64), output 64 dimensions. The three outputs are fused by attention (8 heads) to output 512 dimensions.

[0132] (III) Ultrasonic Feature Extractor The input size is 16×224×224. ConvNeXt V2 uses convnextv2_atto pre-trained weights, with a 768-dimensional output. The 3D convolutional kernel size is 3×3×3, with a stride of 1 and padding of 1, resulting in a 512-dimensional output. The temporal attention has 8 heads, also with a 512-dimensional output.

[0133] (iv) WSI Feature Extractor for Whole-Slide Pathological Images Sixteen patches were downsampled by 40x, each patch measuring 256×256 pixels. The ViT (VisionTransformer) encoder used pre-trained weights with vit_base_patch16_224 (a model with an input size of 224×224 and a single patch size of 16×16), outputting 768 dimensions. Attention-based multi-instance (AMIL) aggregation was employed: the tumor region, stromal region, and lymphocyte region each output 512 dimensions, which were then concatenated and linearly projected to output 512 dimensions.

[0134] (v) Clinical Feature Encoder Category embedding table: nn.Embedding(num_categories (total number of categories), 32). After continuous normalization, it is concatenated with the embeddings. 3-layer MLP: Input → 128 → 64 → 256, ReLU activation, output is linearly projected to 512 dimensions.

[0135] (vi) Cross-modal Transformer fusion module VAE generator: independent encoders for each modality (3-layer MLP), joint encoder (256-dimensional), conditional decoder (3-layer MLP). Cross-attention: 8 heads. =64. Graph convolution: 2 layers, hidden dimension 512, adjacency matrix 5×5, learnable.

[0136] (vii) Spatiotemporal Graph Network Module Spatiotemporal graph construction: temporal edges (added when data exists at adjacent time points), spatial edges (added when the anatomical coordinate distance is <5mm). Adaptive edge weight calculation: .

[0137] Graph convolution propagation: 3 iterations, termination condition Temporal Transformer: 3-layer encoder, 8 heads, 2048-dimensional feedforward, dropout rate 0.1. Transformation Encoder: 3-layer MLP (1024→512→256→512). Spatial Graph Network: 8 nodes, 2-layer graph convolution, 8×8 learnable adjacency matrix. Gated Fusion: Learnable scalar λ.

[0138] (viii) Multi-task prediction head module All four task heads are linear layers (512 → output dimensions). DFS (Disease-Free Survival) and OS (Overall Survival): Output 1 dimension (hazard ratio). Risk stratification: Output 3 dimensions (Softmax). Metastasis and treatment response: Output K dimensions (Sigmoid).

[0139] 3. Model Training Hardware and software environment: 4 NVIDIA A100 SXM 40GB GPUs (total VRAM 160GB), AMD EPYC 7742 64-core CPU, 512GB DDR4 system memory. Software stack: PyTorch 2.0.1, CUDA 11.8, cuDNN 8.7.0, Python 3.10. Mixed-precision training uses torch.cuda.amp (automatic mixed precision), with gradient accumulation steps set to 2 (effective batch size = 32).

[0140] Optimizer: AdamW optimizer is used, betas (hyperparameters) = (0.9, 0.999), epsilon (numerical stability constant) = 10. -8 Initial learning rate 10 -4 Parameter grouping strategy: Encoder bone intervention training parameter learning rate 10. -5 (Reduced by 10 times), learning rate of encoder head and new layer is 10. -4 The learning rate of the fusion module and the spatiotemporal network is 10. -4 Weight decay of 10 -2 This applies only to weight parameters, not to bias and LayerNorm parameters.

[0141] Learning rate scheduling: A cosine annealing scheduling strategy is adopted, with T_max (number of cosine cycle iterations) = 100 (the same as the total number of training iterations) and eta_min (lower learning rate limit) = 10. -6 The first 5 epochs (rounds) use linear warm-up, increasing linearly from eta_min to 10. -4 .

[0142] Training epochs and early stopping: Total training epochs: 200. The early stopping strategy involves monitoring the consistency index of disease-free survival on the validation set. If it does not improve for 10 consecutive epochs (improvement threshold 0.001), training is stopped and the optimal model checkpoint is restored. In actual training, early stopping is triggered on average between epochs 68 and 75.

[0143] Data augmentation: Image modal augmentation includes random horizontal flip (probability 0.5), random rotation ±10° (probability 0.3), random brightness ±15% (probability 0.5), random contrast ±15% (probability 0.5), and Mixup. α =0.2, probability 0.3). Time series enhancement includes equivalent feature perturbation with random time shift ±7 days. Pathological WSI enhancement includes random rotation of 90° / 180° / 270° (probability 0.25 each) and H&E color perturbation (probability 0.5). Clinical feature enhancement includes continuous variables plus Gaussian noise ( σ =0.01 × the standard deviation of this feature).

[0144] Regularization: The Dropout ratio is the default value of the pre-trained model within each encoder, and is uniformly 0.1 for the fusion module and the spatiotemporal module. Gradient clipping uses a global L2 norm threshold of 1.0. KL loss weights. =0.01 (Controls the regularization strength of the latent space in VAE completion). Mode dropout probability. =0.1 (During training, the availability mask of a certain modality is randomly set to 0 to enhance the robustness of VAE completion).

[0145] Validation strategy: A 5-fold cross-validation was employed, with patients stratified. Training and testing were completed independently within each fold, and the final report showed the mean ± standard deviation of the 5-fold test set metrics.

[0146] 4. Verification of technical effectiveness (A) Consistency Index Analysis Figure 15 ) To evaluate the predictive performance of the proposed model, this study systematically compares it with several baseline methods. These baseline methods include VGG-16, ResNet-50, Cox proportional hazards model, random survival forest (RSF), and deep survival model DeepSurv.

[0147] like Figure 15As shown, the system consistency index proposed in this invention reaches 0.812, significantly outperforming baseline models such as VGG-16 (0.716), ResNet-50 (0.742), Cox-PH proportional hazards model (0.698), random survival forest RSF (0.731), and deep survival model DeepSurv (0.756). The error bars represent the 95% confidence interval. This result demonstrates a significant advantage of this invention in terms of the accuracy of prognostic information.

[0148] (B) Analysis of the area under the time-dependent receiver operating characteristic curve (ROC) Figure 16 ) like Figure 16 As shown, time-dependent AUC analysis reveals that the model consistently demonstrates excellent discriminative ability throughout the follow-up period. Specifically, it maintains high predictive performance at the 1-year (AUC=0.86), 3-year (AUC=0.84), and 5-year (AUC=0.78) prediction time points, validating the model's stable discriminative ability across different time windows. These results indicate that the present invention can accurately predict not only short-term prognosis but also maintain good performance in long-term survival prediction.

[0149] (C) Model calibration performance evaluation Figure 17 ) The model calibration performance evaluation results based on the comprehensive Brier score are as follows: Figure 17 As shown, this invention achieved optimal calibration results (IBS=0.142), indicating good consistency between its predicted probabilities and actual observations, outperforming the comparative methods. Good calibration performance means that the risk probabilities output by the model have reliable clinical reference value.

[0150] (D) Subgroup analysis Figure 18 ) To validate the robustness and generalization ability of the model, this study further evaluated its predictive performance in various patient subgroups, covering the overall cohort, hormone receptor status (ER positive / ER negative), disease stage (Stage I–III), HER2+ (human epidermal growth factor receptor 2 positive), and TNBC (triple-negative breast cancer). Figure 18 As shown, the performance of each subgroup was consistent, and the model maintained a stable prediction accuracy in patient groups with different clinical characteristics, demonstrating good universality and robustness, and verifying the applicability of the present invention in real clinical diverse patient groups.

[0151] (E) Ablation test ( Figure 19 ) Ablation experiments quantified the contribution of each input mode to the model performance. For example... Figure 19As shown, the model achieved optimal performance (consistency index = 0.812) when all five modalities (ultrasound, MRI, mammography, pathological examination, and clinical features) were integrated. Removing the ultrasound, MRI, mammography, and pathological examination modalities one by one, or retaining only clinical features, all resulted in varying degrees of performance decline: the performance degradation was most significant after removing the pathology modality, followed by MRI and mammography; the removal of the ultrasound modality had a relatively smaller impact but still made a positive contribution. These results confirm the effectiveness of the multimodal fusion strategy and demonstrate that each modality provides complementary rather than redundant information in prognostic data.

[0152] The above five experiments comprehensively demonstrate the effectiveness and feasibility of the technical solution of the present invention from five dimensions: discrimination ability, time stability, calibration performance, generalization ability, and modal contribution.

[0153] 5. Extended Implementation Method In terms of lightweight implementation, considering the computational power requirements of the five modal Transformer graph convolutional networks when deployed in actual hospitals, large models can be compressed into lightweight models through knowledge distillation, or inference can be accelerated using ONNX or TensorRT to adapt to different clinical deployment environments.

[0154] In terms of simplifying implementation, for clinical scenarios with only partial modal data, the feature extractor corresponding to the missing modality can be skipped, and the available modalities can be directly used for cross-modal fusion and prognostic reference information generation. The present invention can still work normally through the variational autoencoder completion mechanism.

[0155] 6. The clinical characteristic data dictionary is shown in Table 1: Table 1:

Claims

1. A medical data processing system for generating prognostic reference information for breast cancer, characterized in that, The system comprises a four-level architecture: The first level is a multimodal feature extraction module, which uses a dedicated feature extractor for each of the five modalities: mammogram images, breast DCE-MRI images, breast ultrasound images, breast whole slide pathology images, and clinical feature data, to map the raw data of each modality to a feature vector space of the same dimension. The second level is the cross-modal Transformer fusion module, which achieves semantic alignment and fusion of heterogeneous modalities through four sub-processes: missing modality probability completion, cross-modal semantic interaction, dual credibility weighting, and high-order association modeling between modalities, to obtain a global fusion vector. ; The third level is the spatiotemporal graph network module, which receives the global fusion vector. A separate architecture is adopted, which first models the time dimension, then the spatial dimension, and finally performs gated fusion, to obtain spatiotemporal fusion features. ; The fourth level is the multi-task prediction head module, which simultaneously outputs disease-free survival, overall survival, risk stratification, metastasis pattern, and treatment response probability based on shared features, and automatically balances the losses of each task through uncertainty weighting; it also integrates four interpretability mechanisms.

2. The medical data processing system for generating breast cancer prognostic reference information according to claim 1, characterized in that, After receiving the mammogram image, the feature extractor first performs contrast-limited adaptive histogram equalization to enhance local contrast. The enhanced image then extracts features through two parallel paths: the upper path outputs a feature map capturing global contextual information, and the lower path outputs a feature map extracting local details. The outputs of the two paths are concatenated and compressed into a feature vector by an attention pooling layer. After receiving the image, the DCE-MRI feature extractor uses a three-way parallel branch to extract complementary features. Branch 1 processes the original image spatial features and outputs spatiotemporal attention features; Branch 2 processes deep convolutional spatial features and outputs hierarchical spatial features; Branch 3 processes quantitative physiological features and outputs pharmacokinetic features. The outputs of the three branches are weighted and summed by a multi-scale attention fusion module to obtain the feature vector. After receiving an image, the ultrasonic feature extractor first extracts spatial features through a convolutional neighborhood network. Then, it models the dynamic changes between frames by sliding a 3D convolution along the time dimension. Next, it uses a temporal attention mechanism to weight and aggregate the features from different frames, outputting spatiotemporal features. Finally, it compresses the features using global average pooling to obtain the feature vector. The WSI feature extractor for whole-section pathological images receives the image, performs global localization at low magnification to extract the region of interest, and performs patch extraction at high magnification. Then, attention-based multi-instance learning is employed, using learnable attention weights to weighted aggregate features from the tumor area, stromal area, and lymphocyte area respectively. The aggregated results from the three areas are then concatenated and compressed into a feature vector through a linear projection layer. ; The clinical feature encoder receives structured tabular data. Categorical variables are encoded into dense vectors through an embedding layer; continuous variables are standardized to zero mean and unit variance through a normalization layer. All variables are concatenated and input into a fully connected network, and finally mapped into feature vectors through a linear projection layer. .

3. The medical data processing system for generating breast cancer prognostic reference information according to claim 2, characterized in that, The ultrasound feature extractor also includes an elastography auxiliary branch, which is an optional path. The input data of the elastography auxiliary branch is the elastography image of the tissue to be detected. The elastography image is first subjected to preliminary feature extraction and downsampling in the first convolutional layer. After batch normalization and ReLU activation, the output is sent to the second convolutional layer for further extraction of high-level mechanical features and downsampling again. Then, the spatial dimension is compressed by adaptive global average pooling. After flattening, it is mapped to an elastic feature vector through a fully connected layer. The obtained elastic feature vector is expanded by padding zeros on the right side and fused with the feature vector obtained by temporal average pooling of the main branch element by element-wise addition.

4. A medical data processing system for generating breast cancer prognostic reference information according to claim 3, characterized in that, Missing mode probability completion is performed through a cross-modal variational autoencoder submodule. The operation involves: establishing independent variational autoencoders for the five modal features extracted from the multimodal feature extraction module; and obtaining independent latent variables for each modality from the posterior distribution through reparameterization. In the joint encoding stage, a joint encoder is set up to concatenate the independent latent variables of all available modalities and map them into low-dimensional joint latent variables. During the conditional decoding phase, for any missing mode Conditional decoder to combine latent variables and modality Independent latent variables Together as conditional inputs, we reconstruct the feature estimate of the missing mode. ; Cross-modal semantic interaction is performed through a cross-modal multi-head attention interaction submodule, where the following operations are performed: for each modality to its characteristics The query is mapped through three independent linear layers. ,key Sum At the same time, other modalities Mapping to key Sum Then, parallel computation Scaled dot product cross attention with all other modalities; For each pair of modes ,Will and After concatenation, the weights are fed into a two-layer fully connected network and then normalized using Softmax to generate adaptive weights. All attention results are multiplied by learnable adaptive weights. After aggregation by summation nodes, the result is then joined with the original input via residual connections. Adding them together, we get five fused features, which are the attention fusion features.

5. A medical data processing system for generating breast cancer prognostic reference information according to claim 4, characterized in that, The dual credibility weighting is performed through an adaptive weighted fusion submodule, which involves performing a dual-weighted summation of the attention fusion features to obtain the dual-weighted modality features. The dual weighting mechanism is as follows: After the content adaptive weight generator concatenates the attention fusion features of each modality along the feature dimension into a joint vector, a weight vector reflecting the importance of content discrimination of each modality is generated through a fully connected layer and a Softmax function. The modality availability mask is stacked into a 5-dimensional vector through a quality-aware weight generator. After mapping through a multilayer perceptron and a Sigmoid function, a quality score vector reflecting the credibility of the data source of each modality is generated. The two weights are multiplied element by element and normalized. Then, a weighted sum is performed on the attention fusion features of each modality. High-order intermodal correlation modeling is performed through a graph convolutional fusion enhancement submodule. The operations are as follows: First, a modal relationship graph is constructed, with the five modalities as graph nodes. A learnable adjacency matrix is ​​initialized and normalized using Softmax. End-to-end training automatically mines modal collaboration and redundancy relationships. Then, two layers of graph convolution are used, with the attention fusion features of each modality as initial node features. Through multiple rounds of neighborhood aggregation and nonlinear transformation, high-order topological dependencies of the modalities are captured. Finally, the output features are aggregated using the global mean to obtain the graph convolutional enhanced features. Combine it with the double-weighted modal features Element-by-element fusion, followed by a projection layer to obtain the final fused feature. .

6. A medical data processing system for generating breast cancer prognostic reference information according to claim 5, characterized in that, The input to the spatiotemporal graph network module is the patient's longitudinal medical visit sequence. }, each of which For the first The global multimodal fusion features output by the cross-modal Transformer fusion module for each patient visit include the actual time intervals between adjacent visits. Temporal dimension modeling is performed through a temporal branch: first, time location encoding is performed, based on the actual time intervals of each patient's visits, using a learnable linear projection mapping to convert them into time interval embedding vectors. These time embedding vectors are then element-wise added to the corresponding global fusion vectors to obtain the time-encoded feature sequence. This time-encoded feature sequence is fed into a temporal Transformer encoder, which outputs context-enhanced temporal dependency features for each time point. An incremental feature extractor calculates the feature differences between adjacent visit time points. ,based on Tumor growth rate modeling and treatment response modeling were performed separately, and their outputs were spliced ​​and linearly projected before being fused to obtain the temporal dynamic features. .

7. A medical data processing system for generating breast cancer prognostic reference information according to claim 6, characterized in that, The spatiotemporal graph network module models spatial structure through spatial branches. At the corresponding time point of medical visit, tumor imaging sub-regions and regional lymph nodes are used as graph nodes. Connection edges of the graph are constructed based on spatial proximity and anatomical drainage pathways. The resulting graph is represented by a node feature matrix and an adjacency matrix. By using a learnable adjacency matrix and normalizing it with Softmax, and then combining it with a two-layer graph convolutional network for message propagation and node feature aggregation, the spatial structure features are obtained. ; Time dynamic characteristics Spatial structural features Joint encoding is performed using a gated addition unit: in, ∈(0,1) is the gating coefficient, which controls the contribution ratio of the temporal branch and the spatial branch to the final fused feature.

8. A medical data processing system for generating breast cancer prognostic reference information according to claim 7, characterized in that, Spatiotemporal fusion features The global fusion vector corresponding to the time point of visit output by the cross-modal Transformer fusion module. f fused Cross-module fusion is performed using gated addition units to obtain a shared spatiotemporal fusion feature representation. : in, For learnable weight vectors, For bias terms; For learnable scalar gating coefficients, by and After splicing, it is adaptively generated through linear mapping and the Sigmoid function; Multi-task prediction head module in shared spatiotemporal fusion feature representation Based on this, four parallel task-specific prediction heads are set up, and four clinical prediction results are output simultaneously: Cox proportional hazards model branch: The neural Cox proportional hazards model is adopted, and the hazard ratio is output through a linear layer and exponential activation. This branch has two parallel sub-output heads that share the basic hazard function and predict the 1-year / 3-year / 5-year survival probability of disease-free survival and overall survival, respectively. Risk stratification tri-classifier branch: Patients are divided into three risk groups: low risk, medium risk, and high risk through linear layer and Softmax activation; The binary classifier branch for the transfer pattern outputs the predicted probabilities of local transfer and distant transfer after a linear layer and sigmoid activation. pCR prediction classifier branch: outputs the predicted probability of complete pathological remission after linear layer and Sigmoid activation.

9. A medical data processing system for generating breast cancer prognostic reference information according to claim 8, characterized in that, The four-fold interpretability mechanism is as follows: The first level is Grad-CAM attention heatmap visualization: generating a multimodal attention heatmap, highlighting the lesion areas that have the greatest impact on prediction on the original image; extracting the adaptive attention weight matrix in cross-modal fusion to generate an intermodal attention heatmap, showing the intensity and direction of information interaction between each modality; The second level is SHAP attribution analysis: using the SHAP method, the contribution of each modality and each clinical feature to the individual prediction is quantified by bar chart, the Shapley marginal contribution value of each modality feature and each clinical feature under all possible modality combinations is calculated, and the average incremental contribution of each input feature to the final prediction result is quantified. The third level is prototype network case reasoning: by displaying representative cases most similar to the current patient through thumbnails, case-based reasoning is realized, providing clinicians with intuitive analogy references; The fourth level, counterfactual reasoning, presents the impact of changes in key variables on prediction results in the form of text bubbles. By intervening in specific modal feature values ​​and observing the direction and magnitude of changes in prediction results, it generates reasoning explanations and provides quantitative basis for adjusting treatment plans.

10. A medical data processing method for generating prognostic reference information for breast cancer, characterized in that, The method includes: Multimodal feature extraction: For five modalities, namely mammogram images, breast DCE-MRI images, breast ultrasound images, breast whole slide pathology images and clinical feature data, a corresponding dedicated feature extractor is used to map the original data of each modality to a feature vector space of the same dimension. Cross-modal Transformer fusion: Through a four-level sub-process, including missing modality probability completion, cross-modal semantic interaction, dual credibility weighting, and high-order intermodal association modeling, semantic alignment and fusion of heterogeneous modalities are achieved, resulting in a global fusion vector. ; Spatiotemporal graph network fusion: receiving global fusion vector A separate architecture is adopted, which first models the time dimension, then the spatial dimension, and finally performs gated fusion, to obtain spatiotemporal fusion features. ; Multi-task prediction: Based on shared features, it simultaneously outputs disease-free survival, overall survival, risk stratification, metastasis pattern, and treatment response probability, and automatically balances the loss of each task through uncertainty weighting; it also integrates four interpretability mechanisms.