A multimodal fusion method for breast cancer prognosis prediction based on deep learning
Through deep learning technology, the integration of multiple data modes and the fusion of dual-flow attention modes and self-attention pooling mechanisms are used to solve the problem of insufficient accuracy of existing breast cancer prognostic prediction methods, and achieve higher accuracy prognostic prediction and better treatment strategies.
Patent Information
- Application Number
- CN202510454717.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing breast cancer prognosis prediction methods rely on single mode data and cannot effectively capture multi-regional multi-scale features and cross-modal synergistic interactions, resulting in insufficient prediction accuracy and high cost.
The multimodal fusion method based on deep learning is adopted to integrate genomic data, imaging data and pathological data, and through dual-flow attention modal fusion and self-attention pooling mechanisms, multi-regional multi-scale features of breast cancer are captured and complementary information of different modalities are fused.
It improves the accuracy of breast cancer prognosis prediction, can capture tumor characteristics more comprehensively, provides patients with better treatment strategies, and reduces detection costs.
Smart Images

Figure CN119963563B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical artificial intelligence technology, and in particular to a multimodal fusion method for breast cancer prognosis prediction based on deep learning. Background Art
[0002] Breast cancer is the most frequently diagnosed cancer in women worldwide, with high recurrence rates and markedly variable prognoses. Globally, approximately 20–30% of patients with early-stage breast cancer experience recurrence, often resulting in metastatic disease and poor survival outcomes. Among the most common subtypes, invasive ductal carcinoma (IDC) and invasive lobular carcinoma (ILC) have distinct tumor biology, recurrence patterns, and prognostic implications. Existing prognostic prediction methods, such as Oncotype DX and MammaPrint, rely primarily on molecular biomarkers, which estimate the risk of recurrence and benefit of chemotherapy in early-stage patients by evaluating transcriptomic data (gene expression profiles). While these tools have advanced personalized care, their reliance on transcriptomic data has inherent limitations. Transcriptomic profiles capture only a small fraction of tumor heterogeneity and often fail to account for spatial, temporal, and morphological complexity, particularly in IDC and ILC, which exhibit distinct molecular and histological features. Furthermore, the high cost and limited accessibility of such assays restrict their clinical application, especially in resource-constrained settings.
[0003] Advances in artificial intelligence (AI) have expanded prognostic capabilities by leveraging radiology and histopathology imaging data. AI-driven models can automatically perform tumor grading, subtype classification, and risk stratification using dynamic contrast-enhanced MRI (DCE-MRI) or whole slide pathology images (WSIs). However, most existing AI frameworks operate in isolation, analyzing a single modality such as radiology or pathology, thereby ignoring the complementary insights provided by multimodal integration. Radiology, especially through dynamic contrast-enhanced MRI (DCE-MRI), can capture key features of tumor vasculature, such as perfusion dynamics, angiogenesis, and spatial organization of vascular networks, which are key indicators of tumor invasiveness and metastatic potential. At the same time, pathology reveals not only cellular architecture and molecular heterogeneity, but also prognostic markers such as collagen fiber orientation, stromal composition, and immune cell infiltration patterns, which are closely related to breast cancer progression and recurrence. However, separate analysis of these modalities is insufficient to model the synergistic interaction of these features. Furthermore, the lack of spatiotemporal integration across modalities, such as correlation of longitudinal imaging changes with evolving histopathological features, limits the ability to capture the dynamic progression of breast cancer. Summary of the invention
[0004] Purpose of the invention: The technical problem to be solved by the present invention is to address the deficiencies of the prior art and provide a multimodal fusion method for breast cancer prognosis prediction based on deep learning. By integrating genomic data, imaging data and pathological data, the method can accurately predict the survival rate of breast cancer patients. The method can capture the multi-regional and multi-scale characteristics of breast cancer, and through a cross-modal attention mechanism, it can fuse the complementary information of different modalities, thereby providing a more comprehensive prognosis prediction.
[0005] The method of the present invention comprises the following steps:
[0006] Step 1: Preprocess the pathological images of breast cancer patients: read the whole-slice pathological images WSI of the patients, and use the otsu threshold method to detect the tissue area and background; align the rotation, scaling or translation errors between different whole-slice pathological images WSI through the pre-alignment module; process the aligned whole-slice pathological images WSI to create a binary mask to mark the tissue area, and generate 256×256 non-overlapping tissue image blocks at the highest resolution; then use the hierarchical extraction method to extract features from the image blocks;
[0007] Step 2: Preprocess the patient's medical image DCE-MRI, resample the patient's medical image DCE-MRI to 1 mm equal voxel resolution, and then randomly crop 4 sub-volumes of 48×48×3 along the Z axis of the patient's medical image DCE-MRI; pre-align the patient's medical image DCE-MRI using a method based on rigid or non-rigid registration, and construct a registration optimization problem using mutual information MI as a similarity measure to obtain the optimal transformation matrix T MRI :
[0008] ,
[0009] in, For the reference MRI image, is the MRI image to be registered, Indicates that After affine transformation matrix T 1 The image obtained later, Represents mutual information, reflecting the statistical dependence between the two images; through the above optimization process, the optimal transformation matrix T can be obtained MRI , thereby achieving accurate pre-alignment between sequences;
[0010] Step 3: for the consistency of different modalities, dual-stream attention modality fusion is used, and the dual-stream attention modality fusion includes encoder I and encoder P;
[0011] The encoder I and the encoder P both include a regional fusion module, and the regional fusion module includes a window multi-head self-attention mechanism (W-MSA) and a shift window strategy (SW-MSA);
[0012] The whole slice pathology image WSI processes the feature vector sequence through the regional fusion module.
[0013] Medical imaging images DCE-MRI process the original image sub-images through the regional fusion module;
[0014] The regional fusion module in the encoder P also includes patch merging, which groups adjacent image blocks twice along the depth dimension, and then the encoder P gradually reduces the input resolution to fuse multi-region features within the imaging modality;
[0015] The self-attention pooling mechanism is used to capture the cross-modal pathological features output by encoder I and the image features output by encoder P and fuse the cross-modal features:
[0016] Project the 2D features of the whole-slice pathology image WSI processed by encoder I and the 3D features of the medical image DCE-MRI processed by encoder P to the same dimension, and calculate the interaction weight between pathology and image;
[0017] Step 4, making prognosis prediction;
[0018] Step 5: Perform visual interpretation by calculating the integrated gradient (IG) to generate an attention heat map, mapping it to the original whole-slice pathology image WSI, and generating a radiomics feature expression heat map on the medical image DCE-MRI to achieve multi-scale visualization of morphological attributes.
[0019] Step 1 includes: the pre-alignment module uses the feature matching method to construct the following affine transformation matrix T 1 :
[0020] ,
[0021] The parameters a, b, t are obtained by minimizing the feature distance of the overlapping area. x ,t y , a is the image scaling ratio (if a>1, the image is enlarged; if 0<a<1, the image is reduced), and affects the rotation, b is the image shearing degree, and also affects the rotation (if b≠0, the image will tilt along the pixel coordinate XY axis), t x is the displacement of the image on the pixel coordinate X axis, t y is the amount of movement of the image on the Y axis of the pixel coordinate, so as to align the whole slice pathology image WSI.
[0022] Step 1 also includes: extracting a 192-dimensional feature vector for each image block through self-supervised learning using a self-supervised feature extractor HIPT, subdividing the input 256×256 image block into 16×16 image blocks, and extracting fine features at the cell level , extract tissue-level features from 256×256 image blocks ;
[0023] Aggregate adjacent 256×256 image blocks into 2048×2048 image blocks, and then extract macro-scale features in layers ;
[0024] Finally, hierarchical representation aggregation is performed to obtain the features of the whole-slice pathology image WSI. :
[0025] ,
[0026] in, Indicates that the whole-slice pathology image WSI is at the cellular level and the sequence length is 256, i represents the sequence number of the 16×16 image block; j represents the sequence number of the 256×256 image block; k represents the sequence number of the 2048×2048 image block; Represents the output features; N represents the number of 2048×2048 image blocks in the whole-slice pathology image WSI;
[0027] In view of the size differences of image blocks of different patients and different whole-slice pathology images WSI, all images in the batch are uniformly adjusted to the maximum segment size through dynamic filling technology to optimize memory;
[0028] A multi-scale attention mechanism is introduced after each layer of feature extraction. For a given feature map F, the multi-scale attention mechanism The calculation formula is:
[0029] ,
[0030] in is the feature representation at different resolutions, are learnable weights, It is the standard attention calculation module.
[0031] Step 2 includes: the resampling formula is:
[0032] ,
[0033] in, represents the high-resolution image generated after resampling, For low-resolution images, is the resampling operation function, with parameters Control, this function realizes the mapping from low-resolution image to high-resolution image. Represents the parameters that need to be learned in the model or algorithm, is the diffusion noise estimation function, is the step size factor, Represents the diffusion noise estimation function The gradient of the image describes the rate of change of the noise in the image. is the noise compensation term.
[0034] The window multi-head self-attention mechanism can capture the dependency between image blocks in a global scope by parallelly calculating the mutual importance weights between preprocessed image blocks in the window area, that is, it can capture global information.
[0035] The shift window strategy is used to move the window focus by M / 2 pixels along the pixel coordinates X and Y directions, where M represents the number of image blocks in each direction, thereby enhancing feature interaction across window areas and capturing local information;
[0036] The block merging is used to gradually reduce the spatial dimension of the feature map, merging 2×2 (or 3×3) areas each time to reduce the resolution to or After merging two or more image blocks, the feature dimension is improved through splicing or linear transformation, so that the number of channels gradually increases (for example, from the initial number of channels U to 2U and then to 4U).
[0037] In step 3, the encoder I uses a 2D shifted window of size 2×2 to capture local information and a window multi-head self-attention mechanism to capture global information of the pathological modality;
[0038] The encoder P uses a 3D shifted window of size 4×4×4 to capture local information and a multi-head self-attention mechanism to capture global information of the image modality;
[0039] When processing whole-slice pathological images WSI, the improved 2D sliding window and local window multi-head self-attention mechanism are used to fuse local and global information. The image blocks obtained after preprocessing are input, and each image block is first divided into windows. Then, the local features are extracted through the local window multi-head self-attention mechanism. At the same time, the shift window strategy is used to solve the cross-window interaction, thereby indirectly enhancing the global information propagation. The calculation formula is:
[0040] ,
[0041] in, is the final output feature at position i (combining the features of normalization, self-attention and convolution), is the input feature at position i, i is the position index of the pathological image block where the input feature is located, is layer normalization, LocalMSA is a local window multi-head self-attention module, which calculates the attention distribution within the local window, and Conv is a convolution operation;
[0042] For the sub-volume of the image, the encoder P is used, 3D sliding window division is adopted, and the 3D window multi-head self-attention module is used to fuse the local and global information. Block merging is introduced, adjacent small blocks are grouped along the depth direction, and the merging process is repeated to gradually reduce the resolution to achieve multi-region information integration.
[0043] In step 3, the features output by encoder I and encoder P are fused using self-attention pooling. The pathological features output by encoder I are , the image feature output by encoder P is ,in is a real number space, N and M are the number of tokens of encoder I and encoder P respectively (token refers to the basic representation unit obtained after feature extraction, and each token corresponds to a feature vector of a local area or image block), and d is the feature dimension; first, the output features are concatenated, and the concatenated features are obtained using the following formula :
[0044] ,
[0045] Next, it is mapped to query Q, key K and value V through a learnable linear transformation:
[0046] , , ,
[0047] in, is a learnable linear transformation matrix, is the bias term, is the dimension after projection;
[0048] In order to capture the local interactions and multi-scale information between different modalities, two additional terms are introduced in the standard scaled dot product attention, namely, the position bias term and multi-scale weight adjustment ;
[0049] Calculate the bias term at position (i, j) :
[0050] ,
[0051] in, are the location information of the i-th token and the j-th token respectively. is a parameter matrix;
[0052] Calculate the multi-scale weight adjustment at position (i, j) :
[0053] ,
[0054] Among them, G represents the number of scales (that is, how many features are there at different resolutions), and They represent the feature vector of the i-th token at scale g and the feature vector of the j-th token at scale g, respectively. represents element-wise multiplication (Hadamard product), is the learnable weight at scale g, is the importance weight of scale g, which is the bias introduced for information of different resolutions to adjust the importance of cross-modal multi-scale features. This item can be mapped to the weighted item between token pairs after learning a set of weight vectors;
[0055] Attention Matrix The calculation formula is:
[0056] ,
[0057] Wherein, T represents the transposition symbol;
[0058] Calculate the fused output :
[0059] .
[0060] In step 4, the Cox partial negative log-likelihood loss function is constructed to output the survival risk score. The Cox partial negative log-likelihood loss function It is expressed as:
[0061] ,
[0062] Among them, C is the sample set of all events (not censored); R h is the risk set, which means that the event of individual h occurs at time t h The set of individuals who are still at risk (i.e., no event or censoring has occurred and they are still alive) when β is the model parameter; x h is the feature vector of the hth patient sample; x j is the feature vector of the jth patient sample; exp is the natural exponential function.
[0063] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the described method.
[0064] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction is run on a computer, the steps of the method described are executed.
[0065] The present invention has the following beneficial effects:
[0066] The present invention can combine DCE-MRI radiomics with WSI-derived tissue morphological features;
[0067] The present invention can capture multi-scale tumor characteristics, i.e., from macroscopic vascular dynamics to microscopic cellular characteristics;
[0068] The present invention can improve the prognostic accuracy of breast cancer and help formulate better treatment strategies for patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 Flow chart of the method of the present invention.
[0070] Figure 2 It is a structural diagram of the regional fusion module of the present invention. DETAILED DESCRIPTION
[0071] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more clear.
[0072] like Figure 1 As shown, the embodiment of the present invention provides a multimodal fusion method for breast cancer prognosis prediction based on deep learning, comprising the following steps:
[0073] The TCGA dataset was used in the present invention. Figure 1 The TCGA dataset contains WSI, DCE-MRI and labeled data images of 127 breast cancer patients.
[0074] like Figure 1 As shown in the preprocessing steps in , the patient's pathological data and image data are preprocessed respectively. In the example of the present invention, the whole slice pathological image (WSI) of the patient is read and the tissue area and background are detected using the otsu threshold method. For the rotation, scaling or translation errors that may exist between different WSIs, a pre-alignment module is introduced, and the affine transformation matrix is constructed using the feature matching method:
[0075] ,
[0076] The parameters a, b, tx, and ty are obtained by minimizing the feature distance of the overlapping area to align each WSI;
[0077] The pre-aligned WSI is processed to create a binary mask to mark the tissue area, and a 256×256 non-overlapping tissue image patch is generated at the highest resolution; then the feature extraction of the image patch is performed using a hierarchical extraction method.
[0078] After self-supervised learning, the self-supervised feature extractor HIPT is used to extract the 192-dimensional feature vector of each small block. The 256×256 image block is input and subdivided into 16×16 small blocks to extract fine features at the cell level. , process 256×256 image blocks to extract tissue-level features , aggregate adjacent 256×256 blocks into 2048×2048 image blocks to hierarchically extract macro-scale features , extracting a wide range of cell tissue relationships through local and global image comparison learning, and finally performing hierarchical representation aggregation to obtain the features of WSI ; In view of the size differences between different patients and different WSI segments, all images in the batch are uniformly adjusted to the maximum segment size through dynamic filling technology to optimize memory; the specific formula is expressed as:
[0079] ,
[0080] in, Indicates that it is at the cell level and the sequence length is 256, i represents the i-th 16×16 image block; j represents the j-th 256×256 image block; k represents the k-th 2048×2048 image block; represents the output features; N represents the number of 2048×2048 image blocks in WSI.
[0081] In view of the size differences between different patients and different WSI segments, all images in the batch were uniformly adjusted to the maximum segment size through dynamic filling technology to optimize memory.
[0082] Since the patient's image data in the present invention is already DCE-MRI, the preprocessing of the patient's DCE-MRI image data only needs to be resampled to a voxel resolution of 1 mm. In order to better capture local and global information at different scales, a multi-scale (multi-resolution) attention mechanism is introduced after each layer of feature extraction; for a given feature map F, its multi-scale attention is expressed as:
[0083] ,
[0084] in is the feature representation at different resolutions, are learnable weights, It is the standard attention calculation module.
[0085] After preprocessing the data, the processed data is fused in multiple regions and at multiple scales, such as Figure 1As shown in the multi-region multi-scale fusion steps in , for the consistency of different modalities, it is proposed to use dual-stream attention modal fusion, in which the regional fusion module is used to model the multi-scale image features to achieve multi-scale fusion, in which the pathology branch processes the feature vector sequence through a two-dimensional regional fusion module (encoder I), and the radiology branch uses a three-dimensional regional fusion module (encoder P) to process the original image sub-image; in each 3D regional fusion module, there is a small block merging layer to group adjacent image blocks along the depth dimension. After repeating this process twice, the regional fusion module will gradually reduce the input resolution, thereby fusing the multi-region features in the modality; the self-attention pooling mechanism is used to capture cross-modal pathological microscopic features and radiological macroscopic anatomical features and fuse cross-modal features: the 2D features of the pathological image and the 3D features of the imaging image are projected to the same dimension, and the interaction weights of the pathology and the imaging (i.e., the cross-modal mutuality) are calculated. The regional fusion module parameters of the example of the present invention are shown in Table 1:
[0086] Table 1
[0087]
[0088] For the feature fusion of pathology and imaging, the dual-stream attention modality fusion proposed in the present invention uses a self-attention pooling method to fuse the features output by the encoders of the above two modalities. The output pathological features are , the output image features are , where N and M are the number of tokens respectively (token refers to the basic representation unit obtained after feature extraction, and each token corresponds to a feature vector of a local area or image block), and d is the feature dimension; first, the output features are concatenated:
[0089] ,
[0090] Next, it is mapped to query, key, and value through a learnable linear transformation:
[0091] , , ,
[0092] in, is a learnable linear transformation matrix, is the bias term, d' is the dimension after projection;
[0093] In order to capture local interactions and multi-scale information between different modalities, two additional terms are introduced in the standard scaled dot product attention: position bias and multi-scale weight adjustment ;
[0094] Position offset : Calculated based on the relative position information of each token, it can be formally defined as
[0095] ,
[0096] in, are the location information of the i-th and j-th tokens respectively. Usually implemented by a parameter matrix;
[0097] Multi-scale weight adjustment :
[0098] ,
[0099] Among them, G represents the number of scales (that is, how many features are there at different resolutions), and Respectively represent the feature vectors of the i-th and j-th token at scale g, represents element-wise multiplication (Hadamard product), is the learnable weight at scale g, is the importance weight of scale g, which is the bias introduced for information of different resolutions to adjust the importance of cross-modal multi-scale features. This item can be mapped to the weighted item between token pairs after learning a set of weight vectors.
[0100] Therefore, the attention matrix formula is:
[0101] ,
[0102] Use the attention weight matrix and value vector to calculate the fused output:
[0103] ,
[0104] In this way, the feature output of dual-stream attention modality fusion and attention weighting can be obtained.
[0105] Finally, the survival risk score is output by calculating the Cox partial negative log-likelihood loss function for survival prediction. The Cox partial negative log-likelihood loss function is expressed as:
[0106] ,
[0107] Among them, C is the sample set of all events (not censored); R h is the risk set, which means that the event of individual h occurs at time t h The set of individuals who are still at risk (i.e., no event or censoring has occurred and they are still alive) when β is the model parameter; xh is the feature vector of the hth patient sample; x j is the feature vector of the jth patient sample; exp is the natural exponential function.
[0108] By calculating the integrated gradient (IG) to generate an attention heat map, mapping it to the original WSI, and generating a radiomics feature expression heat map on MRI to achieve multi-scale visualization of morphological attributes.
[0109] The patient information of one of the clinical verification cases in the example of the present invention is: female, 47 years old, estrogen receptor positive ER+, human epidermal growth factor receptor 2 negative HER2-, clinical stage T2N1M0 (according to the tumor staging system, T2 indicates that the size or invasion depth of the tumor is medium, N1 indicates that the patient has a lymph node metastasis, and M0 indicates no distant metastasis).
[0110] Input data: Magnetic resonance imaging MRI: the maximum diameter of the tumor is 2.8 cm, the apparent diffusion coefficient of the lymph nodes ADC = 1.1×10⁻³mm² / s; Pathology: a tumor grading system (Scarff-Bloom-Richardson) grade 7 points (3+2+2).
[0111] The results were as follows: Feature extraction: MRI detected ring-shaped enhancement at the edge of the tumor (a sign of poor prognosis, with a weight of 0.71); pathology revealed focal lymphatic invasion (attention score of 0.83); Risk prediction: output algorithm score = 0.62 (high-risk threshold > 0.55), predicted 3-year disease-free survival (DFS) probability: 42% vs actual value 38% (recurrence time 34 months).
[0112] The effect comparison between the present invention and other methods is shown in Table 2.
[0113] Table 2
[0114]
[0115] Based on the above example analysis, the multimodal fusion method for breast cancer prognosis prediction based on deep learning proposed in the present invention can effectively improve the accuracy of breast cancer prognosis prediction. By utilizing the RF module and self-attention pooling mechanism to perform multi-region multi-scale fusion and cross-modal fusion, more comprehensive breast cancer characteristics can be obtained, providing strong support for personalized treatment strategies.
[0116] The present invention provides a multimodal fusion method for breast cancer prognosis prediction based on deep learning. There are many methods and approaches to implement the technical solution. The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the scope of protection of the present invention. All components not specified in this embodiment can be implemented by existing technologies.
Claims
1. A multimodal fusion method for breast cancer prognosis prediction based on deep learning, characterized in that: The following steps are involved: Step 1: Preprocess the pathological images of breast cancer patients: read the whole slice pathological images WSI of the patients, use the otsu threshold method to detect the tissue area and background; align the rotation, scaling or translation errors between different whole slice pathological images WSI through the pre-alignment module; The aligned whole-slice pathology images WSI were processed to create binary masks to mark tissue regions and generate 256×256 non-overlapping tissue image blocks at the highest resolution. The features of the image blocks were then extracted using a hierarchical extraction method. Step 2: Preprocess the patient's medical image DCE-MRI, resample the patient's medical image DCE-MRI to 1 mm equal voxel resolution, and then randomly crop 4 sub-volumes of 48×48×3 along the Z axis of the patient's medical image DCE-MRI; pre-align the patient's medical image DCE-MRI using a method based on rigid or non-rigid registration, and construct a registration optimization problem using mutual information MI as a similarity measure to obtain the optimal transformation matrix T MRI : , in, For the reference MRI image, is the MRI image to be registered, Indicates that The image obtained after affine transformation matrix T1, represents mutual information; Step 3: for the consistency of different modalities, dual-stream attention modality fusion is used, and the dual-stream attention modality fusion includes encoder I and encoder P; The encoder I and the encoder P both include a regional fusion module, and the regional fusion module includes a window multi-head self-attention mechanism and a shift window strategy; The whole-slice pathology image WSI processes the feature vector sequence through an improved regional fusion module. Medical imaging images DCE-MRI process the original image sub-images through the regional fusion module; The regional fusion module in the encoder P also includes block merging, which groups adjacent image blocks twice along the depth dimension, and then the encoder P gradually reduces the input resolution to fuse multi-region features in the imaging modality; The self-attention pooling mechanism is used to capture the cross-modal pathological features output by encoder I and the image features output by encoder P and fuse the cross-modal features: Project the 2D features of the whole-slice pathology image WSI processed by encoder I and the 3D features of the medical image DCE-MRI processed by encoder P to the same dimension, and calculate the interaction weight between pathology and image; Step 4, making prognosis prediction; Step 5: Perform visual interpretation by calculating the comprehensive gradient to generate an attention heat map, mapping it to the original whole-slice pathology image WSI, and generating a radiomics feature expression heat map on the medical image DCE-MRI to achieve multi-scale visualization of morphological attributes.
2. The method according to claim 1, characterized in that Step 1 includes: the pre-alignment module constructs the following affine transformation matrix T1 using a feature matching method: , The parameters a, b, t are obtained by minimizing the feature distance of the overlapping area. x ,t y , a is the image scaling ratio, b is the image shearing degree, t x is the displacement of the image on the pixel coordinate X axis, t y The amount of movement of the image on the Y axis of the pixel coordinate.
3. The method according to claim 2, characterized in that Step 1 also includes: extracting a 192-dimensional feature vector for each image block through self-supervised learning using a self-supervised feature extractor HIPT, subdividing the input 256×256 image block into 16×16 image blocks, and extracting fine features at the cell level , extract tissue-level features from 256×256 image blocks ; Aggregate adjacent 256×256 image blocks into 2048×2048 image blocks, and then extract macro-scale features in layers ; Finally, hierarchical representation aggregation is performed to obtain the features of the whole-slice pathology image WSI. : , in, Indicates that the whole-slice pathology image WSI is at the cellular level and the sequence length is 256, i represents the sequence number of the 16×16 image block; j represents the sequence number of the 256×256 image block; k represents the sequence number of the 2048×2048 image block; Represents the output features; N represents the number of 2048×2048 image blocks in the whole-slice pathology image WSI; In view of the size differences of image blocks of different patients and different whole-slice pathology images WSI, all images in the batch are uniformly adjusted to the maximum segment size through dynamic filling technology; A multi-scale attention mechanism is introduced after each layer of feature extraction. For a given feature map F, the multi-scale attention mechanism The calculation formula is: , in is the feature representation at different resolutions, are learnable weights, It is the standard attention calculation module.
4. The method according to claim 3, characterized in that Step 2 includes: the resampling formula is: , in, represents the high-resolution image generated after resampling, For low-resolution images, is the resampling operation function, is the diffusion noise estimation function, is the step size factor, Represents the diffusion noise estimation function Regarding the gradient of the image, is the noise compensation term.
5. The method according to claim 4, characterized in that In step 3, the window multi-head self-attention mechanism calculates the mutual importance weights between the preprocessed image blocks in the window area in parallel; The shift window strategy is used to move the window focus along the pixel coordinates X and Y directions by M / 2 pixels, where M represents the number of image blocks in each direction; The block merging is used to gradually reduce the spatial dimension of the feature map, and after merging two or more image blocks, the feature dimension is increased by splicing or linear transformation.
6. The method according to claim 5, characterized in that In step 3, the encoder I uses a 2×2 two-dimensional sliding window to capture local information and a window multi-head self-attention mechanism to capture global information of the pathological modality; The encoder P uses a 4×4×4 three-dimensional sliding window to capture local information and a multi-head self-attention mechanism to capture global information of the image modality; When processing whole-slice pathological images WSI, the improved 2D sliding window and local window multi-head self-attention mechanism are used to fuse local and global information. The image blocks obtained after preprocessing are input, and each image block is first divided into windows. Then, the local features are extracted through the local window multi-head self-attention mechanism. At the same time, the shift window strategy is used to solve the cross-window interaction, thereby indirectly enhancing the global information propagation. The calculation formula is: , in, is the final output feature at position i, is the input feature at position i, i is the position index of the pathological image block where the input feature is located, is layer normalization, LocalMSA is a local window multi-head self-attention module, which calculates the attention distribution within the local window, and Conv is a convolution operation; For the sub-volume of the image, the encoder P is used, 3D sliding window partitioning is adopted, and the 3D window multi-head self-attention module is used to fuse the local and global information. Block merging is introduced to group adjacent small blocks along the depth direction, and the merging process is repeated to gradually reduce the resolution and realize multi-region information integration.
7. The method according to claim 6, characterized in that In step 3, the features output by encoder I and encoder P are fused using self-attention pooling. The pathological features output by encoder I are , the image feature output by encoder P is ,in is a real number space, N and M are the number of tokens of encoder I and encoder P respectively, and d is the feature dimension; first, the output features are concatenated, and the concatenated features are obtained using the following formula : , Next, it is mapped to query Q, key K and value V through a learnable linear transformation: , , , in, is a learnable linear transformation matrix, is the bias term, is the dimension after projection; In order to capture the local interactions and multi-scale information between different modalities, two additional terms are introduced in the standard scaled dot product attention, namely, the position bias term and multi-scale weight adjustment ; Calculate the bias term at position (i, j) : , in, are the location information of the i-th token and the j-th token respectively. is a parameter matrix; Calculate the multi-scale weight adjustment at position (i, j) : , Where G represents the number of scales, and They represent the feature vector of the i-th token at scale g and the feature vector of the j-th token at scale g, respectively. represents element-wise multiplication, is the learnable weight at scale g, is the importance weight of scale g; Attention Matrix The calculation formula is: , Wherein, T represents the transposition symbol; Calculate the fused output : 。 8. The method according to claim 7, characterized in that In step 4, the Cox partial negative log-likelihood loss function is constructed to output the survival risk score. The Cox partial negative log-likelihood loss function It is expressed as: , Among them, C is the sample set of all events; R h is the risk set; β is the model parameter; x h is the feature vector of the hth patient sample; x j is the feature vector of the jth patient sample; exp is the natural exponential function.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
Image segmentation method based on full-scale jump connection U-shaped structure
CN116958541A
Multi-scale feature-based triple negative breast cancer immunophenotype prediction method and system
CN117152506A