Multi-modal image fusion breast tumor diagnosis method, device, equipment and medium

By using multimodal image fusion technology, accurate diagnosis of breast tumors has been achieved, solving the stability problem of single-modal image processing, improving the ability to identify complex lesions, and reducing the risk of misdiagnosis.

CN120876398AInactive Publication Date: 2025-10-31YICHANG CENT PEOPLES HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510974440.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing breast tumor identification technologies mostly employ single-modal image processing, which makes it difficult to maintain stable diagnostic performance in complex or atypical lesions. Furthermore, multimodal image fusion methods suffer from semantic inconsistencies between modalities and a lack of deep interaction mechanisms.

Method used

By acquiring multimodal images for spatial registration, extracting modality-sensitive features, and fusing modality features, the system uses global fusion features to predict breast tumors and trains a classifier using a joint loss function to achieve multimodal image fusion diagnosis of breast tumors.

Benefits of technology

It improves the accuracy and stability of breast tumor diagnosis, and is particularly suitable for identifying cases with dense breast tissue or complex structures, reducing the false positive and false negative rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876398A_ABST
    Figure CN120876398A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal image fusion breast tumor diagnosis method and device, equipment and a medium. The method comprises the following steps: acquiring a multi-modal image of a subject, and performing spatial registration processing to obtain a multi-modal image group; the multi-modal image group comprises a molybdenum target image, a B ultrasonic image and an elastic ultrasonic image; performing modal sensitive feature extraction on each modal of the multi-modal image group to obtain molybdenum target image features, B ultrasonic image features and elastic ultrasonic image features; carrying out fusion processing on the molybdenum target image features, the B ultrasonic image features and the elastic ultrasonic image features to obtain global fusion features; performing breast tumor prediction according to the global fusion features to obtain a prediction result; the prediction result comprises the breast tumor focus existence probability and the breast tumor focus type. By adopting the method, cross-modal semantic alignment can be realized while original information of each modal is reserved, and the capability of discriminating a complex breast tumor focus is enhanced through fusion discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of auxiliary medical diagnosis, and in particular relates to a method, device, equipment and medium for diagnosing breast tumors using multimodal image fusion. Background Technology

[0002] With the development of medical imaging technology, early screening and quantitative diagnosis of breast diseases are increasingly dominated by imaging. In clinical practice, various imaging methods, such as mammography, B-mode ultrasound, and elastography, are widely used in the detection and analysis of breast masses. Among them, mammography, with its high resolution and sensitivity to microcalcifications, is the preferred method for breast cancer screening; ultrasound can detect masses in high-density breast tissue, offering advantages such as real-time detection and no radiation; and elastography provides information on tissue stiffness, helping to determine the benign or malignant tendency of the mass.

[0003] Traditional breast tumor identification techniques typically employ single-modality image processing, relying solely on mammography images to assess density, shape, and calcification distribution, or ultrasound images to classify boundary features and internal echoes. While these methods achieve initial screening of breast lesions to some extent, they often struggle to maintain consistent diagnostic performance in complex or atypical lesions due to the varying advantages and disadvantages of each imaging modality.

[0004] In recent years, some studies have attempted to fuse multimodal images through feature concatenation or result voting to construct fusion classification models to improve diagnostic accuracy. However, current fusion methods still suffer from drawbacks such as semantic inconsistencies between modalities and a lack of deep interaction mechanisms. Summary of the Invention

[0005] Therefore, it is necessary to provide a method, device, equipment, and medium for breast tumor diagnosis that can identify breast tumors by fusing multiple modal images, in order to address the above-mentioned technical problems.

[0006] In a first aspect, this application provides a multimodal image fusion-based method for diagnosing breast tumors, including:

[0007] Multimodal images of the subject were acquired and spatially registered to obtain a multimodal image set, which included mammogram images, B-ultrasound images, and elastography images.

[0008] Modality-sensitive features were extracted from each mode of the multimodal image group to obtain features of molybdenum target image, B-ultrasound image, and elastography image;

[0009] The features of mammogram images, B-mode ultrasound images, and elastography images are fused to obtain global fused features.

[0010] Breast tumor prediction is performed based on global fusion features, and the prediction results include the probability of the presence of breast tumor lesions and the type of breast tumor lesions present.

[0011] In one embodiment, multimodal images of the subject are acquired and spatially registered to obtain a multimodal image set, including:

[0012] Using B-ultrasound images as a reference coordinate system, structural feature points were extracted from mammogram images and elastography images respectively to obtain a set of feature points;

[0013] Each feature point in the feature point set is matched with the feature point in the ultrasound image to obtain an initial set of point pairs.

[0014] Mismatched points are removed from the initial point pair set, and affine transformation matrices mapping the molybdenum target image and the elastic ultrasound image to the reference coordinate system are fitted to obtain coarse registration images; the coarse registration images include the molybdenum target coarse registration image and the elastic ultrasound coarse registration image;

[0015] Affine adjustment is performed on the coarsely registered image based on the structural similarity of the registered images to obtain a multimodal image group.

[0016] In one embodiment, the coarsely registered image is affine adjusted based on the structural similarity of the registered images to obtain a multimodal image set, including:

[0017] The mutual information between the coarse registration image and the B-ultrasound image was calculated to obtain the structural similarity between the coarse registration image of the mammogram and the coarse registration image of the elastic ultrasound and the B-ultrasound image, respectively.

[0018] Based on the maximum similarity of each structure as the termination condition, the affine transformation matrix is ​​optimized according to gradient descent to obtain the optimal affine parameters corresponding to the coarse registration image of the molybdenum target and the coarse registration image of elastic ultrasound, respectively.

[0019] The coarse registration images of the molybdenum target and the coarse registration images of elastic ultrasound are adjusted according to the optimal affine parameters to obtain a multimodal image set.

[0020] In one embodiment, modality-sensitive feature extraction is performed on each modality of the multimodal image group to obtain molybdenum target image features, B-ultrasound image features, and elastography image features, including:

[0021] Based on the preset features of interest, modal features are extracted from the mammogram, B-mode ultrasound and elastography images respectively to obtain the hierarchical feature maps of each modality;

[0022] Sensitive and specific regions are emphasized in the hierarchical feature maps of each modality to obtain the enhanced features of each modality;

[0023] The corresponding feature cues are obtained based on the hierarchical feature maps of each modality; the feature cues include location cues, rigidity cues, and boundary cues.

[0024] Based on feature suggestions, cross-modal feature enhancement guidance is performed on the enhancement features of each modality to obtain mammogram image features, B-ultrasound image features, and elastography image features.

[0025] In one embodiment, the features of the mammogram image, the B-mode ultrasound image, and the elastography image are fused to obtain a global fused feature, including:

[0026] Modal complementary representations of features extracted from mammogram images, B-mode ultrasound images, and elastography images;

[0027] Modal complementary representations are residually connected with features from molybdenum target images, B-mode ultrasound images, and elastic ultrasound images. Spatial collaborative recognition capability is enhanced through feature cross paths to obtain global fusion features.

[0028] In one embodiment, breast tumor prediction is performed based on global fusion features to obtain prediction results, including:

[0029] The probability of the presence of breast tumor lesions is obtained based on global fusion features;

[0030] If the probability of the presence of a breast tumor lesion exceeds a preset threshold, the probability of malignancy is calculated based on the global fusion features.

[0031] If the probability of malignancy is greater than the preset malignancy threshold, then the presence of a breast tumor lesion is determined to be a malignant tumor.

[0032] If the probability of malignancy is less than the preset malignancy threshold, then the presence of a breast tumor lesion is determined to be a benign tumor.

[0033] In one embodiment, the method further includes:

[0034] A fully connected perceptron for breast tumor prediction is obtained by constraining the classifier for breast tumor prediction through a predefined joint loss function; the joint loss function includes the consistency embedding loss function and the cross-entropy loss function.

[0035] The joint loss function is obtained using the following formula:

[0036] L total =λL CE +(1-λ)L align

[0037] L CE= -y·log(P) - (1-y)·log(1-P)

[0038]

[0039] Among them, L total For the joint loss function; L CE L is the cross-entropy loss function; align λ is the consistency embedding loss function; y is the weight hyperparameter; P is the probability; i is the i-th training sample; N is the total number of training samples; m and n are the labels of different modalities; φ is the embedding projection function; F i Let be the feature map of the i-th training sample in each modality.

[0040] Secondly, this application also provides a multimodal image fusion-based breast tumor diagnostic device, comprising:

[0041] The multimodal data module is used to acquire multimodal images of the subject and perform spatial registration processing to obtain a multimodal image set;

[0042] The feature extraction module is used to extract modality-sensitive features from each modality of the multimodal image group to obtain features of the molybdenum target image, B-ultrasound image, and elastography image.

[0043] The feature fusion module is used to fuse features from molybdenum target images, B-ultrasound images, and elastography images to obtain global fused features.

[0044] The breast tumor prediction module is used to predict breast tumors based on global fusion features and obtain prediction results. The prediction results include the probability of the presence of breast tumor lesions and the type of breast tumor lesions present.

[0045] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above-described multimodal image fusion methods for diagnosing breast tumors.

[0046] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the above-described multimodal image fusion methods for diagnosing breast tumors.

[0047] The aforementioned multimodal image fusion-based methods, devices, equipment, and media for breast tumor diagnosis fully leverage the complementary advantages of different modalities through spatial registration, sensitive feature extraction, and semantic fusion, achieving a more comprehensive and consistent lesion expression. It distinguishes the physical attributes of each imaging modality and extracts modality-specific regional features, maintaining modal independence while avoiding information ambiguity and feature dilution, thus improving the perception of fine-grained lesion information. Global fusion features are used for final discrimination, ensuring accurate judgment even when a particular modality signal is weak or interfered with by artifacts, thanks to the information provided by other modalities. This approach is particularly suitable for identifying cases with dense breast tissue, complex structures, or atypical symptoms. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating the multimodal image fusion method for breast tumor diagnosis according to the present invention.

[0050] Figure 2 This is a flowchart illustrating the steps of step S101.

[0051] Figure 3 This is a flowchart illustrating the steps of step S102.

[0052] Figure 4 This is a structural diagram of the multimodal image fusion breast tumor diagnostic device of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0054] In one embodiment, such as Figure 1 As shown, a multimodal image fusion method for breast tumor diagnosis is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0055] S101. Acquire multimodal images of the subject and perform spatial registration processing to obtain a multimodal image set; the multimodal image set includes mammogram images, B-ultrasound images and elastography images.

[0056] This study illustrates the acquisition of multimodal breast imaging data from a patient and the spatial registration of these images to construct a structurally consistent, coordinate-corresponding multimodal image set. Multimodal breast imaging includes at least three common clinical examination imaging data: mammography (MAM), B-mode ultrasound (BUS), and elastography (ELA). Mammography images are typically acquired in both anteroposterior (CC) and oblique (MLO) views to present information on breast tissue density distribution and calcification areas; BUS images provide information on lesion boundaries and echogenicity, suitable for high-density breast tissue; and elastography reflects the biomechanical rigidity distribution of local tissues, helping to differentiate between benign and malignant tumors. Because the three modalities originate from different imaging devices and principles, differences exist between the images in scale, viewpoint, and location, necessitating spatial alignment processing. Specifically, the BUS image can be selected as the reference coordinate system. Corner points or local structural features can be extracted from the MAM and ELA images, and preliminary registration can be performed by combining the affine transformation method based on matching points. Then, fine registration can be achieved by further using optimization algorithms based on mutual information or structural similarity, so as to accurately align the three modal images and obtain a spatially corresponding multimodal image group.

[0057] S102. Modality-sensitive features are extracted from each mode of the multimodal image group to obtain the features of the molybdenum target image, the B-ultrasound image, and the elastic ultrasound image.

[0058] In a schematic manner, modal-sensitive features are extracted for each modality in the multimodal image set to retain the most diagnostically valuable features for breast lesions in each modality, constructing a cross-modal comparable semantic feature representation. Modal-sensitive features refer to image attributes that have a strong responsiveness to lesion identification in a particular imaging modality. For example, for mammography images, the imaging is based on the difference in X-ray absorption by breast tissue density, thus showing a high response to structures such as calcification and abnormal mass density. Its sensitive features are mainly manifested in texture features such as regional grayscale distribution, edge sharpness, and density gradient. For ultrasound images, the irregularity of lesion edges and the attenuation or enhancement of posterior echoes are often important criteria for determining benignity or malignancy. Therefore, the sensitive features of ultrasound images include echo boundary morphology, signal distribution non-uniformity, and edge spiculation. Elastic images utilize the difference in tissue response to probe pressure to generate a pseudo-color elastic distribution map. Its sensitive features are mainly reflected in the tissue hardness value of the lesion area, rigidity distribution pattern, and rigidity contrast with surrounding tissues. For the three types of images, lightweight neural network branches with different convolutional structures were constructed for feature extraction. A modality-specific attention mechanism was introduced into the shallow layer of the network to enhance the network's ability to perceive its own sensitive areas, thereby obtaining features of molybdenum target images, ultrasound images, and elasticity images, respectively.

[0059] S103. The features of the mammogram, B-mode ultrasound and elastography images are fused to obtain the global fused features.

[0060] Indicatively, deep fusion processing is performed on the extracted features from molybdenum target images, ultrasound images, and elasticity images to construct a global fusion feature capable of representing cross-modal consistency and complementarity. To achieve high-quality modal fusion, a cross-modal interaction mechanism is introduced, establishing information channels between the three modal features. A multi-head attention mechanism is used to learn the complementary relationships between different modalities. For example, the location cues provided by the molybdenum target image features, the rigidity cues provided by the elasticity image, and the boundary cues provided by the ultrasound can be aligned with each other in the semantic space. A residual connection structure is added during the fusion process to retain feature channels in each modality that are not fully considered but may still be valuable, avoiding information loss during feature fusion. To enhance the collaborative representation capability between modalities, a feature cross-path network can be designed, enabling features to not only be spatially aligned but also to form a consistent discriminative tendency at the semantic level. The fused multi-channel feature map is then subjected to global pooling and dimensionality reduction projection to obtain a global fusion feature vector of uniform dimension.

[0061] S104. Breast tumor prediction is performed based on global fusion features to obtain prediction results; the prediction results include the probability of the presence of breast tumor lesions and the type of breast tumor lesions present.

[0062] Based on the fused feature vectors, the system outputs the diagnostic probability and type classification results for the image group. Illustratively, it first determines whether there are any suspicious breast tumor lesions, and then, if lesions are present, further determines their benign or malignant nature. Specifically, a lightweight discriminant subnetwork performs binary classification on the fused feature vectors, outputting the probability value of tumor presence. When this probability exceeds a preset threshold, a suspected lesion area is considered to exist in the sample. If a lesion is determined to exist, the system proceeds to the second stage, where another classification subnetwork further determines the nature of the lesion, outputting the benign and malignant probabilities. The overall prediction result includes the probability value of lesion presence and the benign / malignant classification result when the lesion is present.

[0063] In the aforementioned multimodal image fusion-based breast tumor diagnosis method, joint registration of mammogram, ultrasound, and elastography images ensures precise spatial alignment of the three modalities, enabling consistent localization of lesion regions. This improves the accuracy of subsequent feature extraction and fusion, reducing misjudgments caused by spatial mismatches. A modality-sensitive feature extraction mechanism is employed, focusing each modality on its most discriminative regions and characteristics in breast tumor identification, such as abnormal density in mammograms, boundary morphology in ultrasound, and rigid distribution in elastography. This preserves key diagnostic clues and improves the overall quality of feature expression. Deep fusion of the three modal features constructs a global fusion feature vector, integrating complementary modal information and uniformly expressing the comprehensive risk characteristics of the lesion region at the semantic level, thereby enhancing identification stability and robustness. The prediction structure employs a two-stage discrimination strategy: first determining the presence of breast tumor lesions, then determining the benign or malignant type of the lesion. This effectively reduces false positive and false negative rates, improving the reliability of the final diagnosis.

[0064] In one embodiment, such as Figure 2 As shown, multimodal images of the subject were acquired and spatially registered to obtain a multimodal image set, including:

[0065] S201. Using the B-ultrasound image as a reference coordinate system, extract structural feature points from the mammogram image and the elastic ultrasound image respectively to obtain a set of feature points.

[0066] This illustration demonstrates the extraction of structural feature points from molybdenum target images and elastic ultrasound images to represent important locations with geometric or edge features. Structural feature points can be extracted using corner detection algorithms based on grayscale changes or local invariant feature extraction algorithms based on scale space. For example, Harris corner detection, SIFT (Scale Invariant Feature Transform), or FAST feature point extraction algorithms are used. Feature points are generally distributed in areas of abrupt edge changes, structural transitions, or significant texture, exhibiting strong stability and good repeatability, and can serve as matching geometric anchor points in cross-modal registration. Through extraction operations, feature point sets are obtained for both the molybdenum target image and the elastic ultrasound image.

[0067] S202. Match each feature point in the feature point set with the feature point in the ultrasound image to obtain an initial set of point pairs.

[0068] Furthermore, the extracted set of molybdenum target feature points is matched with the ultrasound image, and similarly, the feature points of the elastic image are matched with the ultrasound image to obtain two initial set of point pairs. Point pair matching can be based on similarity measures between feature descriptors. For example, Euclidean distance, Hamming distance, or correlation functions can be used, and high-confidence point pairs can be selected through nearest neighbor or bidirectional verification to form the basis for the positional mapping between different modal images and ultrasound images.

[0069] S203. Remove mismatched points from the initial point pair set and fit the affine transformation matrix of the molybdenum target image and the elastic ultrasound image to the reference coordinate system to obtain the coarse registration image; the coarse registration image includes the molybdenum target coarse registration image and the elastic ultrasound coarse registration image.

[0070] Furthermore, to remove unreliable point pairs caused by image noise, texture artifacts, or feature mismatches, a robust RANSAC matching elimination algorithm is used to clean the initial point pair set, retaining point pairs with good spatial consistency. Based on the cleaned point pair set, affine transformation matrices are fitted to map the molybdenum target image and the elastic image to the ultrasound image reference coordinate system, respectively. Affine transformations can include parameters such as translation, rotation, scaling, and shearing, which can maintain structural continuity and perform preliminary alignment of large-scale geometric differences. The original molybdenum target image and elastic image are transformed using this affine transformation to obtain the coarsely registered molybdenum target image and the coarsely registered elastic ultrasound image, thus completing the coarse registration stage.

[0071] S204. Affine adjustment is performed on the coarsely registered image based on the structural similarity of the registered images to obtain a multimodal image group.

[0072] To address the issue of minor misalignments in the detailed structure of coarsely registered images, affine fine-tuning is further achieved through structural similarity optimization to obtain more accurate pixel-level alignment results, ultimately forming a multimodal image group after spatial registration.

[0073] In one embodiment, the coarsely registered image is affine adjusted based on the structural similarity of the registered images to obtain a multimodal image set, including:

[0074] S11. Calculate the mutual information between the coarse registration image and the B-ultrasound image to obtain the structural similarity between the coarse registration image of the molybdenum target and the coarse registration image of elastic ultrasound and the B-ultrasound image, respectively.

[0075] To illustrate, mutual information is used as a structural similarity metric to calculate the mutual information value between the coarsely registered mammogram image and the ultrasound image. And the mutual information value between the elastic ultrasound coarse registration image and the B-mode ultrasound image. Mutual Information As an indicator of the correlation of joint image distributions, mutual information effectively reflects the structural consistency of image content across different modalities; a higher mutual information value indicates better registration. The calculation results serve as a reference indicator for optimizing the objective function.

[0076] S12. Based on the maximum similarity of each structure as the termination condition, the affine transformation matrix is ​​optimized according to gradient descent to obtain the optimal affine parameters corresponding to the coarse registration image of the molybdenum target and the coarse registration image of elastic ultrasound, respectively.

[0077] Mutual information and Maximizing the objective function, an optimization process is constructed, continuously fine-tuning the affine transformation parameters through gradient descent or heuristic search methods. During the optimization process, the transformation parameters include angle, displacement, scaling ratio, etc. After each adjustment, the mutual information index is recalculated. If the mutual information increases, the current transformation is retained; otherwise, it is rolled back, until the index converges or reaches a preset threshold, thus obtaining the optimal affine parameters of the molybdenum target image and the elastic image in the ultrasound image coordinate system.

[0078] S13. Adjust the corresponding coarse registration images of the molybdenum target and the coarse registration images of elastic ultrasound according to each optimal affine parameter to obtain a multimodal image group.

[0079] Furthermore, based on the optimized affine parameters, the coarse registration images of the mammogram target and the elastic coarse registration images are affinely adjusted to obtain a set of registered images that are highly consistent with the ultrasound images in structure and space. The registered multimodal image set has lesion-level spatial alignment capability, ensuring a one-to-one semantic correspondence between different modalities during subsequent feature extraction.

[0080] In one embodiment, such as Figure 3As shown, modality-sensitive feature extraction was performed on each modality of the multimodal image group to obtain features of mammography images, B-mode ultrasound images, and elastography images, including:

[0081] S301. Based on the preset features of interest, modal features are extracted from the mammogram, B-ultrasound, and elastography images respectively to obtain the hierarchical feature maps of each modality.

[0082] In a schematic manner, independent modality coding networks are constructed for three different image modalities, and multi-level feature extraction is performed based on preset key features of interest. This coding network can be a lightweight convolutional neural network structure, progressively extracting texture, structure, and semantic information of the image in a shallow-medium-deep sequence. In the structural design, the coding network for molybdenum images focuses on density and edges as core regions of interest, extracting tissue density gradients and calcification distribution through multi-scale convolutional kernels; the coding network for ultrasound images emphasizes boundary extraction and echo difference perception, with the shallow layer focusing on edge intensity and speckle noise suppression, and the middle layer emphasizing the identification of lesion contour changes and echo morphology; the elastic image emphasizes the identification and differentiation of rigid regions, with the convolutional layers of the coding network used to capture strong response regions and stiffness abrupt boundary changes in the hardness spectrum. Through coding network processing, multi-level feature maps corresponding to the three modalities are obtained, covering a progressively abstract expression from low-level visual information to high-level semantic information.

[0083] S302. Emphasize sensitive and specific regions in the hierarchical feature maps of each modality to obtain the enhanced features of each modality.

[0084] To illustrate, to enhance the model's ability to express its own modal core features, a modality-specific region emphasis mechanism is introduced to enhance sensitive regions of the hierarchical feature map. This mechanism employs a hybrid attention module combining channel attention and spatial attention, adaptively enhancing the response intensity of key regions during forward propagation within each modal branch. For example, the mammography image branch uses channel attention to strengthen the response of calcified and densely structured regions; the ultrasound image branch uses spatial attention to highlight regions with blurred lesion edges and irregular shapes; and the elasticity image branch's attention module focuses on high-rigidity regions in the image and the boundary saliency resulting from the stiffness difference between the image and soft tissue. Through this processing, enhanced feature maps for each modality are obtained, exhibiting stronger activation distributions in modality-sensitive regions compared to the original encoded features, which helps the subsequent fusion model focus on key areas.

[0085] S303. Obtain corresponding feature hints based on the hierarchical feature maps of each modality; feature hints include location hints, rigidity hints, and boundary hints.

[0086] Furthermore, representative modal cue information is calculated and extracted based on the hierarchical feature map as mediating variables for the cross-modal guidance mechanism. Modal cues reflect discriminative clues in a certain modal image that can be used to assist in the discrimination of other modalities. For example, location cues generated from mammograms are used to locate areas where lesions may exist in the overall tissue structure, such as the location of high-density or aggregated structures; rigidity cues generated from elastic images characterize the distribution of tissue stiffness and the degree of local hardening, used to indicate possible malignant areas; boundary cues generated from ultrasound images reflect whether the lesion outline is regular and whether the edges are clear, indirectly expressing the invasiveness of the lesion. The three types of cues are extracted by weighted summation, centrality analysis, or boundary gradient analysis of spatial attention maps and channel activation maps in the mid-layer feature map, thus constructing the three types of modal cue features.

[0087] S304. Based on the feature prompts, perform cross-modal feature enhancement guidance on the enhancement features of each modality to obtain the features of the mammogram, B-mode ultrasound, and elastic ultrasound.

[0088] Indicatively, three types of modal cues are used as information guidance sources to perform cross-modal feature enhancement operations on three types of enhanced feature maps, thereby further improving the perception and discrimination accuracy of each modality for lesion regions. Specifically, a modal mutual guidance mechanism is constructed, employing an attention-based guided enhancement module (Cross-Guided Attention) to embed the cue features of the source modality into the feature map of the target modality as a control factor for feature reweighting. For example, after obtaining the rigidity cue provided by the elastic image, the ultrasound branch can enhance its boundary features in the corresponding region, thereby better identifying the irregular contours corresponding to hard masses; after receiving the boundary cue provided by the ultrasound, the mammography image can correct the contours of the density region, thereby reducing the interference of dense tissue on lesion judgment. Each modality not only retains the high response of its own sensitive region but also receives semantic supplementation from other modalities, achieving cross-modal enhancement at the feature level, and outputting fused guided mammography image features, ultrasound image features, and elastic image features respectively.

[0089] In one embodiment, the features of the mammogram image, the B-mode ultrasound image, and the elastography image are fused to obtain a global fused feature, including:

[0090] S21. Modal complementary representation of features extracted from molybdenum target image features, B-ultrasound image features, and elastography image features.

[0091] This paper illustrates how modal complementary representations are extracted from mammography, ultrasound, and elasticity image features. Specifically, a cross-modal interaction module is constructed, which performs full-channel cross-querying of feature maps from different modalities using a multi-head attention mechanism to extract key information regions that can complement each other across modalities. For example, the mid-level semantic features of each modality are used as Query, Key, and Value inputs, respectively, and their correlation in spatial location and channel response is calculated to form a cross-modal attention weight matrix. This attention weight reflects the semantic support of other modalities in the region of interest of the current modality, thereby converging the originally discrete diagnostic criteria from multiple modalities to the same spatial location. The weighted combination result is the modal complementary representation, representing the information intersection and complementary parts of mammography, ultrasound, and elasticity modalities in lesion identification, providing semantic enhancement support for different modalities.

[0092] S22. Modal complementary representations are residually connected with features of molybdenum target images, B-ultrasound images, and elastic ultrasound images, and spatial collaborative recognition capabilities are enhanced through feature cross paths to obtain global fusion features.

[0093] Schematic, the modal complementary representations are residually connected to the original three-modal feature maps, preserving the original structural information of each modality while introducing semantic increments from other modalities. Residual connections not only mitigate information loss caused by channel dimension reconstruction during fusion but also enhance the network's tolerance to inconsistencies caused by modal heterogeneity. Based on the residual fusion result, a feature cross-path mechanism is further introduced to perform spatial convolutional reconstruction on the fused feature map, enhancing the collaborative recognition capability of multimodal spatial layout. This mechanism, by constructing shared convolutional channels and local reconstruction pathways, allows features with different response regions between modalities to correspond spatially and jointly enhance each other, thereby improving the perception of complex lesion spatial structures. Finally, a global fusion feature vector of uniform length is obtained through global average pooling and dimensionality reduction projection.

[0094] In one embodiment, breast tumor prediction is performed based on global fusion features to obtain prediction results, including:

[0095] S31. Obtain the probability of the presence of breast tumor lesions based on global fusion features.

[0096] In a schematic representation, the global fusion features are input into the first-level discriminant network, namely the lesion detection network, to calculate the probability of the presence of breast tumor lesions in the image. This network can employ a lightweight multilayer perceptron structure, and its output is a continuous value in the range [0,1], representing the probability of the presence of a lesion, denoted as the probability of breast tumor lesion presence. If this probability value is lower than a preset discrimination threshold, a diagnosis of no lesion will be directly returned, thereby effectively filtering out normal images or abnormal images with large errors, avoiding entering the misclassification process of benign or malignant.

[0097] S32. If the probability of the presence of a breast tumor lesion exceeds a preset threshold, the probability value of malignancy is calculated based on the global fusion features.

[0098] If the probability of a lesion's presence exceeds a threshold, the system will automatically enter the second-level judgment process, which classifies the lesion as benign or malignant. Specifically, the globally fused features are still used as input to the benign / malignant classification network. This network structure can share some parameters with the previous-level network or be designed independently. Its output is a malignancy probability value, representing the probability that a suspected lesion in the current image is a malignant tumor. Similarly, this probability value ranges from [0,1], with values ​​closer to 1 indicating a higher confidence level in the system's malignancy assessment.

[0099] S33. If the malignancy probability value is greater than the preset malignancy threshold, then the presence of a breast tumor lesion is determined to be a malignant tumor.

[0100] S34. If the malignancy probability value is less than the preset malignancy threshold, then the presence of a breast tumor lesion is determined to be a benign tumor.

[0101] When the output malignancy probability value is higher than the preset malignancy judgment threshold, the lesion will be identified as a malignant tumor, and a diagnosis of breast malignancy will be returned; when the malignancy probability value is lower than the threshold, the lesion is considered more likely to be benign, and a result of benign breast tumor will be returned. This method combines multimodal image information to output clearly defined diagnostic suggestions under high confidence conditions, assisting doctors in further clinical judgment.

[0102] In one embodiment, the method further includes:

[0103] A fully connected perceptron for breast tumor prediction is obtained by constraining the classifier for breast tumor prediction through a predefined joint loss function; the joint loss function includes the consistency embedding loss function and the cross-entropy loss function.

[0104] The joint loss function is obtained using the following formula:

[0105] L total =λL CE +(1-λ)L align

[0106] L CE = -y·log(P) - (1-y)·log(1-P)

[0107]

[0108] Among them, L total For the joint loss function; L CE L is the cross-entropy loss function; align λ is the consistency embedding loss function; y is the weight hyperparameter; P is the probability; i is the i-th training sample; N is the total number of training samples; m and n are the labels of different modalities; φ is the embedding projection function; F i Let be the feature map of the i-th training sample in each modality.

[0109] Indicatively, the core model of breast tumor diagnosis relies on a predictive classifier based on a fully connected perceptron. This classifier receives the fused multimodal feature vector and outputs the presence or absence of a tumor and its benign or malignant type. To achieve high accuracy and stability in classification, the classifier needs to be trained using a well-structured training process. The training process is guided by the construction of a joint loss function, and the network parameters are continuously updated through an end-to-end backpropagation optimization strategy, allowing the model to gradually converge to its optimal state and achieve practical application value. Specifically, the data used in the training phase comes from a multimodal breast examination dataset of patients with blurred clinical information. Each sample in the dataset includes three image modalities: mammography, ultrasound, and elastography, along with manually annotated lesion boundaries and benign / malignant labels. The original images are first aligned according to a spatial registration process, and three sets of semantic features are extracted by convolutional neural networks in their respective branches. Feature enhancement is then performed through a modal intermodal communication mechanism, and a unified global fused feature representation is constructed in the fusion module. This feature vector is the input data to the classifier during the training phase. During the forward propagation, the fused features are first input into the lesion presence discriminant subnetwork of the first stage, and the output lesion presence probability P is generated. tumor If the true label is "tumor present," then the system proceeds to the second stage, the benign / malignant discrimination network, which outputs the malignancy probability P. malign The entire model structure implements a dual-branch output, corresponding to the training objectives of two sub-tasks respectively.

[0110] Furthermore, to achieve collaborative training of the two sub-tasks mentioned above, a joint loss function consisting of the cross-entropy loss function and the consistency embedding loss function was designed as the optimization objective during network training. The cross-entropy loss function measures the deviation between the model's predictions and the true labels, including the lesion presence task and the benign / malignant classification task. The loss for the lesion presence task is L.tumor =-y tumor ·log(P tumor )-(1-y tumor )·log(1-P tumor ), where y tumor The marker indicates the presence of a real lesion, with a value of 0 or 1. The loss for the benign / malignant classification task is L. malign =-y malign ·log(P malign )-(1-y malign )·log(1-P malign ), where y malign The label represents the true benign or malignant nature, with a value of 0 (benign) or 1 (malignant).

[0111] To improve the consistency of feature representation across different modalities in the semantic space and prevent certain modalities from being masked by the dominant modality during training, a consistency embedding loss function is further introduced as auxiliary supervision. This loss constructs a low-dimensional embedding space to constrain the consistency of representations of molybdenum image features, ultrasound image features, and elasticity image features in the lesion region. Specifically, a feature projection function φ is constructed. m (·), projecting the features of modality m onto the shared embedding space, minimizing the Euclidean distance between the embedding representations of the same lesion region, with the loss form being: Among them, F i This represents the modal feature vector at the i-th location of the lesion region.

[0112] During training, the standard backpropagation algorithm is used for end-to-end optimization of the entire network. In each iteration, the model calculates the prediction results through forward propagation, then calculates the loss function value, and calculates the gradient of the loss with respect to all learnable parameters. Finally, the network weights are updated using the optimization algorithm. The training process employs a Mini-Batch strategy, utilizing batch image data to improve optimization stability and using an Early-Stopping mechanism to prevent overfitting. After completing a certain number of iterations or when the loss function converges, a converged fully connected perceptron classifier is obtained, thus completing the training process.

[0113] The final trained model can be used to judge new samples in the testing phase. By using the fused features obtained after registration, feature extraction and fusion processing of the input image, the model can directly determine the existence and benign or malignant classification of breast tumor lesions, thereby realizing intelligent auxiliary diagnosis of breast tumors.

[0114] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0115] Based on the same inventive concept, this application also provides a breast tumor diagnostic device for implementing the aforementioned multimodal image fusion breast tumor diagnostic method. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the multimodal image fusion breast tumor diagnostic device provided below can be found in the above-described limitations of the multimodal image fusion breast tumor diagnostic method, and will not be repeated here.

[0116] In one exemplary embodiment, such as Figure 4 As shown, a multimodal image fusion-based breast tumor diagnostic device is provided, comprising:

[0117] The multimodal data module 401 is used to acquire multimodal images of the subject and perform spatial registration processing to obtain a multimodal image group.

[0118] Feature extraction module 402 is used to extract modality-sensitive features from each modality of the multimodal image group to obtain molybdenum target image features, B-ultrasound image features and elastic ultrasound image features;

[0119] The feature fusion module 403 is used to fuse the features of the mammogram image, the B-ultrasound image, and the elastography image to obtain the global fused features;

[0120] The breast tumor prediction module 404 is used to predict breast tumors based on global fusion features and obtain prediction results. The prediction results include the probability of the presence of breast tumor lesions and the type of breast tumor lesions present.

[0121] In one embodiment, it further includes:

[0122] The feature point extraction module is used to extract structural feature points from the mammogram image and the elastography image, respectively, using the B-ultrasound image as a reference coordinate system, to obtain a set of feature points;

[0123] The feature point matching module is used to match each feature point in the feature point set with the ultrasound image to obtain an initial set of point pairs.

[0124] The coarse calibration module is used to remove mismatched points from the initial point pair set and fit the affine transformation matrix of the molybdenum target image and the elastic ultrasound image to the reference coordinate system to obtain the coarse registration image; the coarse registration image includes the molybdenum target coarse registration image and the elastic ultrasound coarse registration image;

[0125] The fine calibration module performs affine adjustment on the coarsely registered image based on the structural similarity of the registered images to obtain a multimodal image group.

[0126] In one embodiment, it further includes:

[0127] The similarity module is used to calculate the mutual information between the coarse registration image and the B-ultrasound image, and to obtain the structural similarity between the coarse registration image of the mammogram and the coarse registration image of the elastic ultrasound and the B-ultrasound image, respectively.

[0128] The affine adjustment module is used to optimize the affine transformation matrix based on the maximum similarity of each structure as the termination condition, and obtain the optimal affine parameters corresponding to the coarse registration image of the molybdenum target and the coarse registration image of elastic ultrasound, respectively.

[0129] The fine calibration module is also used to adjust the corresponding coarse registration images of the molybdenum target and the coarse registration images of elastic ultrasound according to each optimal affine parameter, so as to obtain a multimodal image group.

[0130] In one embodiment, it further includes:

[0131] The feature extraction module 402 is also used to extract modal features from the molybdenum target image, B-ultrasound image and elastic ultrasound image respectively according to the preset features of interest, so as to obtain the hierarchical feature map of each modality;

[0132] The feature enhancement module is used to emphasize sensitive and specific regions of the hierarchical feature maps of each modality to obtain enhanced features for each modality;

[0133] The feature suggestion module is used to obtain corresponding feature suggestions based on the hierarchical feature maps of each modality; the feature suggestions include position suggestions, rigidity suggestions, and boundary suggestions.

[0134] The feature cross-modal feature enhancement module is used to guide cross-modal feature enhancement based on feature prompts to obtain features of molybdenum target image, B-ultrasound image and elastic ultrasound image.

[0135] In one embodiment, it further includes:

[0136] A complementary guidance module is used to extract modal complementary representations of features from molybdenum target images, B-mode ultrasound images, and elastography images;

[0137] The feature fusion module 403 is also used to perform residual connection between the modal complementary representation and the features of the molybdenum target image, the B-ultrasound image and the elastic ultrasound image, and to enhance the spatial collaborative recognition capability through the feature cross path to obtain the global fused features.

[0138] In one embodiment, it further includes:

[0139] The breast tumor prediction module 404 is also used to obtain the probability of the presence of breast tumor lesions based on global fusion features;

[0140] The breast tumor prediction module 404 is also used to calculate the malignancy probability value based on global fusion features if the probability of the presence of a breast tumor lesion exceeds a preset threshold.

[0141] The breast tumor prediction module 404 is also used to determine that the breast tumor lesion is malignant if the malignancy probability value is greater than the preset malignancy threshold.

[0142] The breast tumor prediction module 404 is also used to determine that if the malignancy probability value is less than a preset malignancy threshold, the type of breast tumor lesion is benign.

[0143] In one embodiment, it further includes:

[0144] The model training module is used to constrain and train a classifier for breast tumor prediction using a preset joint loss function, resulting in a fully connected perceptron for breast tumor prediction. The joint loss function includes a consistency embedding loss function and a cross-entropy loss function.

[0145] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above method embodiments.

[0146] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0147] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0148] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. A method for diagnosing breast tumors using multimodal image fusion, characterized in that, The method includes: Multimodal images of the subject are acquired and spatially registered to obtain a multimodal image set; the multimodal image set includes mammogram images, B-ultrasound images and elastography images; Modality-sensitive features are extracted from each mode of the multimodal image group to obtain molybdenum target image features, B-ultrasound image features, and elastography image features; The features of the mammogram image, the B-ultrasound image, and the elastic ultrasound image are fused to obtain a global fused feature. Breast tumor prediction is performed based on the global fusion features to obtain prediction results; the prediction results include the probability of the presence of breast tumor lesions and the type of breast tumor lesions present.

2. The method according to claim 1, characterized in that, The process of acquiring multimodal images of the subject and performing spatial registration processing to obtain a multimodal image set includes: Using the B-ultrasound image as a reference coordinate system, structural feature points are extracted from the mammogram and the elastic ultrasound image respectively to obtain a set of feature points; Each feature point in the feature point set is matched with the feature point in the ultrasound image to obtain an initial set of point pairs. Mismatched points are removed from the initial point pair set, and an affine transformation matrix mapping the molybdenum target image and the elastic ultrasound image to the reference coordinate system is fitted to obtain a coarse registration image; the coarse registration image includes the molybdenum target coarse registration image and the elastic ultrasound coarse registration image; The coarsely registered image is affinely adjusted based on the structural similarity of the registered images to obtain the multimodal image group.

3. The method according to claim 2, characterized in that, The step of performing affine adjustment on the coarsely registered image based on the structural similarity of the registered images to obtain the multimodal image group includes: The mutual information between the coarse registration image and the ultrasound image is calculated to obtain the structural similarity between the mammogram coarse registration image and the elastic ultrasound coarse registration image and the ultrasound image, respectively. Based on the maximum similarity of each structure as the termination condition, the affine transformation matrix is ​​optimized according to gradient descent to obtain the optimal affine parameters corresponding to the coarse registration image of the molybdenum target and the coarse registration image of the elastic ultrasound, respectively. The multimodal image group is obtained by adjusting the corresponding coarse registration image of the molybdenum target and the coarse registration image of the elastic ultrasound according to the optimal affine parameters.

4. The method according to claim 1, characterized in that, The modality-sensitive feature extraction of each modality of the multimodal image group yields molybdenum target image features, B-ultrasound image features, and elastography image features, including: Based on preset features of interest, modal features are extracted from the molybdenum target image, the B-ultrasound image, and the elastic ultrasound image to obtain hierarchical feature maps for each modality; Sensitive and specific regions are emphasized in the hierarchical feature maps of each modality to obtain the enhanced features of each modality; Based on the hierarchical feature maps of each modality, corresponding feature cues are obtained; the feature cues include location cues, rigidity cues, and boundary cues; Based on the feature prompts, cross-modal feature enhancement guidance is performed on the enhancement features of each modality to obtain the molybdenum target image features, the B-ultrasound image features, and the elastic ultrasound image features.

5. The method according to claim 1, characterized in that, The process of fusing the features of the mammogram, the ultrasound image, and the elastography image to obtain global fused features includes: Modal complementary representations of the features of the mammogram, the B-mode ultrasound, and the elastic ultrasound are extracted; The modal complementary representation is residually connected with the molybdenum target image features, the B-ultrasound image features, and the elastic ultrasound image features, and the spatial collaborative recognition capability is enhanced through feature cross paths to obtain the global fusion features.

6. The method according to claim 1, characterized in that, The prediction of breast tumors based on the global fusion features, and the resulting prediction, include: The probability of the presence of the breast tumor lesion is obtained based on the global fusion features. If the probability of the presence of the breast tumor lesion exceeds a preset threshold, the malignancy probability value is calculated based on the global fusion feature. If the malignancy probability value is greater than the preset malignancy threshold, then the type of breast tumor lesion is determined to be a malignant tumor. If the malignancy probability value is less than the preset malignancy threshold, then the type of breast tumor lesion is determined to be a benign tumor.

7. The method according to claim 6, characterized in that, The method further includes: A classifier for breast tumor prediction is trained under constraints using a preset joint loss function to obtain a fully connected perceptron for breast tumor prediction; the joint loss function includes a consistency embedding loss function and a cross-entropy loss function. The joint loss function is obtained using the following formula: THE total =λL CE +(1-λ)L align L CE =-y·log(P)-(1-y)·log(1-P) Among them, L total For the joint loss function; L CE L is the cross-entropy loss function; align λ is the consistency embedding loss function; y is the weight hyperparameter; P is the probability; i is the i-th training sample; N is the total number of training samples; m and n are the labels of different modalities; φ is the embedding projection function; F i Let be the feature map of the i-th training sample in each modality.

8. A multimodal image fusion-based breast tumor diagnostic device, characterized in that, The device includes: The multimodal data module is used to acquire multimodal images of the subject and perform spatial registration processing to obtain a multimodal image set; The feature extraction module is used to extract modality-sensitive features from each modality of the multimodal image group to obtain molybdenum target image features, B-ultrasound image features, and elastic ultrasound image features. The feature fusion module is used to fuse the features of the mammogram image, the B-ultrasound image, and the elastic ultrasound image to obtain global fused features. The breast tumor prediction module is used to predict breast tumors based on the global fusion features and obtain prediction results; the prediction results include the probability of the presence of breast tumor lesions and the type of breast tumor lesions present.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Focus identification method based on multi-modal ultrasonic time series data

    CN121304658A

  • Intelligent focus detection and diagnosis system based on multi-modal medical image fusion

    CN121482012A