A method and system for processing medical image data

By introducing attention mechanism and feature alignment algorithm in multimodal medical image data processing, deep fusion of multimodal image features is achieved, and edge monitoring algorithm is used to eliminate pseudo-lesion areas, solving the problems of insufficient feature fusion and inaccurate pseudo-lesion removal, significantly improving the accuracy of image segmentation and lesion positioning.

CN119624978BActive Publication Date: 2025-05-02THE THIRD MEDICAL CENT OF THE CHINESE PEOPLES LIBERATION ARMY GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510168427.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-02
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient feature fusion and inaccurate removal of pseudo-lesion areas in the processing of multimodal medical image data, which affects the accuracy and reliability of image segmentation and lesion positioning.

Method used

By introducing attention mechanisms and feature alignment algorithms, specific features of multimodal images are extracted and deep fusion between modes is achieved. At the same time, the edge monitoring algorithm combined with the preliminary segmentation results were used to eliminate the pseudo-lesion area through multi-scale edge features comparison.

Benefits of technology

It significantly improves the processing accuracy of multimodal image data, ensures the accuracy and robustness of lesion positioning, and solves the problems of insufficient feature fusion and inaccurate pseudo-lesion elimination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119624978B_ABST
    Figure CN119624978B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for processing medical image data, and relates to medical image data processing, including: based on the specific features of each modality, introducing an attention mechanism to calculate the attention weight of each modality, and generating a specific feature map according to the attention weight, and spatially aligning the specific feature maps of different modalities through a feature alignment algorithm to generate an aligned specific feature map, and using a multi-head attention mechanism to identify the correlation between the modalities to generate a comprehensive feature map; based on the comprehensive feature map, pixel classification is performed through UNet to generate a preliminary segmentation result of the lesion area, and an edge monitoring algorithm is used to obtain a multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result, and a pseudo lesion area is eliminated according to the comparison result, and a lesion localization map is finally generated. The present invention eliminates pseudo lesion areas through multi-scale edge feature comparison, thereby ensuring the accuracy and robustness of lesion localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to medical image data processing, and in particular to a method and system for processing medical image data. Background Art

[0002] With the rapid development of medical imaging technology, multimodal medical imaging (such as CT, MRI, and PET) has become an indispensable and important tool in modern medical diagnosis and treatment. These imaging technologies each have unique imaging mechanisms and clinical application values. For example, CT can provide high-resolution anatomical information, MRI can show the contrast characteristics of soft tissues, and PET can reflect metabolic activity and functional information. However, due to the limitations of single-modality imaging, clinical diagnosis often requires the combination of multiple modal imaging to comprehensively evaluate the anatomical, functional, and metabolic characteristics of lesions. In recent years, multimodal medical image processing technology based on deep learning has rapidly emerged. Through feature extraction, fusion, and segmentation, it can efficiently extract disease-related information from multimodal data. However, the fusion processing of multimodal imaging data faces many challenges, including differences in features between modalities, difficulty in spatial alignment, and interference from pseudo-lesion areas. These problems directly affect the accuracy and reliability of image segmentation and lesion localization.

[0003] In the prior art, the processing methods for multimodal image data are mainly focused on feature extraction and modality fusion, but there are still the following shortcomings: On the one hand, most traditional multimodal image fusion methods use simple weighted average or feature splicing methods, which makes it difficult to fully explore the specific features of each modality image and the correlation between modalities, resulting in the fusion result not being accurate enough in describing the details of the lesion area; on the other hand, due to the presence of artifacts, noise and background interference, the existing segmentation technology is insufficient in the accuracy of the lesion edge, and it is easy to misjudge the pseudo-lesion area as the real lesion area, thereby affecting the reliability of the positioning map. Especially in complex multimodal data scenarios, the removal of pseudo-lesion areas has always been a technical difficulty, and the effects of existing methods in edge feature extraction and pseudo-lesion identification still need to be improved. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method for processing medical image data to solve the problems of insufficient fusion of multi-modal specific features and inaccurate removal of pseudo lesion areas.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In the first aspect, the present invention provides a method for processing medical image data, which includes obtaining multimodal medical image data of a patient including CT, MRI and PET, and preprocessing the multimodal image data, inputting the preprocessed multimodal image data into a deep convolutional neural network, extracting specific features of each modality, and generating a preliminary feature map; based on the specific features of each modality, introducing an attention mechanism to calculate the attention weight of each modality, and generating a specific feature map according to the attention weight, and using a feature alignment algorithm to spatially align the specific feature maps of different modalities to generate an aligned specific feature map, and using a multi-head attention mechanism to identify the correlation between the modalities to generate a comprehensive feature map; based on the comprehensive feature map, using UNet to perform pixel classification to generate a preliminary segmentation result of the lesion area, and using an edge monitoring algorithm to obtain a multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result, eliminating pseudo lesion areas according to the comparison results, and finally generating a lesion localization map.

[0008] As a preferred solution of the method for processing medical image data of the present invention, the preprocessing of multimodal image data comprises the following specific steps:

[0009] Use SIFT to spatially align multimodal image data;

[0010] Using intensity normalization and size normalization processing, the pixel values ​​of multimodal imaging data are normalized to a uniform range and the images are adjusted to the same resolution;

[0011] Gaussian filtering method is used to remove background noise and artifacts from multimodal imaging data.

[0012] As a preferred solution of the method for processing medical image data of the present invention, the pre-processed multimodal image data is input into a deep convolutional neural network to extract the specific features of each modality and generate a preliminary feature map. The specific steps are as follows:

[0013] Each modality data in the preprocessed multimodal image data is used as a multi-channel input of a deep convolutional neural network;

[0014] Based on each modality of input data, the first convolutional layer of the deep convolutional neural network combines the ReLU activation function and batch normalization to extract low-level specific features;

[0015] Based on the low-level specific features, the second convolutional layer performs convolution and pooling operations again to extract the specific features of each modality and generate a preliminary feature map for each modality;

[0016] The specific features include density distribution features of CT, soft tissue contrast features of MRI and metabolic activity features of PET.

[0017] As a preferred solution of the method for processing medical image data of the present invention, the specific steps are as follows: based on the specific features of each modality, an attention mechanism is introduced to calculate the attention weight of each modality, and a specific feature map is generated according to the attention weight.

[0018] For the preliminary features of each modality, global average pooling and maximum pooling operations are performed on the channel dimension to generate a spatial attention feature map;

[0019] The spatial attention feature map is convolved and activated with Sigmoid to generate a spatial weight map for each pixel.

[0020] Multiply the spatial weight map with the preliminary feature map pixel by pixel to identify the lesion area;

[0021] For each modality’s specific feature map, global average pooling and maximum pooling operations are performed in the spatial dimension to generate channel-aggregated features;

[0022] The channel aggregation features are processed through convolution operation and Sigmoid activation function to generate the weights of each channel;

[0023] Multiply the channel-aggregated features with the preliminary feature map channel by channel to enhance the feature channels of the lesion area;

[0024] Based on the lesion area and the enhanced lesion area feature channel, nonlinear weighting and normalization are used to calculate the attention weight of each modality, which is expressed as:

[0025] ;

[0026] in, It is The attention weights of each modality, It is The spatial feature map of the modal The weight coefficient of It is The spatial features in the spatial feature map of each modality, It is The weight coefficient of the interaction term between the spatial features of each modality and the channel aggregation features, It is to adjust The weight coefficient of the channel aggregation feature of each mode, It is The channel aggregation characteristics of the modes, is the Sigmoid activation function, is the element-wise multiplication operator;

[0027] The attention weight of each modality It is multiplied layer by layer with the preliminary feature map to generate a specific feature map for each modality.

[0028] As a preferred solution of the method for processing medical image data of the present invention, the specific feature maps of different modalities are spatially aligned by a feature alignment algorithm to generate an aligned specific feature map, and the correlation between modalities is learned by a multi-head attention mechanism to generate a comprehensive feature map. The specific steps are as follows:

[0029] Bilinear interpolation is used to unify the specific feature maps of all modalities to the same spatial resolution, and the specific feature maps of each modality are preliminarily spatially aligned through affine transformation;

[0030] Perform secondary spatial alignment on the specific feature maps after preliminary alignment through STN, and finally generate an aligned specific feature map;

[0031] The aligned specific feature maps are subjected to three sets of linear transformations to generate query vectors, key vectors, and value vectors;

[0032] Based on the query vector, key vector and value vector, the correlation between the modalities is identified, and the expression is:

[0033] ;

[0034] in, It is The query vector of each modality, It is The key vector of the modes, It is The transpose of a mode, is the key vector The dimension of Indicates The mode and The similarity between the modes, is the index variable of the modality, is different from The index variable of another mode;

[0035] Based on the correlation between modalities , construct the complementarity and difference mapping between modalities, and fuse them with the specific feature maps of each modality to finally generate a comprehensive feature map.

[0036] As a preferred solution of the method for processing medical image data of the present invention, wherein: based on the comprehensive feature map, pixel classification is performed on the comprehensive feature map by UNet to generate a preliminary segmentation result of the lesion area, and the specific steps are as follows:

[0037] Based on the comprehensive feature map, the comprehensive feature map is encoded through UNet, the deep semantic information of the comprehensive feature map is gradually extracted, and the spatial resolution in the comprehensive feature map is compressed through sampling operation;

[0038] After encoding and sampling, the comprehensive feature map is upsampled layer by layer to gradually restore the spatial resolution, and the comprehensive feature map is mapped into channels of the target area and background through convolution operations;

[0039] Based on the channels of the target area and the background, the Softmax activation function is used to calculate the probability that each pixel belongs to the lesion area, and the preliminary segmentation result of the lesion area is generated;

[0040] The preliminary segmentation result includes the outline and position of the lesion area in the comprehensive feature map.

[0041] As a preferred solution of the method for processing medical image data of the present invention, the edge monitoring algorithm is used to obtain a multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result, and the pseudo lesion area is eliminated according to the comparison result, and finally a lesion localization map is generated. The specific steps are as follows:

[0042] Use Canny to extract the contour of the lesion area and the boundary of the internal structure from the preliminary segmentation results, and use Gaussian pyramid to generate edge feature maps of different scales;

[0043] Based on edge feature maps of different scales, the local gradient direction features and intensity features of the edge of the lesion area are extracted on the edge feature maps of each scale through multi-resolution filters;

[0044] The local gradient direction features and intensity features of each scale are fused to generate multi-scale edge features of the lesion area;

[0045] Compare the multi-scale edge feature map with the preliminary segmentation result of the lesion area by pixel comparison to obtain the edge coverage H within the lesion area;

[0046] Based on the statistical analysis of the segmentation results and edge feature distribution in the historical data, the edge coverage standard L is defined;

[0047] When H≥L, the current lesion area is marked as the real lesion area;

[0048] When H<L, the current lesion area is marked as a pseudo lesion area;

[0049] The pseudo lesion areas in the preliminary segmentation results are eliminated, and finally a lesion localization map is generated.

[0050] In a second aspect, the present invention provides a medical image data processing system, including a preliminary feature map generation module, a comprehensive feature map generation module and a lesion localization map generation module; the preliminary feature map generation module is used to obtain multimodal medical image data of patients including CT, MRI and PET, and preprocess the multimodal image data, input the preprocessed multimodal image data into a deep convolutional neural network, extract the specific features of each modality, and generate a preliminary feature map; the comprehensive feature map generation module is used to introduce an attention mechanism based on the specific features of each modality to calculate the attention weight of each modality, and generate a specific feature map according to the attention weight, and through a feature alignment algorithm, the specific feature maps of different modalities are spatially aligned to generate an aligned specific feature map, and a multi-head attention mechanism is used to identify the correlation between the modalities to generate a comprehensive feature map; the lesion localization map generation module is used to generate a preliminary segmentation result of the lesion area based on the comprehensive feature map and pixel classification through UNet, and at the same time, an edge monitoring algorithm is used to obtain a multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result, eliminate the pseudo lesion area according to the comparison result, and finally generate a lesion localization map.

[0051] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the method for processing medical image data as described in the first aspect of the present invention is implemented.

[0052] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the method for processing medical image data as described in the first aspect of the present invention is implemented.

[0053] The beneficial effects of the present invention are as follows: by introducing the attention mechanism and feature alignment algorithm, the specific features of multimodal medical images are accurately extracted and deep fusion between modalities is achieved, which significantly improves the processing accuracy of multimodal image data; at the same time, the edge monitoring algorithm is combined with the preliminary segmentation results, and the pseudo lesion area is eliminated through multi-scale edge feature comparison, ensuring the accuracy and robustness of lesion positioning. Overall, the present invention solves the problems of insufficient multimodal feature fusion and inaccurate pseudo lesion elimination in the prior art, and provides more reliable technical support for high-precision medical image segmentation and lesion positioning, as well as clinical diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0055] Figure 1 This is a flowchart of the method for processing medical image data in Example 1.

[0056] Figure 2 Schematic diagram of the medical image data processing system in Example 1. DETAILED DESCRIPTION

[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0058] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0059] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0060] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides a method for processing medical image data, comprising the following steps:

[0061] S1. Obtain the patient's multimodal medical imaging data including CT, MRI and PET, preprocess the multimodal imaging data, input the preprocessed multimodal imaging data into a deep convolutional neural network, extract the specific features of each modality, and generate a preliminary feature map.

[0062] Use SIFT to spatially align multimodal image data;

[0063] Furthermore, the scale-invariant feature transform (SIFT) algorithm is used to process multimodal medical image data. First, stable key feature points are extracted from the multimodal medical image data. These feature points can remain unchanged when the image is rotated, scaled or the illumination changes. Then, by matching the feature points in the medical image data of different modalities, the geometric transformation matrix between the two images is calculated, including translation, rotation and scaling parameters. Finally, the multimodal medical image data is affine transformed according to the geometric transformation matrix to ensure that the medical image data of different modalities are aligned in space.

[0064] Using intensity normalization and size normalization processing, the pixel values ​​of multimodal imaging data are normalized to a uniform range and the images are adjusted to the same resolution;

[0065] Furthermore, the multimodal medical imaging data is subjected to intensity normalization processing, and the pixel grayscale values ​​of different modal medical imaging data are linearly mapped to the same range to eliminate the grayscale differences caused by different imaging mechanisms. Subsequently, the spatial resolution of the multimodal medical imaging data is adjusted to be consistent, and the physical size of the image is unified to ensure that all multimodal medical imaging data have the same spatial distribution characteristics at the pixel scale.

[0066] Gaussian filtering method is used to remove background noise and artifacts from multimodal imaging data.

[0067] Furthermore, the Gaussian filtering method is applied to the multimodal medical image data, and the image data is smoothed through convolution operation, and the smoothing characteristics of the Gaussian filter are used to eliminate the high-frequency noise and artifact interference in the multimodal medical image data. At the same time, the edge information and important structural features in the multimodal medical image data are retained to ensure that the key lesion information is not lost while reducing the noise.

[0068] Each modality data in the preprocessed multimodal image data is used as a multi-channel input of a deep convolutional neural network;

[0069] Based on each modality of input data, the first convolutional layer of the deep convolutional neural network combines the ReLU activation function and batch normalization to extract low-level specific features;

[0070] Furthermore, the preprocessed medical imaging data of each modality is input into a deep convolutional neural network. The first convolutional layer performs a convolution operation on the input image, and extracts the features of the local spatial region in the image through multiple convolution kernels. After the convolution operation, the convolution result is nonlinearly transformed in combination with the ReLU activation function to enhance the significant features in the image and suppress irrelevant information. Subsequently, batch normalization is used to normalize the activated features to reduce the differences in the distribution of data in different batches and ensure the stability and expression ability of the feature data.

[0071] Based on the low-level specific features, the second convolutional layer performs convolution and pooling operations again to extract the specific features of each modality and generate a preliminary feature map for each modality;

[0072] Furthermore, using low-level specific features as input, the second convolutional layer further extracts higher-level features through convolution operations to mine more complex local patterns and edge information in the image data. Subsequently, the pooling operation is used to downsample the feature map after convolution, reducing the spatial resolution of the feature map while retaining the main feature information, reducing the computational complexity and enhancing the feature expression ability of the lesion area. Through this process, the deep convolutional neural network extracts the corresponding specific features from each modality of medical imaging data and generates a preliminary feature map for each modality.

[0073] Specific features include density distribution characteristics of CT, soft tissue contrast characteristics of MRI, and metabolic activity characteristics of PET.

[0074] S2. Based on the specific features of each modality, an attention mechanism is introduced to calculate the attention weight of each modality, and a specific feature map is generated according to the attention weight. The specific feature maps of different modalities are spatially aligned through the feature alignment algorithm to generate an aligned specific feature map. The multi-head attention mechanism is used to identify the correlation between the modalities and generate a comprehensive feature map.

[0075] For the preliminary features of each modality, global average pooling and maximum pooling operations are performed on the channel dimension to generate a spatial attention feature map;

[0076] Furthermore, for the preliminary feature maps of each modality of medical imaging data, global average pooling and maximum pooling operations are performed in the channel dimension. The average distribution characteristics of the pixel values ​​in each feature map are extracted by global average pooling to reflect the importance of the global features, and the maximum value characteristics of the pixel values ​​in the feature map are extracted by maximum pooling to highlight the most significant local information. In this way, each preliminary feature map is globally converged in the channel dimension to generate a spatial attention feature map. The spatial attention feature map can comprehensively reflect the global distribution characteristics and local significant characteristics of the lesion area in the image data, thereby providing a basis for the subsequent feature enhancement of the lesion area. For example, in the preliminary feature map of CT image data, global average pooling extracts the overall trend of density distribution, and maximum pooling focuses on the highest density point in the lesion area to jointly generate a spatial attention feature map.

[0077] The spatial attention feature map is convolved and activated with Sigmoid to generate a spatial weight map for each pixel.

[0078] Furthermore, the generated spatial attention feature map is input into the convolution layer, and the spatial attention feature map is integrated and refined through the convolution operation, so that the global features and local features of medical imaging data of different modalities are further integrated in the spatial dimension. After the convolution operation, the convolution result is nonlinearly mapped using the Sigmoid activation function, and the value of each pixel is normalized to a fixed range, thereby generating a spatial weight map corresponding to each pixel. The spatial weight map reflects the importance of each pixel in the image, giving a higher weight to the lesion area while reducing the influence of the background area. For example, for PET imaging data, the metabolic activity features in the spatial attention feature map are convoluted and activated by Sigmoid, and the generated spatial weight map can highlight the significance of the metabolically active area.

[0079] Multiply the spatial weight map with the preliminary feature map pixel by pixel to identify the lesion area;

[0080] Furthermore, the generated spatial weight map is multiplied pixel by pixel with the corresponding preliminary feature map, and the feature intensity of each pixel in the preliminary feature map is adjusted by the weight value in the spatial weight map. The high-weight area in the spatial weight map will enhance the feature expression of the corresponding pixel in the preliminary feature map, while the low-weight area will suppress its feature value, thereby highlighting the important features of the lesion area and weakening the interference of the background area. For example, for CT image data, the high-weight area in the spatial weight map corresponds to the density distribution characteristics of the lesion. After pixel-by-pixel multiplication, the features of the lesion area in the preliminary feature map are further amplified, while the features of the background area are effectively weakened, thereby achieving accurate identification of the lesion area.

[0081] For each modality’s specific feature map, global average pooling and maximum pooling operations are performed in the spatial dimension to generate channel-aggregated features;

[0082] Furthermore, for the specific feature map of each modality of medical imaging data, global average pooling and maximum pooling operations are performed in the spatial dimension respectively. The average eigenvalue of each channel in the entire spatial range is extracted by global average pooling to reflect the overall response strength of the channel. The maximum eigenvalue in the spatial range of each channel is extracted by maximum pooling to highlight the most significant feature response in the channel. The results of global average pooling and maximum pooling are converged in the channel dimension to generate channel convergence features. Channel convergence features can comprehensively express the global importance and local significance of each channel in the specific feature map. For example, in MRI image data, the soft tissue contrast feature of the specific feature map extracts the overall contrast intensity through global average pooling, and the most significant soft tissue area feature is extracted through maximum pooling. The convergence of the two generates channel convergence features.

[0083] The channel aggregation features are processed through convolution operation and Sigmoid activation function to generate the weights of each channel;

[0084] Furthermore, the generated channel aggregation features are input into the convolution layer, and the results of global average pooling and maximum pooling are integrated through the convolution operation to further extract and refine the importance of each channel. The convolutional features are nonlinearly mapped through the Sigmoid activation function to generate a normalized channel weight distribution. The weight value of each channel reflects the importance of the channel in the specific feature map. For example, for PET imaging data, the convolution operation can aggregate the global and local information of metabolic activity features, and the channel weights generated after Sigmoid activation can highlight the significance of channels related to metabolically active areas.

[0085] Multiply the channel-aggregated features with the preliminary feature map channel by channel to enhance the feature channels of the lesion area;

[0086] Furthermore, the generated channel weights are multiplied by the corresponding preliminary feature maps channel by channel, and the features of each channel are weighted and adjusted by the channel weights. Channels with higher weights will have their feature expressions enhanced, while channels with lower weights will be suppressed, thereby enhancing the significance of the features related to the lesion area at the channel level and weakening irrelevant or noise information. For example, in CT image data, the weights of channels containing lesion density distribution features are higher. After multiplying with the preliminary feature map, the lesion features of these channels are enhanced, while the features of the channels related to the background area are effectively reduced, thereby highlighting the channel characteristics of the lesion area.

[0087] Based on the lesion area and the enhanced lesion area feature channel, nonlinear weighting and normalization are used to calculate the attention weight of each modality, which is expressed as:

[0088] ;

[0089] in, It is The attention weights of each modality, It is The spatial feature map of the modal The weight coefficient of It is The spatial features in the spatial feature map of each modality, It is The weight coefficient of the interaction term between the spatial features of each modality and the channel aggregation features, It is to adjust The weight coefficient of the channel aggregation feature of each mode, It is The channel aggregation characteristics of the modes, is the Sigmoid activation function, is the element-wise multiplication operator;

[0090] It should be noted that, first, the interaction between the spatial feature map and its weight coefficient is used to extract the nonlinear changes of the spatial features through the sine function, and the square value of the interaction term between the spatial features and the channel convergence features is combined to highlight the coupling relationship between the two. Secondly, the dynamic range of the channel convergence feature is captured by the logarithmic change, and the importance of the channel feature is adjusted by combining the cosine function of its weight coefficient. In addition, the difference between the spatial features and the channel features is normalized by the exponential function to balance the influence of the two. Finally, the above calculation results are mapped to a fixed range through the Sigmoid activation function to generate the attention weight of each modality, reflecting the importance of the modality to the feature expression of the lesion area. For example, in MRI image data, the interaction between the soft tissue contrast information in the spatial feature map and the channel convergence feature is nonlinearly weighted, and the attention weight can highlight the modal features with significant soft tissue differences.

[0091] The attention weight of each modality It is multiplied layer by layer with the preliminary feature map to generate a specific feature map for each modality.

[0092] Furthermore, the attention weight of each modality is multiplied layer by layer with the corresponding preliminary feature map, and the features of each layer are weighted and enhanced by the attention weight, highlighting the significant features of the lesion area and suppressing irrelevant information. After layer-by-layer multiplication, the specific feature map of each modality is generated, which can reflect the specific expression of each modality to the lesion area. For example, in PET imaging data, the attention weight will enhance the feature intensity of the metabolically active area and reduce the feature value of the background area, thereby generating a specific feature map with prominent metabolic features.

[0093] Bilinear interpolation is used to unify the specific feature maps of all modalities to the same spatial resolution, and the specific feature maps of each modality are preliminarily spatially aligned through affine transformation;

[0094] Furthermore, the spatial resolution of the specific feature maps of all modalities is adjusted to a uniform scale using the bilinear interpolation method to ensure that the features of different modalities can be compared and fused in the same spatial dimension. Subsequently, the specific feature maps are preliminarily aligned in space through affine transformation to correct the rotation, translation and scaling differences between the data of different modalities so that they are preliminarily matched in the global space. For example, the specific feature map of the PET image is aligned to the spatial resolution of the MRI image so that the metabolic features and structural features are basically overlapped in space.

[0095] Perform secondary spatial alignment on the specific feature maps after preliminary alignment through STN, and finally generate an aligned specific feature map;

[0096] Furthermore, the spatial transformer network STN is used to fine-tune the specific feature map after preliminary alignment to further correct the spatial deviation between modalities in local areas. Through the learned transformation parameters, the spatial transformer network can perform elastic deformation or nonlinear mapping on the local area to ensure that the specific feature maps of different modalities are accurately aligned in the lesion area. The aligned specific feature map finally generated lays the foundation for spatial consistency of the fusion of multimodal features. For example, the spatial transformer network can correct the slight misalignment of the CT image specific feature map and the MRI image specific feature map in the lesion edge area.

[0097] The aligned specific feature maps are subjected to three sets of linear transformations to generate query vectors, key vectors, and value vectors;

[0098] Furthermore, for each modality’s aligned specific feature map, three sets of independent linear transformation operations are performed to generate the corresponding query vector, key vector, and value vector. Linear transformation extracts query features for association calculation, key features for feature matching, and value features for representation in the specific feature map by learning different projection matrices. For example, after linear transformation of the aligned CT image specific feature map, the query vector can capture the density features of the lesion area, the key vector describes the potential feature association pattern, and the value vector retains the main feature information of the lesion.

[0099] Based on the query vector, key vector and value vector, the correlation between the modalities is identified, and the expression is:

[0100] ;

[0101] in, It is The query vector of each modality, It is The key vector of the modes, It is The transpose of a mode, is the key vector The dimension of Indicates The mode and The similarity between the modes, is the index variable of the modality, is different from The index variable of another mode;

[0102] It should be noted that, first, the query vector of the ith modality is dot-producted with the transpose of the key vector of the jth modality to obtain the degree of match between the features of the two modalities. In order to avoid numerical instability caused by the size of the vector dimension, the dot product result is scaled and normalized by dividing it by the square root of the key vector dimension. Then, the normalized result is probability-mapped using the Softmax function to generate the correlation weight between the ith modality and the jth modality. The correlation weight can quantify the semantic or content similarity between the features of the two modalities. For example, between MRI images and PET images, the query vector can represent the characteristics of the soft tissue area in the MRI image, the key vector can represent the characteristics of the metabolically active area in the PET image, and the correlation weight By matching the feature similarities of the two modalities in the lesion area, the degree of correlation between soft tissue and metabolic characteristics is revealed, providing a basis for subsequent multimodal feature fusion.

[0103] Based on the correlation between modalities , construct the complementarity and difference mapping between modalities, and fuse them with the specific feature maps of each modality to finally generate a comprehensive feature map.

[0104] Furthermore, first, according to the correlation matrix, the features of the i-th modality are weighted and aggregated, and the cross-modal information is fused with the features of the j-th modality to generate a complementary map, highlighting the collaborative feature expression of different modalities in the lesion area. At the same time, by analyzing the weight distribution of the correlation matrix, the information of the low-correlation area is extracted, and a difference map is generated to retain the independent and complementary characteristics of each modality. Then, the generated complementary and difference maps are fused with the specific feature maps of each modality, and the inter-modal correlation features are organically combined with the single modality features by element-by-element weighting or splicing. Finally, the global features shared by each modality and the local features unique to the modality are extracted through the fused feature map to generate a comprehensive feature map. For example, in CT images and PET images, the specific feature map of CT images highlights the density distribution characteristics, and the specific feature map of PET images reflects the metabolically active areas. The complementary mapping constructed through the correlation matrix expresses the metabolic areas and soft tissue structures in a coordinated manner, while the differential mapping retains the uniqueness of the density characteristics and metabolic characteristics. The comprehensive feature map generated after fusion can capture the anatomical and functional characteristics of the lesions at the same time.

[0105] S3. Based on the comprehensive feature map, pixel classification is performed through UNet to generate the preliminary segmentation result of the lesion area. At the same time, the edge detection algorithm is used to obtain the multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result. According to the comparison result, the pseudo lesion area is eliminated and finally the lesion localization map is generated.

[0106] Based on the comprehensive feature map, the comprehensive feature map is encoded through UNet, the deep semantic information of the comprehensive feature map is gradually extracted, and the spatial resolution in the comprehensive feature map is compressed through sampling operation;

[0107] Furthermore, the comprehensive feature map is input into the encoder part of UNet, and the deep semantic information in the feature map is extracted through multiple layers of progressively stacked convolution operations. Each layer of convolution is followed by a pooling operation or a stride convolution operation to gradually compress the spatial resolution of the feature map and increase the number of channels, thereby extracting more abstract semantic information. During the encoding process, the initial shallow features retain local detail information, while the deep features aggregate more global lesion area features. For example, for MRI images containing tumors, the encoder can extract the boundary features, morphological features, and global distribution information of lesions in soft tissues layer by layer.

[0108] After encoding and sampling, the comprehensive feature map is upsampled layer by layer to gradually restore the spatial resolution, and the comprehensive feature map is mapped into channels of the target area and background through convolution operations;

[0109] Furthermore, after encoding is completed, the comprehensive feature map is upsampled layer by layer through the decoder part of UNet, and the spatial resolution is gradually restored using bilinear interpolation or deconvolution operations. After upsampling at each layer, the shallow features retained by the corresponding encoder layer are jump-connected with the upsampled features, and then the shallow and deep feature information are fused through convolution operations to ensure that the segmentation result contains both global semantic information and detailed information. In the last layer of decoding, the comprehensive feature map is converted into two channels through convolution operations, representing the features of the target area (lesion) and the background area respectively. For example, for tumor segmentation tasks, the feature map restored by the decoder can clearly express the boundary and location of the tumor while suppressing the noise information in the background.

[0110] Based on the channels of the target area and the background, the Softmax activation function is used to calculate the probability that each pixel belongs to the lesion area, and the preliminary segmentation result of the lesion area is generated;

[0111] Furthermore, the target area channel and background channel output by the decoder are normalized using the Softmax activation function to normalize the value of each pixel, generating a probability distribution of each pixel belonging to the lesion area and the background area. Subsequently, based on the probability map of the target area, a preliminary segmentation result of the lesion area is generated by setting a threshold or directly selecting the category with the maximum probability. This segmentation result can mark the pixel position of the lesion area. For example, in PET images, the Softmax calculation can mark high-probability pixels in metabolically active areas as lesions, and low-probability pixels as background, thereby generating a preliminary segmentation map of the tumor area.

[0112] The preliminary segmentation results include the contour and location of the lesion area in the comprehensive feature map.

[0113] Use Canny to extract the contour of the lesion area and the boundary of the internal structure from the preliminary segmentation results, and use Gaussian pyramid to generate edge feature maps of different scales;

[0114] Furthermore, the preliminary segmentation results are processed by the Canny edge detection algorithm. First, a smoothing filter is used to suppress noise, and then the image gradient is calculated to detect the significant edges of the lesion area, including the boundaries of the contours and internal structures. Next, the edge pixels are accurately located through non-maximum suppression and double threshold segmentation to obtain a clear boundary of the lesion area. On this basis, the Gaussian pyramid is used to perform multi-scale downsampling on the preliminary segmentation results to generate multiple edge feature maps of different resolutions, so that the features of the lesion edge can be captured at different scales. For example, for tumor segmentation results, the Gaussian pyramid can generate edge feature maps from small scales with rich details to global views.

[0115] Based on edge feature maps of different scales, the local gradient direction features and intensity features of the edge of the lesion area are extracted on the edge feature maps of each scale through multi-resolution filters;

[0116] Furthermore, on the edge feature map of each scale, a multi-resolution filter is used to calculate the local gradient direction and intensity features of each edge pixel. The gradient direction feature represents the main direction of the edge in the local area, while the gradient intensity feature represents the significance and change amplitude of the edge. By extracting gradient information at different scales, the global trend and local details of the edge of the lesion can be captured. For example, on a smaller-scale edge feature map, the filter can highlight the subtle boundaries inside the tumor, while on a larger-scale feature map, the direction and intensity features of the overall contour of the tumor can be extracted.

[0117] The local gradient direction features and intensity features of each scale are fused to generate multi-scale edge features of the lesion area;

[0118] Furthermore, the gradient direction features and intensity features of each scale are fused scale by scale, and the edge information at different resolutions is integrated into the same feature space. In the fusion process, the small-scale detail features are retained, and the large-scale global edge features are highlighted, generating multi-scale edge features that can fully reflect the edge information of the lesion area. For example, for the lesion area of ​​CT images, the small-scale gradient features can accurately depict the local boundary details of the lesion, while the large-scale gradient features can describe the overall shape and contour of the lesion. The fused multi-scale edge features can simultaneously reflect the details and global information of the lesion.

[0119] Compare the multi-scale edge feature map with the preliminary segmentation result of the lesion area by pixel comparison to obtain the edge coverage H within the lesion area;

[0120] Furthermore, the multi-scale edge feature map is compared pixel by pixel with the preliminary segmentation result of the lesion area. First, in the preliminary segmentation result of the lesion area, it is determined whether each pixel belongs to the lesion area. Then, the number of pixels with edge response in the multi-scale edge feature map is counted among the pixels in the lesion area, which is recorded as the number of edge coverage pixels. Next, the edge coverage rate H in the lesion area is obtained by dividing the number of edge coverage pixels by the total number of pixels in the lesion area. The edge coverage rate H in the lesion area represents the proportion of the lesion area that is accurately covered by the edge features. For example, in the tumor segmentation task, if the lesion area contains 100 pixels in the preliminary segmentation result, and 80 pixels in the multi-scale edge feature map have edge responses in the lesion area, then the edge coverage rate H in the lesion area is 80%.

[0121] Based on the statistical analysis of the segmentation results and edge feature distribution in the historical data, the edge coverage standard L is defined;

[0122] When H≥L, the current lesion area is marked as the real lesion area;

[0123] When H<L, the current lesion area is marked as a pseudo lesion area;

[0124] The pseudo lesion areas in the preliminary segmentation results are eliminated, and finally a lesion localization map is generated.

[0125] It should be noted that the lesion localization map reflects the spatial distribution, significant features and differences between the lesion area and the surrounding tissue in the medical image. It can also reveal the possible location, shape and range of the lesion area, and reflect the characteristic differences between the lesion and the background area in terms of structure, intensity or texture. For example, in MRI images, the lesion localization map can reflect the location of the abnormal signal area and its significance in the image, providing guidance for accurate segmentation and diagnosis.

[0126] The present embodiment also provides a medical image data processing system, including: a preliminary feature map generation module, a comprehensive feature map generation module and a lesion localization map generation module; the preliminary feature map generation module is used to obtain multimodal medical image data of a patient including CT, MRI and PET, and preprocess the multimodal image data, input the preprocessed multimodal image data into a deep convolutional neural network, extract the specific features of each modality, and generate a preliminary feature map; the comprehensive feature map generation module is used to introduce an attention mechanism based on the specific features of each modality to calculate the attention weight of each modality, and generate a specific feature map according to the attention weight, and spatially align the specific feature maps of different modalities through a feature alignment algorithm to generate an aligned specific feature map, and use a multi-head attention mechanism to identify the correlation between the modalities to generate a comprehensive feature map; the lesion localization map generation module is used to generate a preliminary segmentation result of the lesion area based on the comprehensive feature map and pixel classification through UNet, and at the same time use an edge monitoring algorithm to obtain a multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result, eliminate the pseudo lesion area according to the comparison result, and finally generate a lesion localization map.

[0127] This embodiment also provides a computer device suitable for the case of a method for processing medical image data, comprising: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to implement the method for processing medical image data proposed in the above embodiment.

[0128] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0129] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for processing medical image data proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0130] In summary, the present invention introduces an attention mechanism and a feature alignment algorithm to accurately extract specific features of multimodal medical images and achieve deep fusion between modalities, significantly improving the processing accuracy of multimodal image data; at the same time, it utilizes an edge monitoring algorithm in combination with preliminary segmentation results to eliminate pseudo-lesion areas through multi-scale edge feature comparison, thereby ensuring the accuracy and robustness of lesion localization. Overall, the present invention solves the problems of insufficient multimodal feature fusion and inaccurate pseudo-lesion elimination in the prior art, and provides more reliable technical support for high-precision medical image segmentation and lesion localization, as well as clinical diagnosis and treatment.

[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for processing medical image data, characterized in that: include, Obtain multimodal medical imaging data of patients including CT, MRI and PET, preprocess the multimodal imaging data, input the preprocessed multimodal imaging data into a deep convolutional neural network, extract specific features of each modality, and generate a preliminary feature map; Based on the specific features of each modality, an attention mechanism is introduced to calculate the attention weight of each modality, and a specific feature map is generated according to the attention weight. The specific feature maps of different modalities are spatially aligned through the feature alignment algorithm to generate an aligned specific feature map. The multi-head attention mechanism is used to identify the correlation between the modalities and generate a comprehensive feature map. Based on the comprehensive feature map, pixel classification is performed through UNet to generate the preliminary segmentation result of the lesion area. At the same time, the edge detection algorithm is used to obtain the multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result. The pseudo lesion area is eliminated according to the comparison results, and finally the lesion localization map is generated.

2. The method for processing medical image data according to claim 1, wherein: The multimodal image data is preprocessed, and the specific steps are as follows: Use SIFT to spatially align multimodal image data; Using intensity normalization and size normalization processing, the pixel values ​​of multimodal imaging data are normalized to a uniform range and the images are adjusted to the same resolution; Gaussian filtering method is used to remove background noise and artifacts from multimodal imaging data.

3. The method for processing medical image data according to claim 2, wherein: The preprocessed multimodal image data is input into a deep convolutional neural network to extract the specific features of each modality and generate a preliminary feature map. The specific steps are as follows: Each modality data in the preprocessed multimodal image data is used as a multi-channel input of a deep convolutional neural network; Based on each modality of input data, the first convolutional layer of the deep convolutional neural network combines the ReLU activation function and batch normalization to extract low-level specific features; Based on the low-level specific features, the second convolutional layer performs convolution and pooling operations again to extract the specific features of each modality and generate a preliminary feature map for each modality; The specific features include density distribution features of CT, soft tissue contrast features of MRI and metabolic activity features of PET.

4. The method for processing medical image data according to claim 3, wherein: Based on the specific features of each modality, the attention mechanism is introduced to calculate the attention weight of each modality, and a specific feature map is generated according to the attention weight. The specific steps are as follows: For the preliminary features of each modality, global average pooling and maximum pooling operations are performed on the channel dimension to generate a spatial attention feature map; The spatial attention feature map is convolved and activated with Sigmoid to generate a spatial weight map for each pixel. Multiply the spatial weight map with the preliminary feature map pixel by pixel to identify the lesion area; For each modality’s specific feature map, global average pooling and maximum pooling operations are performed in the spatial dimension to generate channel-aggregated features; The channel aggregation features are processed through convolution operation and Sigmoid activation function to generate the weights of each channel; Multiply the channel-aggregated features with the preliminary feature map channel by channel to enhance the feature channels of the lesion area; Based on the lesion area and the enhanced lesion area feature channel, nonlinear weighting and normalization are used to calculate the attention weight of each modality, which is expressed as: ; in, It is The attention weights of each modality, It is The spatial feature map of the modal The weight coefficient of It is The spatial features in the spatial feature map of each modality, It is The weight coefficient of the interaction term between the spatial features of each modality and the channel aggregation features, It is to adjust The weight coefficient of the channel aggregation feature of each mode, It is The channel aggregation characteristics of the modes, is the Sigmoid activation function, is the element-wise multiplication operator; The attention weight of each modality It is multiplied layer by layer with the preliminary feature map to generate a specific feature map for each modality.

5. The method for processing medical image data according to claim 4, wherein: The feature alignment algorithm is used to spatially align the specific feature maps of different modalities to generate aligned specific feature maps, and the multi-head attention mechanism is used to learn the correlation between modalities to generate a comprehensive feature map. The specific steps are as follows: Bilinear interpolation is used to unify the specific feature maps of all modalities to the same spatial resolution, and the specific feature maps of each modality are preliminarily spatially aligned through affine transformation; Perform secondary spatial alignment on the specific feature maps after preliminary alignment through STN, and finally generate an aligned specific feature map; The aligned specific feature maps are subjected to three sets of linear transformations to generate query vectors, key vectors, and value vectors; Based on the query vector, key vector and value vector, the correlation between the modalities is identified, and the expression is: ; in, It is The query vector of each modality, It is The key vector of the modes, It is The transpose of a mode, is the key vector The dimension of Indicates The mode and The similarity between the modes, is the index variable of the modality, is different from The index variable of another mode; Based on the correlation between modalities , construct the complementary and difference mapping between modalities, and fuse them with the specific feature maps of each modality to finally generate a comprehensive feature map.

6. The method for processing medical image data according to claim 5, wherein: Based on the comprehensive feature map, pixel classification is performed on the comprehensive feature map through UNet to generate a preliminary segmentation result of the lesion area. The specific steps are as follows: Based on the comprehensive feature map, the comprehensive feature map is encoded through UNet, the deep semantic information of the comprehensive feature map is gradually extracted, and the spatial resolution in the comprehensive feature map is compressed through sampling operation; After encoding and sampling, the comprehensive feature map is upsampled layer by layer to gradually restore the spatial resolution, and the comprehensive feature map is mapped into channels of the target area and background through convolution operations; Based on the channels of the target area and the background, the Softmax activation function is used to calculate the probability that each pixel belongs to the lesion area, and the preliminary segmentation result of the lesion area is generated; The preliminary segmentation result includes the outline and position of the lesion area in the comprehensive feature map.

7. The method for processing medical image data according to claim 6, wherein: The edge monitoring algorithm is used to obtain a multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result, and the pseudo lesion area is eliminated according to the comparison result, and finally a lesion localization map is generated. The specific steps are as follows: Use Canny to extract the contour of the lesion area and the boundary of the internal structure from the preliminary segmentation results, and use Gaussian pyramid to generate edge feature maps of different scales; Based on edge feature maps of different scales, the local gradient direction features and intensity features of the edge of the lesion area are extracted on the edge feature maps of each scale through multi-resolution filters; The local gradient direction features and intensity features of each scale are fused to generate multi-scale edge features of the lesion area; Compare the multi-scale edge feature map with the preliminary segmentation result of the lesion area by pixel comparison to obtain the edge coverage H within the lesion area; Based on the statistical analysis of the segmentation results and edge feature distribution in the historical data, the edge coverage standard L is defined; When H≥L, the current lesion area is marked as the real lesion area; When H<L, the current lesion area is marked as a pseudo lesion area; The pseudo lesion areas in the preliminary segmentation results are eliminated, and finally a lesion localization map is generated.

8. A system for processing medical image data, based on the method for processing medical image data according to any one of claims 1 to 7, characterized in that: It includes a preliminary feature map generation module, a comprehensive feature map generation module and a lesion localization map generation module; A preliminary feature map generation module is used to obtain multimodal medical imaging data of patients including CT, MRI and PET, and preprocess the multimodal imaging data, input the preprocessed multimodal imaging data into a deep convolutional neural network, extract the specific features of each modality, and generate a preliminary feature map; The comprehensive feature map generation module is used to introduce an attention mechanism to calculate the attention weight of each modality based on the specific features of each modality, and generate a specific feature map according to the attention weight. The specific feature maps of different modalities are spatially aligned through a feature alignment algorithm to generate an aligned specific feature map, and the multi-head attention mechanism is used to identify the correlation between the modalities to generate a comprehensive feature map. The lesion localization map generation module is used to generate a preliminary segmentation result of the lesion area based on the comprehensive feature map and pixel classification through UNet. At the same time, the edge detection algorithm is used to obtain the multi-scale edge feature map of the lesion area and compare it with the preliminary segmentation result. The pseudo lesion area is eliminated according to the comparison result, and finally the lesion localization map is generated.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for processing medical image data according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for processing medical image data according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Liver image recognition method and device based on graph neural network

    CN116934754A

  • Fundus pseudo focus detection method

    CN118115466A