Medical image multi-modal feature extraction method based on multi-dimensional convolution transformation

Through the combination of multi-dimensional convolution transformation and PHER optimization algorithm, the problem of insufficient information integration and adaptability in multimodal medical image feature extraction is solved, efficient feature extraction and deep fusion are achieved, and accuracy and adaptability are significantly improved.

CN120071035AInactive Publication Date: 2025-05-30YONGZHOU CENT HOSPITAL
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510139748.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate complementary information in multimodal medical images, and it is not adaptable in high noise and small sample scenarios, resulting in low feature extraction accuracy.

Method used

Using a multi-dimensional convolution transformation method, convolution operations are performed on spatial dimensions and modal dimensions through a multi-dimensional convolution kernel, combined with the PHER optimization algorithm, dynamically adjust the feature weights and feature fusion weights between modals, and deep fusion of multi-modal features is achieved through the fusion network.

Benefits of technology

It significantly reduces information conflicts between modes, improves the expressiveness of the model when dealing with complex lesion areas, enhances the adaptability to high-noise and small-sample scenes, and improves the accuracy of multimodal medical imaging feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071035A_ABST
    Figure CN120071035A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image multi-modal feature extraction method based on multi-dimensional convolution transformation. The method comprises the following steps: S1, acquiring multi-modal medical image data; s2, obtaining initial feature mapping of each mode; s3, multi-modal feature mapping after alignment is obtained; s4, obtaining an optimized multi-modal feature map; s5, performing deep fusion on the optimized multi-modal feature mapping by using a fusion network, and performing information complementation and weight adaptive distribution on each modal feature mapping to obtain fused multi-modal feature representation; s6, carrying out regularization processing on the fused multi-modal characteristic representation so as to inhibit characteristic noise and redundant information, and obtaining regularized optimized multi-modal characteristic representation; and S7, outputting the regularized optimized multi-modal characteristic representation. According to the method, information conflicts among modals can be remarkably reduced, and the expressive force of the model in processing a complex focus area is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical imaging technology, and in particular, to a method for extracting multi-modal features of medical images based on multi-dimensional convolution transformation. Background Art

[0002] With the rapid development of medical imaging technology, multi-modal imaging technologies such as CT, MRI, and PET have been widely used in clinical diagnosis. These imaging technologies can provide multi-dimensional information on the internal structure, function, and lesions of patients' bodies, which is of great significance for the early diagnosis of diseases and the formulation of precise treatment plans. However, due to significant differences in physical properties, imaging principles, and resolutions among different modal images, how to effectively integrate complementary information in multi-modal medical images has become an important research topic in the field of current medical image analysis.

[0003] Currently, the extraction of multi-modal medical image features mainly relies on traditional image processing algorithms and some deep learning models. For example, methods based on manually designed features extract texture and shape features and fuse them. However, such methods usually have difficulty in processing complex inter-modal correlation information and have limited feature expression capabilities. With the development of deep learning technology, convolutional neural networks have been increasingly widely used in medical image processing. Some studies have attempted to extract and fuse features of multi-modal medical images through two-dimensional or three-dimensional convolutional networks. However, these existing technologies still have some prominent problems in practical applications:

[0004] First of all, existing technologies are difficult to make full use of complementary information in multi-modal medical images. Due to significant spatial differences and information redundancy between modal data, existing methods are prone to information loss during inter-modal alignment and fusion, and cannot effectively capture deep-level correlation features between modalities. Secondly, existing deep learning methods have insufficient adaptability to medical image data with high noise and small samples. Medical image data often has noise, blurring, and uneven distribution, and the scarcity of labeled data further limits the training effect of the model, resulting in low feature extraction accuracy.

[0005] In summary, existing technologies have significant defects in inter-modal information integration, adaptability to high-noise and small-sample scenarios, and parameter optimization. The problems of existing technologies limit the effect of multi-modal medical image feature extraction technology in practical clinical applications. There is an urgent need for an innovative method that can efficiently align and fuse multi-modal features and adapt to the characteristics of high noise and small samples in medical images to solve the above problems. Summary of the Invention

[0006] An object of the present invention is to propose a method for extracting multi-modal features of medical images based on multi-dimensional convolution transformation. The present invention can significantly reduce information conflicts between modalities and improve the performance of the model in processing complex lesion regions.

[0007] A method for extracting multimodal features of medical images based on multidimensional convolution transformation according to an embodiment of the present invention comprises the following steps:

[0008] S1. Acquire multimodal medical image data, where the multimodal medical image data includes medical image data of at least two different imaging modalities, and perform spatial alignment, intensity standardization, and denoising on the multimodal medical image data to obtain preprocessed multimodal medical image data;

[0009] S2. Multidimensional convolution transform performs convolution operation on multimodal medical imaging data in spatial dimension and modality dimension through multidimensional convolution kernel to obtain preliminary feature mapping of each modality;

[0010] S3. Performing spatial alignment and channel alignment on the preliminary feature maps, aligning the preliminary feature maps of each modality in the spatial dimension and the channel dimension to obtain an aligned multimodal feature map;

[0011] S4. Optimizing the aligned multimodal feature mapping based on the Puller optimization algorithm. The Puller optimization algorithm adaptively optimizes the aligned multimodal feature mapping by global search and local search by dynamically iteratively adjusting feature weights and multidimensional convolution kernel parameters to obtain an optimized multimodal feature mapping.

[0012] S5. Using the fusion network to deeply fuse the optimized multimodal feature mapping, perform information complementation and weight adaptive allocation on each modality feature mapping, and obtain a fused multimodal feature representation;

[0013] S6. performing regularization processing on the fused multimodal feature representation to suppress feature noise and redundant information, and obtaining a regularized optimized multimodal feature representation;

[0014] S7. Output the regularized optimized multimodal feature representation for medical image analysis tasks such as disease classification, lesion segmentation, or lesion detection.

[0015] Optionally, the S1 includes:

[0016] S11. Acquire multimodal medical image data, including medical image data of at least two different imaging modalities, namely, first modality medical image data and second modality medical image data, wherein the first modality medical image data includes CT image data D CT (x, y, z), the second modality medical image data includes MRI image data D MRI (x,y,z);

[0017] S12. Perform preliminary alignment processing on the first-modal medical image data and the second-modal medical image data according to the spatial resolution and voxel size of the first-modal medical image data and the second-modal medical image data;

[0018] S13. Perform precise alignment on the first-modal medical image data and the second-modal medical image data through the calibration point matching method to generate the aligned first-modal medical image data and the aligned second-modal medical image data;

[0019] S14. Sample the aligned first-modal medical image data and the second-modal medical image data, and uniformly adjust them to the target resolution, where the target resolution is preset according to the requirements of the actual application scenario;

[0020] S15. Perform normalization processing on the first-modal medical image data and the second-modal medical image data after adjusting the resolution, so that their pixel intensity values are distributed within the range of [-1, 1];

[0021] S16. Output the first-modal medical image data and the second-modal medical image data as multi-modal medical image data.

[0022] Optionally, the S3 includes:

[0023] S21. Perform a chunking operation on the first-modal medical image data and the second-modal medical image data after normalization processing, and divide the input data into several overlapping local sub-regions;

[0024] S22. Define a multi-dimensional adaptive convolution kernel W m,n (k x , k y , k z , α, β), where k x , k y , k z is the size of the convolution kernel in the spatial dimension, α and β respectively represent the dynamic adjustment parameters of the multi-dimensional adaptive convolution kernel for the mutual information between modalities and the local noise intensity, m represents the input channel index, corresponding to the number of input feature channels for each modality, and n represents the output channel index, characterizing the types of features extracted by convolution;

[0025] S23. According to the mutual information intensity and noise distribution between modalities within the local sub-region, update the convolution kernel parameters α and β in real time, adjust the convolution kernel weights, and optimize the cross-modal feature extraction ability:

[0026]

[0027] where, MI represents the mutual information between modalities, σnoise represents the local noise intensity, Vol(x ′ ,y ′ ,z ′ ) represents the volume of the local sub-region;

[0028] S24. Perform a convolution operation on the first-modal medical image data and the second-modal medical image data using the updated adaptive convolution kernel to calculate the modality-specific preliminary feature map. The preliminary feature map of the first modality is:

[0029]

[0030] where M represents the number of input feature channels, K x ,K y ,K z represents the receptive field range of the convolution kernel in each dimension;

[0031] Perform a convolution operation on the second-modal medical image data to obtain the preliminary feature map of the second modality:

[0032]

[0033] Optionally, the S4 includes:

[0034] S31. Align the first-modal preliminary feature map and the second-modal preliminary feature map in the spatial dimension, and calculate the transformation matrix T spatial in the spatial dimension. The transformation matrix is calculated based on the geometric correspondence and spatial deviation of the two-modal data. According to the transformation matrix T spatial perform a spatial transformation on the first-modal preliminary feature map and the second-modal preliminary feature map to generate the spatially aligned first-modal feature map F ′ CT (x ′ ,y ′ ,z ′ ) and the second-modal feature map F ′ MRI (x ′ ,y ′ ,z ′ );

[0035] S33. Align the spatially aligned first-modal feature map and the second-modal feature map in the channel dimension, and define the modality-specific channel weighting matrix W channel , where W channel,CT and W channel,MRI respectively represent the channel weights of the first modality and the second modality. The size of the weighting matrix is the same as the number of input feature channels. Reallocate the feature channels of each modality through a weighting operation:

[0036] F ″ CT (x ′ ,y ′ ,z ′ ) = W channel,CT ·F ′ CT (x ′ ,y ′ ,z ′ );

[0037] F ″ MRI (x ′ ,y ′ ,z ′ ) = W channel,MRI ·F ′ MRI (x ′ ,y ′ ,z ′ );

[0038] Wherein, F ″ CT (x ′ ,y ′ ,z ′ ) and F ″ MRI (x ′ ,y ′ ,z ′ ) respectively represent the first modal feature map and the second modal feature map after channel alignment;

[0039] S34. Combine the first modal feature map and the second modal feature map after spatial alignment and channel alignment to generate the finally aligned multimodal feature map F aligned (x ′ ,y ′ ,z ′ ):

[0040] F aligned (x ′ ,y ′ ,z ′ ) = γ CT ·F ″ CT (x ′ ,y ′ ,z ′ ) + γ MRI ·F ″ MRI (x ′ ,y ′ ,z ′ );

[0041] Among them, γ CT and γ MRI are the feature fusion weights of the first modality and the second modality respectively.

[0042] Optionally, the S5 includes:

[0043] S41. Receive the aligned multimodal feature map F aligned (x ′ ,y ′ ,z ′ ), initialize the particle position of the Puller optimization algorithm Particle Speed Individual optimal position The global optimal position G best , as well as inertia weight w, individual learning factor c 1 , group learning factor c 2 , particle positions and velocities are initialized to random distribution;

[0044] S42. Define the optimization objective function J(θ), which integrates the mutual information between modalities, feature extraction accuracy and noise suppression capability:

[0045]

[0046] Among them, θ represents the multidimensional convolution kernel parameters and modal weights, is the loss function, which measures the accuracy of feature extraction, MI is the mutual information between modalities, which measures the correlation between modal features, and σ noise is the local noise intensity in the feature map, λ 1 and λ 2 is the balance coefficient, which is used to adjust the contribution of different objectives to the optimization process;

[0047] S43. Dynamically iterate and adjust the particle swarm through the Puller optimization algorithm to update the speed and position of each particle:

[0048]

[0049] Among them, t is the current iteration number, r 1 and r 2 is a random number in the range [0,1], and Respectively represent the position and velocity of the particle after the current iteration;

[0050] Update individual optimal positions and global optimal positions:

[0051]

[0052] S44. In the iterative process, according to the optimized parameter θ* , dynamically adjust the weights of the convolutional kernel and the weights of modal fusion, and is the optimized result of the modal weights, and the optimized multi-dimensional convolutional kernel weights are used to recalculate the feature map;

[0053] S45. Re-convolve the aligned multi-modal feature maps using the optimized parameters, and generate the optimized multi-modal feature map F optimized (x ′ , y ′ , z ′ ):

[0054]

[0055] Optionally, the S6 includes:

[0056] S51. Define the structure of the fusion network through the optimized multi-modal feature map F optimized (x ′ , y ′ , z ′ ), including multiple convolutional modules, an attention mechanism module, and a feature reallocation module;

[0057] S52. In the first stage of the fusion network, use the multiple convolutional modules to perform a preliminary fusion operation on the optimized multi-modal feature map, and input the optimized multi-modal feature map F optimized (x ′ , y ′ , z ′ ) into the multiple convolutional modules to extract cross-modal joint features and output the preliminary fusion feature map F conv (x ′ , y ′ , z ′ ):

[0058] F conv (x ′ , y ′ , z ′ ) = Conv fusion (F optimized (x ′ , y ′ , z ′ ));

[0059] Among them, Conv fusion represents the multiple convolutional operations in the fusion network, which are used to extract the deep features of multi-modal;

[0060] S53. In the second stage of the fusion network, the attention mechanism module is used to adjust the weights of the preliminary fusion feature maps, dynamically allocate the weights of each modality feature, and define the modality attention weights α CT and α MRI , which respectively represent the feature importance of the first modality and the second modality. Through global pooling and the weight calculation formula, the modality attention weights are dynamically calculated:

[0061]

[0062] The preliminary fusion feature maps are weighted and adjusted using the attention weights to generate attention-optimized feature maps:

[0063] F att (x ′ ,y ′ ,z ′ ) = α CT ·F conv,CT (x ′ ,y ′ ,z ′ ) + α MRI ·F conv,MRI (x ′ ,y ′ ,z ′ );

[0064] S54. In the third stage of the fusion network, the feature reassignment module is used to perform deep fusion on the attention-optimized feature map F att (x ′ ,y ′ ,z ′ ), combine the complementary information between modalities and the spatial distribution characteristics, and generate the fused multi-modal feature representation F fused (x ′ ,y ′ ,z ′ ):

[0065] F fused (x ′ ,y ′ ,z ′ ) = Reassign(F att (x ′ ,y ′ ,z ′ ));

[0066] Among them, Reassign represents the feature reassignment operation, which is used to further optimize the complementary characteristics between modalities.

[0067] The beneficial effects of the present invention are:

[0068] (1) The present invention effectively solves the problem of difficulty in aligning and fusing multimodal image features in the prior art by designing a multidimensional adaptive convolution kernel and combining inter-modal mutual information and a dynamic adjustment mechanism for local noise. The multidimensional convolution transform not only extracts features in the spatial dimension, but also realizes the deep capture of inter-modal complementary information in the modal dimension. The adaptive multidimensional convolution kernel can automatically adjust weights and parameters according to the characteristics of the modal data, thereby significantly improving the utilization rate of modal mutual information in the feature extraction process.

[0069] (2) The present invention introduces the Puller optimization algorithm in the feature optimization process to dynamically adjust the key parameters in the feature extraction process, including the weights of the multi-dimensional convolution kernel and the weight allocation strategy for inter-modal feature fusion, thereby significantly improving the adaptability of the model in high noise and small sample scenarios. The Puller optimization algorithm dynamically balances between global search and local search, and by defining a comprehensive objective function, takes modal mutual information maximization and noise suppression as optimization goals, effectively suppressing noise interference in feature extraction and avoiding model overfitting in small sample scenarios.

[0070] (3) The present invention realizes the deep fusion of multimodal features by constructing a fusion network, including a multi-layer convolution module, an attention mechanism module and a feature reallocation module. Based on the initially extracted optimization features, the fusion network dynamically adjusts the weights of different modal features using the attention mechanism so that the model can focus on the key lesion area, and further optimizes the complementary characteristics between modalities through the feature reallocation module, thereby improving the global representation ability of the fusion features. Different from the existing simple modal feature superposition method, the method of the present invention can significantly reduce the information conflict between modalities and improve the expressiveness of the model when processing complex lesion areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0072] Figure 1 This is a flowchart of a medical image multimodal feature extraction method based on multidimensional convolution transformation proposed by the present invention. DETAILED DESCRIPTION

[0073] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0074] refer to Figure 1 , a medical image multimodal feature extraction method based on multidimensional convolution transformation, comprising the following steps:

[0075] S1. Acquire multimodal medical image data, where the multimodal medical image data includes medical image data of at least two different imaging modalities, and perform spatial alignment, intensity standardization, and denoising on the multimodal medical image data to obtain preprocessed multimodal medical image data;

[0076] S2. Multidimensional convolution transform performs convolution operation on multimodal medical imaging data in spatial dimension and modality dimension through multidimensional convolution kernel to obtain preliminary feature mapping of each modality;

[0077] S3. Performing spatial alignment and channel alignment on the preliminary feature maps, aligning the preliminary feature maps of each modality in the spatial dimension and the channel dimension to obtain an aligned multimodal feature map;

[0078] S4. Optimizing the aligned multimodal feature mapping based on the Puller optimization algorithm. The Puller optimization algorithm adaptively optimizes the aligned multimodal feature mapping by global search and local search by dynamically iteratively adjusting feature weights and multidimensional convolution kernel parameters to obtain an optimized multimodal feature mapping.

[0079] S5. Use the fusion network to deeply fuse the optimized multimodal feature maps, perform information complementation and weight adaptive allocation on each modality feature map, and obtain the fused multimodal feature representation;

[0080] S6. performing regularization processing on the fused multimodal feature representation to suppress feature noise and redundant information, and obtaining a regularized optimized multimodal feature representation;

[0081] S7. Output the regularized optimized multimodal feature representation for medical image analysis tasks such as disease classification, lesion segmentation, or lesion detection.

[0082] In this implementation, S1 includes:

[0083] S11. Acquire multimodal medical image data, including medical image data of at least two different imaging modalities, namely, first modality medical image data and second modality medical image data, wherein the first modality medical image data includes CT image data D CT (x, y, z), the second modality medical image data includes MRI image data D MRI (x,y,z);

[0084] S12. Preliminarily aligning the first modality medical image data and the second modality medical image data according to the spatial resolution and voxel size of the first modality medical image data and the second modality medical image data;

[0085] S13. Precisely align the first-modal medical image data and the second-modal medical image data through the calibration point matching method to generate the aligned first-modal medical image data and the aligned second-modal medical image data;

[0086] S14. Sample the aligned first-modal medical image data and the second-modal medical image data and uniformly adjust them to the target resolution, which is preset according to the requirements of the actual application scenario;

[0087] S15. Perform normalization processing on the first-modal medical image data and the second-modal medical image data after adjusting the resolution so that their pixel intensity values are distributed within the range of [-1, 1];

[0088] S16. Output the normalized first-modal medical image data and the second-modal medical image data as the multi-modal medical image data.

[0089] In this embodiment, S3 includes:

[0090] S21. Perform a chunking operation on the normalized first-modal medical image data and the second-modal medical image data, and divide the input data into several overlapping local sub-regions;

[0091] S22. Define a multi-dimensional adaptive convolution kernel W m,n (k x , k y , k z , α, β), where k x , k y , k z is the size of the convolution kernel in the spatial dimension, α and β respectively represent the dynamic adjustment parameters of the multi-dimensional adaptive convolution kernel for the inter-modal mutual information and the local noise intensity, m represents the input channel index, corresponding to the number of input feature channels of each modality, and n represents the output channel index, characterizing the types of features extracted by convolution;

[0092] S23. According to the inter-modal mutual information intensity and the noise distribution within the local sub-region, update the convolution kernel parameters α and β in real time, adjust the convolution kernel weights, and optimize the cross-modal feature extraction ability:

[0093]

[0094] where MI represents the inter-modal mutual information, σ noise represents the local noise intensity, and Vol(x ′ , y ′ , z ′ ) represents the volume of the local sub-region;

[0095] S24. Perform a convolution operation on the first-modal medical image data and the second-modal medical image data using the updated adaptive convolution kernel to calculate the modality-specific preliminary feature maps. The preliminary feature map of the first modality is:

[0096]

[0097] where M represents the number of input feature channels, and K x , K y , K z represents the receptive field range of the convolution kernel in each dimension;

[0098] Perform a convolution operation on the second-modal medical image data to obtain the preliminary feature map of the second modality:

[0099]

[0100] In this embodiment, S4 includes:

[0101] S31. Align the first-modal preliminary feature map and the second-modal preliminary feature map in the spatial dimension, and calculate the transformation matrix T spatial in the spatial dimension of the two modality feature maps. The transformation matrix is calculated based on the geometric correspondence and spatial deviation of the two modality data. According to the transformation matrix T spatial perform a spatial transformation on the first-modal preliminary feature map and the second-modal preliminary feature map to generate the spatially aligned first-modal feature map F ′ CT (x ′ , y ′ , z ′ ) and the second-modal feature map F ′ MRI (x ′ , y ′ , z ′ );

[0102] S33. Align the spatially aligned first-modal feature map and the second-modal feature map in the channel dimension, and define the modality-specific channel weighting matrix W channel , where W channel,CT and W channel,MRI represent the channel weights of the first modality and the second modality respectively. The size of the weight matrix is the same as the number of input feature channels. Reallocate the feature channels of each modality through a weighting operation:

[0103] F ″ CT (x ′ , y ′ , z ′ ) = Wchannel,CT ·F ′ CT (x ′ ,y ′ ,z ′ );

[0104] F ″ MRI (x ′ ,y ′ ,z ′ )=W channel,MRI ·F ′ MRI (x ′ ,y ′ ,z ′ );

[0105] Among them, F ″ CT (x ′ ,y ′ ,z ′ ) and F ″ MRI (x ′ ,y ′ ,z ′ ) respectively represent the first-modal feature map and the second-modal feature map after channel alignment;

[0106] S34. Combine the first-modal feature map and the second-modal feature map after spatial alignment and channel alignment to generate the finally aligned multi-modal feature map F aligned (x ′ ,y ′ ,z ′ ):

[0107] F aligned (x ′ ,y ′ ,z ′ )=γ CT ·F ″ CT (x ′ ,y ′ ,z ′ )+γ MRI ·F ″ MRI (x ′ ,y ′ ,z ′ );

[0108] Among them, γ CT and γ MRI are respectively the feature fusion weights of the first modality and the second modality.

[0109] In this implementation, S5 includes:

[0110] S41. Receive the aligned multimodal feature map F aligned (x ′ ,y ′ ,z ′ ), initialize the particle position of the Puller optimization algorithm Particle Speed Individual optimal position The global optimal position G best , as well as inertia weight w, individual learning factor c 1 , group learning factor c 2 , particle positions and velocities are initialized to random distribution;

[0111] S42. Define the optimization objective function J(θ), which integrates the mutual information between modalities, feature extraction accuracy and noise suppression capability:

[0112]

[0113] Among them, θ represents the multidimensional convolution kernel parameters and modal weights, is the loss function, which measures the accuracy of feature extraction, MI is the mutual information between modalities, which measures the correlation between modal features, and σ noise is the local noise intensity in the feature map, λ 1 and λ 2 is the balance coefficient, which is used to adjust the contribution of different objectives to the optimization process;

[0114] S43. Dynamically iterate and adjust the particle swarm through the Puller optimization algorithm to update the speed and position of each particle:

[0115]

[0116] Among them, t is the current iteration number, r 1 and r 2 is a random number in the range [0,1], and Respectively represent the position and velocity of the particle after the current iteration;

[0117] Update individual optimal positions and global optimal positions:

[0118]

[0119] S44. In the iterative process, according to the optimized parameter θ * , dynamically adjust the convolution kernel weights and modality fusion weights, and The modal weight optimization result is used to recalculate the feature map after the optimization of the multi-dimensional convolution kernel weights.

[0120] S45. Re - convolve the aligned multi - modal feature maps using the optimized parameters, and generate the optimized multi - modal feature map F by combining the modal weights. optimized (x ′ ,y ′ ,z ′ ):

[0121]

[0122] In this embodiment, S6 includes:

[0123] S51. Define the structure of the fusion network through the optimized multi - modal feature map F optimized (x ′ ,y ′ ,z ′ ), including multiple convolutional modules, an attention mechanism module, and a feature re - distribution module;

[0124] S52. In the first stage of the fusion network, use the multiple convolutional modules to perform a preliminary fusion operation on the optimized multi - modal feature map, and input the optimized multi - modal feature map F optimized (x ′ ,y ′ ,z ′ ) into the multiple convolutional modules to extract cross - modal joint features and output the preliminary fusion feature map F conv (x ′ ,y ′ ,z ′ ):

[0125] F conv (x ′ ,y ′ ,z ′ ) = Conv fusion (F optimized (x ′ ,y ′ ,z ′ ));

[0126] Among them, Conv fusion represents the multiple convolutional operations in the fusion network, which are used to extract deep - level features of multi - modalities;

[0127] S53. In the second stage of the fusion network, use the attention mechanism module to adjust the weights of the preliminary fusion feature map, dynamically allocate the weights of each modal feature, and define the modal attention weights α CT and α MRI , which respectively represent the feature importance of the first modality and the second modality, and dynamically calculate the modal attention weights through global pooling and the weight calculation formula:

[0128]

[0129] The preliminary fusion feature map is weighted and adjusted using the attention weights to generate an attention-optimized feature map:

[0130] F att (x ′ ,y ′ ,z ′ ) = α CT ·F conv,CT (x ′ ,y ′ ,z ′ ) + α MRI ·F conv,MRI (x ′ ,y ′ ,z ′ );

[0131] S54. In the third stage of the fusion network, the feature reassignment module is used to perform deep fusion on the attention-optimized feature map F att (x ′ ,y ′ ,z ′ ), combining the complementary information between modalities and the spatial distribution characteristics to generate the fused multi-modal feature representation F fused (x ′ ,y ′ ,z ′ ):

[0132] F fused (x ′ ,y ′ ,z ′ ) = Reassign(F att (x ′ ,y ′ ,z ′ ));

[0133] Among them, Reassign represents the feature reassignment operation, which is used to further optimize the complementary characteristics between modalities.

[0134] Example 1:

[0135] On November 12, 2024, a 67-year-old male patient was admitted to a certain tertiary hospital. The patient presented with persistent abdominal pain, weight loss, and decreased appetite. The doctor suspected that he might have gastric cancer with metastasis and decided to perform a combined CT and MRI examination for further diagnosis. The CT image resolution of the patient was 512×512×300, mainly used to display the details of anatomical structures; the MRI image resolution was 256×256×150, used to observe soft tissue lesions and functional characteristics.

[0136] After the acquisition of the imaging data is completed, the radiologist analyzes the patient's images through the medical imaging multi-modal feature extraction system of the present invention. The system first receives the CT and MRI image data of the patient and automatically performs normalization processing to adjust the CT image resolution to 256×256×150 to match the MRI data. Subsequently, the system performs spatial alignment on the data of the two modalities through a geometric correction algorithm based on landmark matching to ensure the consistency of the data between modalities.

[0137] In the multi-dimensional convolution transform feature extraction stage, the system detects that the density in the CT image of the patient's stomach area is abnormally increased, and the density range is about 120–130 HU, which is initially judged as a tumor-related feature; the MRI image shows diffusion restriction in the corresponding area, indicating the possible presence of malignant lesions. By dynamically adjusting the weights of the multi-dimensional convolution kernels, the system further extracts the cross-modal deep features of this area and enhances the complementarity of the CT and MRI features through the modal mutual information maximization algorithm.

[0138] The system performs deep fusion on the optimized multi-modal features, focuses on the lesion area through the attention mechanism, automatically suppresses the feature weights of other normal tissues, and the system generates a three-dimensional fusion image of the patient's stomach tumor, clearly showing the spatial distribution of the lesion and its relationship with the surrounding organs. The imaging analysis report indicates that the tumor size is 3.2 cm×2.8 cm×2.5 cm, is closely related to the gastric wall, and may invade the pancreas and duodenum.

[0139] During the feature extraction and analysis process, the system simultaneously records the calculation results and processing time of each step. For example, the spatial alignment of the CT and MRI images takes about 2.1 seconds, the multi-dimensional convolution feature extraction takes 6.5 seconds, the fusion feature generation takes 3.8 seconds, and the total processing time is 12.4 seconds. The system also lists the key indicators of feature extraction in the final diagnosis report: the modal mutual information score is 0.87, the Dice coefficient is 93.4%, and the overlap rate with high-quality annotated data is relatively high.

[0140] After the analysis is completed, the system automatically generates a diagnosis report and submits the fused three-dimensional image and analysis results to the clinician. The report details the size, location, invasion range of the tumor, and the comprehensive analysis results of the features between modalities suggest that the tumor has a high malignant risk, and it is recommended to further perform endoscopic examination and pathological biopsy to clarify the diagnosis. Based on the multi-modal fusion feature report provided by the system, the doctor quickly formulates a personalized diagnosis and treatment plan, avoiding misdiagnosis or missed diagnosis caused by insufficient single-modal imaging data.

[0141] This example demonstrates the practical application capability of the method of the present invention in complex clinical scenarios. Through dynamic optimization and multimodal feature fusion, the efficiency and accuracy of image analysis are significantly improved, providing doctors with more intuitive and comprehensive diagnostic support. In the analysis of high-noise data and complex lesion areas, the system shows obvious advantages, providing important technical support for the realization of precision medicine.

[0142] The present invention effectively solves the problem of difficulty in aligning and fusing multimodal image features in the prior art by designing a multidimensional adaptive convolution kernel, combining inter-modal mutual information and a dynamic adjustment mechanism of local noise. The multidimensional convolution transformation not only extracts features in the spatial dimension, but also realizes the deep capture of inter-modal complementary information in the modal dimension. The adaptive multidimensional convolution kernel can automatically adjust weights and parameters according to the characteristics of the modal data, thereby significantly improving the utilization rate of modal mutual information in the feature extraction process.

[0143] The present invention introduces the Puller optimization algorithm in the feature optimization process to dynamically adjust the key parameters in the feature extraction process, including the weights of the multidimensional convolution kernel and the weight allocation strategy for inter-modal feature fusion, thereby significantly improving the adaptability of the model in high noise and small sample scenarios. The Puller optimization algorithm dynamically balances global search and local search, and by defining a comprehensive objective function, takes modal mutual information maximization and noise suppression as optimization goals, effectively suppressing noise interference in feature extraction and avoiding model overfitting in small sample scenarios.

[0144] The present invention realizes the deep fusion of multimodal features by constructing a fusion network, including a multi-layer convolution module, an attention mechanism module and a feature reallocation module. Based on the initially extracted optimization features, the fusion network dynamically adjusts the weights of different modal features using the attention mechanism so that the model can focus on the key lesion area, and further optimizes the complementary characteristics between modalities through the feature reallocation module, thereby improving the global representation ability of the fusion features. Different from the existing simple modal feature superposition method, the method of the present invention can significantly reduce the information conflict between modalities and improve the expressiveness of the model when processing complex lesion areas.

[0145] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for extracting multimodal features from medical images based on multidimensional convolution transformation, characterized in that: The steps include: S1. Acquire multimodal medical image data, where the multimodal medical image data includes medical image data of at least two different imaging modalities, and perform spatial alignment, intensity standardization, and denoising on the multimodal medical image data to obtain preprocessed multimodal medical image data; S2. Multidimensional convolution transform performs convolution operation on multimodal medical imaging data in spatial dimension and modality dimension through multidimensional convolution kernel to obtain preliminary feature mapping of each modality; S3. Perform spatial alignment and channel alignment on the preliminary feature maps, align the preliminary feature maps of each modality in the spatial dimension and the channel dimension, and obtain an aligned multimodal feature map; S4. Optimizing the aligned multimodal feature mapping based on the Puller optimization algorithm. The Puller optimization algorithm adaptively optimizes the aligned multimodal feature mapping by global search and local search by dynamically iteratively adjusting feature weights and multidimensional convolution kernel parameters to obtain an optimized multimodal feature mapping. S5. Using the fusion network to deeply fuse the optimized multimodal feature mapping, perform information complementation and weight adaptive allocation on each modality feature mapping, and obtain a fused multimodal feature representation; S6. performing regularization processing on the fused multimodal feature representation to suppress feature noise and redundant information, and obtaining a regularized optimized multimodal feature representation; S7. Output the regularized optimized multimodal feature representation for medical image analysis tasks such as disease classification, lesion segmentation, or lesion detection.

2. The method for extracting multimodal features of medical images based on multidimensional convolution transformation according to claim 1, characterized in that: The S1 includes: S11. Acquire multimodal medical image data, including medical image data of at least two different imaging modalities, namely, first modality medical image data and second modality medical image data, wherein the first modality medical image data includes CT image data D CT (x, y, z), the second modality medical image data includes MRI image data D MRI (x,y,z); S12. Preliminarily aligning the first modality medical image data and the second modality medical image data according to the spatial resolution and voxel size of the first modality medical image data and the second modality medical image data; S13. Accurately aligning the first modality medical image data and the second modality medical image data by a calibration point matching method to generate aligned first modality medical image data and aligned second modality medical image data; S14. Sampling the aligned first modality medical image data and the second modality medical image data, and uniformly adjusting them to a target resolution, where the target resolution is pre-set according to the requirements of an actual application scenario; S15. performing normalization processing on the first modality medical image data and the second modality medical image data after adjusting the resolution so that the pixel intensity values ​​thereof are distributed within the range of [-1, 1]; S16. Output the normalized first modality medical image data and second modality medical imaging data As multimodal medical imaging data.

3. The method for extracting multimodal features of medical images based on multidimensional convolution transformation according to claim 1, characterized in that: The S3 includes: S21. performing a block operation on the normalized first modality medical image data and the second modality medical image data, dividing the input data into a plurality of overlapping local sub-regions; S22. Define multi-dimensional adaptive convolution kernel W m,n (k x ,k y ,k z ,α,β), where k x ,k y ,k z is the size of the convolution kernel in the spatial dimension, α and β represent the dynamic adjustment parameters of the multidimensional adaptive convolution kernel for the mutual information between modalities and the local noise intensity, respectively, m represents the input channel index, corresponding to the number of input feature channels of each modality, and n represents the output channel index, which represents the type of features extracted by convolution; S23. Update the convolution kernel parameters α and β in real time according to the mutual information strength and noise distribution between the modalities in the local sub-region, adjust the convolution kernel weights, and optimize the cross-modal feature extraction capability: Among them, MI represents the mutual information between modes, σ noise represents the local noise intensity, Vol(x ′ ,y ′ ,z ′ ) represents the volume of the local sub-region; S24. Perform a convolution operation on the first modality medical image data and the second modality medical image data using the updated adaptive convolution kernel to calculate a modality-specific preliminary feature map, where the preliminary feature map of the first modality is: Among them, M represents the number of input feature channels, K x ,K y ,K z Represents the receptive field range of the convolution kernel in each dimension; Perform convolution operation on the second modality medical image data to obtain the preliminary feature map of the second modality:

4. The method for extracting multimodal features of medical images based on multidimensional convolution transformation according to claim 1, characterized in that: The S4 includes: S31. Align the first modal preliminary feature map and the second modal preliminary feature map in the spatial dimension, and calculate the transformation matrix T of the two modal feature maps in the spatial dimension. spatial The transformation matrix is ​​calculated based on the geometric correspondence and spatial deviation of the two modal data. According to the transformation matrix T spatial Perform spatial transformation on the first modality preliminary feature map and the second modality preliminary feature map to generate the spatially aligned first modality feature map F ′ CT (x ′ ,y ′ ,z ′ ) and the second modality feature map F ′ MRI (x ′ ,y ′ ,z ′ ); S33. Align the spatially aligned first modal feature map and the second modal feature map in the channel dimension, and define a modality-specific channel weighting matrix W channel , where W channel,CT and W channel,MRI They represent the channel weights of the first mode and the second mode respectively. The size of the weight matrix is ​​consistent with the number of input feature channels. The feature channels of each mode are redistributed through weighted operations: F ″ CT (x ′ ,y ′ ,z ′ )=W channel,CT ·F ′ CT (x ′ ,y ′ ,z ′ ); F ″ MRI (x ′ ,y ′ ,z ′ )=W channel,MRI ·F ′ MRI (x ′ ,y ′ ,z ′ ); Among them, F ″ CT (x ′ ,y ′ ,z ′ ) and F ″ MRI (x ′ ,y ′ ,z ′ ) respectively represent the first modality feature map and the second modality feature map after channel alignment; S34. Combining the first modal feature map and the second modal feature map after spatial alignment and channel alignment to generate a final aligned multimodal feature map F aligned (x ′ ,y ′ ,z ′ ): F aligned (x ′ ,y ′ ,z ′ )=γ CT ·F ″ CT (x ′ ,y ′ ,z ′ )+γ MRI ·F ″ MRI (x ′ ,y ′ ,z ′ ): Among them, γ CT and γ MRI are the feature fusion weights of the first modality and the second modality respectively.

5. The method for extracting multimodal features of medical images based on multidimensional convolution transformation according to claim 1, characterized in that: The S5 includes: S41. Receive the aligned multimodal feature map F aligned (x ′ ,y ′ ,z ′ ), initialize the particle position of the Puller optimization algorithm Particle Speed Individual optimal position The global optimal position G best , as well as the inertia weight w, individual learning factor c1, group learning factor c2, the particle position and velocity are initialized to random distribution; S42. Define the optimization objective function J(θ), which integrates the mutual information between modalities, feature extraction accuracy and noise suppression capability: Among them, θ represents the multidimensional convolution kernel parameters and modal weights, is the loss function, which measures the accuracy of feature extraction, MI is the mutual information between modalities, which measures the correlation between modal features, and σ noise is the local noise intensity in the feature map, λ1 and λ2 are balance coefficients used to adjust the contribution of different objectives to the optimization process; S43. Dynamically iterate and adjust the particle swarm through the Puller optimization algorithm to update the speed and position of each particle: Where t is the current iteration number, r1 and r2 are random numbers in the range [0,1]. and Respectively represent the position and velocity of the particle after the current iteration; Update individual optimal positions and global optimal positions: S44. In the iterative process, according to the optimized parameter θ * , dynamically adjust the convolution kernel weights and modality fusion weights, and The modal weight optimization result is used to recalculate the feature map after the optimization of the multi-dimensional convolution kernel weights. S45. Use the optimized parameters to reconvolve the aligned multimodal feature map and generate the optimized multimodal feature map F by combining the modal weights optimized (x ′ ,y ′ ,z ′ ):

6. The method for extracting multimodal features of medical images based on multidimensional convolution transformation according to claim 1, characterized in that: The S6 includes: S51. Through the optimized multimodal feature map F optimized (x ′ ,y ′ ,z ′ ) Define the structure of the fusion network, including multi-layer convolution modules, attention mechanism modules and feature redistribution modules; S52. In the first stage of the fusion network, a multi-layer convolution module is used to perform a preliminary fusion operation on the optimized multimodal feature map F optimized (x ′ ,y ′ ,z ′ ) is input into the multi-layer convolution module to extract cross-modal joint features and output the preliminary fusion feature map F conv (x ′ ,y ′ ,z ′ ): F conv (x ′ ,y ′ ,z ′ )=Conv fusion (F optimized (x ′ ,y ′ ,z ′ )); Among them, Conv fusion Represents the multi-layer convolution operation in the fusion network, which is used to extract multi-modal deep features; S53. In the second stage of the fusion network, the attention mechanism module is used to adjust the weight of the preliminary fusion feature map, dynamically assign the weight of each modal feature, and define the modal attention weight α CT and α MRI , respectively represent the feature importance of the first modality and the second modality. The modality attention weight is dynamically calculated through global pooling and weight calculation formula: Use the attention weights to perform weighted adjustment on the preliminary fused feature map to generate an attention optimized feature map: F att (x ′ ,y ′ ,z ′ )=α CT ·F conv,CT (x ′ ,y ′ ,z ′ )+α MRI ·F conv,MRI (x ′ ,y ′ ,z ′ ): S54. In the third stage of the fusion network, the feature redistribution module is used to optimize the attention feature map F att (x ′ ,y ′ ,z ′ ) is deeply fused, combining the complementary information between modalities and spatial distribution characteristics to generate a fused multimodal feature representation F fused (x ′ ,y ′ ,z ′ ): F fused (x ′ ,y ′ ,z ′ )=Reassign(F att (x ′ ,y ′ ,z ′ )); Among them, Reassign represents the feature redistribution operation, which is used to further optimize the complementary characteristics between modalities.

Citation Information

Cited By

  • Head focus multi-level automatic positioning method based on multi-modal image cooperation

    CN120931699A

  • Multi-modal data fusion pituitary adenoma invasion behavior characteristic modeling method and system

    CN121075674A