Medical image fusion method and device based on pre-trained large model

By using pre-trained large models in medical image fusion, the problem of difficult CT and MRI image registration fusion is solved, and high-precision image fusion is achieved.

CN120164066APending Publication Date: 2025-06-17LONGWOOD VALLEY MEDICAL TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510233741.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Registration and fusion of CT images and MRI images is difficult to achieve high accuracy.

Method used

Using a medical image fusion method based on a pre-trained large model, the fusion image to be fused is obtained by obtaining the MRI and CT images to be fused and inputting them into the image fusion model composed of a first coding structure, a second coding structure, a pre-trained large model structure, a fusion module and a decoded structure to obtain the fusion image.

Benefits of technology

High-precision registration and fusion of CT and MRI images are realized, and the accuracy and efficiency of image fusion are improved through pre-trained large models' semantic understanding ability and multimodal feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164066A_ABST
    Figure CN120164066A_ABST
Patent Text Reader

Abstract

The invention provides a medical image fusion method and device based on a pre-training large model. The method comprises the steps of obtaining a to-be-fused MRI image and a to-be-fused CT image; inputting the to-be-fused CT image and the to-be-fused MRI image into a trained image fusion model to obtain a fused image; the image fusion model is composed of a first coding structure, a second coding structure, a pre-training large model structure, a fusion module and a decoding structure. In the application, the pre-trained large model is embedded into the image fusion module, and high-precision registration and fusion of the CT and MRI images are realized through the semantic understanding ability and multi-modal feature extraction of the pre-trained large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical image recognition technology. Specifically, it relates to a medical image fusion method and device based on a pre-trained large model. Background Art

[0002] CT images are good at showing bones and calcified tissues, and MRI images are good at showing soft tissues and lesion areas. However, the fusion of the two requires high-precision multi-modal image registration, which is currently very difficult. Summary of the Invention

[0003] The problem solved by this application is that it is difficult to achieve high precision in the registration and fusion of CT images and MRI images.

[0004] To solve the above problems, the first aspect of this application provides a medical image fusion method based on a pre-trained large model, including:

[0005] Obtain the MRI image to be fused and the CT image to be fused;

[0006] Input the CT image to be fused and the MRI image to be fused into a trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module, and a decoding structure.

[0007] The second aspect of this application provides a medical image fusion device based on a pre-trained large model, which includes:

[0008] An image acquisition module, which is used to obtain the MRI image to be fused and the CT image to be fused;

[0009] An image fusion module, which is used to input the CT image to be fused and the MRI image to be fused into a trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module, and a decoding structure.

[0010] The third aspect of this application provides an electronic device, which includes: a memory and a processor;

[0011] The memory is used to store programs;

[0012] The processor, coupled to the memory, is used to execute the program for:

[0013] Obtain the MRI image to be fused and the CT image to be fused;

[0014] Input the CT image to be fused and the MRI image to be fused into the trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module, and a decoding structure.

[0015] The fourth aspect of this application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the above-mentioned medical image fusion method based on a pre-trained large model.

[0016] In this application, a pre-trained large model is embedded in the image fusion module, and through the semantic understanding ability and multi-modal feature extraction of the pre-trained large model, high-precision registration and fusion of CT and MRI images are realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of a medical image fusion method based on a pre-trained large model according to an embodiment of this application;

[0018] Figure 2 It is a model architecture diagram of a medical image fusion method based on a pre-trained large model according to an embodiment of this application;

[0019] Figure 3 It is an architecture diagram of the first encoding structure of a medical image fusion method based on a pre-trained large model according to an embodiment of this application;

[0020] Figure 4 It is an architecture diagram of the decoding structure of a medical image fusion method based on a pre-trained large model according to an embodiment of this application;

[0021] Figure 5 It is an architecture diagram of the fusion module of a medical image fusion method based on a pre-trained large model according to an embodiment of this application;

[0022] Figure 6 It is an architecture diagram of the attention module of a medical image fusion method based on a pre-trained large model according to an embodiment of this application;

[0023] Figure 7 It is a structural block diagram of a medical image fusion device based on a pre-trained large model according to an embodiment of this application;

[0024] Figure 8 It is a structural block diagram of an electronic device according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the above objects, features, and advantages of the present application more apparent and understandable, the following provides a detailed description of the specific embodiments of the present application with reference to the accompanying drawings. Although the accompanying drawings show exemplary embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0026] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present application should have the ordinary meaning understood by those skilled in the art to which the present application belongs.

[0027] In view of the above problems, the present application provides a new medical image fusion solution based on a pre-trained large model, which can embed the pre-trained large model into the image fusion module to eliminate the problem that it is difficult to achieve high-precision registration and fusion of CT images and MRI images.

[0028] The embodiments of the present application provide a medical image fusion method based on a pre-trained large model. The specific solution of this method is Figures 1 - 6 as shown. This method can be executed by a medical image fusion device based on a pre-trained large model, and this medical image fusion device based on a pre-trained large model can be integrated in electronic devices such as computers, servers, computers, server clusters, and data centers. Combining Figure 1 、 Figure 2 as shown, it is a flowchart of a medical image fusion method based on a pre-trained large model according to an embodiment of the present application; wherein, the medical image fusion method based on a pre-trained large model includes:

[0029] S101, obtain the MRI image to be fused and the CT image to be fused;

[0030] In the present application, due to the different imaging principles of CT and MRI images, there are differences in their gray levels, resolutions, and spatial positions, and direct fusion will result in information loss or distortion.

[0031] In the present application, CT images are good at showing bones and calcified tissues, and MRI images are good at showing soft tissues and lesion areas.

[0032] S102, input the CT image to be fused and the MRI image to be fused into the trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module, and a decoding structure.

[0033] In the present application, the pre-trained large model is embedded in the image fusion module, and through the semantic understanding ability and multi-modal feature extraction of the pre-trained large model, high-precision registration and fusion of CT and MRI images are achieved.

[0034] In this application, by introducing a pre-trained large model structure and a multi-encoding-decoding architecture, the following technical effects are achieved:

[0035] High-precision feature extraction and fusion: The first encoding structure and the second encoding structure respectively extract the features of CT and MRI images. The pre-trained large model structure further enhances the feature expression ability, and the fusion module realizes multi-modal feature complementarity.

[0036] In this application, the pre-trained large model structure (such as DeepSeek) has strong generalization ability and can adapt to the image fusion requirements of different parts and tasks.

[0037] In this application, through the transfer learning ability of the pre-trained large model, the training time and the consumption of computing resources are reduced.

[0038] In this application, the decoding structure reconstructs the fused features into a high-resolution and high-fidelity fused image, while retaining the bone density information of CT and the soft tissue contrast of MRI.

[0039] In one implementation, in combination with Figure 2 as shown in, in S102, inputting the CT image to be fused and the MRI image to be fused into the trained image fusion model to obtain a fused image, including:

[0040] Inputting the CT image to be fused into the first encoding structure to obtain a first encoding;

[0041] Inputting the MRI image to be fused into the second encoding structure to obtain a second encoding;

[0042] Inputting the first encoding and the second encoding into the pre-trained large model structure to obtain a pre-trained encoding;

[0043] Inputting the first encoding, the second encoding and the pre-trained encoding into the fusion module to obtain a fused feature map;

[0044] Inputting the fused feature map into the decoding structure to obtain a fused image.

[0045] In this application, the first encoding structure and the second encoding structure respectively process CT and MRI images, ensuring the independent extraction of the features of the two modalities, and ensuring the alignment of the features of the two modalities in the encoding stage through independent extraction.

[0046] In this application, a pre-trained large model (such as DeepSeek) is introduced as a feature enhancement module to improve the generalization ability and feature expression ability of the model.

[0047] In this application, the fusion module dynamically fuses CT and MRI features through an attention mechanism or adaptive weight allocation, achieving multi-modal information complementarity and dynamic fusion of multi-modal features, and solving the problem of fixed fusion weights in traditional methods.

[0048] In this application, the entire model (encoding structure, pre-trained large model, fusion module, decoding structure) can be trained end-to-end to optimize the fusion effect.

[0049] In one implementation, an intermediate module is further provided between the fusion module and the decoding structure. After the fusion feature map of the fusion module is input into the intermediate module, an adjusted fusion feature map is obtained; then the adjusted fusion feature map is input into the decoding structure.

[0050] In one implementation, the structure and processing process of the intermediate module include:

[0051] The fusion feature map is divided into blocks to obtain independent blocks;

[0052] For each independent block, first neighborhood blocks and second neighborhood blocks with different spacings are obtained;

[0053] Based on the independent block and the first neighborhood blocks, a first feature block is generated;

[0054] Based on the independent block and the second neighborhood blocks, a second feature block is generated;

[0055] The first feature block and the second feature block are subjected to feature compression to obtain a compressed block;

[0056] All independent blocks are traversed, and an adjusted fusion feature map is generated based on the obtained compressed blocks.

[0057] In this application, dividing the fusion feature map into blocks means dividing the fusion feature map into corresponding feature map blocks through a checkerboard; among them, the feature map blocks can be at the pixel level (that is, each pixel is a feature map block), or at other levels, and the specific division depends on the actual processing situation.

[0058] In this application, the feature map is divided into blocks of the same size using a sliding window or a fixed step size.

[0059] It should be noted here that if the fusion feature map is a two-dimensional feature map, it is directly divided into a checkerboard, and each grid is a feature map block; if the fusion feature map is a three-dimensional feature map, a plane is selected for checkerboard division, and each grid is a strip-shaped grid with a lot of depth (the depth is the depth of the three-dimensional feature map), and this strip-shaped grid is a feature map block.

[0060] Preferably, in the present application, each feature patch is 100 - 1000 pixels, so as to perform more feature calculations between local regions on the basis of ensuring the generation accuracy and reducing the computational amount.

[0061] In the present application, one feature patch is selected as the independent patch, and the adjacent feature patches above, below, to the left, and to the right of the independent patch are the first neighborhood patches; the feature patches separated by one grid above, below, to the left, and to the right of the independent patch are the second neighborhood patches. The distances between the first neighborhood patches and the second neighborhood patches and the independent patch are different.

[0062] In the present application, neighborhood information is extracted for each independent patch to capture the local structure.

[0063] In the present application, the first feature patch is generated by using the independent patch and its first neighborhood patches to generate a local feature representation. Specifically, it can be: the independent patch and the first neighborhood patches are processed through a convolutional layer and an attention layer to obtain the first feature patch.

[0064] In the present application, the specific structure and specific parameters of the convolutional layer and the attention layer can be obtained according to the training data or determined according to the actual situation.

[0065] It should be noted that in the present application, there are four first neighborhood patches and multiple first feature patches.

[0066] In the present application, the independent patch and the first neighborhood patches are processed through a convolutional layer and an attention layer to obtain the first feature patch. The specific process is: the independent patch and the four neighborhood patches are concatenated together to form a multi-channel input, and the convolutional layer is used to extract features from the concatenated patch; the self-attention mechanism or the channel attention mechanism is used to enhance the important features, calculate the attention weights, and weight the output of the convolutional layer to enhance the important features; the output of the attention layer is split into multiple feature patches, and each feature patch corresponds to the processing result of the independent patch and at least one neighborhood patch.

[0067] In the present application, the second feature patch is generated by using the independent patch and its second neighborhood patches to generate a more extensive local feature representation. Its specific generation process is the same as that of the first feature patch, except that the parameters of the convolutional layer and the attention layer are different.

[0068] In the present application, the generated feature patches are compressed into a more compact representation to reduce the computational amount and retain the key information. Pooling operations (such as max pooling or average pooling) or fully connected layers are used for feature compression.

[0069] In this way, through compression, multiple first feature patches and second feature patches are compressed into one compressed patch, which corresponds to the independent patch in size and position and is used to replace the independent patch. All feature patches are replaced by compressed patches to obtain the adjusted fused feature map.

[0070] In this application, by means of traversal, each feature map block of the fused feature map is traversed to obtain the corresponding compressed block.

[0071] In this application, for the feature map blocks / independent blocks near the edge, their first neighborhood blocks and second neighborhood blocks are incomplete. At this time, the first neighborhood blocks and second neighborhood blocks in the relative positions are copied for completion. For example, if the first neighborhood block above the independent block does not exist, the first neighborhood block below is copied and used as the block above.

[0072] In this application, through completion, the processing accuracy of the edge-adjacent feature map blocks is greatly improved.

[0073] In this application, the intermediate module captures the similarity relationship between local regions, thereby enhancing the feature representation.

[0074] In one implementation, combined Figure 3 as shown, the input of the CT image to be fused into the first encoding structure to obtain the first encoding includes:

[0075] Multiply the CT image to be fused by the downsampling coefficient to obtain a coefficient feature map;

[0076] Perform convolution processing on the coefficient feature map with two groups of convolutions in sequence to obtain a convolution feature map;

[0077] Add the convolution feature map and the coefficient feature map to obtain the first encoding.

[0078] In this application, adding the feature maps means adding the feature maps element by element to fuse the information.

[0079] In this application, when performing convolution processing on the coefficient feature map with two groups of convolutions in sequence, each group of convolutions is a 1×1 convolution, a 3×3 convolution, and a 1×1 convolution connected in sequence.

[0080] In this application, the downsampling coefficient is the weight coefficient used for scaling the size or resolution of the original image; it may also be a specific scaling factor obtained according to other network structures or multi-scale analysis (such as pyramids, hierarchical convolutions, etc.).

[0081] In this application, the downsampling coefficient has the same or similar dimensions as the image to be fused to ensure that the multiplication operation can be completed pixel by pixel (or voxel by voxel).

[0082] In this application, local weighting or information screening of the original image is performed through the downsampling coefficient, so that the subsequent convolutional network focuses on important regions or suppresses irrelevant background noise.

[0083] In this application, the resolution of the input image is reduced through the downsampling coefficient, the amount of calculation is reduced, and the feature extraction efficiency is improved.

[0084] In this application, two groups of convolutional operations extract features at different levels, enhancing the feature expression ability.

[0085] In this application, by adding the convolutional feature map and the coefficient feature map, high-frequency information is retained and information loss is avoided.

[0086] In this application, through downsampling coefficient optimization, double convolutional structure and feature fusion mechanism, problems such as computational efficiency, feature expression ability and information loss in CT image coding are solved, and efficient and high-precision feature extraction is achieved.

[0087] In this application, the first coding structure and the second coding structure have the same structure but different parameters.

[0088] In this application, using the same network structure facilitates maintenance and unified management of the model architecture; at the same time, through independent parameter training, it can more flexibly meet the requirements of multi-modal fusion or multi-stage fusion while ensuring consistency and modularity.

[0089] In one implementation, combined Figure 4 As shown, inputting the fusion feature map into the decoding structure to obtain the fused image includes:

[0090] Multiplying the fusion feature map by the upsampling coefficient to obtain the upsampled feature map;

[0091] Performing convolutional processing on the upsampled feature map with two groups of convolutions in sequence to obtain the convolutional feature map;

[0092] Adding the convolutional feature map and the coefficient feature map to obtain the fused image.

[0093] In this application, when performing convolutional processing on the coefficient feature map with two groups of convolutions in sequence, each group of convolutions is a 1×1 convolution, 3×3 convolution, and 1×1 convolution connected in sequence.

[0094] In this application, the decoding structure is similar to the encoding structure. This symmetric design helps the feature maps to be spatially aligned and reduces information loss; it can ensure the consistency and integrity of information during the encoding and decoding processes.

[0095] In this application, the symmetric design simplifies the network structure, reduces the training difficulty, and improves the convergence speed and stability of the model.

[0096] In this application, the upsampling coefficient can be regarded as the weight matrix (or interpolation kernel weight) used when performing spatial scale expansion (such as magnifying in the width and height directions) on the feature map, and can be represented as a group of parameterized convolutional kernels or fixed interpolation coefficients in actual implementation.

[0097] In this application, through element-wise multiplication or a calculation process with transposed convolution, the fused feature map is enlarged to the required resolution to obtain an upsampled feature map.

[0098] In this application, while maintaining high-level semantic information, the decoding structure gradually restores spatial details and fuses them with high-semantic features, thereby achieving accurate reconstruction or output of the image.

[0099] In one implementation, combined with Figure 5 As shown, inputting the first encoding, second encoding, and pre-trained encoding into the fusion module to obtain a fused feature map includes:

[0100] Adding the first encoding, second encoding, and pre-trained encoding to obtain an added graph;

[0101] Performing 1×1 convolution, 1×3 convolution, and 3×1 convolution on the first encoding in sequence to obtain a first graph;

[0102] Performing 1×1 convolution, 3×1 convolution, and 1×3 convolution on the second encoding in sequence to obtain a second graph;

[0103] Performing 1×1 convolution, 3×1 convolution, and 1×3 convolution on the pre-trained encoding in sequence to obtain a third graph;

[0104] Performing average pooling on the added graph to obtain a pooled graph;

[0105] Concatenating the first graph, second graph, third graph, and pooled graph and then performing 1×1 convolution to obtain a convolved graph;

[0106] Adding the convolved graph and the added graph to obtain a fused feature map.

[0107] In this application, by directly superimposing the features from three sources through element-wise addition, a comprehensive representation is obtained, which helps subsequent branches quickly access and utilize the feature information of all encoding branches.

[0108] In this application, 1×1 convolution is mainly used for channel-wise transformation or channel compression / dilation, which can adjust the number of output channels or fuse information between channels without changing the spatial resolution.

[0109] In this application, 1×3 convolution and 3×1 convolution decompose the convolution kernel in the horizontal direction (1×3) and vertical direction (3×1). Compared with directly using 3×3 convolution, this decomposition can reduce the number of parameters, improve the operation efficiency, and also bring a more flexible feature extraction effect.

[0110] In this application, the process of the second encoding is similar to that of the first encoding, but the order of the convolutional kernels is different. This order difference can diversely extract spatial features in different direction orders, enhancing the network's ability to capture horizontal / vertical edges, textures, or shapes.

[0111] In this application, the addition graph serves as the input to the fourth branch here. The low-resolution or global information obtained through pooling the addition graph provides supplementation for subsequent fusion, avoiding focusing only on local details while lacking overall background or statistical information.

[0112] In this application, the outputs of the aforementioned four branches (the first graph, the second graph, the third graph, and the pooled graph) are concatenated along the channel dimension; the concatenated feature map will contain diverse feature channels such as different-direction convolutions, different coding sources, and different resolutions / global information.

[0113] In this application, after concatenation, the number of channels increases significantly. Channel compression or fusion is performed through 1×1 convolutions to integrate multi-channel information into a more compact and usable feature representation.

[0114] In this application, the convolutional graph and the addition graph are added together. Adopting the idea of residual connection, the original addition feature (the addition graph) is retained, avoiding possible feature loss or excessive deformation after deep convolution.

[0115] In this application, residual connection enables the network to be more easily optimized during training and avoids gradient vanishing or explosion. And it can perform information fusion: containing both the fusion results of multi-branch convolutions and the native information of the directly added graph.

[0116] In this application, by decomposing convolutional kernels such as 1×3 and 3×1, the flexibility of the network to extract horizontal / vertical details is enhanced, reducing the number of parameters while still being able to capture the receptive field of a conventional 3×3 convolution.

[0117] In this application, three types of features, namely "the first encoding", "the second encoding", and "the pre-trained encoding", are processed simultaneously to fuse multiple feature sources; then an "average pooling" branch is added to enhance global information.

[0118] In this application, the addition operation is introduced into the structure multiple times, which not only ensures the effective transmission of features and gradients but also helps the network converge more easily and remain stable in the deep structure.

[0119] In one implementation, an attention module is further provided after the first encoding structure, the second encoding structure, and the pre-trained large model structure to capture the attention features of the first encoding, the second encoding, and the pre-trained encoding.

[0120] Among them, the attention module is adjacent to the first encoding structure, the second encoding structure, and the pre-trained large model structure, and the outputs of the first encoding structure, the second encoding structure, and the pre-trained large model structure are the inputs of the attention module.

[0121] In one implementation, in combination with Figure 6 as shown, the processing process of the attention module includes:

[0122] Split the fused feature map into a first feature map and a second feature map; the local parameters of the first feature map and the second feature map are different;

[0123] Perform 1×1 convolution and sampling processing on the first feature map for different branches respectively to obtain a first sampled feature map and a second sampled feature map;

[0124] Perform 1×1 convolution on the second feature map to obtain a convolutional feature map;

[0125] After feature extraction on the second sampled feature map, merge it with the downsampled convolutional feature map to obtain a merged feature map;

[0126] After merging the merged feature map and the first sampled feature map, perform upsampling to obtain an adjusted fused feature map.

[0127] In this application, split according to the channel dimension or branch dimension into a first feature map and a second feature map

[0128] In this application, "different local parameters" means that the convolution kernels or processing operations applied to these two feature maps are not exactly the same to obtain more diverse feature expressions.

[0129] In this application, splitting helps to perform independent attention analysis and feature extraction in two paths respectively, improving the network's discrimination ability for key regions or channels.

[0130] In this application, perform 1×1 convolution and sampling processing on the first feature map for different branches respectively to obtain a first sampled feature map and a second sampled feature map; perform 1×1 convolution and sampling processing on the first feature map for the first branch to obtain the first sampled feature map; perform 1×1 convolution and sampling processing on the first feature map for the second branch to obtain the second sampled feature map. The structures of the first branch and the second branch are the same, but the parameters are different.

[0131] In this application, set multiple parallel branches for the first feature map, so that the original channel information can be decomposed and reorganized.

[0132] In this application, relative to the first feature map, the second feature map may be different in the number of channels or spatial scale.

[0133] In this application, a 1×1 convolution is used to perform channel integration or transformation on the second feature map to generate a new "convolution feature map". Compared with the multi-branch processing of the first feature map, only one branch is processed here for subsequent merging.

[0134] In this application, local convolution or attention operation is performed on the sampled second feature map again to strengthen the key regions therein. Preferably, a self-attention or channel-attention module can be inserted at this step (to replace feature extraction) to mine the most important feature channels or spatial positions.

[0135] In this application, the "convolution feature map" obtained in the third step has a relatively large dimension and needs to be downsampled to align with the sampled second feature map, and then merged (concatenated or added) in the spatial or channel dimension.

[0136] In this application, this merging fuses the focus points of the two-way features, and the complementary extracted features can be superimposed after being aligned at the same scale.

[0137] In this application, the "merged feature map" and the "sampled first feature map" are merged again (element-wise addition or concatenation) to complete the aggregation of multi-branch information.

[0138] In this application, the sampled first feature map may retain more original resolution or detailed features, while the merged feature map stores the attention information from the deep fusion of the second feature map and the sampled second feature map. Combining the two can obtain a more complete and refined feature representation.

[0139] In this application, finally, an upsampling operation is performed to restore the fused feature map to the required spatial resolution (or align with the previous stage).

[0140] In this application, the first and second feature maps respectively undergo convolution, sampling, and attention enhancement in different branches to extract detailed and global information from multiple perspectives.

[0141] In this application, multiple merging operations enable the network to better fuse attention features of different scales and modalities, and finally restore an appropriate spatial resolution through upsampling to ensure that the output takes into account both high semantics and details.

[0142] The embodiment of this application provides a medical image fusion device based on a pre-trained large model, which is used to execute a medical image fusion method based on a pre-trained large model described above. The following provides a detailed description of the medical image fusion device based on a pre-trained large model.

[0143] As Figure 7As shown, the medical image fusion device based on a pre-trained large model includes:

[0144] An image acquisition module 101 for acquiring the MRI image to be fused and the CT image to be fused;

[0145] An image fusion module 102 for inputting the CT image to be fused and the MRI image to be fused into a trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module, and a decoding structure.

[0146] In one implementation, the image fusion module 102 is further configured to:

[0147] Input the CT image to be fused into the first encoding structure to obtain a first encoding; input the MRI image to be fused into the second encoding structure to obtain a second encoding; input the first encoding and the second encoding into the pre-trained large model structure to obtain a pre-trained encoding; input the first encoding, the second encoding, and the pre-trained encoding into the fusion module to obtain a fused feature map; input the fused feature map into the decoding structure to obtain a fused image.

[0148] In one implementation, the image fusion module 102 is further configured to:

[0149] Multiply the CT image to be fused by a downsampling coefficient to obtain a coefficient feature map; perform convolution processing on the coefficient feature map with two groups of convolutions in sequence to obtain a convolution feature map; add the convolution feature map and the coefficient feature map to obtain a first encoding.

[0150] In one implementation, the image fusion module 102 is further configured to:

[0151] Multiply the fused feature map by an upsampling coefficient to obtain an upsampled feature map; perform convolution processing on the upsampled feature map with two groups of convolutions in sequence to obtain a convolution feature map; add the convolution feature map and the coefficient feature map to obtain a fused image.

[0152] In one implementation, the image fusion module 102 is further configured to:

[0153] Add the first encoding, the second encoding, and the pre-trained encoding to obtain an added graph; perform 1×1 convolution, 1×3 convolution, and 3×1 convolution on the first encoding in sequence to obtain a first graph; perform 1×1 convolution, 3×1 convolution, and 1×3 convolution on the second encoding in sequence to obtain a second graph; perform 1×1 convolution, 3×1 convolution, and 1×3 convolution on the pre-trained encoding in sequence to obtain a third graph; perform average pooling on the added graph to obtain a pooled graph; splice the first graph, the second graph, the third graph, and the pooled graph and then perform 1×1 convolution to obtain a convolution graph; add the convolution graph and the added graph to obtain a fused feature map.

[0154] In one implementation, after the first encoding structure, the second encoding structure, and the pre-trained large model structure, an attention module is further provided to capture the attention features of the first encoding, the second encoding, and the pre-trained encoding.

[0155] In one implementation, the image fusion module 102 is further configured to:

[0156] Split the fused feature map into a first feature map and a second feature map; the local parameters of the first feature map and the second feature map are different; perform 1×1 convolution and sampling processing on the first feature map in different branches to obtain a first sampled feature map and a second sampled feature map respectively; perform 1×1 convolution on the second feature map to obtain a convolutional feature map; after feature extraction on the second sampled feature map, merge it with the downsampled convolutional feature map to obtain a merged feature map; after merging the merged feature map and the first sampled feature map, perform upsampling to obtain an adjusted fused feature map.

[0157] A medical image fusion device based on a pre-trained large model provided in the above embodiments of the present application and a medical image fusion method based on a pre-trained large model provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0158] The internal functions and structures of a medical image fusion device based on a pre-trained large model are described above. As Figure 8 shown, in practice, the medical image fusion device based on a pre-trained large model can be implemented as an electronic device, including: a memory 301 and a processor 303.

[0159] The memory 301 can be configured to store programs.

[0160] In addition, the memory 301 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0161] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0162] The processor 303 is coupled to the memory 301 and is configured to execute the programs in the memory 301 for:

[0163] Obtain the MRI image to be fused and the CT image to be fused;

[0164] Input the CT image to be fused and the MRI image to be fused into the trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module, and a decoding structure.

[0165] In one embodiment, the processor 303 is further configured to:

[0166] Input the CT image to be fused into the first encoding structure to obtain a first encoding; input the MRI image to be fused into the second encoding structure to obtain a second encoding; input the first encoding and the second encoding into the pre-trained large model structure to obtain a pre-trained encoding; input the first encoding, the second encoding, and the pre-trained encoding into the fusion module to obtain a fused feature map; input the fused feature map into the decoding structure to obtain a fused image.

[0167] In one embodiment, the processor 303 is further configured to:

[0168] Multiply the CT image to be fused by the downsampling coefficient to obtain a coefficient feature map; perform convolutional processing on the coefficient feature map with two sets of convolutions in sequence to obtain a convolutional feature map; add the convolutional feature map and the coefficient feature map to obtain a first encoding.

[0169] In one embodiment, the processor 303 is further configured to:

[0170] Multiply the fused feature map by the upsampling coefficient to obtain an upsampled feature map; perform convolutional processing on the upsampled feature map with two sets of convolutions in sequence to obtain a convolutional feature map; add the convolutional feature map and the coefficient feature map to obtain a fused image.

[0171] In one embodiment, the processor 303 is further configured to:

[0172] Add the first encoding, the second encoding, and the pre-trained encoding to obtain an added graph; perform 1×1 convolution, 1×3 convolution, and 3×1 convolution on the first encoding in sequence to obtain a first graph; perform 1×1 convolution, 3×1 convolution, and 1×3 convolution on the second encoding in sequence to obtain a second graph; perform 1×1 convolution, 3×1 convolution, and 1×3 convolution on the pre-trained encoding in sequence to obtain a third graph; perform average pooling on the added graph to obtain a pooled graph; splice the first graph, the second graph, the third graph, and the pooled graph and then perform 1×1 convolution to obtain a convolutional graph; add the convolutional graph and the added graph to obtain a fused feature map.

[0173] In one embodiment, an attention module is further provided after the first encoding structure, the second encoding structure, and the pre-trained large model structure to capture the attention features of the first encoding, the second encoding, and the pre-trained encoding.

[0174] In one embodiment, the processor 303 is further configured to:

[0175] Split the fused feature map into a first feature map and a second feature map; the local parameters of the first feature map and the second feature map are different; perform 1×1 convolution and sampling processing on different branches of the first feature map respectively to obtain a sampled first feature map and a sampled second feature map; perform 1×1 convolution on the second feature map to obtain a convolved feature map; after feature extraction on the sampled second feature map, merge it with the downsampled convolved feature map to obtain a merged feature map; after merging the merged feature map and the sampled first feature map, perform upsampling to obtain an adjusted fused feature map.

[0176] In this application, Figure 8 only some components are schematically shown, which does not mean that the electronic device only includes Figure 8 the components shown.

[0177] The electronic device provided in this embodiment and a medical image fusion method based on a pre-trained large model provided in an embodiment of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0178] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0179] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0180] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in the block or blocks.

[0182] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0183] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (Flash RAM). Memory is an example of computer-readable media.

[0184] This application also provides a computer-readable storage medium corresponding to a medical image fusion method based on a pre-trained large model provided in the foregoing embodiment. A computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it will execute a medical image fusion method based on a pre-trained large model provided in any of the foregoing embodiments.

[0185] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0186] The computer-readable storage medium provided by the above embodiments of the present application and a medical image fusion method based on a pre-trained large model provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0187] It should be noted that a large number of specific details are set forth in the specification provided herein. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0188] It should also be noted that the term "comprising", "including", or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0189] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A medical image fusion method based on a pre-trained large model, characterized in that: include: Acquire an MRI image to be fused and a CT image to be fused; The CT image to be fused and the MRI image to be fused are input into a trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module and a decoding structure.

2. The medical image fusion method based on a pre-trained large model according to claim 1, characterized in that: The step of inputting the to-be-fused CT image and the to-be-fused MRI image into a trained image fusion model to obtain a fused image comprises: Inputting the CT image to be fused into a first coding structure to obtain a first code; Inputting the MRI image to be fused into a second coding structure to obtain a second code; Inputting the first code and the second code into the pre-trained large model structure to obtain a pre-trained code; Inputting the first code, the second code and the pre-trained code into a fusion module to obtain a fusion feature map; The fused feature map is input into the decoding structure to obtain the fused image.

3. The medical image fusion method based on a pre-trained large model according to claim 2, characterized in that: The step of inputting the CT image to be fused into a first coding structure to obtain a first code includes: Multiply the CT image to be fused by the downsampling coefficient to obtain a coefficient feature map; Convolve the coefficient feature map with the two groups of convolutions in sequence to obtain a convolution feature map; The convolution feature map and the coefficient feature map are added to obtain the first code.

4. The medical image fusion method based on a pre-trained large model according to claim 2, characterized in that: The step of inputting the fused feature map into a decoding structure to obtain a fused image includes: Multiply the fused feature map by the upsampling coefficient to obtain the upsampled feature map; Convolve the upsampled feature map with the two groups of convolutions in sequence to obtain a convolution feature map; Add the convolution feature map and the coefficient feature map to get the fused image.

5. The medical image fusion method based on a pre-trained large model according to claim 2, characterized in that: The first code, the second code and the pre-trained code are input into the fusion module to obtain a fusion feature map, including: Adding the first code, the second code and the pre-trained code to obtain an addition graph; Perform 1×1 convolution, 1×3 convolution, and 3×1 convolution on the first code in sequence to obtain the first image; Perform 1×1 convolution, 3×1 convolution, and 1×3 convolution on the second code in sequence to obtain the second image; The pre-trained code is sequentially subjected to 1×1 convolution, 3×1 convolution, and 1×3 convolution to obtain the third image; Perform average pooling on the added image to obtain a pooled image; The first image, the second image, the third image and the pooled image are concatenated and then subjected to 1×1 convolution to obtain a convolution image. Add the convolution map and the addition map to get the fused feature map.

6. The medical image fusion method based on a pre-trained large model according to any one of claims 2 to 5, characterized in that: An attention module is also provided after the first encoding structure, the second encoding structure and the pre-trained large model structure to capture the attention features of the first encoding, the second encoding and the pre-trained encoding.

7. The medical image fusion method based on a pre-trained large model according to claim 6, characterized in that: The processing process of the attention module includes: Splitting the fused feature map into a first feature map and a second feature map; the local parameters of the first feature map and the second feature map are different; Perform 1×1 convolution and sampling processing of different branches on the first feature map to obtain a sampling 1 feature map and a sampling 2 feature map respectively; Perform 1×1 convolution on the second feature map to obtain a convolution feature map; After feature extraction is performed on the sampled second feature map, it is merged with the downsampled convolution feature map to obtain a merged feature map; The merged feature map and the sampled feature map are merged and then upsampled to obtain the adjusted fused feature map.

8. A medical image fusion device based on a pre-trained large model, characterized in that: include: An image acquisition module, which is used to acquire the MRI image to be fused and the CT image to be fused; The image fusion module is used to input the CT image to be fused and the MRI image to be fused into the trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module and a decoding structure.

9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program to: Acquire an MRI image to be fused and a CT image to be fused; The CT image to be fused and the MRI image to be fused are input into a trained image fusion model to obtain a fused image; the image fusion model is composed of a first encoding structure, a second encoding structure, a pre-trained large model structure, a fusion module and a decoding structure.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by the processor to implement a medical image fusion method based on a pre-trained large model as described in any one of claims 1-7.

Citation Information

Cited By

  • Crohn disease intestinal fibrosis condition early warning method based on large model driving

    CN121121240A

  • Crohn's disease intestinal fibrosis condition early warning method based on large model driving

    CN121121240B