An optimized image synthesis method and system for multimodal magnetic resonance imaging

By using modal-specific encoder, ViT backbone, Mamba module, dual-scale feature fusion device and dual-stream fusion device in multimodal magnetic resonance imaging image synthesis, problems such as loss of detail and instability of training are solved, high-quality image synthesis is achieved, and diagnostic accuracy and resource utilization efficiency are improved.

CN119693249BActive Publication Date: 2025-06-10FUDAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510191666.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art has problems such as loss of detail, unstable GAN training, insufficient global structure and context consistency in multimodal magnetic resonance imaging image synthesis, affecting image quality and diagnostic accuracy.

Method used

Modal-specific encoder is used to retain the features of each mode, combine the ViT backbone and Mamba module to solve the shortcomings of CNN in long-distance dependence processing, and introduce a dual-scale feature fusion device and a dual-stream fusion module to optimize the feature aggregation module to enhance training stability.

Benefits of technology

Effectively avoid details loss, improve the global structural consistency of the image and the multi-scale feature fusion ability, significantly improve image quality and diagnostic value, reduce dependence on additional MRI scans, and save resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693249B_ABST
    Figure CN119693249B_ABST
Patent Text Reader

Abstract

The present invention proposes an optimized image synthesis method and system for multimodal magnetic resonance imaging, belonging to the technical field of image synthesis, including: constructing a multimodal image synthesis model, extracting first features from each modality of multimodal magnetic resonance imaging data; extracting context information from the first features to obtain second features; extracting information of different scales in the image from the second features to obtain front-half stage features and back-half stage features; jump-connecting the front-half stage features to the back-half stage features for integration to obtain multi-scale features; cascading and fusing the multi-scale features of each modality to obtain fused features; synthesizing a target modality image according to the fused features. The present invention obtains more complete image data through the optimized image synthesis method, which is helpful for accurate clinical diagnosis and scientific research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image synthesis, and particularly relates to an optimized image synthesis method and system for multimodal magnetic resonance imaging. Background Art

[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Magnetic Resonance Imaging (MRI) is a non-invasive and non-radiative imaging technology widely used in medical diagnosis, which can provide high-resolution soft tissue images, especially outstanding in the imaging of soft tissues such as the brain, spine, muscles, and joints. However, despite its important position in the medical field, there are still many challenges. First, MRI equipment is expensive and has high maintenance costs, which limits its popularization in some medical institutions, especially in resource-scarce areas. Second, MRI scans often take a long time, which may cause discomfort to some patients, especially those who cannot stay still, thereby affecting the image quality. In addition, patients with metal implants or medical devices may encounter safety risks when undergoing MRI scans. Moreover, due to noise interference, motion artifacts, and limitations in imaging resolution, MRI images may be incomplete or of poor quality in some clinical applications, affecting the accuracy of diagnosis.

[0004] The prior art discloses the Hi-Net network, which includes a modality-specific network, a multimodal fusion network, and a multimodal synthesis network. Among them, the modality-specific network is used to learn the feature representation of each modality, and the multimodal fusion network adaptively fuses the features of different modalities through a hybrid fusion block, and finally generates the target modality image through a Generative Adversarial Network (GAN). However, the Hi-Net network has several problems in practical applications. First, when fusing different modality information, the detailed features of a specific modality may be lost, which will affect the accuracy of the synthesized image. Second, the training process of GAN is prone to instability, resulting in the generated images being not diverse enough or not realistic enough, thus affecting the image quality.

[0005] The prior art discloses the CACR-Net model, which mainly includes a confidence-guided aggregation module (CGA) and a cross-modal refinement module (CMR). CGA adaptively aggregates the target modality image generated by multiple source modalities through confidence maps, while CMR further mines the correlation information between the multimodal input and the target modality image to optimize the target modality image. Although the model performs well in multimodal MR image synthesis, there are still some potential problems. First, the accuracy of the confidence map is crucial to the quality of the target image. If the confidence map cannot accurately reflect the similarity between the generated image and the real image, it may lead to the loss of image details and structural integrity. Secondly, when the cross-modal refinement module fuses multimodal features, if the fusion is insufficient, it may cause the synthesized image to lack details or have incomplete structures in some areas, affecting the diagnostic value of the image.

[0006] The prior art discloses the ResViT model, which combines an encoder, an information bottleneck, and a decoder, and uses convolutional layers to capture local features. The information bottleneck contains an aggregated residual Transformer module to combine the advantages of Transformer and CNN to achieve high-quality medical image synthesis. Although the model performs well in local feature extraction and global context modeling, it also has several obvious problems. First, CNN performs poorly in dealing with long-range spatial dependencies, which may cause defects in the global structure and context consistency of the synthesized image, and the computational burden is heavy. Secondly, although the Transformer can capture long-range contextual information, its positioning ability is weak, which may result in inaccurate details. Summary of the invention

[0007] In order to overcome the deficiencies of the above-mentioned prior art, the present invention proposes an optimized image synthesis method and system for multimodal magnetic resonance imaging, designs a modality-specific encoder to process the data of each modality separately, retains the features of different modalities, and avoids detail loss; adopts the ViT backbone and Mamba module to solve the problem of CNN processing long-distance dependency differences and reduce the computational burden. In addition, the dual-scale feature fuser can better capture information of different scales, enhance the expression ability of the network, and effectively reduce the problem of detail loss. Finally, by optimizing the feature aggregation module and the dual-stream fusion module, the present invention indirectly enhances the stability of the training process, improves the quality of the generated image, and avoids the problem of unstable GAN training.

[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0009] In a first aspect, a method for optimizing image synthesis of multimodal magnetic resonance imaging is disclosed, comprising:

[0010] Acquiring multimodal magnetic resonance imaging data;

[0011] Construct a multi-modal image synthesis model and synthesize the target modal image from the multi-modal magnetic resonance imaging data;

[0012] Among them, the multi-modal image synthesis model extracts the first features from each modality of the multi-modal magnetic resonance imaging data; extracts context information from the first features to obtain the second features; extracts information at different scales in the image from the second features to obtain the first half-stage features and the second half-stage features; jumps and connects the first half-stage features to the second half-stage features to integrate and obtain multi-scale features; cascades and fuses the multi-scale features of each modality to obtain fused features; synthesizes the target modal image according to the fused features.

[0013] In a second aspect, an optimized image synthesis system for multi-modal magnetic resonance imaging is disclosed, including:

[0014] A data acquisition module configured to: acquire multi-modal magnetic resonance imaging data;

[0015] A multi-modal image synthesis module configured to: construct a multi-modal image synthesis model and synthesize the target modal image from the multi-modal magnetic resonance imaging data;

[0016] Among them, the multi-modal image synthesis model extracts the first features from each modality of the multi-modal magnetic resonance imaging data; extracts context information from the first features to obtain the second features; extracts information at different scales in the image from the second features to obtain the first half-stage features and the second half-stage features; jumps and connects the first half-stage features to the second half-stage features to integrate and obtain multi-scale features; cascades and fuses the multi-scale features of each modality to obtain fused features; synthesizes the target modal image according to the fused features.

[0017] In a third aspect, an electronic device is disclosed, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above-mentioned optimized image synthesis method for multi-modal magnetic resonance imaging are completed.

[0018] In a fourth aspect, a computer-readable storage medium is disclosed, which is used to store computer instructions. When the computer instructions are executed by the processor, the steps of the above-mentioned optimized image synthesis method for multi-modal magnetic resonance imaging are completed.

[0019] Compared with the prior art, the beneficial effects of the present invention are:

[0020] The method proposed in the present invention can preserve the features of each modality by using a specific encoder for each modality, avoiding the loss of details during the fusion process. Compared with traditional Convolutional Neural Networks (CNNs), the Transformer model captures high-level information of each modality better through its global context awareness ability, providing an accurate and complete input representation for subsequent multi-modal image synthesis.

[0021] The SSM layer proposed in the present invention enhances the global understanding of the image by effectively capturing long-range context information, making up for the deficiency of traditional CNNs in long-range dependence modeling. This enables the network to better maintain the overall structure and consistency of the image during the synthesis process, reducing the loss of details in the image, which is particularly significant for complex anatomical structures.

[0022] Through the fusion of multi-scale features, the Dual-Fuser in the present invention can better process detail information at different scales in the image, improving the expression ability of details and structures. The high-resolution branch can capture key information of a specific modality more precisely through more downsampling stages, while the low-resolution branch can better capture global features. Finally, more details of specific modalities are retained through skip connections, effectively enhancing the quality and diagnostic value of the image.

[0023] The present invention introduces a Dual-Scale Feature Fuser (Dual-Fuser). The network design is similar to the U-Net structure, including multiple downsampling stages, and uses Swin Transformer layers for feature fusion. After the output feature map undergoes bilinear interpolation to form a high-resolution image, multi-scale feature fusion is performed.

[0024] The dual-stream structure proposed in the present invention can effectively process the spatial information and channel information of the feature map, suppressing less useful features and retaining more valid information. The spatial stream can enhance the expression of spatial features by calculating the spatial attention map, while the channel stream captures the relationship between channels through global average pooling, enabling the network to perform more accurate feature selection and suppress irrelevant features. This structure improves the efficiency of feature fusion, making the finally synthesized image more realistic and rich in details.

[0025] The FA module proposed in the present invention effectively fuses the features of multi-modal images through global and local attention mechanisms, enhancing the complementary information between different modalities. Global attention enables the network to focus on the global features of the image, while local attention helps to extract the detailed parts. This fusion method improves the quality of the image, making the synthesized image highly accurate and consistent in both overall structure and detail presentation.

[0026] The present invention optimizes resource utilization: By synthesizing the missing modalities, it effectively reduces the dependence on additional MRI scans, saves scanning time and equipment costs, thereby enabling more efficient utilization of medical resources.

[0027] The present invention enhances image usability: In traditional methods, the missing modalities may affect clinical judgment and the comprehensiveness of research. By synthesizing multi-modal images, doctors can obtain more complete image data, which helps in making more accurate clinical diagnoses and scientific research.

[0028] The present invention improves diagnostic efficiency: This method reduces the risk of misdiagnosis caused by image noise and artifacts through image quality improvement. Higher-quality synthesized images help doctors make quick decisions in complex medical conditions, improving diagnostic efficiency.

[0029] The present invention reduces the burden on patients: For some patients who cannot remain stationary for a long time or due to medical condition limitations cannot undergo a comprehensive scan, the image synthesis technology of the present invention can provide high-quality image support without increasing the burden on patients, ensuring the accuracy of diagnostic results.

[0030] The present invention promotes the expansion of medical data: The synthesized multi-modal images can provide more sample data for medical research and education, promote the expansion of the dataset, contribute to conducting more extensive clinical experiments and scientific research work, and drive the further development of the medical imaging field.

[0031] Advantages of additional aspects of the present invention will be partially given in the following description, partially will become apparent from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0033] Figure 1 It is the overall framework diagram of the optimized image synthesis method for multi-modal magnetic resonance imaging described in Embodiment 1 of the present invention.

[0034] Figure 2 It is the schematic diagram of the two-stream fusion module described in Embodiment 1 of the present invention.

[0035] Figure 3 It is the visualization result of the T1 and T2 to FLAIR synthesis task performed by the multi-modal image synthesis model described in Embodiment 1 of the present invention.

[0036] Figure 4 It is the visualization result of the T1 and FLAIR to T2 synthesis task performed by the multi-modal image synthesis model described in Embodiment 1 of the present invention.

[0037] Figure 5 Visualization results of the T2 and FLAIR synthetic T1 task for the multimodal image synthesis model described in Embodiment 1 of the present invention. Detailed implementation manners

[0038] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0039] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present invention.

[0040] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0041] Embodiment 1

[0042] In one or more implementation manners, an optimized image synthesis method for multimodal magnetic resonance imaging is disclosed, including the following steps:

[0043] Step 1, obtain magnetic resonance imaging (MRI) of different modalities, including MRI T1, MRI T2, and MRI FLAIR. When the FLAIR modality is missing, obtain the T1 and T2 modalities as inputs; when the T1 modality is missing, obtain the T2 and FLAIR modalities as inputs; when the T2 modality is missing, obtain the T1 and FLAIR modalities as inputs, so as to perform the synthesis task of the missing modality.

[0044] Step 2, construct a multimodal image synthesis model, as Figure 1 shown, the model includes an encoder, a Mamba module, a dual-scale feature fuser, a two-stream fusion module, and a feature aggregation module connected in sequence. Specifically:

[0045] Step 2-1, construct an encoder, and extract first features of each modality of the multimodal magnetic resonance images through the encoder.

[0046] In order to maintain the integrity of different modality data, by designing a modality-specific encoder and inputting each modality, the specific encoder first assigns each input modality image to an independent branch. In this embodiment, the modality-specific encoder includes two branches, and each branch uses a ViT (Vision Transformer) backbone architecture to process images.

[0047] Following the ViT method, the input is first segmented into a series of non-overlapping patches, each patch is regarded as a sequence, and is converted into a vector through an embedding layer to be consistent with the input specification of the Transformer. A vector sequence is obtained through an ordered sorted set of multiple vectors. This serialization process enables the present invention to utilize the Transformer model to process image data and overcome the limitations of traditional convolutional neural networks.

[0048] Subsequently, these vector sequences are input into the Transformer encoder for additional feature extraction and processing. This modality-specific encoder structure effectively captures the features of different modality data and maintains its integrity, providing a more accurate input representation for subsequent tasks.

[0049] Among them, the main difference between the specific encoder and the ordinary encoder lies in its design purpose and processing method. The ordinary encoder usually processes all input data through a single structure, while the specific encoder designs independent branches for each input modality (such as T1, T2, FLAIR). Each branch specifically processes the images of one modality to ensure that the features of each modality can be effectively extracted and retained. This encoder is designed with a two-branch structure, and each branch consists of a ViT backbone, which can effectively process different modality data. The input data of each modality will undergo serialization processing and feature extraction through the Transformer encoder.

[0050] Step 2-2: Construct a Mamba module, and obtain the second feature by extracting context information of the first feature through the Mamba module;

[0051] The Mamba block contains an SSM layer to capture long-range context information. There are two branches in the Mamba block. First, the input feature map is divided into pieces of size according to formula (1) to generate a sequence which are respectively input into the two branches;

[0052] (1)

[0053] where is the number of blocks; and are the height and width of the input feature map respectively; is the size of each small block after the feature map is divided into small pieces.

[0054] In the first branch, the sequence is linearly embedded to calculate the gating variable, as shown in the following formula:

[0055] (2)

[0056] Among them, is the gating variable; represents the activation function; is the linear transformation. In the second branch, the sequence undergoes linear embedding, is processed using depthwise convolution, and then projected through the SSM layer, as shown in the following equation:

[0057] (3)

[0058] Among them, is the feature representation processed by linear transformation, depthwise separable convolution, and the state space model; is the depthwise separable convolution; is the SSM layer.

[0059] In the above formula, the SSM layer is implemented based on the selective state space sequence model. Initially, the patch sequence is extended by performing selective scans in all directions on the 2D feature map. Then the extended sequence is processed through state space modeling and merged back together. During this process, a discrete state space model is established.

[0060] (4)

[0061] (5)

[0062] Among them, represents the hidden state, represents the input sequence, represents the output sequence, represents the sequence index, , , , are the learnable SSM parameters, and , , and , where represents the dimension of the state space. Finally, the gating variable obtained from the first branch and the feature obtained from the second branch are multiplied, and the linear projection of the multiplication result is combined with the input sequence residual:

[0063] (6)

[0064] Among them, is the output feature of the Mamba module, is the linear transformation, is the per-pixel multiplication operation.

[0065] The Mamba module contains a selective state space model (SSM) layer that extends the sequence by selectively scanning a 2D feature map and processes the extended sequence using state space modeling.

[0066] Step 2-3: Construct a dual-scale feature fuser to extract information at different scales in the image, obtaining the front-half stage features and the back-half stage features.

[0067] To better capture information at different scales in the image and enhance the model's understanding and representation capabilities, the present invention introduces an innovative dual-scale feature fuser (Dual-Fuser) aimed at fully utilizing the output of the Mamba module to build a powerful function for a specific modality, and its structure is as Figure 1 shown in (b) of the figure.

[0068] First, the output features of the Mamba module are used as low-resolution feature maps , and at the same time bilinear interpolation is applied to generate higher-resolution feature maps . Next, the low-resolution feature maps and the high-resolution feature maps are input into separate branches of the dual-scale fuser, and each branch has several downsampling stages. Each stage includes a Swin Transformer layer and a patch merging unit, which are applied to reduce the resolution of the feature map while increasing its dimension to form a hierarchical feature representation.

[0069] It should be noted that the high-resolution branch has more downsampling stages than the low-resolution branch, which enhances the network's ability to capture features. To retain key details of a specific modality in the final output, the present invention uses a patch expansion operation and a Swin Transformer layer in the later stage to enhance the feature resolution. In the dual-scale feature fuser, the downsampling stage is the front-half stage of the dual-scale fuser, and the back-half stage is the upsampling. By introducing an innovative twin-stream fusion module (TSF), skip connections are used to retain more multi-scale information of a specific modality, and this design allows the Dual-Fuser to more flexibly and effectively capture the relationships between different scales during feature fusion.

[0070] Step 2-4: Construct a twin-stream fusion module (TSF) to integrate multi-scale features by skipping the front-half stage features of the dual-scale feature fuser to the back-half stage features;

[0071] To better integrate the features extracted in the first half of the Dual-Fuser, a Two-Stream Fusion (TSF) module is proposed. The feature maps extracted from the first half of the dual-scale feature fuser through multiple Swin transformers and block mergers are used as the features in the first half stage, and the feature information extracted from the second half of the dual-scale feature fuser through multiple Swin transformers and block expansions is used as the features in the second half stage. The features in the first half stage are passed through the TSF module to the features in the second half stage, effectively integrating multi-scale features.

[0072] As an implementation, this embodiment adopts the operation of combining Swin Transformer with block merging and block expansion. The low-resolution feature map and the high-resolution feature map are respectively input into the Swin Transformer and block merging branch, and sequentially pass through multiple Swin Transformer and block merging combination modules, which are referred to as the first half of the dual-scale feature fuser above. The second half of the dual-scale feature fuser is composed of multiple Swin Transformer and block expansion combinations. The output of the Swin Transformer and block merging passes through the TSF module and is then concatenated with the output of the Swin Transformer and block expansion in the second half of the dual-scale feature fuser and input into the next Swin Transformer and block expansion module to achieve the skip connection between the first half stage and the second half stage.

[0073] The TSF module consists of two special data streams, namely the spatial stream and the channel stream. Specifically, as Figure 2 shown, the multi-scale features extracted in the first half stage are represented as and . The features after connecting them are input into the pointwise convolution operation to generate the multi-scale feature map . Then, they enter the spatial stream and channel stream stages respectively. The TSF module is designed to suppress less useful features and only allow more informative features to pass through.

[0074] (1) Channel stream (CF): This branch uses compression and excitation operations to utilize the relationship between channels in the convolutional feature map. Starting from the feature map , global average pooling is performed in the spatial dimension to capture the global context, generating the feature descriptor . Then, it undergoes two convolutional layers and an activation function Sigmoid during the excitation process. The output of the CF branch is obtained by rescaling the original feature map using the activation value .

[0075] (2)Spatial Flow (SF): This branch is designed to utilize the spatial relationships within the convolutional features. The purpose of SF is to create a spatial attention map for recalibrating the input feature map of the feature map. To construct the spatial attention map, the SF branch first performs global average pooling and max pooling along the channel dimension of the feature map respectively, and then combines the results to form a new feature map . The mapping subsequently passes through a convolutional layer, followed by activation, resulting in a spatial attention mapping . Then this mapping is used to adjust the feature map , generating the output of the SF branch.

[0076] The outputs of the channel flow branch and the spatial flow branch are respectively multiplied with the multi-scale feature map and then concatenated, and after per-pixel convolution, the output of the two-stream fusion module (TSF) is obtained.

[0077] Step 2 - 5: Construct the Feature Aggregation Module (FA), and through the feature aggregation module, the multi-scale features of each modality are cascaded and fused to obtain the fused features;

[0078] Although the two-stream fusion module (TSF) has successfully integrated the multi-resolution information of a specific modality, a mechanism is still needed to manage the multi-scale information between different modalities. Specifically, for the FA module, first, the multi-scale features from different modalities are fused by the dual-fuser module and integrated through cascading. Then, the FA module is divided into two branches: one branch fuses features from a global attention perspective, and the other from a local attention perspective. Both branches undergo per-pixel convolution processing. In the final step, the features from the two branches are combined again and multiplied with the initial features per-pixel to produce the final fused features.

[0079] The corresponding formula is as follows:

[0080] (7)

[0081] (8)

[0082] Among them, represents the finally fused features, represents the intermediate features before entering the activation function, GA represents the global branch of the FA module, LA represents the local branch of the FA module, and represent the multi-scale features of different modalities, is the activation function, It is a pixel-by-pixel addition operation. Through this method, the FA module effectively integrates features from different modalities while retaining the importance of multi-scale information. This design enables the model of the present invention to better understand and process the complex relationships between different modality data, thereby improving the performance of the model in various tasks.

[0083] The FA module effectively fuses the features of multi-modal images through global and local attention mechanisms, enhancing the complementary information between different modalities. Global attention enables the network to focus on the global features of the image, while local attention helps to extract the detailed parts. This fusion method improves the quality of the image, making the synthesized image have high accuracy and consistency in both overall structure and detail presentation.

[0084] Step 2-6: The output of the feature aggregation module generates the target modality image after passing through the Mamba module.

[0085] Step 3: Input the obtained multi-modal image into the constructed multi-modal image synthesis model to obtain the target modality.

[0086] The optimized image synthesis method for multi-modal MRI provided by the present invention effectively compensates for the problems of image loss or quality degradation caused by factors such as scanning limitations, insufficient time, or patient discomfort by fusing the MRI image information of different modalities. Compared with traditional methods, this image synthesis method not only avoids the time and economic costs of additional scans but also can significantly improve the integrity and quality of the image, thereby effectively enhancing the application effect and accuracy of MRI in clinical diagnosis.

[0087] Embodiment 2

[0088] In one or more embodiments, an optimized image synthesis system for multi-modal magnetic resonance imaging is disclosed, specifically including:

[0089] A data acquisition module, which is configured to: acquire multi-modal magnetic resonance imaging data;

[0090] A multi-modal image synthesis module, which is configured to: construct a multi-modal image synthesis model and synthesize the multi-modal magnetic resonance imaging data to obtain a target modality image;

[0091] Among them, the multi-modal image synthesis model extracts first features for each modality of the multi-modal magnetic resonance imaging data; extracts context information from the first features to obtain second features; extracts information of different scales in the image from the second features to obtain front-half stage features and rear-half stage features; jumps and connects the front-half stage features to the rear-half stage features to integrate and obtain multi-scale features; cascades and fuses the multi-scale features of each modality to obtain fused features; and synthesizes a target modality image according to the fused features.

[0092] Embodiment 3

[0093] This embodiment provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above-mentioned optimized image synthesis method for multimodal magnetic resonance imaging are completed.

[0094] Embodiment 4

[0095] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned optimized image synthesis method for multimodal magnetic resonance imaging are completed.

[0096] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operation steps are executed on the computer or other programmable device to generate computer-implemented processing. Therefore, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0099] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0100] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for optimizing image synthesis of multimodal magnetic resonance imaging, characterized in that: include: Acquiring multimodal magnetic resonance imaging data; Construct a multimodal image synthesis model and synthesize multimodal magnetic resonance imaging data to obtain the target modality image; Among them, the multimodal image synthesis model extracts the first feature of each modality of the multimodal magnetic resonance imaging data; extracts context information from the first feature to obtain the second feature; extracts information of different scales in the image from the second feature to obtain the first half stage feature and the second half stage feature; jumps the first half stage feature to the second half stage feature integration to obtain the multi-scale feature; cascades and fuses the multi-scale features of each modality to obtain the fused feature; synthesizes the target modality image according to the fused feature; The multimodal image synthesis model includes: an encoder, a Mamba module and a dual-scale feature fuser; The encoder includes two branches, the input of each branch is magnetic resonance imaging data of a modality, and the output of each branch is the first feature of the modality. The encoder assigns each input modality image to an independent branch, and each branch uses the Vision Transformer backbone architecture to process the image; the first features of each modality are respectively subjected to the Mamba module to extract context information to obtain the second features of each modality; the second features of each modality and the feature maps obtained by interpolating the second features are respectively input into separate branches of the dual-scale feature fuser to extract the first-half stage features and the second-half stage features, specifically, by using the Swin Transformer layer and the patch merging unit in the downsampling stage, and using the patch expansion operation and the Swin Transformer layer in the upsampling stage to construct a dual-scale feature fuser to respectively extract the first-half stage features and the second-half stage features.

2. The method for optimizing image synthesis of multimodal magnetic resonance imaging according to claim 1, characterized in that: The Mamba module includes two branches. In the first branch, the sequence generated by the input feature map segmentation is linearly embedded to calculate the gated variable; in the second branch, the sequence undergoes linear embedding, is processed using deep convolution, and then projected through the SSM layer; specifically: In the first branch, the sequence is linearly embedded to calculate the gating variable as shown below: in, is the gating variable; represents the activation function; is a linear transformation; In the second branch, the sequence undergoes linear embedding, processed using depthwise convolution, and then projected through an SSM layer as shown below: in, It is the feature representation after processing by linear transformation, depthwise separable convolution and state-space model; It is a depth-wise separable convolution; It is the SSM layer; Finally, the gated variable obtained by the first branch And the features obtained by the second branch Perform a multiplication operation and combine the linear projection of the multiplication result with the input sequence residual to obtain the second feature: in, Output features for the Mamba module, is a linear transformation, It is a pixel-by-pixel multiplication operation.

3. The method for optimizing image synthesis of multimodal magnetic resonance imaging according to claim 1, characterized in that: The multimodal image synthesis model also includes: a dual-stream fusion module; The dual-stream fusion module consists of two data streams, namely a channel stream and a spatial stream. The input of the dual-stream fusion module is the first half stage features and the second half stage features. The first half stage features and the second half stage features are connected and the multi-scale feature maps obtained by point-by-point convolution operation enter the spatial stream and the channel stream respectively. The outputs of the spatial stream and the channel stream are connected and convolved to obtain multi-scale features.

4. The method for optimizing image synthesis of multimodal magnetic resonance imaging according to claim 3, characterized in that: The spatial stream performs global average pooling and maximum pooling on the multi-scale feature map along the channel dimension, combines the pooling results to form a new feature map, and then passes through a convolution layer and an activation function to obtain a spatial attention map.

5. The method for optimizing image synthesis of multimodal magnetic resonance imaging according to claim 4, characterized in that: The channel flow performs global average pooling on the multi-scale feature map in the spatial dimension to generate a feature descriptor, which then undergoes two convolutional layers and an activation function.

6. The method for optimizing image synthesis of multimodal magnetic resonance imaging according to claim 1, characterized in that: The multimodal image synthesis model further includes: a feature aggregation module; The feature aggregation module cascades and integrates the multi-scale features of each modality, and fuses the features from the global attention perspective and the local attention perspective respectively to obtain the global attention feature and the local attention feature, and then combines the global attention feature and the local attention feature again and multiplies them pixel by pixel with the multi-scale features of different modalities to generate the final fused feature.

7. An optimized image synthesis system for multimodal magnetic resonance imaging, characterized in that: include: A data acquisition module, configured to: acquire multimodal magnetic resonance imaging data; A multimodal image synthesis module is configured to: construct a multimodal image synthesis model, and synthesize multimodal magnetic resonance imaging data to obtain a target modality image; Among them, the multimodal image synthesis model extracts the first feature of each modality of the multimodal magnetic resonance imaging data; extracts context information from the first feature to obtain the second feature; extracts information of different scales in the image from the second feature to obtain the first half stage feature and the second half stage feature; jumps the first half stage feature to the second half stage feature integration to obtain the multi-scale feature; cascades and fuses the multi-scale features of each modality to obtain the fused feature; synthesizes the target modality image according to the fused feature; The multimodal image synthesis model includes: an encoder, a Mamba module and a dual-scale feature fuser; The encoder includes two branches, the input of each branch is magnetic resonance imaging data of a modality, and the output of each branch is the first feature of the modality. The encoder assigns each input modality image to an independent branch, and each branch uses the Vision Transformer backbone architecture to process the image; the first features of each modality are respectively subjected to the Mamba module to extract context information to obtain the second features of each modality; the second features of each modality and the feature maps obtained by interpolating the second features are respectively input into separate branches of the dual-scale feature fuser to extract the first-half stage features and the second-half stage features, specifically, by using the Swin Transformer layer and the patch merging unit in the downsampling stage, and using the patch expansion operation and the Swin Transformer layer in the upsampling stage to construct a dual-scale feature fuser to respectively extract the first-half stage features and the second-half stage features.

8. An electronic device, characterized in that: The invention comprises a memory and a processor and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the optimized image synthesis method for multi-modal magnetic resonance imaging according to any one of claims 1 to 6 is completed.

9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, complete the optimized image synthesis method for multimodal magnetic resonance imaging as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Brain tumor image segmentation method based on multi-scale convolution and Mama structure

    CN118447244A