A medical image fusion method, system, device and medium based on space-frequency domain feature interaction analysis
By employing a spatial-frequency domain feature interaction analysis method, and utilizing the Restormer image encoder and spatial-frequency co-fusion block, the problem of insufficient utilization of frequency domain information in existing medical image fusion is solved, achieving medical image fusion effects with greater detail preservation and global semantics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2026-04-07
AI Technical Summary
Existing medical image fusion methods are insufficient in utilizing frequency domain information, making it difficult to fully utilize frequency domain information. Furthermore, deep learning-based methods struggle to simulate long-term dependencies and global contexts in complex medical scenarios, neglecting fine-grained local details.
A method based on spatial-frequency domain feature interaction analysis is adopted. Through the Restormer image encoder and spatial-frequency co-fusion block, combined with the frequency interaction module and spatial compensation module, features of structural modality and functional modality images are extracted, and fused medical images are generated by training with a hybrid loss function.
It achieves full utilization of frequency domain information, improves the detail preservation ability and global semantic understanding of medical image fusion, and generates more accurate fused images.
Smart Images

Figure CN121121374B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing technology, and in particular to a medical image fusion method, system, device and medium based on spatial-frequency domain feature interaction analysis. Background Technology
[0002] The purpose of multimodal medical image fusion is to combine complementary information from different imaging sources to generate a comprehensive fused image. MRI, CT, and ultrasound can provide structural information such as lesion location, size, and changes in surrounding tissues, while PET, fMRI, and SPECT can display detailed functional information. Therefore, the effective interaction of functional and structural cues in medical images can help clinicians perform more accurate disease analysis and monitoring.
[0003] In recent years, the success of deep learning in computer vision has also driven the rapid development of medical image fusion. Existing medical image fusion methods can be broadly classified into three categories: methods based on convolutional neural networks (CNNs), methods based on Transformers, and methods based on GANs. CNN-based methods primarily rely on convolutional operations to capture local texture and spatial structure from source images. These methods typically employ an encoder-decoder architecture to extract features from different modalities and fuse them together through concatenation, addition, or attention mechanisms. Representative works, such as DenseFuse, U2Fusion, and variants based on U-Net, have shown promising results in preserving important details from both modalities. However, due to their limited receptive field, CNNs often struggle to model long-term dependencies and global context, which are crucial for accurate structural alignment in complex medical scenarios. Therefore, some researchers have introduced Transformers to establish global dependencies and better preserve contextual information. However, due to patch-based tagging, most methods may ignore fine-grained local details. Furthermore, Transformers typically have a large number of parameters and require substantial data for training, which is challenging in the field of medical imaging. GAN-based methods often employ adversarial training strategies to improve perceptual quality and structural consistency. For example, FusionGAN and TCGAN utilize modality-aware discriminators and fusion generators to produce visually believable fusion results. Although deep learning-based methods have made significant progress in multimodal medical image fusion, they primarily focus on spatial domain modeling while neglecting the rich potential of the frequency domain. Summary of the Invention
[0004] To address the aforementioned issues, this application provides a medical image fusion method, system, device, and medium based on spatial-frequency domain feature interaction analysis, thereby resolving the problem of fully utilizing frequency domain information in the fusion process of medical images.
[0005] According to the first aspect of this application, a medical image fusion method based on spatial-frequency domain feature interaction analysis is provided, the method comprising:
[0006] Obtain the medical image pair to be fused, which includes structural modality images and functional modality images; convert the functional modality images to the YCbCr color space, extract the Y channel as the channel to be fused for the functional modality images, and retain the Cb and Cr channels;
[0007] Two Restormer-based image encoders are used to extract features from the structural modality image and the functional modality image to be fused channels, respectively, to obtain multi-layer structural modality features and multi-layer functional modality features to be fused channels; the two image encoders are bridged by at least two spatial-frequency co-fusion blocks, which include a frequency interaction module and a spatial compensation module;
[0008] Based on the space-frequency co-fusion block, the structural modal features of each layer are analyzed. and the channel features to be fused for each functional mode. The fusion process is performed to obtain the fusion characteristics;
[0009] The fusion features are input into the decoder, which generates a fused Y-channel image through at least three upsampling operations and corresponding convolutional layer processing. The fused Y-channel image is then merged with the retained Cb and Cr channels in the YCbCr color space and converted to the target color space to obtain the fused medical image.
[0010] A medical image fusion model is formed by combining two image encoders and decoders bridged by at least two spatial-frequency co-fusion blocks. The medical image fusion model is trained based on a hybrid loss function and a medical dataset. The fusion of structural modality images and functional modality images is achieved based on the trained medical image fusion model.
[0011] Furthermore, based on the space-frequency co-fusion block, the structural modal features of each layer are analyzed. and the channel features to be fused for each functional mode. The fusion process yields fusion features, including:
[0012] Through the frequency interaction module and Perform discrete Fourier transforms on each, and obtain phase components amplitude component and phase components amplitude component Amplitude gating map W generated based on modal guided gating unit. A Phase-gated graph W P WA according to Generate, W P according to Generate; utilize W A right and By performing element-wise weighted summation, the fused amplitude components are obtained. Using W P right and Perform element-wise weighted summation to obtain the fused phase components. right and Performing the inverse discrete Fourier transform yields the frequency domain fused features. Will Downsampled output with adjacent lower-level spatial-frequency co-fusion block By concatenating the data, a global information representation can be obtained.
[0013] Calculated using the space compensation module and The absolute difference is used to obtain the spatial texture difference; the spatial texture difference is then input into the spatial attention unit. Generate a spatial attention map; combine the spatial attention map with... Multiply element by element, then multiply by each element. Adding them together yields the functional compensation features. and structural compensation features Will Input channel attention units respectively Obtain the functional characteristics after channel optimization and channel-optimized structural features Will and The concatenated data, after convolution, activation, and pooling processes, yields the fused features of the current layer.
[0014] Furthermore, based on the modal guided gating unit, the amplitude gating map W is generated using the following formula. A Phase-gated graph W P :
[0015]
[0016] in, This represents the sigmoid function, and Con3 represents a 3×3 convolution. Represents the ReLU activation function;
[0017] Using W A right and By performing element-wise weighted summation, the fused amplitude components are obtained. And using W P right and Perform element-wise weighted summation to obtain the fused phase components. The calculation process is expressed as follows:
[0018]
[0019] Here, ⊙ represents element-wise multiplication.
[0020] Furthermore, on and Performing the inverse discrete Fourier transform yields the frequency domain fused features. The calculation process is expressed as follows:
[0021]
[0022] in, Indicates the inverse discrete Fourier transform;
[0023] Will Downsampled output with adjacent lower-level spatial-frequency co-fusion block By concatenating the data, a global information representation can be obtained. The calculation process is expressed as follows:
[0024]
[0025] in, Represents the feature integration function. This indicates a splicing operation.
[0026] Furthermore, the functional compensation feature is calculated using the following formula. and structural compensation features
[0027]
[0028] Here, ⊙ represents element-wise multiplication.
[0029] Furthermore, the optimized functional characteristics of the channel are calculated using the following formula. and channel-optimized structural features
[0030]
[0031] Will and The concatenated data, after convolution, activation, and pooling processes, yields the fused features of the current layer. The calculation process is expressed as follows:
[0032]
[0033] in, This indicates a processing function that includes 3×3 convolution, LeakyReLU activation, and pooling. This indicates a splicing operation.
[0034] Furthermore, the hybrid loss function is expressed as:
[0035]
[0036] Among them, table Show the mixed loss function, Indicates strength loss. This represents the maximum gradient loss, where α represents the weighting factor.
[0037] The strength loss is calculated using the following formula:
[0038]
[0039] Where max(,) represents the operation of extracting the maximum value of two images pixel by pixel, H represents the height of the image, W represents the width of the image, and I represents the height of the image. F I represents the fused image. V Represents structural modal images, I I Represents functional modality images;
[0040] The maximum gradient loss is calculated using the following formula:
[0041]
[0042] in, This represents the Sobel gradient operator.
[0043] According to the second technical solution of this application, a medical image fusion system based on spatial-frequency domain feature interaction analysis is provided, the system comprising:
[0044] The image preprocessing module is configured to acquire medical image pairs to be fused, including structural modality images and functional modality images; convert the functional modality images to the YCbCr color space, extract the Y channel as the channel to be fused for the functional modality images, and retain the Cb and Cr channels;
[0045] The feature extraction module is configured to use two Restormer-based image encoders to extract features from the structural modality image and the functional modality image to be fused channels, respectively, to obtain multi-layer structural modality features and multi-layer functional modality features to be fused channels; the two image encoders are bridged by at least two spatial-frequency co-fusion blocks, which include a frequency interaction module and a spatial compensation module;
[0046] The feature fusion module is configured to perform feature fusion based on a space-frequency co-fusion block for each layer of structural modal features. and the channel features to be fused for each functional mode. The fusion process is performed to obtain the fusion characteristics;
[0047] The feature decoding module is configured to input the fused features into the decoder. The decoder generates a fused Y-channel image through at least three upsampling operations and corresponding convolutional layer processing. The fused Y-channel image is then merged with the retained Cb and Cr channels in the YCbCr color space and converted to the target color space to obtain the fused medical image.
[0048] The model training module is configured to combine two image encoders and decoders bridged by at least two spatial-frequency co-fusion blocks to form a medical image fusion model. The medical image fusion model is trained based on a hybrid loss function and a medical dataset. The fusion of structural modality images and functional modality images is achieved based on the trained medical image fusion model.
[0049] According to the third technical solution of this application, an electronic device is provided, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the above-described method.
[0050] According to the fourth technical solution of this application, a non-transitory computer-readable storage medium storing instructions is provided, which performs the above method when the instructions are executed by a processor.
[0051] The medical image fusion methods, systems, devices, and media based on spatial-frequency domain feature interaction analysis according to the various schemes of this application have at least the following technical effects:
[0052] To address the issue of insufficient utilization of frequency domain information, this application designs a general medical image pairing scheme and validates it on a medical dataset. To address the problem of insufficient interaction between spatial and frequency information, this application designs a spatial-frequency co-fusion block to bridge the two image encoders in each layer. The frequency interaction module uses Discrete Fourier Transform to extract phase and amplitude components from the two modes and uses modality-guided gating units to aggregate them individually within each mode to explore global semantics. To extract local details, the spatial compensation module formulates an adaptive refinement mechanism in the spatial domain to compensate for texture differences between different modes, thereby achieving more detailed medical image fusion effects.
[0053] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0054] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0055] Figure 1 A flowchart illustrating a medical image fusion method based on spatial-frequency domain feature interaction analysis, provided for embodiments of this application;
[0056] Figure 2 This application provides a network framework diagram for implementing medical image fusion in its embodiments.
[0057] Figure 3 A framework diagram of the frequency interaction module provided in an embodiment of this application;
[0058] Figure 4 A framework diagram of the space compensation module provided in an embodiment of this application;
[0059] Figure 5 This is a structural diagram of a medical image fusion system based on spatial-frequency domain feature interaction analysis, provided in an embodiment of this application. Detailed Implementation
[0060] To enable those skilled in the art to better understand the technical solution of this application, the application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0061] This application provides a medical image fusion method based on temperature measurement data and deep learning, utilizing spatial-frequency domain feature interaction analysis. Please refer to... Figure 1 This is a flowchart of a medical image fusion method based on spatial-frequency domain feature interaction analysis provided in an embodiment of this application. The method includes steps S10-S50.
[0062] S10: Obtain the medical image pair to be fused, which includes structural modality images and functional modality images; convert the functional modality images to the YCbCr color space, extract the Y channel as the channel to be fused for the functional modality images, and retain the Cb and Cr channels.
[0063] It is important to note that structural modality images and functional modality images are two distinct types of medical imaging. The core function of structural modality images is to provide anatomical information such as the location, size, and surrounding tissue changes of lesions. Common image types include MRI (Magnetic Resonance Imaging) images and CT (Computed Tomography) images. The core function of functional modality images is to display information about the functional activity of tissues or organs. Common image types include PET (Positron Emission Tomography), SPECT (Single Photon Emission Tomography), and fMRI (Functional Magnetic Resonance Imaging) images.
[0064] Step S10 is the image preprocessing step, used to extract the fusion channels of the functional modality image. In this embodiment, taking the MRI-PET fusion task as an example, given an MRI image and a PET image, the PET image is converted from the RGB color space to the YCbCr color space, and the Y channel is extracted for fusion.
[0065] S20: Two Restormer-based image encoders are used to extract features from the structural modality image and the functional modality image to be fused channels, respectively, to obtain multi-layer structural modality features and multi-layer functional modality features to be fused channels; the two image encoders are bridged by at least two spatial-frequency co-fusion blocks, which include a frequency interaction module and a spatial compensation module.
[0066] It should be noted that the method of extracting corresponding features based on the Restormer image encoder is an existing technology, so it will not be elaborated on in detail here.
[0067] A key improvement of this application is the addition of a spatial-frequency co-fusion block to the existing image encoder to achieve the fusion of two medical images. The process of obtaining the fusion features will be described in detail in the subsequent step S30.
[0068] S30: Based on the space-frequency co-fusion block, the structural modal features of each layer are analyzed. and the channel features to be fused for each functional mode. The fusion process is performed to obtain the fusion characteristics.
[0069] This embodiment uses an MRI-PET fusion task as an example. Figure 2The diagram shows a network framework for medical image fusion provided in this embodiment. The network first processes MRI images and Y-channel extracted PET images using a Restormer-based image encoder, generating multi-layer features: multi-layer structural modal features and multi-layer functional modal features to be fused. Subsequently, in each encoder layer, MRI and PET features are integrated through an SFCF (Spatial-Frequency Co-fusion Block). This SFCF block consists of a FI (Frequency Interaction) module for exploring global contextual information and an SC (Spatial Compensation) module for compensating for fine-grained spatial details. The FI module extracts the amplitude and phase components of the multi-layer features using Discrete Fourier Transform and adaptively integrates them using a modality-guided gating unit. The fusion result is then concatenated with the output of the adjacent lower-layer SFCF block to produce a global representation. To supplement local texture information, the SC module takes the global representation, MRI multi-layer features, and multi-layer features obtained from the Y-channel PET image as input, aggregating their complementary spatial cues in the spatial domain through a dedicated thinning mechanism to obtain fused features. By connecting SFCF blocks from bottom to top, the network can gradually balance image modalities and achieve complementary information fusion.
[0070] In some embodiments, step S30 can be specifically implemented through the following steps S301-S302 to achieve feature fusion.
[0071] S301: Through the frequency interaction module and Perform discrete Fourier transforms on each, and obtain phase components amplitude component and phase components amplitude component Amplitude gating map W generated based on modal guided gating unit. A Phase-gated graph W P W A according to Generate, W P according to Generate; utilize W A right By performing element-wise weighted summation, the fused amplitude components are obtained. Using W P right and Perform element-wise weighted summation to obtain the fused phase components. right and Performing the inverse discrete Fourier transform yields the frequency domain fused features. Will Downsampled output with adjacent lower-level spatial-frequency co-fusion block By concatenating the data, a global information representation can be obtained.
[0072] S302: Calculated via the space compensation module and The absolute difference is used to obtain the spatial texture difference; the spatial texture difference is then input into the spatial attention unit. Generate a spatial attention map; combine the spatial attention map with... Multiply element by element, then multiply by each element. Adding them together yields the functional compensation features. and structural compensation features Will Input channel attention units respectively Obtain the functional characteristics after channel optimization and channel-optimized structural features Will and The concatenated data, after convolution, activation, and pooling processes, yields the fused features of the current layer.
[0073] It should be noted that step S301 is the data processing flow of the frequency interaction module, and S302 is the data processing flow of the spatial compensation module.
[0074] Taking the MRI-PET fusion task as an example, such as Figure 3 and Figure 4 As shown below, the formulas will be used to explain in detail how the frequency interaction module and the spatial compensation module work together to achieve feature fusion.
[0075] The frequency interaction module aims to aggregate global contextual information from the source images into the fused image. For example... Figure 3 As shown, firstly, Fourier transforms are performed on the features of MRI and PET images to obtain their amplitude and phase components. Then, they are integrated from the two modalities using two independent sets of modality-guided gating units. Considering that PET images typically have rich functional textures that are closely related to amplitude in the frequency domain, amplitude gating maps are generated using the amplitude of PET features. Similarly, since MRI images emphasize anatomical structures closely related to the phase component, phase gating maps are generated using the phase of MRI features. Specifically, the modality-guided gating units process the amplitude of PET features and the phase components of MRI images through 3×3 convolution and ReLU activation functions, respectively, before performing 3×3 convolution to improve nonlinear representation. Then, an amplitude gating map W is generated using the sigmoid function. A Phase-gated graph W P for:
[0076]
[0077] in, This represents the sigmoid function, and Con3 represents a 3×3 convolution. This represents the ReLU activation function.
[0078] Then use the output gating mapping W A and W P By fusing the amplitude and phase components from the two modes through element-weighted summation, the fused amplitude and phase representations, i.e., the fused amplitude components, are obtained. and fusion phase components
[0079]
[0080] Here, ⊙ represents element-wise multiplication.
[0081] The inverse discrete Fourier transform is applied to transform the fused amplitude and phase components back into the spatial domain:
[0082]
[0083] at last, Downsampled output of adjacent lower-level SFCF blocks Connect these elements to enrich the current layer with fine-grained cues, thereby generating a global information representation:
[0084]
[0085] in, This represents the inverse discrete Fourier transform.
[0086] To enhance the global information representation with detailed textures from multiple layers of features, a spatial compensation module is used in the thinning mechanism. Specifically, such as... Figure 4 As shown, the spatial compensation module calculates the absolute difference between MRI and PET features to capture spatial texture differences. Subsequently, a spatial attention unit is applied. This simulates spatial dependencies. The output attention map is then multiplied by a global information representation to compensate for complex textures, and the result is combined with the original input to preserve original details.
[0087]
[0088] Here, ⊙ represents element-wise multiplication.
[0089] To generate more information-rich feature representations, spatially compensated features are processed through channel attention units. Integrating with the original fusion features, this unit utilizes inter-channel relationships for complementary learning to obtain... and
[0090]
[0091] Then, and The convolutional layers, consisting of 3×3 convolutions, LeakyReLU activation, and pooling layers, are used to generate fused features.
[0092]
[0093] in, This indicates a processing function that includes 3×3 convolution, LeakyReLU activation, and pooling. This indicates a splicing operation.
[0094] Through bottom-up spatial compensation modules, the network can progressively achieve multimodal feature fusion at different scales. After obtaining the final fused features, they are input into the decoder to generate a fused Y-channel image. In this embodiment, the decoder consists of three upsampling operations and a convolutional layer. Finally, the fused Y-channel image is combined with the Cb and Cr channels of the original PET image to generate the final fused image in the YCbCr color space.
[0095] S40: Input the fusion features into the decoder. The decoder generates a fused Y-channel image through at least three upsampling operations and corresponding convolutional layer processing. The fused Y-channel image is then merged with the retained Cb and Cr channels in the YCbCr color space and converted to the target color space to obtain the fused medical image.
[0096] In this embodiment, please refer to Figure 2 As shown, a lightweight decoder can be used to convert the fused features back to the original image space to obtain the fusion result of the Y channel. This fusion result is then merged with the Cb and Cr channels of the original PET image in the channel dimension to obtain the final fused image. The original image space can be used as the target color image; for example, if the original image space is RGB, it is converted to RGB. In practical implementation, the target color space can be determined based on the image space of the desired output image. This embodiment is merely an example and does not constitute a limitation of this application.
[0097] S50: A medical image fusion model is formed by combining two image encoders and decoders bridged by at least two spatial-frequency co-fusion blocks. The medical image fusion model is trained based on a hybrid loss function and a medical dataset. The fusion of structural modal images and functional modal images is achieved based on the trained medical image fusion model.
[0098] It should be noted that the core data processing flow of the medical image fusion model has been described above and will not be repeated here. The training method can employ gradient descent, and the medical dataset used can be from publicly available datasets or datasets constructed from multiple medical images legally and compliantly obtained in medical settings. This embodiment focuses on designing a hybrid loss function to ensure that the fused image retains both structural information and texture details. Therefore, this invention proposes a hybrid loss function from the perspectives of texture and structure. It consists of two parts: strength loss and maximum gradient loss
[0099]
[0100] Where α is the weighting factor, which is set to 5 based on experience.
[0101] The method used to emphasize and highlight the target by maximizing the response intensity of the target region during the fusion process is determined by the following formula:
[0102]
[0103] Where max(,) represents the operation of extracting the maximum value of two images pixel by pixel, H represents the height of the image, W represents the width of the image, and I represents the height of the image. F I represents the fused image. V Represents structural modal images, I I Represents functional modality images;
[0104] The aim is to preserve the edge information present in the two source images, determined by the following formula:
[0105]
[0106] in, This represents the Sobel gradient operator.
[0107] The feasibility and progressiveness of the proposed method will be further illustrated below with specific calculation examples.
[0108] This embodiment uses the Harvard Medical School dataset to evaluate the performance of the proposed method, and divides the data into three fusion tasks based on modality: MRI-CT, MRI-PET, and MRI-SPECT. During the dataset segmentation phase, 147, 215, and 285 image pairs are randomly selected for training the MRI-CT, MRI-PET, and MRI-SPECT fusion tasks, respectively. The remaining 37, 54, and 72 image pairs for each task are used for testing. Considering limited computational resources, the training image size is adjusted to 96×96. During the inference phase, the query image pairs are directly input into the well-trained fusion model to generate the fused image.
[0109] The medical image fusion model constructed in this embodiment is implemented in the PyTorch deep learning framework and runs on the Ubuntu 18.04 operating system. The network model uses the Adam optimizer. During the training phase, the proposed network is optimized using the Adam optimizer with an initial learning rate of 0.0001, which is reduced by a factor of 0.5 every 5000 iterations. The network is trained for a total of 100,000 iterations with a batch size of 4.
[0110] This embodiment uses spatial frequency (SF), which is widely recognized in the field of medical image fusion, to measure network performance. Generally, a higher SF indicates better performance.
[0111] Experimental results show that the SF index of the medical image fusion model is 34.493 on the MRI-CT fusion task, 33.429 on the MRI-PET fusion task, and 19.763 on the MRI-SPECT fusion task.
[0112] Another aspect of this application provides a medical image fusion system based on spatial-frequency domain feature interaction analysis, such as... Figure 5 The diagram shown is a structural diagram of a medical image fusion system based on spatial-frequency domain feature interaction analysis provided in an embodiment of this application. The medical image fusion system based on spatial-frequency domain feature interaction analysis includes:
[0113] Image preprocessing module 501 is configured to acquire a pair of medical images to be fused, the medical image pair including a structural modality image and a functional modality image; convert the functional modality image to the YCbCr color space, extract the Y channel as the channel to be fused in the functional modality image, and retain the Cb channel and Cr channel;
[0114] The feature extraction module 502 is configured to use two Restormer-based image encoders to extract features from the structural modality image and the functional modality image to be fused channels, respectively, to obtain multi-layer structural modality features and multi-layer functional modality features to be fused channels; the two image encoders are bridged by at least two spatial-frequency co-fusion blocks, which include a frequency interaction module and a spatial compensation module;
[0115] Feature fusion module 503 is configured to perform feature fusion based on a space-frequency co-fusion block for each layer of structural modal features. and the channel features to be fused for each functional mode. The fusion process is performed to obtain the fusion characteristics;
[0116] The feature decoding module 504 is configured to input the fused features into the decoder. The decoder generates a fused Y-channel image through at least three upsampling operations and corresponding convolutional layer processing. The fused Y-channel image is then merged with the retained Cb and Cr channels in the YCbCr color space and converted to the target color space to obtain a fused medical image.
[0117] The model training module 505 is configured to combine two image encoders and decoders bridged by at least two spatial-frequency co-fusion blocks to form a medical image fusion model. The medical image fusion model is trained based on a hybrid loss function and a medical dataset. The fusion of structural modality images and functional modality images is achieved based on the trained medical image fusion model.
[0118] It should be noted that the medical image fusion device based on spatial-frequency domain feature interaction analysis provided in the above embodiments and the medical image fusion method based on spatial-frequency domain feature interaction analysis provided in the foregoing embodiments belong to the same concept. The specific way in which each module and unit performs operations has been described in detail in the method embodiments, and will not be repeated here.
[0119] Another aspect of this application provides an electronic device, including: a controller; and a memory for storing one or more programs, which, when executed by the controller, perform the methods described in the various embodiments above.
[0120] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.
[0121] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.
[0122] According to one aspect of the embodiments of this application, a computer system is also provided, including a central processing unit (CPU), which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from storage into random access memory (RAM), such as performing the methods described above. Various programs and data required for system operation are also stored in the RAM. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0123] For example, a computer system includes a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or loaded from storage into random access memory (RAM), such as executing the methods described in the above embodiments. The RAM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0124] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN (Local Area Network) cards and modems. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.
[0125] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs various functions defined in the system of this application.
[0126] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0128] The module units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.
[0129] The above embodiments are only used to illustrate this application and are not intended to limit this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of this application. Therefore, all equivalent technical solutions also fall within the scope of this application, and the patent protection scope of this application should be defined by the claims.
Claims
1. A medical image fusion method based on spatial-frequency domain feature interaction analysis, characterized in that, The methods include: Acquire medical image pairs to be fused, including structural modality images and functional modality images; The functional modality image is converted to the YCbCr color space, the Y channel is extracted as the channel to be fused in the functional modality image, and the Cb and Cr channels are retained. Two Restormer-based image encoders were used to extract features from the channels to be fused in the structural modality image and the functional modality image, respectively, to obtain multi-layer structural modality features and multi-layer functional modality channel features to be fused. Two image encoders are bridged by at least two spatial-frequency co-fusion blocks, which include a frequency interaction module and a spatial compensation module. Based on the space-frequency co-fusion block, the structural modal features of each layer are analyzed. and the channel features to be fused for each functional mode. The fusion process is performed to obtain the fusion characteristics; The fusion features are input into the decoder, which generates a fused Y-channel image through at least three upsampling operations and corresponding convolutional layer processing. The fused Y-channel image is then merged with the retained Cb and Cr channels in the YCbCr color space and converted to the target color space to obtain the fused medical image. A medical image fusion model is formed by combining two image encoders and decoders bridged by at least two spatial-frequency co-fusion blocks. The medical image fusion model is trained based on a hybrid loss function and a medical dataset. The fusion of structural modality images and functional modality images is achieved based on the trained medical image fusion model. Based on the space-frequency co-fusion block, the structural modal features of each layer are analyzed. and the channel features to be fused for each functional mode. The fusion process yields fusion features, including: Through the frequency interaction module and Perform discrete Fourier transforms on each, and obtain phase components , amplitude component and phase components , amplitude component Amplitude gating maps are generated based on modal guided gating units. and phase-gated graph ,in according to generate, according to Generate; utilize right and By performing element-wise weighted summation, the fused amplitude components are obtained. ;use right and Perform element-wise weighted summation to obtain the fused phase components. ;right and Performing the inverse discrete Fourier transform yields the frequency domain fused features. ;Will Downsampled output with adjacent lower-level spatial-frequency co-fusion block By concatenating the data, a global information representation can be obtained. ; Calculated using the space compensation module and The absolute difference is used to obtain the spatial texture difference; the spatial texture difference is then input into the spatial attention unit. Generate a spatial attention map; combine the spatial attention map with... Multiply element by element, then multiply by each element. , Adding them together yields the functional compensation features. and structural compensation features ;Will , Input channel attention units respectively The functional characteristics after channel optimization are obtained. and channel-optimized structural features ;Will and The concatenated data, after convolution, activation, and pooling processes, yields the fused features of the current layer. .
2. The method according to claim 1, characterized in that, Based on modal guided gating elements, amplitude gating maps are generated using the following formula. and phase-gated graph : in, This represents the sigmoid function. This represents a 3×3 convolution. Represents the ReLU activation function; use right and By performing element-wise weighted summation, the fused amplitude components are obtained. and utilization right and Perform element-wise weighted summation to obtain the fused phase components. The calculation process is expressed as follows: in, This indicates element-wise multiplication.
3. The method according to claim 1, characterized in that, right and Performing the inverse discrete Fourier transform yields the frequency domain fused features. The calculation process is expressed as follows: in, Indicates the inverse discrete Fourier transform; Will Downsampled output with adjacent lower-level spatial-frequency co-fusion block By concatenating the data, a global information representation can be obtained. The calculation process is expressed as follows: in, Represents the feature integration function. This indicates a splicing operation.
4. The method according to claim 1, characterized in that, The function compensation feature is calculated using the following formula. and structural compensation features : in, This indicates element-wise multiplication.
5. The method according to claim 3, characterized in that, The optimized functional characteristics of the channel are calculated using the following formula. and channel-optimized structural features : Will and The concatenated data, after convolution, activation, and pooling processes, yields the fused features of the current layer. The calculation process is expressed as follows: in, This indicates a processing function that includes 3×3 convolution, LeakyReLU activation, and pooling. This indicates a splicing operation.
6. The method according to any one of claims 1 to 5, characterized in that, The mixed loss function is expressed as: Among them, table Show the mixed loss function, Indicates strength loss. This represents the maximum gradient loss. Indicates the weighting factor; The strength loss is calculated using the following formula: in, This represents the operation of extracting the maximum value of two images pixel by pixel. Indicates the height of the image. Indicates the width of the image. This represents the merged image. Represents structural modal images. Represents functional modality images; The maximum gradient loss is calculated using the following formula: in, This represents the Sobel gradient operator.
7. A medical image fusion system based on spatial-frequency domain feature interaction analysis, characterized in that, The system includes: The image preprocessing module is configured to acquire medical image pairs to be fused, the medical image pairs including structural modality images and functional modality images; The functional modality image is converted to the YCbCr color space, the Y channel is extracted as the channel to be fused in the functional modality image, and the Cb and Cr channels are retained. The feature extraction module is configured to use two Restormer-based image encoders to extract features from the structural modality image and the functional modality image to be fused channels, respectively, to obtain multi-layer structural modality features and multi-layer functional modality features to be fused channels; the two image encoders are bridged by at least two spatial-frequency co-fusion blocks, which include a frequency interaction module and a spatial compensation module; The feature fusion module is configured to perform feature fusion based on a space-frequency co-fusion block for each layer of structural modal features. and the channel features to be fused for each functional mode. The fusion process yields fusion features, including: Through the frequency interaction module and Perform discrete Fourier transforms on each, and obtain phase components , amplitude component and phase components , amplitude component Amplitude gating maps are generated based on modal guided gating units. and phase-gated graph ,in according to generate, according to Generate; utilize right and By performing element-wise weighted summation, the fused amplitude components are obtained. ;use right and Perform element-wise weighted summation to obtain the fused phase components. ;right and Performing the inverse discrete Fourier transform yields the frequency domain fused features. ;Will Downsampled output with adjacent lower-level spatial-frequency co-fusion block By concatenating the data, a global information representation can be obtained. ; Calculated using the space compensation module and The absolute difference is used to obtain the spatial texture difference; the spatial texture difference is then input into the spatial attention unit. Generate a spatial attention map; combine the spatial attention map with... Multiply element by element, then multiply by each element. , Adding them together yields the functional compensation features. and structural compensation features ;Will , Input channel attention units respectively The functional characteristics after channel optimization are obtained. and channel-optimized structural features ;Will and The concatenated data, after convolution, activation, and pooling processes, yields the fused features of the current layer. ; The feature decoding module is configured to input the fused features into the decoder. The decoder generates a fused Y-channel image through at least three upsampling operations and corresponding convolutional layer processing. The fused Y-channel image is then merged with the retained Cb and Cr channels in the YCbCr color space and converted to the target color space to obtain the fused medical image. The model training module is configured to combine two image encoders and decoders bridged by at least two spatial-frequency co-fusion blocks to form a medical image fusion model. The medical image fusion model is trained based on a hybrid loss function and a medical dataset. The fusion of structural modality images and functional modality images is achieved based on the trained medical image fusion model.
8. An electronic device, characterized in that, Electronic devices include: Memory, used to store computer programs; A processor for executing a computer program to implement the method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing instructions, characterized in that, When the instructions are executed by the processor, the method according to any one of claims 1 to 6 is performed.