Image synthesis method and device based on multi-scale structural similarity and long-range dependence
By introducing self-attention mechanism and multi-scale structural similarity loss function in medical image synthesis, the problem that convolutional neural networks cannot use global information and synthesized images to blur in high-frequency areas is solved, which significantly improves the quality and edge clarity of image synthesis.
Patent Information
- Application Number
- CN202111404670.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The existing medical image synthesis method has the problem that convolutional neural networks cannot use global information and synthesize target images to blur in high-frequency areas.
By introducing a self-attention mechanism into a fully convolutional neural network, long-range dependence between image regions is established, and a hybrid loss function is designed to combine multi-scale structural similarity and L1 loss to improve the image synthesis quality of the network.
It effectively overcomes the disadvantage that convolutional neural networks cannot utilize global information, improves the ability of image synthesis networks to solve image long-range and multi-level dependencies, and significantly improves the edge clarity and visual effect of synthetic images.
Smart Images

Figure CN114187217B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and pattern recognition, and in particular to an image synthesis method and device based on multi-scale structural similarity and long-range dependence. Background Art
[0002] Medical imaging with different resolutions and modalities provides supplementary information about human tissues, which is of great reference value for clinical diagnosis. Medical image synthesis evaluates the quality of synthesized images by comparing the differences between synthesized images and real images. Existing medical image synthesis methods are mainly divided into three categories: 1) atlas registration-based methods; 2) learning-based methods; and 3) deep learning-based methods.
[0003] The effect of the atlas-based registration method depends on the quality of the original image data. This method is very sensitive to the registration accuracy and segmentation accuracy, and there may be deviations in the synthesis of some fine parts.
[0004] The main goal of learning-based methods is to establish a nonlinear mapping between the source image and the target image, including random forests, dictionary learning, sparse representation and other methods. However, this method requires extracting the features of the source image first and then establishing a mapping of the target image. Therefore, the quality of the generated image depends on the artificially constructed features and the quality of the source image based on the extracted features.
[0005] The deep learning method mainly uses convolutional neural network (CNN), but because the convolution operator has a local receptive field, a big defect of the CNN-based network is that it can only process long-distance dependencies after passing through multiple convolutional layers, but this may prevent long-term dependency learning for various reasons. Generally, this problem is solved by increasing the network depth or increasing the size of the convolution kernel to increase the network's representation ability. However, this will not only increase the network parameters and lead to low computational efficiency, but also increase the difficulty of network optimization when long-term dependencies gradually propagate over multiple layers. Summary of the invention
[0006] Based on adversarial learning technology, this paper proposes an image synthesis method based on multi-scale structural similarity and long-range dependency, aiming to solve the shortcomings of convolutional neural networks that cannot utilize global information and the blurring phenomenon of synthesized target images in high-frequency areas. The self-attention mechanism is introduced into the fully convolutional neural network to model the long-range dependency of images. By establishing a non-local model, the network's ability to solve long-range and multi-level dependencies of images is improved. The designed hybrid loss function is used to improve the clarity of edge details, thereby improving the quality of the synthesized image.
[0007] The image synthesis method comprises the following steps:
[0008] S1, obtaining an image training data set and a target image 1 corresponding to the image training data set;
[0009] S2, preprocessing the training image data to obtain a source image patch;
[0010] S3. Establish a generative adversarial network, set the number of network layers, the number of training batches, and the hybrid loss function, introduce a self-attention mechanism after the first layer of convolution operation, and obtain an image synthesis network with global information expression capabilities;
[0011] S4. Input the source image patch into the image synthesis network, and output the synthesis result of the target image one.
[0012] Furthermore, the specific steps to obtain an image synthesis network with global information expression capability are as follows:
[0013] S31. A fully convolutional neural network model is used to establish a generative adversarial network, which consists of a generator network and a discriminator network.
[0014] S32. Introduce the self-attention mechanism into the generator network to establish long-range dependencies between image regions, and thus obtain an image synthesis network with the ability to express global information.
[0015] Furthermore, the specific steps in step S31 are:
[0016] Input the source image patch into a generator network in a generative adversarial network, and output a second target image, wherein the second target image is an image with different resolution and modality;
[0017] Determine a mixed loss function in a generator network according to a pixel value of each pixel in the target image one and a corresponding pixel value in the target image two;
[0018] According to the hybrid loss function, the target image 2 is compared with the corresponding pixel points in the target image 1, and the corresponding hybrid loss function values are compared to obtain a generated target image when the hybrid loss function is minimized;
[0019] Obtaining the probability that the generated target image belongs to the target image one or the target image two through the discriminator network, performing an exponential operation or a logarithmic operation on the probability to obtain a discriminant loss to train the discriminator network, and using the discriminant loss as a generation loss value of the generator network, and further training the generator network to improve the performance of the generator network;
[0020] The training samples in the image training data set are input into the trained generator network batch by batch, and a new training result image is output. According to the new training result image, the generator gradient is obtained by back propagation, and the parameters of the generator network are updated according to the gradient to obtain a generative adversarial network that minimizes the hybrid loss function.
[0021] Furthermore, the generator network is a fully convolutional neural network.
[0022] Furthermore, the discriminator network is a convolutional neural network.
[0023] Furthermore, the hybrid loss function is a MS-SSIM+L1 hybrid loss function.
[0024] Furthermore, the expression of the hybrid loss function is:
[0025]
[0026] Among them, L Mix represents the mixed loss function, L MS_SSIM represents the multi-scale structural similarity loss function, L l1 represents the L1 loss function, α is a hyperparameter, indicating the allocation to the loss L MS_SSIM The weight size, The standard deviation is σ G The subscript G is the Gaussian distribution parameter, and M represents the corresponding scale of the target image at different sampling times.
[0027] An image synthesis device based on multi-scale structural similarity and long-range dependence, the image synthesis device comprising:
[0028] An acquisition module, used to acquire an image training data set and a target image 1 corresponding to the image training data set;
[0029] A preprocessing module, used for preprocessing the image training data set to obtain a source image patch;
[0030] The construction module is used to build a generative adversarial network and set the number of network layers, the number of training batches, and the hybrid loss function. The self-attention mechanism is introduced after the first layer of convolution operation to obtain an image synthesis network with global information expression capabilities.
[0031] A synthesis module is used to input the source image patch into the above-mentioned image synthesis network that integrates multi-scale structural similarity and long-range dependency, and output the synthesis result of the target image one.
[0032] Furthermore, in the construction module, the generative adversarial network consists of two parts: a generator network and a discriminator network. The generator network is a fully convolutional neural network, and the discriminator network is a convolutional neural network.
[0033] Furthermore, the construction module also includes a network training process, which is based on the training samples in the image training data set, and a generative adversarial network that minimizes the mixed loss function through a back-propagation algorithm in a supervised manner, introduces a self-attention mechanism, and obtains an image synthesis network with global information expression capabilities.
[0034] The technical solution provided by the present invention has the following beneficial effects: (1) The present invention overcomes the shortcoming that convolutional neural networks cannot utilize global information, introduces the self-attention mechanism into fully convolutional neural networks to model the long-range dependencies of images, and improves the network's ability to solve long-range and multi-level dependencies of images by establishing a non-local model.
[0035] (2) The present invention takes into account the property of multi-scale structural similarity that conforms to the human visual system, combines it with the mean absolute error, and designs a new hybrid loss function. This effectively improves the edge clarity of the synthesized target image without changing any network model structure, providing better visual effects.
[0036] (3) The present invention constructs a new network model to synthesize medical images without relying on manual annotation. The new method can make full use of the acquired source image data and synthesize the required target image data step by step in adversarial training. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0038] Figure 1 A flowchart for performing medical image synthesis based on multi-scale structural similarity and long-range dependency of the present invention;
[0039] Figure 2 This is a network structure diagram of medical image synthesis based on multi-scale structural similarity and long-range dependency of the present invention;
[0040] Figure 3 A sample schematic diagram of source image data (3T MRI or T1MRI) and target image data (7T MRI or T2 MRI) of the present invention;
[0041] Figure 4 is a schematic diagram of the long-range dependent expression module of the present invention;
[0042] Figure 5 This is a structural diagram of the medical image synthesis device that integrates multi-scale structural similarity and long-range dependence in the present invention. DETAILED DESCRIPTION
[0043] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, specific embodiments of the present invention are now described in detail with reference to the accompanying drawings.
[0044] Please refer to Figure 1 The present invention proposes a method and device for medical image synthesis by integrating multi-scale structural similarity and long-range dependency, comprising the following steps:
[0045] S1, obtaining an image training data set and a target image 1 corresponding to the image training data set;
[0046] S2, preprocessing the image data set to obtain source image patches; the preprocessing is to divide the source image in the image data set into a plurality of source image patches;
[0047] S3. Establish a generative adversarial network based on a fully convolutional neural network, set the number of network layers, the number of training batches, and the hybrid loss function, introduce a self-attention mechanism after the first layer of convolution operation, establish long-range dependencies between image regions, and obtain an image synthesis network with global information expression capabilities;
[0048] Step S3 specifically includes:
[0049] S31. A fully convolutional neural network model is used to establish a generative adversarial network, which consists of a generator network (G) and a discriminator network (D). The game between the discriminator and the generator is used to improve the performance of the G network and the D network.
[0050] Step S31 is specifically as follows:
[0051] The preprocessed source image patch is input into the generator network in the generative adversarial network, and the trained target image patch is output, wherein the target image patch is an image with different image resolution and modality generated based on the preprocessed source image patch, and these target image patches constitute target image 2. The source image patch is obtained by dividing the source image into several small blocks, each of which corresponds to a source image patch, and then these small blocks are sent to the network for training to avoid insufficient video memory when directly inputting the training due to the large size of the entire image, and thus the inability to train.
[0052] According to the pixel value of each pixel point in the target image one and the corresponding pixel value in the target image two, it is determined that the mixed loss function in the generator network is the MS-SSIM+L1 mixed loss function.
[0053] According to the hybrid loss function, the target image 2 generated by the generator network is compared with the corresponding pixel points in the target image 1, and the corresponding hybrid loss function values are compared to obtain the generated target image when the hybrid loss function is minimized.
[0054] The probability that the generated target image belongs to the target image one or the target image two is judged by the discriminator network, the probability is subjected to exponential operation or logarithmic operation to obtain the discriminant loss to train the discriminator network, and the discriminant loss is used as the generation loss value of the generator network to further train the generator network to improve the performance of the generator network.
[0055] The training samples in the image training data set are input into the trained generator network batch by batch, and a new training result image is output. According to the new training result image, the generator gradient is obtained by back propagation, and the parameters of the generator are updated according to the gradient to obtain a generative adversarial network that minimizes the hybrid loss function.
[0056] S32. Introduce the self-attention mechanism into the generator network G to obtain an image synthesis network with the ability to express global information.
[0057] Please refer to Figure 2 , use multi-scale structural similarity and long-range dependency to synthesize the adversarial network structure of 7T MRI, adopt adversarial learning strategy, fully convolutional neural network (FCN) as the generator network, convolutional neural network (CNN) as the discriminator of the network, G network and D network are trained simultaneously and constitute a dynamic "game process", and the entire network structure continuously improves the expression ability of the two networks through error back propagation.
[0058] Please refer to Figure 3 , the input to the generator network G is the source image patch (3T MRI or T1MRI), which is 5×64×64 in size, and the target image patch (7T MRI or T2 MRI) is estimated, which is 1× 64×64 in size. That is, the input 3T MRI image is used to obtain a 7T MRI image with higher resolution, or the input T1 MRI image is used to obtain a T2 For MRI images, the convolution kernel size of the generator network G is 3×3, the stride is 1, the padding is 1, the number of convolution layer filters is 64, 128, 256, and 512 respectively, and each convolution layer is followed by a batch normalization layer (BN) and a relu activation layer. The idea of residual learning is adopted, and the residual map in the last layer is directly learned by adding the input to the last layer through a long skip connection to solve the gradient vanishing problem and insufficient memory problem that may be caused by a too deep network. The self-attention mechanism fusion is implemented after the first layer of convolution operation. A self-attention module (Self-Attention) is added after the first layer of convolution operation to modulate the original response generated by the convolution kernel receptive field, that is, to establish a new global receptive field around the receptive field, and then send the new global receptive field to the subsequent convolution layer for more refined feature learning.
[0059] The discriminator network D adopts the CNN architecture in the neural network, including three convolutional layers with kernel sizes of 9×9, 5×5, and 5×5, a batch normalization layer, a relu activation layer, a maximum pooling layer, and a self-attention layer. A self-attention mechanism is added after the last convolution layer of the discriminator network to solve the long-range dependencies in the network. There are three fully connected layers. The number of convolutional layer filters of these three fully connected layers is 32, 64, and 512, respectively. The number of output nodes in the fully connected layer is 512, 64, and 1.
[0060] Please refer to Figure 4 , Figure 4 Schematic diagram of the long-range dependency expression module, using L in the generator network MS_SSIM +L l1 The mixed loss function is used to help the network update parameters. The final mixed loss function expression is shown in formula (1):
[0061]
[0062] Among them, L Mix represents the mixed loss function, L MS_SSIM represents the multi-scale structural similarity loss function, L l1 represents the L1 loss function, α is a hyperparameter, indicating the allocation to the loss L MS_SSIM The weight size, The standard deviation is σ G A Gaussian filter is constructed, where the subscript G is a Gaussian distribution parameter, M represents the scale of the target image 1 at different sampling times, M=1 represents the original size of the target image 1 at the first sampling, M=2 represents 1 / 2 of the original size of the target image 1 at the second sampling, and M=3 represents 1 / 4 of the original size of the target image 1 at the third sampling, and so on to obtain images with multiple resolutions.
[0063] The self-attention mechanism is used to establish long-range dependencies between images to improve the ability of neural networks to solve long-distance and multi-level dependencies between images. The core of self-attention is the transpose product operation of the matrix. The expression of the self-attention feature image is shown in formula (2):
[0064]
[0065] Among them, Att(o j ) represents the self-attention feature image, exp() represents the exponential function with the natural constant e as the base, i = 1, 2, ..., N, x i , x j denote the i-th pixel and j-th pixel of the input image respectively, h(xj ) represents the input signal at any position j, S(x i , x j )=f(x i ) T g(x j ), which represents the matrix f(x i ) and the matrix g(x j ) can represent the correlation weight between the current position i and any position j in the image, ∑ j exp(S(x i , x j )) represents j S(x i , x j ) after exp().
[0066] The feature map after the convolution operation is obtained by three 1×1×1 convolutions to obtain three matrices f(x i )=W f x i , g(x j )=W g x j and h(x j )=W h x j , W f ∈R C×C , W g ∈R C×C , W h ∈R C×C is the weight matrix obtained during training.
[0067] S(x i , x j ) is passed through the softmax function to normalize the original weights. The softmax function formula is shown in (3):
[0068]
[0069] As learning changes during training to allow the network to gradually learn more non-local cues from local cues, the final image matrix z i It contains both the information of the original image (i.e., target image 1) and the relationship between it and the pixels of other surrounding images. Its expression is shown in formula (4):
[0070] z i =α1·Att(o j )+x i (4)
[0071] Among them, α1 is a hyperparameter, which is set to 0 at the beginning of training.
[0072] Multi-scale structural similarity is integrated into the neural network to help network training. Structural similarity measures the similarity between two images by using the mean as an estimate of brightness (L(x, y)), the standard deviation as an estimate of contrast (C(x, y)), and the covariance as a measure of structural similarity (S(x, y)). The product of the three is defined as structural similarity, that is, structural similarity SSIM(x, y) = [αL(x, y)*βC(x, y)*γS(x, y)], where a, β, and γ are used to adjust the weights of each component. After setting α, β, and γ to 1, the following formula (5) is obtained:
[0073]
[0074] where u x is the mean of the target image x, u y is the mean value of the target image y, is the variance of x, is the variance of y, σ xy is the covariance of x and y, c1, c2, c3 are all constants.
[0075] Constantly increase the width and height of target image 1 x and target image 2 y by 2 M-1 The image is downsampled (reduced) by a factor, M represents the different scales of the target image 1 at different sampling times, M=1 represents the original size of the target image 1 at the first sampling, M=2 represents 1 / 2 of the original size of the target image 1 at the second sampling, and M=3 represents 1 / 4 of the original size of the target image 1 at the third sampling, and so on to obtain images with multiple resolutions, and then these images are evaluated by SSIM in turn, and the multi-scale structural similarity is obtained by fusing these SSIMs into one value in some way, L M , C J , S J They correspond to the brightness, contrast, and structural similarity in SSIM. For simplicity, α can be J =β J =γ J , The specific formula is shown in (6):
[0076]
[0077] in, α J , β J , γ J Used to adjust L M , C J , S J The weight of .
[0078] S4. Input the source image (3T MRI or T1 MRI) into the above-mentioned medical image synthesis network that integrates multi-scale structural similarity and long-range dependency, and output the synthesis result of the target image (7T MRI or T2 MRI).
[0079] refer to Figure 5 In order to implement the method for synthesizing medical images by fusing multi-scale structural similarity and long-range dependency, a specific embodiment of the present invention further provides a medical image synthesis device fusing multi-scale structural similarity and long-range dependency, the image synthesis device comprising:
[0080] An acquisition module, used to acquire an image training data set and a target image 1 corresponding to the image training data set;
[0081] A preprocessing module, used for preprocessing the image training data set to obtain a source image patch;
[0082] A construction module is used to establish a generative adversarial network, set the number of network layers, the number of training batches and the hybrid loss function, introduce a self-attention mechanism after the first layer of convolution operation, and obtain an image synthesis network with global information expression capability; the generative adversarial network consists of a generator network and a discriminator network, the generator network is a fully convolutional neural network, and the discriminator network is a convolutional neural network; the construction module also includes a network training process, based on the training samples in the image training data set, the generative adversarial network when the hybrid loss function is minimized by the back propagation algorithm, introduces a self-attention mechanism, and obtains an image synthesis network with global information expression capability;
[0083] The synthesis module inputs the source image (3T MRI or T1MRI) into the medical image synthesis network that integrates multi-scale structural similarity and long-range dependency, and outputs the synthesis result of the target image (7T MRI or T2 MRI).
[0084] The present invention establishes a fully convolutional network structure with the ability to model long-range dependencies, and integrates the self-attention mechanism into it to solve the weakness of traditional convolutional networks that can only perceive local information; through an adversarial learning strategy, the nonlinear mapping from source images to target images is better modeled; a new multi-scale structural similarity loss function is used to help us better update the network model parameters while keeping the image edges clear; long residual units and automatic context models are used to further improve the learning efficiency of image global information and improve the synthesis quality of the target image.
[0085] The examples of the present invention are described above in conjunction with the accompanying drawings, but the present invention does not take precedence over the above-mentioned specific embodiments. The above-mentioned specific embodiments are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, all of which are within the protection of the present invention.
Claims
1. An image synthesis method based on multi-scale structural similarity and long-range dependence, characterized in that: The image synthesis method comprises the following steps S1, obtaining an image training data set and a target image 1 corresponding to the image training data set; S2, preprocessing the training image data to obtain a source image patch; S3. Establish a generative adversarial network, set the number of network layers, the number of training batches, and the hybrid loss function, introduce the self-attention mechanism after the first layer of convolution operation, and obtain an image synthesis network with global information expression capabilities; the specific steps are as follows: S31. Establish a generative adversarial network, which consists of a generator network and a discriminator network. The specific steps are: Input the source image patch into a generator network in a generative adversarial network, and output a second target image, wherein the second target image is an image with different resolution and modality; Determine a mixed loss function in a generator network according to a pixel value of each pixel in the target image one and a corresponding pixel value in the target image two; According to the hybrid loss function, the target image 2 is compared with the corresponding pixel points in the target image 1, and the corresponding hybrid loss function values are compared to obtain a generated target image when the hybrid loss function is minimized; Obtaining the probability that the generated target image belongs to the target image one or the target image two through the discriminator network, performing an exponential operation or a logarithmic operation on the probability to obtain a discriminant loss to train the discriminator network, and using the discriminant loss as a generation loss value of the generator network, and further training the generator network to improve the performance of the generator network; Inputting the training samples in the image training data set into the trained generator network batch by batch, outputting new training result images, obtaining the generator gradient by back propagation according to the new training result images, updating the parameters of the generator network according to the gradient, and obtaining a generative adversarial network when minimizing the hybrid loss function; S32. Introduce the self-attention mechanism in the generator network to establish long-range dependencies between image regions, thereby obtaining an image synthesis network with global information expression capabilities; S4. Input the source image patch into the image synthesis network, and output the synthesis result of the target image one.
2. The image synthesis method based on multi-scale structural similarity and long-range dependence according to claim 1, characterized in that: The generator network is a fully convolutional neural network.
3. The image synthesis method based on multi-scale structural similarity and long-range dependence according to claim 1, characterized in that: The discriminator network is a convolutional neural network.
4. The image synthesis method based on multi-scale structural similarity and long-range dependence according to claim 1, characterized in that: The hybrid loss function is a MS-SSIM+L1 hybrid loss function.
5. The image synthesis method based on multi-scale structural similarity and long-range dependence according to claim 1, characterized in that: The expression of the mixed loss function is: in, represents the mixed loss function, represents the multi-scale structural similarity loss function, express L 1 loss function, is a hyperparameter, indicating the allocation to loss The weight size, The standard deviation is Gaussian filter, subscript G is the Gaussian distribution parameter, and M represents the corresponding scale of the target image at different sampling times.
6. An image synthesis device based on multi-scale structural similarity and long-range dependence, the device implementing the image synthesis method based on multi-scale structural similarity and long-range dependence as claimed in any one of claims 1 to 5, characterized in that: The image synthesis device comprises: An acquisition module, used to acquire an image training data set and a target image 1 corresponding to the image training data set; A preprocessing module, used for preprocessing the image training data set to obtain a source image patch; The construction module is used to build a generative adversarial network and set the number of network layers, the number of training batches, and the hybrid loss function. The self-attention mechanism is introduced after the first layer of convolution operation to obtain an image synthesis network with global information expression capabilities. The synthesis module is used to input the source image patch into the above-mentioned multi-scale structural similarity and long-range dependency image synthesis network, and output the synthesis result of the target image one.
7. The image synthesis device based on multi-scale structural similarity and long-range dependence according to claim 6, characterized in that: In the building module, the generative adversarial network consists of two parts: a generator network and a discriminator network. The generator network is a fully convolutional neural network, and the discriminator network is a convolutional neural network.
8. The image synthesis device based on multi-scale structural similarity and long-range dependence according to claim 6, characterized in that: The construction module also includes a network training process, which is based on the training samples in the image training data set, and a generative adversarial network that minimizes the mixed loss function in a supervised manner through a back-propagation algorithm, introduces a self-attention mechanism, and obtains an image synthesis network with global information expression capabilities.