Cross-modal medical image synthesis method, system, terminal and storage medium

By constructing a generative adversarial network model, learning and generating feature mapping relationships across modal medical images, the problem that cross-modal medical image synthesis results in the prior art cannot effectively represent human tissue edge information, achieving higher signal-to-noise ratio and clearer edges.

CN114240753BActive Publication Date: 2025-06-27SHENZHEN PING AN MEDICAL HEALTH TECHNOLOGY SERVICES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111551447.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-06-27
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

The existing cross-modal medical image synthesis method cannot effectively represent the edge information of human tissues with limited paired data, resulting in low signal-to-noise ratio and blurred edges.

Method used

A generative adversarial network model including a generator and a discriminator is constructed. The generator learns the feature mapping relationship between the first modal medical image and the second modal medical image, generates the synthesized second modal medical image, and distinguishes true and false through the discriminator, and constructs different loss functions for image synthesis training.

Benefits of technology

By adding pairs of data, the generated synthetic images can better represent edge information of human tissue, improve signal-to-noise ratio, reduce edge blur, and obtain more reliable multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114240753B_ABST
    Figure CN114240753B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross-modal medical image synthesis method, system, terminal and storage medium. The method includes constructing a generative adversarial network model including a generator and a discriminator; wherein, the generator takes a real first-modal medical image as input, learns the feature mapping relationship between the real first-modal medical image and the real second-modal medical image, and generates a synthesized second-modal medical image according to the feature mapping relationship, and then splices the synthesized second-modal medical image and the real second-modal medical image with the real first-modal medical image respectively to output a first image pair and a second image pair; the discriminator takes the first image pair and the second image pair as input, respectively performs true / false discrimination on them, and outputs a true / false discrimination result; different loss functions are constructed for the generator and the discriminator respectively to perform image synthesis training on the generative adversarial network model. This method can generate multi-modal data with high reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of medical image processing, and particularly relates to a cross-modal medical image synthesis method, system, terminal, and storage medium. Background Art

[0002] With the development of science and technology, there are various ways to obtain medical images, and different modalities of medical images have different advantages and disadvantages. For example, Magnetic Resonance Imaging (MRI) has no radiation to the human body, shows soft tissue structures clearly, and can obtain rich diagnostic information, but the acquisition time is long and artifacts are easily generated; Positron Emission Computed Tomography (PET) can make early diagnoses of diseases through functional changes in tissues in the lesion area, but it is expensive and has low image resolution. Studies have shown that morphological or functional abnormalities caused by diseases in the human body are often manifested in various aspects, and the information obtained by a single-modal imaging device often cannot comprehensively reflect the complex characteristics of diseases. Clinically, collecting different modalities of medical images simultaneously requires a large amount of time and financial resources. Therefore, how to use existing modalities of medical images to accurately synthesize the required modality of images through computer technology has been a research direction in recent years.

[0003] Although existing cross-modal synthesis methods have achieved good results, due to the complex spatial structure of medical images, these synthesis results still cannot well represent the edge information of human tissues, and there are problems such as low signal-to-noise ratio and blurred edges. This reduces the synthesis effect of multi-modal images of specific subjects in the case of limited paired data.

[0004] Therefore, how to improve the synthesis effect of multi-modal images of specific subjects in the case of limited paired data is an urgent problem to be solved. Summary of the Invention

[0005] Based on this, in view of the above problems, it is necessary to provide a cross-modal medical image synthesis method, including:

[0006] Constructing a generative adversarial network model including a generator and a discriminator; wherein, the generator takes a real first-modal medical image as input, learns the feature mapping relationship between the real first-modal medical image and the real second-modal medical image, and generates a synthesized second-modal medical image according to this feature mapping relationship. Then, the synthesized second-modal medical image and the real second-modal medical image are respectively spliced with the real first-modal medical image to output a first image pair and a second image pair;

[0007] The discriminator takes the first image pair and the second image pair as inputs, discriminates their authenticity respectively, and outputs the authenticity discrimination results; and

[0008] Construct different loss functions for the generator and the discriminator respectively to perform image synthesis training on the generative adversarial network model.

[0009] The above cross-modal medical image synthesis method first constructs a generative adversarial network model including a generator and a discriminator, then controls the generator to map the features between the first-modal medical image and the second-modal medical image, and generates a synthesized second-modal medical image according to the feature mapping relationship. Then, the synthesized second-modal medical image and the real second-modal medical image are respectively stitched with the real first-modal medical image to output a first image pair and a second image pair. On the one hand, it can increase paired data and make the synthesis effect better. On the other hand, since the generated synthesized second-modal medical image is obtained according to the feature mapping relationship, and at the same time, the first image pair and the second image pair also pass through the discriminator for authenticity discrimination, and different loss functions are respectively constructed for the generator and the discriminator to perform image synthesis training on the generative adversarial model. Therefore, the final obtained result can be made more reliable. That is to say, the synthesis method of the present application is constructed based on 3D CGAN, which can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, and there are problems such as low signal-to-noise ratio and blurred edges.

[0010] In a possible embodiment, the generator adopts a U-Net network structure, which includes an encoder and a decoder with symmetric network structures;

[0011] The generation step of the synthesized second-modal medical image includes:

[0012] Through the feature extraction operation of multi-layer convolution of the encoder, the feature map of the real first-modal medical image is output;

[0013] The decoder performs multi-layer deconvolution operations on the feature map output by the encoder, and performs multiple stitching operations on the generated feature map and the feature map of the same size at the corresponding position of the encoder, and finally outputs the target reconstructed image, which is the synthesized second-modal medical image.

[0014] In a possible embodiment, the synthesis method further includes:

[0015] Regarding each pixel in the feature map as a random variable, calculate the paired covariance between the pixels;

[0016] Selectively enhance or weaken the value of each pixel according to the calculated paired covariance.

[0017] In a possible embodiment, the encoder includes a convolutional module layer, a batch normalization layer, and an activation layer;

[0018] Among them, the number of the convolutional module layers is seven, and the second to fifth convolutional module layers are hybrid dilated convolutional module layers, and the rest are fully convolutional layers.

[0019] In a possible embodiment, the hybrid dilated convolutional module layer includes six 3D convolutional layers of 3×3×3, and the dilation rate of each 3D convolutional layer is set to a zigzag structure;

[0020] Each of the convolutional layers is respectively denoted as convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, convolutional layer 5, and convolutional layer 6; among them, convolutional layer 1 and convolutional layer 3, convolutional layer 2 and convolutional layer 5, and convolutional layer 4 and convolutional layer 6 are respectively connected through a residual structure;

[0021] The dilation rates of convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, convolutional layer 5, and convolutional layer 6 are 1, 2, 5, 1, 2, and 5 respectively.

[0022] In a possible embodiment, the discriminator includes six convolutional layers, a batch normalization layer, and an activation layer; among them, each of the convolutional layers is respectively denoted as convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, convolutional layer 5, and convolutional layer 6; among them, convolutional layer 1 and convolutional layer 4, and convolutional layer 2 and convolutional layer 6 are respectively connected through a residual structure.

[0023] In a possible embodiment, the real first-modal medical image includes a CT image or an MRI image; the synthesized second-modal medical image includes a SPECT image or a PET image.

[0024] In a possible embodiment, the loss function MAE of the generator is set as:

[0025]

[0026] Among them, m is the model batch size, y i is the true value, is the predicted value.

[0027] In a possible embodiment, the loss function MSE of the discriminator is set as:

[0028]

[0029] Among them, m is the model batch size, y i is the true value, is the predicted value.

[0030] Based on the same inventive concept, this application also provides a cross-modal medical image synthesis system, including:

[0031] A model construction module configured to construct a generative adversarial network model including a generator and a discriminator; wherein, the generator takes a real first-modal medical image as input, learns the feature mapping relationship between the real first-modal medical image and the real second-modal medical image, and generates a synthetic second-modal medical image according to this feature mapping relationship. Then, the synthetic second-modal medical image and the real second-modal medical image are respectively stitched with the real first-modal medical image to output a first image pair and a second image pair;

[0032] The discriminator takes the first image pair and the second image pair as input, respectively discriminates their authenticity, and outputs the authenticity discrimination result;

[0033] A model training module for respectively constructing different loss functions for the generator and the discriminator to perform image synthesis training on the generative adversarial network model.

[0034] In the above cross-modal medical image synthesis system, by setting the model construction module to construct a generative adversarial network model including a generator and a discriminator, then controlling the generator to learn the feature mapping relationship between the first-modal medical image and the second-modal medical image, and generating a synthetic second-modal medical image according to this feature mapping relationship. Then, the synthetic second-modal medical image and the real second-modal medical image are respectively stitched with the real first-modal medical image to output a first image pair and a second image pair. On the one hand, it can increase paired data and make the synthesis effect better; on the other hand, since the generated synthetic second-modal medical image is obtained according to the feature mapping relationship, and at the same time, the first image pair and the second image pair are also discriminated for authenticity through the discriminator. Then, by setting the model training module to respectively construct different loss functions for the generator and the discriminator to perform image synthesis training on the generative adversarial model, the final obtained result can be made more reliable. That is to say, the synthesis system of this application is constructed based on 3D CGAN, which can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, such as low signal-to-noise ratio and blurred edges.

[0035] Based on the same inventive concept, this application also provides a terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method described in any one of the foregoing items.

[0036] The above-mentioned terminal, since it has a processor that can be used to execute the foregoing cross-modal medical image synthesis method, therefore, the beneficial effects produced by this method naturally apply to the terminal of the present application.

[0037] Based on the same inventive concept, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can be used to execute any one of the foregoing methods.

[0038] The above-mentioned computer-readable storage medium, since the computer program stored thereon can be used to execute the foregoing cross-modal medical image synthesis method when executed by a processor, therefore, the beneficial effects produced by this method naturally apply to the computer-readable storage medium of the present application. Description of the Drawings

[0039] Figure 1 It is a schematic flowchart of the cross-modal medical image synthesis method in an embodiment;

[0040] Figure 2 It is a framework diagram of the generative adversarial network model in an embodiment;

[0041] Figure 3 For Figure 2 the model framework diagram of the generator part in;

[0042] Figure 4 For Figure 3 the schematic diagram of the sub-structure of structure 212 in;

[0043] Figure 5 For Figure 2 the model framework diagram of the discriminator part in;

[0044] Figure 6 It is a schematic diagram of the modules of the cross-modal medical image synthesis system in an embodiment. Detailed Embodiments

[0045] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided so that the disclosure of the present invention can be understood more thoroughly and comprehensively.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0047] With the development of science and technology, there are various ways to acquire medical images, and different modalities of medical images have different advantages and disadvantages. For example, Magnetic Resonance Imaging (MRI) has no radiation to the human body, shows soft tissue structures clearly, and can obtain rich diagnostic information, but the acquisition time is long and artifacts are easily generated; Positron Emission Computed Tomography (PET) can make early diagnoses of diseases based on the functional changes of tissues in the lesion area, but it is expensive and has low image resolution. Research has shown that the morphological or functional abnormalities caused by diseases in the human body are often manifested in various aspects, and the information obtained by a single-modal imaging device often cannot comprehensively reflect the complex characteristics of diseases. Moreover, clinically collecting medical images of different modalities simultaneously requires a large amount of time and financial resources. Therefore, how to use existing modalities of medical images and accurately synthesize the required modality of images through computer technology has been a research direction in recent years.

[0048] Although existing cross-modal synthesis methods have achieved good results, due to the complex spatial structure of medical images, these synthesis results still cannot well represent the edge information of human tissues, and there are problems such as low signal-to-noise ratio and blurred edges. This reduces the synthesis effect of multi-modal images of specific subjects in the case of limited paired data.

[0049] Based on this, this application hopes to provide a new solution to solve the technical problems described above, and its specific composition will be elaborated in detail in the subsequent embodiments.

[0050] According to the first aspect of the present invention, as Figure 1 shown, this application provides a cross-modal medical image synthesis method, which may include steps S100 - S300.

[0051] Step S100, constructing a generative adversarial network model including a generator and a discriminator; wherein, the generator takes a real first-modal medical image as input, learns the feature mapping relationship between the real first-modal medical image and the real second-modal medical image, and generates a synthesized second-modal medical image according to this feature mapping relationship. Then, the synthesized second-modal medical image and the real second-modal medical image are respectively spliced with the real first-modal medical image to output a first image pair and a second image pair;

[0052] Step S200, the discriminator takes the first image pair and the second image pair as input, respectively makes true or false discrimination on them, and outputs the true or false discrimination results; and

[0053] Step S300: Construct different loss functions for the generator and the discriminator respectively to perform image synthesis training on the generative adversarial network model.

[0054] The above cross-modal medical image synthesis method first constructs a generative adversarial network model including a generator and a discriminator, then controls the generator to map the features between the first-modal medical image and the second-modal medical image, and generates a synthesized second-modal medical image according to this feature mapping relationship. Then, the synthesized second-modal medical image and the real second-modal medical image are respectively stitched with the real first-modal medical image to output a first image pair and a second image pair. On the one hand, it can increase paired data and make the synthesis effect better; on the other hand, since the generated synthesized second-modal medical image is obtained according to the feature mapping relationship, and at the same time, the first image pair and the second image pair are also discriminated as true or false through the discriminator, and different loss functions are respectively constructed for the generator and the discriminator to perform image synthesis training on the generative adversarial model. Therefore, the finally obtained result can be more reliable. That is to say, the synthesis method of the present application is constructed based on 3D CGAN, which can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data, so as to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, and there are problems such as low signal-to-noise ratio and blurred edges.

[0055] Specifically, reference can be made to Figure 2 , the real first-modal medical image SP1 of the present application can be image data of size 256*256*256*3. For the convenience of description, the real first-modal medical image SP1 will be simplified to the first-modal medical image SP1 hereinafter. The first-modal medical image SP1 can include CT images or MRI images. At the same time, the first-modal medical image SP1 can be selected from a first-modal medical image set, which is a set of images obtained by collecting images of multiple reference objects in the first modality. For example, the first modality can be MRI, and the multiple reference objects can be a certain organ of multiple people, such as the hearts of multiple people. In this case, the first-modal medical image set is a set of multiple cardiac MRI images obtained by collecting the hearts of multiple people using MRI. The above describes the first modality by taking MRI as an example and multiple reference objects by taking the hearts of multiple people as an example, but it should be understood that the present disclosure is not limited thereto. The first modality can also be various other modalities such as CT, PET, SPECT, etc., and the multiple reference objects can also be various other reference objects such as the kidneys of multiple people, the bones of multiple people, etc.

[0056] Further, the true second-modal medical image SP2 of the present application can also be image data of size 256*256*256*3. For ease of description, the true second-modal medical image SP2 will be simplified to the second-modal medical image SP2 hereinafter. The second-modal medical image SP2 can include SPECT images or PET images. At the same time, the second-modal medical image SP2 can be selected from a second-modal medical image set, which is a set of images obtained by collecting images of multiple reference objects in the second modality. For example, the second modality can be PET, and the multiple reference objects can be a certain organ of multiple individuals, such as the hearts of multiple individuals. In this case, the second-modal medical image set is a set of multiple cardiac PET images obtained by collecting the hearts of multiple individuals using PET. The above describes the second modality by taking PET as an example and multiple reference objects by taking the hearts of multiple individuals as an example. However, it should be understood that the present disclosure is not limited thereto. The second modality can also be various other modalities such as CT, MRI, SPECT, etc., and the multiple reference objects can also be various other reference objects such as the kidneys of multiple individuals, the bones of multiple individuals, etc. It can be understood that no matter what modalities the first-modal medical image SP1 and the second-modal medical image SP2 are specifically selected, they should both be image data obtained for the same sample.

[0057] For ease of description, in the following embodiments, the first-modal medical image SP1 is an MRI image, and the second-modal medical image SP2 is a PET image for illustration.

[0058] In a possible embodiment, refer to Figure 2 、 Figure 3 and Figure 4 continuously. The generator 20 of the present application can adopt a U-Net network structure, which is a 3D structure. It can include an encoder 210 and a decoder 220 with a symmetric network structure. The U-Net model is designed based on a fully convolutional network with skip connections. Its main idea is to design an encoder and a decoder with a symmetric network structure so that they have the same number and size of feature maps, and combine the corresponding feature maps of the encoder and the decoder through skip connections, which can maximize the retention of feature information during the downsampling process, thereby improving the efficiency of feature expression. The MRI image and the PET image come from the same sample, and they share a large amount of primary feature information. Therefore, the U-Net model is very suitable for complex feature mapping between two-modal images.

[0059] Further, the generation steps of the synthesized second-modal medical image SS can include the following sub-steps:

[0060] Output the feature map of the true first-modal medical image through the feature extraction operation of multi-layer convolution of the encoder;

[0061] The decoder performs multi-layer deconvolution operations on the feature map output by the encoder, and performs multiple splicing operations on the generated feature map and the feature map of the same size at the corresponding position of the encoder, and finally outputs the target reconstructed image, that is, the synthesized second-modal medical image SS.

[0062] Specifically, the synthesized second-modal medical image SS can be a SPEC image or a PET image. At the same time, the synthesized second-modal medical image SS should be of the same type and for the same object as the real second-modal medical image SP2.

[0063] Furthermore, the aforementioned first image pair can be formed by splicing the first-modal medical image SP1 and the second-modal medical image SP2, and the second image pair can be formed by splicing the first-modal medical image SP1 and the synthesized second-modal medical image SS. It can be understood that the composition of the first image pair and the second image pair can also be exchanged, which will not be elaborated here.

[0064] In a possible embodiment, reference can be made to Figure 3 , the encoder 210 may include a convolutional module layer, a batch normalization layer (not shown in the figure) and an activation layer (not shown in the figure); among them, the batch normalization layer is also denoted as BN, and the activation layer is also denoted as ReLU. It can be understood that for the batch normalization layer BN and the activation layer ReLU, reference can be made to the description of the prior art, which is not the focus of this application and will not be elaborated further here.

[0065] Among them, the number of the convolutional module layers is seven, and the seven convolutional module layers are respectively denoted as convolutional module layer 211, convolutional module layer 212, convolutional module layer 213, convolutional module layer 214, convolutional module layer 215, convolutional module layer 216 and convolutional module layer 217. And the second to fifth of the convolutional module layers are hybrid dilated convolutional module layers, that is Figure 2 in, the convolutional module layers 212, 213, 214, 215 with thickened borders, and the rest are ordinary 3×3×3 3D fully convolutional layers.

[0066] In a possible embodiment, auxiliary reference can be made to Figure 4 , taking the hybrid dilated convolutional layer 212 as an example, the hybrid dilated convolutional module layer 212 may include 6 3×3×3 3D convolutional layers, and the dilation rate of each 3D convolutional layer is set to a zigzag structure.

[0067] Specifically, each of the convolutional layers is denoted as convolutional layer 1 (2121), convolutional layer 2 (2122), convolutional layer 3 (2123), convolutional layer 4 (2124), convolutional layer 5 (2125), and convolutional layer 6 (2126); among them, convolutional layer 1 (2121) and convolutional layer 3 (2123), convolutional layer 2 (2122) and convolutional layer 5 (2125), and convolutional layer 4 (2124) and convolutional layer 6 (2126) are respectively connected through residual structures;

[0068] The dilation rates of convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, convolutional layer 5, and convolutional layer 6 are 1, 2, 5, 1, 2, and 5 respectively.

[0069] In this application, the hybrid dilated convolution module layer is only placed in the middle four layers of the encoder 210, which not only avoids the network grid problem, but also reduces the network parameters and training time to a certain extent while ensuring sufficient extraction of feature information to improve the generation quality.

[0070] In a possible embodiment, continue to refer to Figure 4 , the decoder 220 mainly reconstructs the final output from the feature maps compressed by the encoder 210. According to the foregoing description, the network structures of the decoder 220 and the encoder 210 are symmetric. Therefore, as shown in the figure, the decoder 220 of this application also consists of 7 transposed convolution module layers, a batch normalization layer (not shown in the figure), and an activation layer (not shown in the figure). Among them, the batch normalization layer is also denoted as BN, and the activation layer is also denoted as ReLU. It can be understood that for the batch normalization layer BN and the activation layer ReLU, reference can be made to the description of the prior art, which is not the focus of this application and will not be further elaborated here.

[0071] Furthermore, the seven transposed convolution module layers are respectively denoted as transposed convolution module layer 221, transposed convolution module layer 222, transposed convolution module layer 223, transposed convolution module layer 224, transposed convolution module layer 225, transposed convolution module layer 226, and transposed convolution module layer 227; each transposed convolution module layer is composed of 3 2*2*2 3D convolutional layers. At the same time, transposed convolution module layer 221 and transposed convolution module layer 223 are connected through a residual structure.

[0072] In a possible embodiment, continue to refer to Figure 5 , the discriminator 220 is an 8-layer 3D fully convolutional network, which may include 6 convolutional layers, a batch normalization layer (317), and an activation layer (318); among them, each of the convolutional layers is denoted as convolutional layer 1 (311), convolutional layer 2 (312), convolutional layer 3 (313), convolutional layer 4 (314), convolutional layer 5 (315), and convolutional layer 6 (316); among them, convolutional layer 1 (311) and convolutional layer 4 (314), and convolutional layer 2 (312) and convolutional layer 6 (316) are respectively connected through residual structures.

[0073] Further, as Figure 5 shown, the convolutional layer 6 adopts global average pooling (GAP), the seventh layer adopts a 1×1×1 convolutional kernel, and finally the Sigmoid activation function is used to determine whether the first image pair and the second image pair output by the generator belong to real images or generated images. It can be understood that the principle of the discriminator for discriminating the authenticity of the input images can be understood with reference to the prior art, and the present application will not elaborate further herein.

[0074] Furthermore, to prevent network overfitting, the present application also adds a dropout operation after the ReLu activation layer in the generator 20, and the related value is set to 0.5. Finally, the encoded and decoded feature information is used to obtain the synthesized second-modal medical image (PET image) through the Tanh activation function.

[0075] Specifically, the manner in which the generator 20 synthesizes the real MRI image into the corresponding PET image through the encoder 210 and the decoder 220 can be understood with reference to Figure 3 and Figure 4 and will not be elaborated herein.

[0076] In a possible embodiment, to further improve the synthesis quality, the synthesis method of the present application may further include:

[0077] Regarding each pixel in the feature map as a random variable, calculating the pairwise covariance between the pixels;

[0078] Selectively enhancing or weakening the value of each pixel according to the calculated pairwise covariance.

[0079] That is to say, the present application introduces a self-attention mechanism between the encoder 210 and the decoder 220. The self-attention mechanism regards each pixel in the feature map as a random variable, calculates the pairwise covariance between all pixels, and enhances or weakens the value of each predicted pixel according to the similarity between each predicted pixel and other pixels in the image, that is, selectively amplifies the more valuable feature channels and suppresses the useless feature channels, that is, amplifies the weights of relevant features and suppresses the weights of irrelevant features. Further eliminating the interference brought by irrelevant features and noise in the skip connection, highlighting the key features in the residual structure connection, so as to better capture the key information of the MRI image.

[0080] In a possible embodiment, the loss function MAE of the generator is set as:

[0081]

[0082] Among them, m is the model batch size, y i is the true value, is the predicted value.

[0083] In a possible embodiment, the loss function MSE of the discriminator is set to:

[0084]

[0085] Among them, m is the model batch size, y i is the true value, is the predicted value.

[0086] Using the loss function to train the generative adversarial network model can improve the quality of synthetic training.

[0087] According to the second aspect of the present application, reference may be made to Figure 6 , the present application also provides a cross-modal medical image synthesis system, which may include a model construction module 2 and a model training module 3. The image data input by the input unit 1 is transmitted to the model construction module 2 and the model training module 3 and then output via the output unit 4.

[0088] Among them, the model construction module 2 is configured to construct a generative adversarial network model including a generator and a discriminator; among them, the generator takes a real first-modal medical image as input, learns the feature mapping relationship between the first-modal medical image and the real second-modal medical image, and generates a synthetic second-modal medical image according to the feature mapping relationship, and outputs a pair of true and false medical images; the pair of true and false medical data images is formed by splicing the synthetic second-modal medical image and the real second-modal medical image with the real first-modal medical image respectively;

[0089] The discriminator takes the pair of true and false medical data images as input, discriminates its true and false, and outputs a discrimination result;

[0090] The model training module 3 is used to construct different loss functions for the generator and the discriminator respectively to perform image synthesis training on the generative adversarial network model.

[0091] The above cross-modal medical image synthesis system constructs a generative adversarial network model including a generator and a discriminator by setting a model construction module 2. Then, it controls the feature mapping relationship between the first-modal medical image and the second-modal medical image by the generator, and generates a synthesized second-modal medical image according to this feature mapping relationship. Then, the synthesized second-modal medical image and the real second-modal medical image are respectively stitched with the real first-modal medical image to output a first image pair and a second image pair. On the one hand, it can increase paired data and make the synthesis effect better. On the other hand, since the generated synthesized second-modal medical image is obtained according to the feature mapping relationship, and at the same time, the first image pair and the second image pair are also discriminated as true or false through the discriminator. Then, a model training module 3 is set to construct different loss functions for the generator and the discriminator respectively to perform image synthesis training on the generative adversarial model. Therefore, the finally obtained result can be made more reliable. That is to say, the synthesis system of the present application is constructed based on 3D CGAN, which can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data, so as to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, such as low signal-to-noise ratio and blurred edges.

[0092] According to the third aspect of the present invention, a terminal is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it can be used to execute any one of the methods in the above embodiments of the present invention.

[0093] Optionally, the memory is used to store programs; the memory may include volatile memory (English: Volatile Memory), such as random access memory (English: Random-Access Memory, abbreviation: RAM), such as static random access memory (English: Static Random-Access Memory, abbreviation: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviation: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: Non-Volatile Memory), such as flash memory (English: Flash Memory). The memory is used to store computer programs (such as application programs and functional modules for implementing the above methods), computer instructions, etc. The above computer programs, computer instructions, etc. can be stored in partitions in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.

[0094] The above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.

[0095] A processor, configured to execute the computer program stored in the memory to implement each step in the method related to the above embodiments. For specific details, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0096] The processor and the memory can be of independent structures or integrated structures. When the processor and the memory are of independent structures, the memory and the processor can be coupled and connected through a bus.

[0097] Since the above terminal includes a processor configured to execute the method described in any of the foregoing embodiments, and for this cross-modal medical image synthesis method, by first constructing a generative adversarial network model including a generator and a discriminator, then controlling the generator for the feature mapping relationship between the first-modal medical image and the second-modal medical image, and generating a synthesized second-modal medical image according to this feature mapping relationship, and then splicing the synthesized second-modal medical image and the real second-modal medical image with the real first-modal medical image respectively to output a first image pair and a second image pair. On the one hand, it can increase paired data and make the synthesis effect better; on the other hand, since the generated synthesized second-modal medical image is obtained according to the feature mapping relationship, and at the same time, the first image pair and the second image pair are also discriminated as true or false through the discriminator, and different loss functions are respectively constructed for the generator and the discriminator to perform image synthesis training on the generative adversarial model. Therefore, the finally obtained result can be more reliable. That is to say, the synthesis method of the present application is constructed based on 3D CGAN, which can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data, so as to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, such as low signal-to-noise ratio and blurred edges.

[0098] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it can be used to execute the method described in any of the above embodiments of the present invention.

[0099] The above computer-readable storage medium, since the computer program stored thereon can be used to execute the cross-modal medical image synthesis method described in any of the foregoing embodiments when executed by a processor. First, a generative adversarial network model including a generator and a discriminator is constructed, and then the generator is controlled to map the features between the first-modal medical image and the second-modal medical image, and a synthesized second-modal medical image is generated according to the feature mapping relationship. Then, the synthesized second-modal medical image and the real second-modal medical image are respectively stitched with the real first-modal medical image to output a first image pair and a second image pair. On the one hand, paired data can be increased, making the synthesis effect better; on the other hand, since the synthesized second-modal medical image is obtained according to the feature mapping relationship, and at the same time, the first image pair and the second image pair are also discriminated as true or false through the discriminator, and different loss functions are respectively constructed for the generator and the discriminator to perform image synthesis training on the generative adversarial model. Therefore, the finally obtained result can be made more reliable. That is to say, the synthesis method of the present application is constructed based on 3D CGAN, which can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data, so as to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, and there are problems such as low signal-to-noise ratio and blurred edges.

[0100] The cross-modal medical image synthesis method and system provided in the above embodiments of the present invention, wherein the system includes modules corresponding to the steps of the method. First, a generative adversarial network model including a generator and a discriminator is constructed, and then the generator is controlled to map the features between the first-modal medical image and the second-modal medical image, and a synthesized second-modal medical image is generated according to the feature mapping relationship and true and false medical image pairs are output. On the one hand, paired data can be increased, making the synthesis effect better; on the other hand, since the synthesized second-modal medical image is obtained according to the feature mapping relationship, and at the same time, these true and false medical image pairs are also discriminated as true or false through the discriminator, and different loss functions are respectively constructed for the generator and the discriminator to perform image synthesis training on the generative adversarial model. Therefore, the finally obtained result can be made more reliable. That is to say, the synthesis method of the present application is constructed based on 3D CGAN, which can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data, so as to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, and there are problems such as low signal-to-noise ratio and blurred edges.

[0101] The cross-modal medical image synthesis method and system provided in the above embodiments of the present invention first construct a generative adversarial network model including a generator and a discriminator, and then control the generator to map the features between the first-modal medical image and the second-modal medical image, and generate a synthesized second-modal medical image according to the feature mapping relationship, and output a pair of real and fake medical images. On the one hand, it can increase the paired data and make the synthesis effect better; on the other hand, since the generated synthesized second-modal medical image is obtained according to the feature mapping relationship, and at the same time, these pairs of real and fake medical images are also discriminated for authenticity through the discriminator, and different loss functions are constructed for the generator and the discriminator respectively to train the generative adversarial model for image synthesis. Therefore, the final obtained result can be made more reliable. That is to say, the synthesis method of the present application is constructed based on 3DCGAN, and can make full use of the spatial structure information of multi-modal medical images to generate highly reliable multi-modal data, so as to solve the problems that the existing synthesis results cannot well represent the edge information of human tissues, and there are problems such as low signal-to-noise ratio and blurred edges.

[0102] It should be noted that the steps in the method provided by the present invention can be implemented by corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the method to implement the composition of the system. That is, the embodiments in the method can be understood as preferred examples for constructing the system, which will not be elaborated here.

[0103] Those skilled in the art know that in addition to implementing the system, device and their respective modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to make the system, device and their respective modules provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to implement the same program. Therefore, the system, device and their respective modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as both software programs for implementing the method and the structure within the hardware component.

[0104] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0105] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. A cross-modal medical image synthesis method, characterized in that, Including: Constructing a generative adversarial network model including a generator and a discriminator; wherein, the generator takes a real first-modal medical image as input, learns the feature mapping relationship between the real first-modal medical image and the real second-modal medical image, and generates a synthetic second-modal medical image according to this feature mapping relationship. Then, the synthetic second-modal medical image and the real second-modal medical image are respectively spliced with the real first-modal medical image to output a first image pair and a second image pair; The discriminator takes the first image pair and the second image pair as input, respectively discriminates their authenticity, and outputs the authenticity discrimination result; and Constructing different loss functions for the generator and the discriminator respectively to perform image synthesis training on the generative adversarial network model; Among them, the generator adopts a U-Net network structure, which includes an encoder and a decoder with symmetric network structures. The generation steps of the synthetic second-modal medical image include: Through the feature extraction operation of multi-layer convolution of the encoder, output the feature map of the real first-modal medical image; The decoder performs multi-layer transposed convolution operations on the feature map output by the encoder, and performs multiple splicing operations on the generated feature map and the feature map of the same size at the corresponding position of the encoder, and finally outputs the target reconstructed image, which is the synthetic second-modal medical image.

2. The cross-modal medical image synthesis method according to claim 1, wherein Also including: Regarding each pixel in the feature map as a random variable, calculating the paired covariance between the pixels; Selectively enhancing or weakening the value of each pixel according to the calculated paired covariance.

3. The cross-modal medical image synthesis method according to claim 1, characterized in that The encoder includes a convolutional module layer, a batch normalization layer, and an activation layer; Among them, the number of convolutional module layers is seven, and the second to fifth convolutional module layers are hybrid dilated convolutional module layers, and the rest are fully convolutional layers.

4. The cross-modal medical image synthesis method according to claim 3, wherein The hybrid dilated convolutional module layer includes 6 3D convolutional layers of 3×3×3, and the dilation rate of each 3D convolutional layer is set to a zigzag structure; Each of the convolutional layers is respectively denoted as convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, convolutional layer 5, and convolutional layer 6; among them, convolutional layer 1 and convolutional layer 3, convolutional layer 2 and convolutional layer 5, and convolutional layer 4 and convolutional layer 6 are respectively connected through a residual structure; The dilation rates of convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, convolutional layer 5, and convolutional layer 6 are 1, 2, 5, 1, 2, and 5 respectively.

5. The cross-modal medical image synthesis method according to claim 1, wherein The discriminator includes 6 convolutional layers, a batch normalization layer, and an activation layer; among them, each of the convolutional layers is respectively denoted as convolutional layer 1, convolutional layer 2, convolutional layer 3, convolutional layer 4, convolutional layer 5, and convolutional layer 6; among them, convolutional layer 1 and convolutional layer 4, and convolutional layer 2 and convolutional layer 6 are respectively connected through a residual structure.

6. The cross-modal medical image synthesis method according to any one of claims 1-5, characterized in that, The real first-modal medical image includes a CT image or an MRI image; the synthetic second-modal medical image includes a SPECT image or a PET image.

7. The cross-modal medical image synthesis method according to any one of claims 1-5, characterized in that The loss function MAE of the generator is set as: where m is the model batch size, is the true value, is the predicted value.

8. The cross-modal medical image synthesis method according to any one of claims 1-5, characterized in that, The loss function MSE of the discriminator is set as: where m is the model batch size, is the true value, is the predicted value.

9. A cross-modal medical image synthesis system, characterized in that, Including: A model construction module, configured to construct a generative adversarial network model including a generator and a discriminator; wherein, the generator takes a real first-modal medical image as input, learns the feature mapping relationship between the real first-modal medical image and the real second-modal medical image, and generates a synthetic second-modal medical image according to this feature mapping relationship. Then, the synthetic second-modal medical image and the real second-modal medical image are respectively stitched with the real first-modal medical image to output a first image pair and a second image pair. Among them, the generator adopts a U-Net network structure, which includes an encoder and a decoder with symmetric network structures. The generation step of the synthetic second-modal medical image includes: through the feature extraction operation of multi-layer convolution of the encoder, outputting the feature map of the real first-modal medical image; the decoder performs multi-layer transposed convolution operations on the feature map output by the encoder, and performs multiple stitching operations on the generated feature map and the feature map of the same size at the corresponding position of the encoder, and finally outputs a target reconstructed image, which is the synthetic second-modal medical image; The discriminator takes the first image pair and the second image pair as input, respectively performs true / false discrimination on them, and outputs the true / false discrimination results; A model training module, used to construct different loss functions for the generator and the discriminator respectively, so as to perform image synthesis training on the generative adversarial network model.

10. A terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it can be used to execute the method described in any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it can be used to execute the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Cross-modal medical image registration method and device

    CN111862174A

  • Priori guidance type network for multi-task medical image synthesis

    CN112669247A