A medical image cross-modal synthesis method, system and readable storage medium
By combining the generative adversarial network model of the visual Transformer and the convolution operator, efficient and accurate cross-modal synthesis of MRI images to CT images is achieved, which solves the problem of insufficient expression of contextual features in the existing technology and improves the accuracy and efficiency of image conversion.
Patent Information
- Application Number
- CN202210942137.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-08-08
AI Technical Summary
Existing cross-modal synthesis methods for medical images cannot effectively express the contextual characteristics of images, resulting in a lack of dependencies between long-distance voxels, which affects the estimation accuracy.
A generative adversarial network model combining visual Transformer and convolution operator is adopted to perform cross-modal synthesis of MRI images to CT images through joint processing of channel attention features and global attention features. The joint attention residual processing module is used to perform downsampling and upsampling operations to collaboratively preserve local and global context.
It improves the estimation accuracy and conversion efficiency of cross-modal synthesis of medical images, overcomes the limitations of contextual feature expression in existing technologies, and enhances the dependencies between distant voxels.
Smart Images

Figure CN115311183B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of medical image optimization, and in particular to a method, system, and readable storage medium for cross-modal synthesis of medical images. Background Art
[0002] Medical imaging plays a vital role in the diagnosis and treatment of various diseases. Because the insights provided by different imaging modalities are often complementary, more than one imaging modality is often involved in clinical decision-making. For example, magnetic resonance (MR) images are widely used in clinical diagnosis and cancer monitoring because they are obtained through noninvasive imaging protocols and provide excellent soft-tissue contrast. However, MR images do not provide the electron density information that computed tomography (CT) images do, which is crucial for dose calculation in radiotherapy treatment planning. CT has the advantage of providing both electron and physical density information on tissues, which is essential for radiotherapy dose planning in cancer patients. On the other hand, radiation exposure during CT acquisition may also increase the risk of secondary cancers, particularly in young patients. Magnetic resonance imaging (MRI) provides excellent soft-tissue contrast. Compared to CT, MRI is also safer and does not involve any radiation; often, some conditions require two or more imaging studies for an accurate diagnosis. Obtaining a set of different clinical images is a time-consuming and expensive process, making it unaffordable for most patients. Therefore, cross-modality synthesis of medical images, such as MR to CT synthesis, is desirable for many diagnostic and therapeutic purposes.
[0003] The above observations reflect a common dilemma: it is desirable to obtain certain forms of medical images, but in practice these are not feasible. To this end, a system that can synthesize images of interest from different sources (e.g., image modalities and acquisition protocols) can be of great benefit. It can provide the highly demanding imaging data for certain clinical uses without incurring the additional cost / risk of performing real acquisitions.
[0004] Generative adversarial networks (GANs) [Goodfellow I et al. Generative adversarial nets, 2014], a generative approach based on deep learning, have become increasingly popular for medical image synthesis. In several studies [Emami H et al. Generating synthetic CTs from magnetic resonance images using generative adversarial networks, 2018] and [Kazemifar S et al. MRI-only brain radiotherapy: Assessing the dosimetric accuracy of synthetic CT images generated using a deep learning approach, 2019], GANs have been used to exploit nonlinear correlations between modalities and estimate realistic synthetic CT images. Various GAN variants have demonstrated remarkable performance in medical image synthesis. Despite their powerful capabilities, previous learning-based synthesis models have been largely based on convolutional architectures, using compact filters to extract local image features. By leveraging correlations between small neighborhoods of image pixels, this inductive bias reduces the number of model parameters and facilitates learning. However, it also limits the expression of contextual features that reflect long-term spatial dependencies [X. Wang et al. Non-local neural networks, 2018]. Medical images contain contextual relationships between healthy and pathological tissues. For example, bones in the skull or cerebrospinal fluid in the ventricles are widely distributed in spatially adjacent or separated brain regions, leading to dependencies between long-range voxels. Although pathological tissue has less regular anatomical basis, its spatial distribution (such as location, number, shape) can still show disease-specific patterns. In principle, this can be improved by comprehensively capturing priors of these relationships. The Visual Transformer (VIT) is promising to achieve this goal because the attention operation of learning contextual features can improve sensitivity to long-range interactions and focus on key image regions to improve generalization to atypical anatomy. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide a method, system and readable storage medium for cross-modal synthesis of medical images, which can avoid limiting the expression of contextual characteristics and improve estimation accuracy.
[0006] The present application also provides a method for cross-modal synthesis of medical images, comprising the following steps:
[0007] Determining a pair of medical images corresponding to the same part, the pair of medical images comprising an MRI image and a CT image that are matched for training;
[0008] Performing cross-modal image synthesis on the MRI image based on the constructed initial generative adversarial network model to obtain a corresponding CT image, wherein the initial generative adversarial network model includes a joint attention residual processing module for jointly extracting channel attention features and global attention features from the image to determine corresponding joint attention features;
[0009] Model training is performed based on the matched medical image pairs. During the training process, the input MRI images are downsampled and upsampled in combination with the joint attention residual processing module to collaboratively preserve local and global context, and a target generative adversarial network model is obtained at the end of training;
[0010] The acquired MRI image to be processed is input into the target generative adversarial network model to obtain the corresponding CT image synthesis result.
[0011] In a second aspect, an embodiment of the present application further provides a medical image cross-modal synthesis system, the system comprising an image processing module, a model building module, a model training module, and an image synthesis module, wherein:
[0012] The image processing module is used to determine a pair of medical images corresponding to the same part, wherein the pair of medical images includes an MRI image and a CT image that are matched for training;
[0013] The model construction module is used to perform cross-modal image synthesis on the MRI image based on the constructed initial generative adversarial network model to obtain a corresponding CT image, wherein the initial generative adversarial network model includes a joint attention residual processing module for combining channel attention features and global attention features extracted from the image to determine corresponding joint attention features;
[0014] The model training module is used to perform model training based on the matched medical image pairs. During the training process, the input MRI image is downsampled and upsampled in combination with the joint attention residual processing module to collaboratively preserve local and global context, and a target generative adversarial network model is obtained at the end of training;
[0015] The image synthesis module is used to input the acquired MRI image to be processed into the target generative adversarial network model to obtain the corresponding CT image synthesis result.
[0016] In a third aspect, an embodiment of the present application further provides a readable storage medium, which includes a medical image cross-modal synthesis method program. When the medical image cross-modal synthesis method program is executed by a processor, the steps of a medical image cross-modal synthesis method as described in any one of the above items are implemented.
[0017] From the above, it can be seen that the embodiments of the present application provide a method, system and readable storage medium for cross-modal synthesis of medical images, which utilize the context sensitivity of the fitted visual Transformer, the precision of the convolution operator and the realism of adversarial learning to establish a corresponding network model for cross-modal synthesis of medical images, so that conversion can be performed between cross-modal medical image data based on the network model, overcoming the above-mentioned limitations of the expression of contextual characteristics in the existing convolutional neural network technology, resulting in a lack of dependence between long-distance voxels, thereby improving image conversion efficiency and estimation accuracy.
[0018] Other features and advantages of the present application will be described in the following description and, in part, will become apparent from the description or be understood by practicing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 A flowchart of a medical image cross-modal synthesis method provided in an embodiment of the present application;
[0021] Figure 2 Schematic diagram of the structure of the joint attention residual processing module;
[0022] Figure 3 Schematic diagram of the structure of the channel attention block;
[0023] Figure 4 Schematic diagram of the network structure for downsampling and upsampling the input target MRI image in combination with the joint attention residual processing module;
[0024] Figure 5 A schematic diagram of the structure of a medical image cross-modal synthesis system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0027] Please refer to Figure 1 , Figure 1 This is a flowchart of a method for cross-modal synthesis of medical images in some embodiments of the present application. This method is described using a computer device (the computer device may be a terminal or a server, and the terminal may be, but is not limited to, various personal computers, laptops, smartphones, tablet computers, and portable wearable devices. The server may be a standalone server or a server cluster consisting of multiple servers) as an example, and includes the following steps:
[0028] Step S100 : determining a medical image pair corresponding to the same part, where the medical image pair includes a matched MRI image and a CT image for training.
[0029] In step S200, the MRI image is cross-modally synthesized based on the constructed initial generative adversarial network model to obtain a corresponding CT image. The initial generative adversarial network model includes a joint attention residual processing module for jointly extracting channel attention features and global attention features from the image to determine the corresponding joint attention features.
[0030] In step S300, model training is performed based on the matched medical image pairs. During the training process, the input MRI image is downsampled and upsampled in combination with the joint attention residual processing module to collaboratively preserve local and global contexts, and at the end of the training, the target generative adversarial network model is obtained.
[0031] In step S400 , the acquired MRI image to be processed is input into the target generative adversarial network model to obtain the corresponding CT image synthesis result.
[0032] From the above, it can be seen that a cross-modal synthesis method of medical images provided in an embodiment of the present application utilizes the context sensitivity of the fitted visual Transformer, the precision of the convolution operator, and the realism of adversarial learning to establish a corresponding network model for cross-modal synthesis of medical images, so that conversion can be performed between cross-modal medical image data based on the network model, overcoming the above-mentioned limitations of the expression of contextual characteristics in the existing convolutional neural network technology, resulting in a lack of dependence between long-distance voxels, thereby improving image conversion efficiency and estimation accuracy.
[0033] In one embodiment, in step S100, determining a target medical image pair corresponding to the same part includes:
[0034] Step S1001 : obtaining a pair of medical images to be processed corresponding to the same part, wherein the pair of medical images to be processed includes a matched initial MRI image and an initial CT image.
[0035] In step S1002, the initial MRI image and the initial CT image are preprocessed according to a preset preprocessing method to obtain matching MRI images and CT images for training, wherein the preprocessing method includes at least one of an N4 bias correction method and an image denoising method.
[0036] Step S1003 : determining a medical image pair corresponding to the same part based on the matched MRI image and CT image used for training.
[0037] It should be noted that before model training on MRI and CT data (i.e., the aforementioned medical image pairs to be processed), these data must first undergo preprocessing, such as N4 bias correction and denoising. Subsequently, the preprocessed data is accurately registered and sliced into corresponding 2D datasets to improve data registration accuracy.
[0038] In one embodiment, please refer to Figure 2 The joint attention residual processing module consists of a channel attention block and multiple cascaded Swin Transformer blocks, where:
[0039] The channel attention block consists of a pooling layer, a first fully connected layer, a dot product operation layer, a second fully connected layer, and a sigmoid activation layer connected in sequence.
[0040] For details, please refer to Figure 3Unlike general channel attention blocks, which only extract statistical information of input feature channels through one pooling operation, the current embodiment extracts statistical information of input feature channels based on a pooling layer that includes two pooling operations, namely, a maximum pooling operation and an average pooling operation.
[0041] In one embodiment, the pooling layer further sequentially undergoes the steps of fully connected layer, dot product operation, fully connected layer, and sigmiod activation, and obtains channel attention features based on the statistical information of the extracted input feature channels. The obtained channel attention features are further multiplied with the input global attention features to obtain the required output features, allowing the model to better adapt to obtain channel statistics.
[0042] The pooling layer consists of a maximum pooling branch layer for extracting statistical information of the input feature channels through a maximum pooling operation, and an average pooling branch layer for extracting statistical information of the input feature channels through an average pooling operation.
[0043] Specifically, commonly used pooling methods include maximum pooling and mean pooling. According to relevant theories, feature extraction errors mainly come from two aspects:
[0044] (1) The variance of the estimated value increases due to the limited neighborhood size;
[0045] (2) The error in the convolutional layer parameters causes the estimated mean to shift.
[0046] Generally speaking, max pooling can reduce the first and second errors mentioned above, preserving more background and texture information of the image. It is similar to mean pooling and, in a local sense, follows the max-pooling principle.
[0047] In one embodiment, the maximum pooling kernel size is typically 2×2. For very large inputs, the kernel size may be set to 4×4. However, larger kernel sizes can significantly reduce the signal size and may result in excessive information loss. Generally, non-overlapping pooling windows perform best.
[0048] In one embodiment, in step S300, when extracting channel attention features from the input MRI image based on the channel attention block, the method includes:
[0049] Step S3001: extract the channel attention feature P from the input target MRI image using the following formula: C :
[0050]
[0051] Among them, x Tis the global attention feature extracted from the input target MRI image, δ1 is the first fully connected layer, and δ2 is the second fully connected layer; is the sigmoid activation function.
[0052] In one embodiment, in step S300, when combining the channel attention features and the global attention features extracted from the input target MRI image to determine the joint attention features of the image, the method includes:
[0053] Step S3002: Determine the joint attention feature P of the image using the following formula: SRTB :
[0054] P SRTB =Conv(P STB (x T )+P C (x T ));
[0055] Among them, x T is the global attention feature extracted from the input target MRI image, P STB is a cascade block consisting of the multiple cascaded Swin Transformer blocks, P C is the channel attention block, and Conv is the convolutional layer.
[0056] For details, please refer to Figure 2 , the above cascade block is composed of six Swin Transformer blocks cascaded. In the current embodiment, the computer device inputs the feature x T After the above channel attention block P C The channel attention feature is extracted. The channel attention feature is also connected to the output of the cascade block to determine the joint attention feature of the image based on the extracted channel attention feature and the global attention feature.
[0057] In one embodiment, in step S300, the input target MRI image is downsampled and upsampled in combination with the joint attention residual processing module, including: performing four downsampling operations and corresponding four upsampling operations on the input target MRI image to maintain the original size of the image, wherein the downsampling operation includes a maximum pooling operation, and the upsampling operation includes a deconvolution operation; when performing the first sampling, the output features obtained by downsampling are jump-connected to the initial input features of the upsampling; when performing the second, third or fourth sampling, the output features obtained by downsampling are passed through the joint attention residual processing module and then connected to the initial input features of the upsampling.
[0058] Specifically, the computer device can implement a downsampling operation on the target MRI image through a preset maximum pooling layer, and implement an upsampling operation on the target MRI image through a preset deconvolution layer.
[0059] Among them, you can refer to Figure 4 When performing the first sampling, the computer device jumps the target output feature output by the maximum pooling layer corresponding to the downsampling operation to the initial input feature of the deconvolution operation layer corresponding to the upsampling operation.
[0060] It should be noted that the target output features of the down-sampling output of the 2nd, 3rd and 4th layers will be processed by the joint attention residual processing module (i.e. Figure 4 RSTB shown in the figure), and then connected to the initial input features of the deconvolution operation layer corresponding to the upsampling operation.
[0061] In this way, it can effectively overcome the existing convolutional neural network technology, which limits the expression of contextual features and leads to a lack of dependence between long-distance voxels, thereby improving the accuracy of data prediction.
[0062] In one embodiment, in step S300, model training is performed based on the target medical image pair, and upon completion of the training, a target generative adversarial network model is obtained, including:
[0063] Step S3003, constructing a target loss function, the target loss function includes generating an adversarial loss function, a mean absolute error loss function determined based on the mean absolute error between the target CT image and the corresponding CT image synthesis result, and a frequency loss function determined based on the frequency difference between the target CT image and the corresponding CT image synthesis result.
[0064] Specifically, the computer device may construct a corresponding loss function based on the deviation between the target CT image and the corresponding CT image synthesis result.
[0065] When multiple types of loss functions are involved, a total loss function can be determined based on the weighted sum of these multiple loss functions. Subsequently, model optimization is performed based on this total loss function, with optimization targets including, but not limited to, adjustments to network parameters and network structure, etc., which are not limited in this embodiment of the present application.
[0066] In step S3004, during the training process, the model is optimized by the gradient descent method based on the target loss function, and when it is determined that the preset training end condition is met, the target generative adversarial network model is obtained.
[0067] Specifically, in a machine learning algorithm, when minimizing the loss function, the gradient descent method can be used to iteratively solve the problem, thereby obtaining the minimized loss function and model parameter values.
[0068] In one embodiment, if the maximum value of the loss function needs to be solved, the gradient ascent method can be used for iterative calculation. It should be noted that the gradient descent method and the gradient ascent method can be converted into each other. For example, when the minimum value of the loss function f(ω) needs to be solved, the gradient descent method can be used for iterative solution. However, in fact, it is also possible to solve the maximum value of the loss function -f(ω) in reverse, and at this time, the gradient ascent method comes in handy.
[0069] In one embodiment, the above-mentioned training end conditions include but are not limited to the target loss function approaching 0, reaching the maximum number of iterations, etc., which are not limited in this embodiment of the present application.
[0070] In one embodiment, the calculation formula of the generative adversarial loss function LcGAN includes:
[0071] LcGAN=E x,y [logD(x,y)]+E x,y [log(1-D(x, G(x)))];
[0072] Where x∈R N is the input target MRI image, y∈R N is the target CT image paired with the target MRI image; E x,y [*] is the expected value of the distribution function; D(x,y) is the matching between the target MRI image and the corresponding target CT image; G(x) is the CT image synthesis result output by the model.
[0073] Specifically, D(x,y) can be further understood as a discriminator for identifying the degree of pairing between the target MRI image x and the corresponding generated CT image y. G(x) can be further understood as a generator for generating the corresponding CT image based on the input target MRI image x.
[0074] It can be understood that in the above formula, logD(x,y) is the probability that the discriminator judges real data as real data, and log(1-D(x,G(x))) is the probability that the discriminator judges the false data generated by the generator as the opposite of real data, that is, the probability that the false data is still judged as false data.
[0075] Adversarial networks, on the other hand, merely propose a network structure. Generally speaking, they utilize two models: a generative model and a discriminative model. The discriminative model determines whether a given image is real (from a dataset), while the generative model's task is to create an image that looks realistic. Initially, both models are untrained. They are then trained adversarially, with the generative model generating an image to deceive the discriminative model, which then determines whether the image is real or fake. Ultimately, as the two models train, their capabilities grow, eventually reaching a steady state.
[0076] The calculation formula of the mean absolute error loss function L1 includes:
[0077] L1=E x,y ||yG(x)||1;
[0078] Here, ||.||1 represents the L1 norm.
[0079] Specifically, the mean absolute error (MAE) function is a commonly used regression loss function. It is the sum of the absolute values of the differences between the target and predicted values, representing the average magnitude of the error in the predictions, regardless of the direction of the error. Compared to the mean error function, the MAE function, because the deviations are absolute, does not cause positive and negative offsets. Therefore, the MAE function can better reflect the actual state of the prediction error and improve model training accuracy.
[0080] Frequency loss function L fre The calculation formula includes:
[0081] L fre =E x,y [||y l -G(x) l ||+||y h -G(x) h ||];
[0082] Among them, y l is the low-frequency information of the target MRI image, G(x) l is the low-frequency information of the CT image synthesis result; y h is the high-frequency information of the target MRI image, G(x) h It is the high-frequency information of the CT image synthesis result.
[0083] Specifically, the computer device can use a Gaussian kernel function to filter out high-frequency features to retain low-frequency information, namely:
[0084]
[0085] Where [i, j] represents the spatial position in the image; σ 2 represents the variance, where the variance increases proportionally with the Gaussian kernel size.
[0086] In one embodiment, y l For example, in determining the low-frequency information y in image y l When , the computer device can use the Gaussian kernel for convolution processing, and its calculation method can refer to the following formula:
[0087] y l [i, j] = ∑ m ∑ n K[m,n]·y[i+m,j+n];
[0088] Among them, y l is the low-frequency information extracted from the image y, [i, j] represents the spatial position in the image; m and n are the indicators of the two-dimensional Gaussian kernel function, and K[m,n] is the Gaussian kernel function.
[0089] In one embodiment, the computer device may filter out low-frequency information y from the image y. l By further extracting high-frequency information y from image y h That is: y h =yy l .
[0090] In this way, during the model training process, the computer equipment uses loss functions such as frequency constraint and adversarial constraint loss functions to perform fusion training on the target generative adversarial network model, so that it can achieve higher quality synthesis effects and improve image synthesis quality compared with ordinary generative adversarial network models.
[0091] Please refer to Figure 5 The present application discloses a medical image cross-modality synthesis system 500, which includes an image processing module 501, a model building module 502, a model training module 503, and an image synthesis module 504, wherein:
[0092] The image processing module 501 is used to determine a medical image pair corresponding to the same part, where the medical image pair includes a matching MRI image and a CT image for training.
[0093] The model construction module 502 is used to perform cross-modal synthesis of MRI images based on the constructed initial generative adversarial network model to obtain corresponding CT images. The initial generative adversarial network model includes a joint attention residual processing module for jointly extracting channel attention features and global attention features from the image to determine the corresponding joint attention features.
[0094] The model training module 503 is used to perform model training based on matched medical image pairs. During the training process, the input MRI image is downsampled and upsampled in combination with the joint attention residual processing module to collaboratively preserve local and global contexts, and at the end of the training, a target generative adversarial network model is obtained.
[0095] The image synthesis module 504 is used to input the acquired MRI image to be processed into the target generative adversarial network model to obtain the corresponding CT image synthesis result.
[0096] In one embodiment, the above modules are also used to implement the method in any optional implementation of the above embodiment.
[0097] From the above, it can be seen that a medical image cross-modal synthesis system disclosed in the present application utilizes the context sensitivity of the fitted visual Transformer, the precision of the convolution operator, and the realism of adversarial learning to establish a corresponding network model for cross-modal synthesis of medical images, so that conversion between cross-modal medical image data can be performed based on the network model, overcoming the above-mentioned limitations of the expression of contextual characteristics in the existing convolutional neural network technology, resulting in a lack of dependence between long-distance voxels, thereby improving image conversion efficiency and estimation accuracy.
[0098] The present application provides a readable storage medium, wherein when the computer program is executed by a processor, the method in any optional implementation of the above embodiment is executed. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.
[0099] The above-mentioned readable storage medium utilizes the context sensitivity of the fitted visual Transformer, the precision of the convolution operator, and the realism of adversarial learning to establish a corresponding network model for cross-modal synthesis of medical images, so that conversion between cross-modal medical image data can be performed based on the network model, overcoming the limitations of the above-mentioned existing convolutional neural network technology on the expression of contextual characteristics, resulting in a lack of dependence between long-distance voxels, thereby improving image conversion efficiency and estimation accuracy.
[0100] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0101] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0102] Furthermore, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0103] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.
[0104] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A medical image cross-modal synthesis method, characterized in that: The following steps are involved: Acquire a pair of medical images to be processed corresponding to the same part, the pair of medical images to be processed comprising a matched initial MRI image and an initial CT image; Preprocessing the initial MRI image and the initial CT image according to a preset preprocessing method to obtain a matching MRI image and CT image for training, wherein the preprocessing method includes at least one of an N4 bias correction method and an image denoising method; determining a medical image pair corresponding to the same part based on the matched MRI image and CT image for training, the medical image pair including the matched MRI image and CT image for training; Performing cross-modal image synthesis on the MRI image based on the constructed initial generative adversarial network model to obtain a corresponding CT image, wherein the initial generative adversarial network model includes a joint attention residual processing module for jointly extracting channel attention features and global attention features from the image to determine corresponding joint attention features; Model training is performed based on the matched medical image pairs. During the training process, the input MRI images are downsampled and upsampled in combination with the joint attention residual processing module to collaboratively preserve local and global context, and a target generative adversarial network model is obtained at the end of training; Inputting the acquired MRI image to be processed into the target generative adversarial network model to obtain the corresponding CT image synthesis result; The joint attention residual processing module consists of a channel attention block and multiple cascaded Swin Transformer blocks, where: The channel attention block consists of a pooling layer, a first fully connected layer, a dot product operation layer, a second fully connected layer, and a sigmoid activation layer connected in sequence; The pooling layer is composed of a maximum pooling branch layer for extracting statistical information of the input feature channel through a maximum pooling operation, and an average pooling branch layer for extracting statistical information of the input feature channel through an average pooling operation.
2. The method according to claim 1, characterized in that When extracting channel attention features from an input MRI image based on the channel attention block, the method includes: The channel attention feature P is extracted from the input target MRI image using the following formula: C : Among them, x T is the global attention feature extracted from the input target MRI image, δ1 is the first fully connected layer, and δ2 is the second fully connected layer; is the sigmoid activation function.
3. The method according to claim 1, characterized in that When jointly extracting the channel attention features and the global attention features from the input target MRI image to determine the joint attention features of the image, the method includes: Determine the joint attention feature P of the image through the following formula SRTB : P SRTB =Conv(P STB (x T )+P C (x T )); Among them, x T is the global attention feature extracted from the input target MRI image, P STB is a cascade block consisting of the multiple cascaded SwinTransformer blocks, P C is the channel attention block, and Conv is the convolutional layer.
4. The method according to claim 1, wherein The step of performing downsampling and upsampling operations on the input target MRI image in combination with the joint attention residual processing module includes: Performing four downsampling operations and corresponding four upsampling operations on the input target MRI image to maintain the original size of the image, wherein the downsampling operation includes a maximum pooling operation and the upsampling operation includes a deconvolution operation; When performing the first sampling, the output features obtained by downsampling are jump-connected to the initial input features of upsampling; When performing the second, third or fourth sampling, the output features obtained by downsampling are passed through the joint attention residual processing module and then connected to the initial input features of the upsampling.
5. The method according to claim 1, wherein The model training is performed based on the target medical image pair, and upon completion of the training, a target generative adversarial network model is obtained, comprising: Constructing a target loss function, the target loss function including at least one of a generative adversarial loss function, a mean absolute error loss function determined based on a mean absolute error between a target CT image and a result of synthesizing the corresponding CT image, and a frequency loss function determined based on a frequency difference between the target CT image and a result of synthesizing the corresponding CT image; During the training process, based on the target loss function, the model is optimized by the gradient descent method, and when it is determined that the preset training end condition is met, the target generative adversarial network model is obtained.
6. The method according to claim 5, characterized in that The calculation formula of the generative adversarial loss function LcGAN includes: LcGAN=E x,y [logD(x,y)]+E x,y [log(1-D(x,G(x)))]; Where x∈R N is the input target MRI image, y∈R N is the target CT image paired with the target MRI image; x,y [*] is the expected value of the distribution function; D(x,y) is the matching between the identified target MRI image and the corresponding target CT image; G(x) is the CT image synthesis result output by the model; The calculation formula of the mean absolute error loss function L1 includes: L1=E x,y ||yG(x)||1; Among them, ||·||1 represents the L1 norm; The frequency loss function L fre The calculation formula includes: L fre =E x,y [||y l -G(x) l ||+||y h -G(x) h ||]; Among them, y l is the low-frequency information of the target MRI image, G(x) l is the low-frequency information of the CT image synthesis result; y h is the high-frequency information of the target MRI image, G(x) h It is the high-frequency information of the CT image synthesis result.
7. A medical image cross-modal synthesis system, characterized in that: The system includes an image processing module, a model building module, a model training module, and an image synthesis module, wherein: The image processing module is configured to obtain a pair of medical images to be processed corresponding to the same part, the pair of medical images to be processed comprising a matched initial MRI image and an initial CT image, preprocess the initial MRI image and the initial CT image respectively according to a preset preprocessing method to obtain matched MRI images and CT images for training, wherein the preprocessing method comprises at least one of an N4 bias correction method and an image denoising method, and determine a pair of medical images corresponding to the same part based on the matched MRI image and CT image for training, the pair of medical images comprising the matched MRI image and CT image for training; The model construction module is used to perform cross-modal image synthesis on the MRI image based on the constructed initial generative adversarial network model to obtain a corresponding CT image, wherein the initial generative adversarial network model includes a joint attention residual processing module for combining channel attention features and global attention features extracted from the image to determine corresponding joint attention features; The model training module is used to perform model training based on a pair of matched medical images. During the training process, the input MRI image is downsampled and upsampled in combination with the joint attention residual processing module to collaboratively preserve local and global contexts, and a target generative adversarial network model is obtained at the end of the training. The joint attention residual processing module is composed of a channel attention block and multiple cascaded Swin Transformer blocks, wherein: the channel attention block is composed of a pooling layer, a first fully connected layer, a dot product operation layer, a second fully connected layer and a sigmoid activation layer connected in sequence, and the pooling layer is composed of a maximum pooling branch layer for extracting statistical information of the input feature channel through a maximum pooling operation, and an average pooling branch layer for extracting statistical information of the input feature channel through an average pooling operation; The image synthesis module is used to input the acquired MRI image to be processed into the target generative adversarial network model to obtain the corresponding CT image synthesis result.
8. A readable storage medium, characterized in that: The readable storage medium includes a medical image cross-modality synthesis method program, and when the medical image cross-modality synthesis method program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Super-resolution reconstruction algorithm for medical imaging
CN110717856A
Image synthesis method based on dynamic self-attention generative adversarial network
CN113379655A