Cross-modal medical image generation method, system, equipment and medium

The cross-modal medical image generation model is constructed through a circular generation adversarial network, which solves the problem of low accuracy and matching of CT images to MRI images, achieves higher quality image generation, and improves the clarity and detail retention of MRI images.

CN120355802APending Publication Date: 2025-07-22SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510351382.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technical methods for converting CT images into MRI images have problems such as poor image global information capture capability and insufficient detail retention, resulting in low accuracy and matching of generated MRI images.

Method used

Using a cross-modal medical image generation method based on a circular generation adversarial network, a cross-modal medical image generation model is constructed, and the CT image data is converted into pseudo-MRI image data using the first generator and the second generator, and optimized through the first discriminator and the second discriminator. Combining structural consistency loss, perceptual loss, adversarial loss and identity loss functions, the accuracy and matching of image conversion are improved.

Benefits of technology

Effectively capture and generate complex mapping relationships in organizational structures between different modal images, improve the clarity and detail performance of MRI images, enhance the matching of visual features and structure, and improve the accuracy and quality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355802A_ABST
    Figure CN120355802A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal medical image generation method, system and device and a medium. The method comprises the steps of obtaining sample image data including CT image data and MRI image data; constructing a cross-modal medical image generation model based on the cyclic generative adversarial network; a cross-modal medical image generation model is trained according to the sample image data, a target cross-modal medical image generation model is obtained, in the training process, the first generator is configured to convert a first feature set of the extracted CT image data into pseudo MRI image data, and the pseudo MRI image data is used as a target cross-modal medical image generation model; a second generator configured to convert a second feature set extracted from the pseudo MRI image data into pseudo CT image data; and inputting each acquired CT image to be converted into the target cross-modal medical image generation model, and constructing a target MRI image data set according to an output result of the target cross-modal medical image generation model. According to the method, the model feature extraction capability can be improved, and the accuracy of the output image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a cross-modal medical image generation method, system, device and medium. Background Art

[0002] Adaptive Radiation Therapy (ART) is a treatment method that adjusts the radiation plan in real time according to the physical changes of patients during radiotherapy. It can improve the treatment accuracy, irradiate tumors more precisely, and at the same time reduce the damage to surrounding healthy tissues. When the position or shape of the tumor changes, such as the tumor shrinks or the patient's body position changes, the doctor will adjust the treatment plan accordingly.

[0003] CT-Only adaptive radiation therapy refers to the use of only the CT imaging technology carried by the linear accelerator to monitor the changes of tumors and surrounding tissues in real time during CT-Linac-based treatment, so as to dynamically adjust the treatment plan without other imaging means. However, due to the low soft tissue resolution of CT images, it may affect the accurate definition of the boundaries of organs at risk and target areas, bringing certain challenges to the treatment accuracy.

[0004] Image-guided radiation therapy using magnetic resonance imaging (MRI) is a new technology that has been widely studied and developed in recent years. The advantage of high soft tissue resolution of MRI imaging can help doctors more accurately locate tumors and organs at risk. However, MRI examinations are not only expensive and require longer scanning and imaging times. In addition, MRI examinations need to be carried out in a relatively enclosed space, which also makes this technology inapplicable to patients with epilepsy, nerve irritation and claustrophobia, and does not have universality.

[0005] Converting CT images collected by CT-Linac into MRI images to assist in delineating tumors and organs at risk can reduce treatment time and costs, improve soft tissue contrast, and enhance the accuracy and efficacy of adaptive radiotherapy. The existing technical methods for converting CT images into MRI images face problems of poor global information capture ability of images and insufficient detail retention, resulting in low accuracy and matching degree of the generated MRI images.

[0006] Therefore, how to improve the accuracy and matching degree of converting CT images into MRI images has become a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention

[0007] The present invention provides a cross-modal medical image generation method, system, device and medium to solve the technical problem of how to improve the accuracy and matching degree of converting CT images into MRI images, and achieve the effect of capturing and generating the complex mapping relationship of tissue structures between different modal medical images, and improving the accuracy and matching degree of the synthesized MRI images.

[0008] In a first aspect, the present invention provides a cross-modal medical image generation method, the method comprising:

[0009] Obtain sample image data, wherein the sample image data includes CT image data of a first modality and MRI image data of a second modality;

[0010] Construct a cross-modal medical image generation model based on a cyclic generative adversarial network;

[0011] Train the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, a first generator is configured to convert a first feature set of the extracted CT image data into pseudo-MRI image data, and a second generator is configured to convert a second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes the first generator and the second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features of different scales, and first tissue and organ global features;

[0012] In the actual image generation process, input each collected CT image to be converted into the target cross-modal medical image generation model, and construct a target MRI image data set according to the output result of the target cross-modal medical image generation model.

[0013] Preferably, the target cross-modal medical image generation model further includes a first discriminator and a second discriminator;

[0014] The first discriminator is configured to perform authenticity discrimination based on a second feature set extracted from the MRI image data and the pseudo-MRI image data to obtain a discrimination result;

[0015] The second discriminator is configured to perform authenticity discrimination based on a second feature set extracted from the CT image data and the pseudo-CT image data to obtain a discrimination result. The second feature set at least includes: second tissue and organ local features of different scales and second tissue and organ global features.

[0016] Preferably, the training of the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model includes:

[0017] Perform a block operation on the CT image data to obtain CT image blocks, input the CT image blocks into the first generator, and obtain pseudo-MRI image blocks corresponding to the CT image data;

[0018] Perform an affine transformation on the MRI image data to obtain MRI image patches, and input the MRI image patches and the pseudo-MRI image patches into the first discriminator to generate a first discrimination result;

[0019] Optimize the first generator and the first discriminator based on the first discrimination result;

[0020] Input the pseudo-MRI image patches into the second generator to obtain pseudo-CT image patches;

[0021] Input the pseudo-CT image patches and the CT image patches into the second discriminator to generate a second discrimination result;

[0022] Optimize the second generator and the second discriminator based on the second discrimination result;

[0023] Based on the optimized first generator, second generator, first discriminator, and second discriminator, obtain the trained target cross-modal medical image generation model.

[0024] Preferably, the step of inputting the CT image patches into the first generator to obtain pseudo-MRI image patches corresponding to the CT images includes:

[0025] Input the CT image patches into the shallow feature extraction module of the first generator to extract shallow features of different levels of the CT image patches, generate several first feature maps with different resolutions, extract the smallest resolution feature map among the several first feature maps, and output the remaining several first feature maps as tissue organ shallow feature maps; the shallow feature extraction module includes at least several downsampling convolutional blocks and several residual convolutional blocks; the tissue organ shallow feature maps include at least the tissue organ detail features and the edge features;

[0026] Input the smallest resolution feature map into the deep feature extraction module of the first generator to extract deep features of the smallest resolution feature map and generate tissue organ deep feature maps of different sizes; the deep feature extraction module includes at least several variable-size window attention blocks; the tissue organ deep feature maps include at least the first tissue organ local features and the first tissue organ global features at different scales;

[0027] Input the tissue organ shallow feature maps and the tissue organ deep feature maps into the upsampling module of the first generator for feature fusion, and output pseudo-MRI image patches.

[0028] Preferably, the step of inputting the MRI image patches and the pseudo-MRI image patches into the first discriminator to generate a first discrimination result includes:

[0029] Input the MRI image patch and the pseudo-MRI image patch into the first discriminator, and the feature extraction convolutional block of the first discriminator extracts the tissue and organ detail information of the MRI image patch and the pseudo-MRI image patch; the first discriminator has a U-shaped structure;

[0030] Input the tissue and organ detail information into the global perception channel focusing module of the first discriminator. The multi-scale attention pooling layer of the global perception channel focusing module performs pooling operations on the tissue and organ detail information at different scales to obtain the second local tissue and organ features at different scales. The global average pooling layer of the global perception channel focusing module performs average pooling operations on each channel of the tissue and organ detail information to obtain the second global tissue and organ feature;

[0031] Fuse the second global tissue and organ features at different scales and the second local tissue and organ features, and then calculate the adaptive attention weight to obtain the attention weight map;

[0032] Link the attention weight map with the MRI image patch and the pseudo-MRI image patch respectively, and then perform upsampling processing to obtain the true / false discrimination feature of each pixel;

[0033] Generate a first discrimination result based on the true / false discrimination feature.

[0034] Preferably, the total objective function of the cross-modal medical image generation model includes a structure consistency loss function, a perception loss function, an adversarial loss function, and an identity loss function;

[0035] The total objective function is expressed as:

[0036] L total =L GAN +L identity +λ scycle ·L scycle +λ perc ·L perc

[0037] where L GAN represents the adversarial loss function term, L identity represents the identity loss function term, L scycle represents the structure consistency loss function term, L perc represents the perception loss function term, λ scycle 、λ perc represent weight parameters to balance the structure consistency loss function term and the perception loss function term;

[0038] The structure consistency loss function is:

[0039]

[0040] Among them, L1 is the cycle consistency loss, which is represented by the pixel difference between the real image x and the pseudo-image y. λ is the weight parameter, μ x is the average value of x, μ y is the average value of y, δ x is the variance of x, δ y is the variance of y, δ xy is the covariance of x and y, and c1 and c2 are constants; the real image includes the CT image data and the MRI image data, and the pseudo-image includes the pseudo-MRI image block and the pseudo-CT image block;

[0041] The perceptual loss function is:

[0042]

[0043] Among them, j is a specific layer in the cross-modal medical image generation model, is the feature extraction function, G x is the pseudo-image, x target is the real image;

[0044] The adversarial loss function is:

[0045] L GAN (G MRI ,G CT ,D MRI ,D CT ) = E[log(D MRI (I MRI ))] + E[1 - log(D MRI (G MRI (I CT )))] + E[log(D CT (I CT ))] + E[1 - log(D CT (G CT (G MRI (I CT ))))]

[0046] Among them, G MRI represents the first generator, G CT represents the second generator, D MRI represents the first discriminator, D CT represents the second discriminator, I MRI represents the MRI image block, I CT represents the CT image block;

[0047] The identity loss function is:

[0048] Lidentity(G CT ,G MRI ) = E[||G CT (I MRI )||] + E[||G MRI (I CT ) - I CT ||].

[0049] Preferably, constructing the target MRI image dataset according to the output result of the target cross-modal medical image generation model includes:

[0050] Performing similarity verification on the output result of the target cross-modal medical image generation model, and generating the target MRI image dataset based on the output result that passes the similarity verification.

[0051] In a second aspect, the present invention also provides a cross-modal medical image generation system for implementing the above-mentioned cross-modal medical image generation method. The system includes: a sample image data acquisition unit, a model construction unit, a model training unit, and a model application unit;

[0052] The sample image data acquisition unit is used to obtain sample image data, where the sample image data includes CT image data of the first modality and MRI image data of the second modality;

[0053] The model construction unit: is used to construct a cross-modal medical image generation model based on a cyclic generative adversarial network;

[0054] The model training unit: is used to train the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, the first generator is configured to convert the first feature set of the extracted CT image data into pseudo-MRI image data, and the second generator is configured to convert the second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes the first generator and the second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features of different scales, and first tissue and organ global features;

[0055] The model application unit: during the actual image generation process, inputs each to-be-converted CT image collected into the target cross-modal medical image generation model, and constructs a target MRI image dataset according to the output result of the target cross-modal medical image generation model.

[0056] In a third aspect, the present invention further provides a computer device, which includes a memory, a processor, and a transceiver, and they are connected through a bus; the memory is used to store a set of computer program instructions and data, and transmit the stored data to the processor, and the processor executes the program instructions stored in the memory to execute the sensitive device failure probability evaluation method described above.

[0057] In a fourth aspect, the present invention further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is run, the sensitive device failure probability evaluation method described above is implemented.

[0058] The present application provides a sensitive device failure probability evaluation method, system, device, and medium. Compared with the prior art, the beneficial effects of the embodiments of the present application are as follows:

[0059] The cross-modal medical image generation method disclosed in the present application can more effectively capture and generate the complex mapping relationship of the tissue structure between different modal images. It can not only focus on the local features of the image, but also consider the global context information, thereby improving the feature extraction ability; it can fuse global information and channel information, strengthen the ability to capture features of different scales, add a structural similarity index to the cycle consistency loss function to enhance the model's ability to retain lesion regions and anatomical structure details, and by introducing a perceptual loss function, encourage the generator to generate an output similar to the real image in the perceptual space, so that the target MRI image is closer to the real MRI image in terms of clarity and detail performance, and has a better match with the real image in terms of visual features and structure. Description of the Drawings

[0060] Figure 1 is a schematic diagram of the steps of a cross-modal medical image generation method provided by a preferred embodiment of the present invention;

[0061] Figure 2 is a schematic diagram of the structure of a target cross-modal medical image generation model provided by a preferred embodiment of the present invention;

[0062] Figure 3 is a schematic diagram of the structure of the improved first generator and second generator provided by a preferred embodiment of the present invention;

[0063] Figure 4 is a schematic diagram of the structure of the improved first discriminator and second discriminator provided by a preferred embodiment of the present invention;

[0064] Figure 5 is a schematic diagram of the structure of a cross-modal medical image generation system provided by a preferred embodiment of the present invention;

[0065] Figure 6 It is the internal structure diagram of the computer device in the embodiment of the present invention. Detailed implementation manners

[0066] The following will specifically clarify the implementation manners of the present invention in conjunction with the accompanying drawings. The provided embodiments are only for illustrative purposes and should not be construed as limiting the present invention. The included drawings are for reference and illustration only and do not constitute a limitation on the protection scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. In the description of the present invention, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.

[0067] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected to" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are only for illustrative purposes and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation on the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0068] In the description of the present invention, it should be noted that unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0069] Please refer to Figure 1 , in the embodiment of the present invention, a cross-modal medical image generation method is provided, and the method includes:

[0070] S1. Obtain sample image data, where the sample image data includes CT image data of the first modality and MRI image data of the second modality; using patients from different hospitals as data samples, collect CT images and MRI images at different times of the data samples, use the CT images as the image data of the first modality, use the MRI images as the image data of the second modality, and construct sample image data with the CT image data of the first modality and the MRI image data of the second modality.

[0071] S2. Build a cross-modal medical image generation model based on the Cycle Generative Adversarial Network; the Cycle Generative Adversarial Network (CycleGAN) consists of two generators and two discriminators, which are respectively used to transform images from one domain to another. Among them, one generator can transform images in domain X (such as human faces) into images in domain Y (such as cat faces), while the other generator can transform images in domain Y back into domain X. The goal of the discriminator is to accurately distinguish real images from the images generated by the generator, and the two discriminators are respectively used to determine whether the input image belongs to domain X and domain Y.

[0072] For medical CT images and MRI images, especially head and neck CT images and MRI images, they contain multiple different anatomical organs, their structures are complex and adjacent to each other, and they may exhibit high heterogeneity in terms of size, shape, density, and signal intensity, which makes it difficult for the generator to capture the consistent features of these tissues and organs. Moreover, the CT images and MRI images of patients are not scanned and generated at the same time, and it is impossible to ensure completely consistent scanning conditions, which will cause the CT images and MRI images to not fully match due to changes in scanning time, body position, body organ movement, and / or filling state.

[0073] In view of this, in an embodiment of the present application, a cross-modal medical image generation model is built based on the Cycle Generative Adversarial Network, as Figure 2 shown, the cross-modal medical image generation model includes a first generator, a second generator, a first discriminator, and a second discriminator.

[0074] S3. Train the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, the first generator is configured to convert the first feature set of the extracted CT image data into pseudo-MRI image data, and the second generator is configured to convert the second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes the first generator and the second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features and first tissue and organ global features at different scales, and also include a first discriminator corresponding to the first generator and a second discriminator corresponding to the second generator; during the model training process, first perform a chunking operation on the CT image data to obtain CT image chunks I CT , perform an affine transformation on the MRI image data to obtain MRI image chunks I MRI . Among them, the first generator is responsible for converting the chunked CT image chunks I CT into pseudo-MRI image chunks I' that are similar to real MRI images in terms of structure and content MRI , and the second generator is responsible for converting the pseudo-MRI image chunks I' generated by the first generator MRI into pseudo-CT image chunks I' that are similar to real CT image chunks I CT in terms of structure and content CT . The first discriminator is used to distinguish whether the pseudo-MRI image chunks I' MRI are real MRI image chunks I MRI or images generated by the first generator, and provide feedback to the first generator to help the first generator improve its parameters so that the generated images are more accurate. The second discriminator is used to distinguish whether the pseudo-CT image chunks I' CT are real CT image chunks or images generated by the second generator, and provide feedback to the second generator to help the second generator improve its parameters so that the generated images are more accurate. Through the design of the dual generators and dual discriminators, the complex mapping relationship of tissue structures between different modality images can be captured and generated more effectively.

[0075] Specifically, input the CT image chunks I CT into the first generator to obtain pseudo-MRI image chunks I' corresponding to the CT image data MRI , and input the MRI image chunks I MRI and the pseudo-MRI image chunks I' MRIInput the first discriminator to generate the first discrimination result. Based on the first discrimination result, optimize the first generator and the first discriminator according to the structural consistency loss function and the adversarial loss function. If the first discrimination result is true, it indicates that the generation quality of the first generator is very good. At this time, the first generator can be adjusted to optimize its performance for other characteristic image patches. For the first discriminator, it is a misjudgment, which means that the first discriminator fails to successfully distinguish the pseudo-MRI image patches and needs to be trained to update the parameters to enhance its discrimination ability. If the first discrimination result is false, it indicates that the first discriminator's judgment is successful, and the quality of the pseudo-MRI images generated by the first generator needs to be improved. Update the parameters of the first generator to make the pseudo-MRI image patches output by the first generator more realistic, so as to better "deceive" the first discriminator.

[0076] Furthermore, input the pseudo-MRI image patches into the second generator to obtain pseudo-CT image patches. Input the pseudo-CT image patches and the CT image patches into the second discriminator to generate the second discrimination result. Based on the second discrimination result, optimize the second generator and the second discriminator according to the structural consistency loss function and the adversarial loss function. If the second discrimination result is true, it indicates that the generation quality of the second generator is very good. At this time, the second generator can be adjusted to optimize its performance for other characteristic image patches. For the second discriminator, it is a misjudgment, which means that the second discriminator fails to successfully distinguish the pseudo-CT image patches and needs to be trained to update the parameters to enhance its discrimination ability. If the second discrimination result is false, it indicates that the second discriminator's judgment is successful, and the quality of the pseudo-CT images generated by the second generator needs to be improved. Update the parameters of the second generator to make the pseudo-CT image patches output by the second generator more realistic, so as to better "deceive" the second discriminator. Based on the optimized first generator, second generator, first discriminator, and second discriminator, obtain the trained target cross-modal medical image generation model.

[0077] In order to enhance the quality of image generation of different modalities and improve the ability to identify tissues and organs, in one embodiment of the present application, the first generator and the second generator are improved. The original generator in CycleGAN is usually based on the architecture of ResNet (Residual Network), which includes three parts: encoder, residual block and decoder. Encoder part (downsampling): convolution and downsampling operations are performed on the input image step by step to compress the spatial information of the high-resolution image into low-resolution, high-semantic information. Generally, a 7×7 convolution kernel is first used for the initial feature extraction, and then two 3×3 convolution operations are used with a stride of 2. Downsampling operations are performed each time, and the size of the feature map is gradually halved. The residual block is the core part of the original generator, which is mainly used to extract the mapping relationship between the input domain and the target domain, and increase the performance of the generator. Each residual block consists of two 3×3 convolution layers, and the number of residual blocks is generally 6 to 9 layers. Residual connection (ResidualConnection) is used to retain the content information between the input and output. The decoder performs deconvolution or transposed convolution operations on the downsampled features to restore the feature map to the same spatial size as the input image, and uses two 3×3 transposed convolutions with a stride of 2 to achieve layer-by-layer amplification of the spatial size. Finally, a 7×7 convolution kernel is used to generate the final output image, which has the same size as the input image. Although the original generator can effectively extract the deep semantic features of the image, it ignores the capture of shallow features to a certain extent, resulting in the loss of details and edge information of tissues and organs in the image.

[0078] like Figure 3As shown, the improved first generator and second generator of the present application at least include a shallow feature extraction module, a deep feature extraction module, and an upsampling module. The shallow feature extraction module at least includes a downsampling convolutional block and a residual convolutional block, specifically composed of a 7×7 convolutional block, three 3×3 residual convolutional blocks, and two 3×3 downsampling convolutional blocks. Taking the first generator as an example, the specific processing process is described as follows. The 7×7 convolutional block extracts features from the input image data. The downsampling convolutional block is used to reduce the size of the feature map generated by the 7×7 convolutional block and increase the feature depth. At the same time, the residual convolutional block aggregates texture information and semantic information at different levels to enhance the representation ability of the cross-modal medical image generation model for image features. Three first feature maps with different resolutions are obtained through the shallow feature extraction module. The smallest-resolution first feature map is input into the deep feature module, and the remaining several first feature maps are output as tissue and organ shallow feature maps, which at least include tissue and organ detail features and edge features. The deep feature extraction module at least includes a variable-sized window attention block (VSA). Organs are included in medical CT images and MRI images and are distributed in continuous spatial regions, and there are dependencies between these organs in distant voxels. In order to generate high-quality target MRI images, it is necessary to capture context information reflecting long-range dependencies. The variable-sized window attention block (VSA block) focuses on different parts of the input features under different window sizes, so as to be able to capture context information in different ranges and enhance the cross-modal medical image generation model's ability to capture local and global features. Considering the depth and complexity of the network, in an embodiment of the present application, the number of variable-sized window attention blocks is 6, so that the cross-modal medical image generation model can capture the features of tissues and organs in the image at different scales and more comprehensively understand the context information in the image. The deep feature extraction module also includes six 3×3 convolutional blocks, which are arranged alternately with the variable-sized window attention blocks. For the variable-sized window attention block, if its input is x, it traverses the LayerNorm (layer normalization) layer, and then the VSA block calculates the multi-head attention of the variable-sized window. Finally, the calculation result and the input x are superimposed to obtain the feature x'. The whole process can be expressed as follows:

[0079] x' = VSA(LN(x)) + x

[0080] Where LN is the processing of the LayerNorm layer, VSA is the multi-head attention processing of the variable-sized window. Further, the MLP (multi-layer perceptron) is also used for residual connection, which is expressed as follows:

[0081] x” = MLP(LN(x')) + x'

[0082] The deep feature extraction module extracts deep feature maps of tissues and organs, which at least include local features of the first tissue and organ at different scales and global features of the first tissue and organ.

[0083] Finally, the shallow feature map of the tissue and organ and the deep feature map of the tissue and organ are input into the upsampling module of the first generator for feature fusion to output a pseudo-MRI image patch. The upsampling module can use a conventional upsampling convolutional block.

[0084] The improved first generator and second generator in this application can fully extract the shallow semantic features and deep semantic features of the image, capture the features of tissues and organs in the image at different scales, and more comprehensively understand the context information in the image, enhancing the model's ability to capture local and global features.

[0085] The task of the discriminator is to distinguish between real images and pseudo-images generated by the generator. In the cross-modal medical image generation model of this application, the first discriminator is used to determine whether the pseudo-MRI image is real or generated by the first generator. If the discrimination result is that the pseudo-MRI image is real, it means that the first discriminator misjudges the pseudo-MRI image as a real image, indicating that the generation quality of the first generator is very good. At this time, the first generator can be adjusted to optimize its performance for other characteristic images. For the first discriminator, misjudgment means that the first discriminator fails to successfully distinguish the pseudo-MRI image, and it needs to be trained to update its parameters to enhance its discrimination ability. If the discrimination result is that the pseudo-MRI image is fake, it means that the first discriminator has distinguished the pseudo-MRI image as a forged image, indicating that the generation quality of the first generator is not good. At this time, the parameters of the first generator are updated to make the image patch output by the first generator more realistic, so as to better "deceive" the first discriminator. For the first discriminator, the task has been correctly completed and no parameter adjustment is required. The logical relationship between the second discriminator and the second generator is similar to that between the first discriminator and the first generator, except that the discrimination object and the image generation object are different.

[0086] Traditional discriminators usually consist of multiple downsampling convolutional layers and introduce fully connected layers at appropriate scales to flatten the feature vectors and increase the number of parameters to achieve the discrimination of the entire image. Although this design can capture the global structural information in the image, it is not sufficient to capture image details and local features.

[0087] In an embodiment of this application, as Figure 4As shown, the first discriminator and the second discriminator are designed as U-shaped structures. Through the skip connection mechanism, the detailed information in the encoder directly flows to the decoder, and a global perception channel focusing module (GPCF) is designed to further improve the performance of the first discriminator and the second discriminator. The real images and fake images input into the first discriminator and the second discriminator first undergo feature extraction through the feature extraction convolutional block to capture the local features and the detailed information of tissues and organs in the images. The ReLU activation function is used to increase the non-linear expression ability of the model, so as to learn more complex features. Subsequently, the local features and the detailed information of tissues and organs generated by the feature extraction convolutional block are input into the global perception channel focusing module (GPCF) that integrates multi-scale attention pooling (MAP) and global average pooling (GAP). MAP performs pooling operations on the local features and the detailed information of tissues and organs at different scales, captures the local feature representations of tissues and organs in the images at different scales, enables the model to understand the multi-scale context of the images, and obtains the second global features of tissues and organs at different scales. GAP performs average pooling operations on each channel of the local features and the detailed information of tissues and organs to generate a feature representation containing global information and obtains the second global features of tissues and organs. The output of MAP is added to the result of GAP to generate a feature representation that contains both local details and global context, enhancing the model's ability to capture features at different scales. The fused feature map passes through a one-dimensional convolutional layer with an adaptive convolution kernel to capture the dependencies between channels and generate an attention weight map. The attention weight map is converted into attention weights through the sigmoid activation function, and the generated attention weights are applied to the feature maps of the original input real images and fake images to adjust the attention weights of the feature maps of the real images and fake images, making the discriminator pay more attention to the features that have an important impact on discriminating real and fake images. After the attention-adjusted feature maps of the real images and fake images are subjected to max pooling, they pass through a series of upsampling transposed convolutional blocks to gradually restore to the spatial dimension of the original images and contain the detailed information required for discriminating real and fake, obtaining the real and fake discrimination features for each pixel. Finally, based on the real and fake discrimination features, the discrimination results are output through the output layer.

[0088] The improved first discriminator and second discriminator in this application can aggregate information at different scales, enabling the model to capture the global information of the images, and integrating the global information and channel information to strengthen the ability to capture features at different scales. The discriminative results of real and fake images for each pixel are output through the output layer, effectively improving the accuracy and quality of the generated images and enhancing the performance of the discriminator.

[0089] The loss functions of the cyclic generative adversarial network mainly include the cyclic consistency loss function, the adversarial loss function, and the identity loss function.

[0090] The cycle consistency loss function is to ensure that the main features of the original input image remain unchanged. After a cross - domain transformation is completed and then reversed back to the original input image, a similar result should be obtained. That is, when a CT image is converted into a pseudo - MRI image and then converted back to a pseudo - CT image, the resulting pseudo - CT image should be the same as the original CT image. The constraint of the cycle consistency loss function helps prevent the cross - modal medical image generation model from learning meaningless or incorrect relationships and ensures that the converted pseudo - CT image still has the key attributes of the input CT image.

[0091] In an embodiment of the present application, the cycle consistency loss is optimized. On the basis of the cycle consistency loss, the Structural Similarity Index Measure (SSIM) is added to obtain the structural consistency loss function. The cycle consistency loss itself does not consider the visual structure of the image and the semantic information of the content. Although it ensures the consistency of the image during the conversion process, however, it has limitations in retaining the details of the lesion area and anatomical structure, which are decisive for medical diagnosis and treatment. The SSIM loss is an index for measuring the similarity of the visual effects of images and can evaluate the similarity of the brightness, contrast, and structural information between the real image and the pseudo - image. Combining the SSIM loss with the cycle consistency loss as the structural consistency loss can simultaneously control the pixel - level accuracy and visual effects of the generated pseudo - image, thereby generating a more realistic and detail - rich pseudo - image to enhance the ability of the cross - modal medical image generation model in retaining the details of the lesion area and anatomical structure. The structural consistency loss function of the present application is expressed as follows:

[0092] L scycle = L1+λ(1 - SSIM)

[0093] Where L1 is the cycle consistency loss, represented by the pixel difference between the real image x and the pseudo - image y, and λ is the weight parameter.

[0094] For the real image x and the pseudo - image y, the SSIM loss is expressed as follows:

[0095]

[0096] Where μ x is the average value of x, μ y is the average value of y, δ x is the variance of x, δ y is the variance of y, δ xy is the covariance of x and y, and c1 and c2 are constants.

[0097] In an embodiment of the present application, a Perceptual Loss is further introduced to measure the difference between the generated pseudo-image and the reference image in the feature space. In this way, the generator is encouraged to produce an output similar to the reference image in the perceptual space, so that the generated pseudo-image is closer to the reference image in terms of clarity and detail performance, and has a better match with the reference image in terms of visual features and structure. The reference image in the present application is MRI image data. The present application uses the VGG19 model pre-trained based on the ImageNet image database to construct a perceptual loss module, extracts high-level features for supervised learning to obtain a perceptual loss function, and the perceptual loss function is expressed as:

[0098]

[0099] where j is a specific layer in the cross-modal medical image generation model, is a feature extraction function, G x is the pseudo-image, x target is the real image.

[0100] By summing up the perceptual losses of multiple layers, the detail fidelity and overall semantic consistency between the pseudo-image and the reference image can be balanced.

[0101] The generative adversarial loss function aims to make the pseudo-image generated by the generator able to deceive the discriminator so that it cannot distinguish between the real image and the pseudo-image, and promotes the competition and progress between the generator and the discriminator during the training process, so as to achieve the goal of generating high-quality pseudo-images. The adversarial loss function is expressed as follows:

[0102] L GAN (G MRI ,G CT ,D MRI ,D CT )=E[log(D MRI (I MRI ))]+E[1-log(D MRI (G MRI (I CT )))]

[0103] +E[log(D CT (I CT ))]+E[1-log(D CT (G CT (G MRI (I CT ))))]

[0104] where G MRI represents the first generator, G CT represents the second generator, D MRI represents the first discriminator, DCT Denotes the second discriminator, I MRI Denotes an MRI image patch, I MRI Denotes a CT image patch, E denotes taking the expected value.

[0105] The objective of the identity loss function is to minimize the difference between the generated fake image and the original input image, thereby ensuring the consistency between the input and output images. The identity loss function is expressed as follows:

[0106] Lidentity(G CT , G MRI ) = E[||G CT (I MRI )||] + E[||G MRI (I CT ) - I CT ||]

[0107] The structural consistency loss function, perceptual loss function, adversarial loss function, and identity loss function all have their respective advantages, and can improve the clarity of the generated fake MRI image data based on aspects such as the structural similarity, high-level semantic features, and detailed textures of medical images. Fusing the above various loss functions into the total loss objective function of the cross-modal medical image generation model can improve the similarity between the complex structures and detailed textures of the fake MRI image data generated by the cross-modal medical image generation model and the MRI image data, as well as the accuracy of the generated fake MRI image data. The total loss objective function is expressed as:

[0108] L total = L GAN + L identity + λ scycle · L scycle + λ perc · L perc

[0109] Where λ scycle , λ perc Denote weight parameters, used to balance the structural consistency loss function term and the perceptual loss function term.

[0110] During the training process, optimize the parameters of the cross-modal medical image generation model based on the structural consistency loss function, perceptual loss function, adversarial loss function, and identity loss function to obtain the target cross-modal medical image generation model.

[0111] S4. During the actual image generation process, input each of the collected CT images to be converted into the target cross-modal medical image generation model, and construct a target MRI image dataset according to the output result of the target cross-modal medical image generation model; in the actual application of image generation, perform a blocking operation on each of the collected CT images to be converted to generate corresponding CT image blocks to be converted, and input the corresponding CT image blocks to be converted into the first generator of the target cross-modal medical image generation model to obtain corresponding pseudo-MRI image blocks; input the corresponding pseudo-MRI image blocks into the second generator of the target cross-modal medical image generation model to obtain corresponding pseudo-CT image blocks; at this time, input the pseudo-CT image blocks and the corresponding CT image blocks to be converted into the second discriminator of the target cross-modal medical image generation model to generate a real-time discrimination result to judge the image generation quality of the first generator and the second generator. If the real-time discrimination result is that the pseudo-CT image block is false, perform dynamic optimization on the first generator and the second generator of the target cross-modal medical image generation model until the real-time discrimination result is that the pseudo-CT image block to be converted is true. At this time, output the target MRI image dataset with the corresponding pseudo-MRI image blocks. During the process of generating the cross-modal medical image to be converted, using the second discriminator to perform dynamic optimization on the first generator and the second generator can not only make the target MRI image set output each time more accurate, but also optimize the target cross-modal medical image generation model in real time according to the update of the CT image, ensuring that the target cross-modal medical image generation model can adapt to the update of the CT technology.

[0112] In a preferred embodiment of the present application, when constructing the target MRI image dataset according to the output result of the target cross-modal medical image generation model, it is also necessary to perform similarity verification on the output result of the target cross-modal medical image generation model, and generate the target MRI image dataset based on the output result that passes the similarity verification. The similarity verification calculates the variance of the peak signal-to-noise ratio of different MRI image data in the output MRI image dataset, and compares the variance with a preset threshold, and generates the target MRI image dataset according to the comparison result. If the variance is less than the threshold, it indicates that the generated MRI image quality is good, and the similarity verification is passed, and the MRI image dataset that passes the similarity verification is output as the target MRI image dataset.

[0113] In a preferred embodiment of the present invention, sample image data is obtained, wherein the sample image data includes CT image data of a first modality and MRI image data of a second modality; a cross-modal medical image generation model is constructed based on a cyclic generative adversarial network; the cross-modal medical image generation model is trained according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, the first generator is configured to convert the first feature set of the extracted CT image data into pseudo-MRI image data, and the second generator is configured to convert the second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes a first generator and a second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features of different scales, and first tissue and organ global features. In the actual image generation process, each to-be-converted CT image collected is input into the target cross-modal medical image generation model, and a target MRI image data set is constructed according to the output result of the target cross-modal medical image generation model. The cross-modal medical image generation method disclosed in this application can more effectively capture and generate the complex mapping relationship of tissue structures between different modality images. It can not only focus on the local features of the images, but also take into account the global context information, thereby improving the feature extraction ability; it can fuse global information and channel information, strengthen the ability to capture features of different scales, add a structural similarity index to the cyclic consistency loss function to enhance the model's ability to retain lesion regions and anatomical structure details, and by introducing a perceptual loss function, encourage the generator to generate an output similar to the real image in the perceptual space, so that the target MRI image is closer to the real MRI image in terms of clarity and detail performance, and has a better matching degree with the real image in terms of visual features and structure.

[0114] Correspondingly, as Figure 5 shown, based on a cross-modal medical image generation method, an embodiment of the present invention further provides a cross-modal medical image generation system for implementing the cross-modal medical image generation method disclosed in the embodiment of the present invention, including: a sample image data acquisition unit 1, a model construction unit 2, a model training unit 3, and a model application unit 4;

[0115] The sample image data acquisition unit 1 is configured to obtain sample image data, wherein the sample image data includes CT image data of a first modality and MRI image data of a second modality;

[0116] The model construction unit 2 is configured to construct a cross-modal medical image generation model based on a cyclic generative adversarial network;

[0117] The model training unit 3 is configured to train the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, the first generator is configured to convert the first feature set of the extracted CT image data into pseudo-MRI image data, and the second generator is configured to convert the second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes the first generator and the second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features at different scales, and first tissue and organ global features.

[0118] The model application unit 4 is configured to input each collected CT image to be converted into the target cross-modal medical image generation model during the actual image generation process, and construct a target MRI image data set according to the output result of the target cross-modal medical image generation model.

[0119] For the specific limitations of a cross-modal medical image generation system, reference may be made to the above limitations of a cross-modal medical image generation method, which will not be elaborated here. Those of ordinary skill in the art can realize that, in combination with the various modules and steps described in the embodiments disclosed in the present invention, they can be implemented in hardware, software, or a combination of both. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0120] As Figure 6 shown, a computer device provided in an embodiment of the present invention includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps in the cross-modal medical image generation embodiment as described above, such as Figure 1 the steps S1 to S4 described therein.

[0121] Those skilled in the art can understand that the schematic Figure 6 is only an example of a computer device and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine some components, or different components. For example, the computer device may further include input and output devices, network access devices, a bus, etc.

[0122] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits.

[0123] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, the image playback function, etc.); the data storage area can store the data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.

[0124] Among them, if the modules integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0125] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above-described embodiment methods, it can be completed by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0126] Correspondingly, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. Among them, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the steps in the cross-modal medical image generation in the above-described embodiment, such as Figure 1 the steps S1 to S4 described therein.

[0127] In summary, a cross-modal medical image generation method, system, device, and medium provided by embodiments of the present application solve the technical problem of how to improve the accuracy and matching degree of converting CT images into MRI images. The method includes: obtaining sample image data, where the sample image data includes CT image data of a first modality and MRI image data of a second modality; constructing a cross-modal medical image generation model based on a cyclic generative adversarial network; training the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, a first generator is configured to convert a first feature set of the extracted CT image data into pseudo-MRI image data, and a second generator is configured to convert a second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes a first generator and a second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features of different scales, and first tissue and organ global features; in the actual image generation process, each to-be-converted CT image collected is input into the target cross-modal medical image generation model, and a target MRI image data set is constructed according to the output result of the target cross-modal medical image generation model. The cross-modal medical image generation method disclosed in the present application can more effectively capture and generate complex mapping relationships of tissue structures between different modality images. It can not only focus on local features of images but also consider global context information, thereby improving the feature extraction ability; it can fuse global information and channel information to strengthen the ability to capture features of different scales, add a structural similarity index to the cyclic consistency loss function to enhance the model's ability to retain lesion regions and anatomical structure details, and by introducing a perceptual loss function, encourage the generator to produce outputs similar to real images in the perceptual space, so that the target MRI images are closer to real MRI images in terms of clarity and detail performance, and have a better match with real images in terms of visual features and structures.

[0128] Each embodiment in this specification is described in a progressive manner. For parts that are the same or similar in each embodiment, they can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0129] The above-described embodiments merely represent several preferred embodiments of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patented application. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present application, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the protection scope of the claims described above.

Claims

1. A cross-modal medical image generation method, characterized in that The method includes: Obtaining sample image data, where the sample image data includes CT image data of a first modality and MRI image data of a second modality; Constructing a cross-modal medical image generation model based on a cyclic generative adversarial network; Training the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, a first generator is configured to convert a first feature set of the extracted CT image data into pseudo-MRI image data, and a second generator is configured to convert a second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes the first generator and the second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features of different scales, and first tissue and organ global features; During the actual image generation process, input each to-be-converted CT image collected into the target cross-modal medical image generation model, and construct a target MRI image data set according to the output result of the target cross-modal medical image generation model.

2. The cross-modal medical image generation method according to claim 1, wherein The target cross-modal medical image generation model further includes a first discriminator and a second discriminator; The first discriminator is configured to perform authenticity discrimination based on the second feature set extracted from the MRI image data and the pseudo-MRI image data to obtain a discrimination result; The second discriminator is configured to perform authenticity discrimination based on the second feature set extracted from the CT image data and the pseudo-CT image data to obtain a discrimination result. The second feature set at least includes: second tissue and organ local features of different scales and second tissue and organ global features.

3. The cross-modal medical image generation method according to claim 2, wherein The training the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model includes: Performing a blocking operation on the CT image data to obtain CT image blocks, inputting the CT image blocks into the first generator to obtain pseudo-MRI image blocks corresponding to the CT image data; Performing an affine transformation on the MRI image data to obtain MRI image blocks, inputting the MRI image blocks and the pseudo-MRI image blocks into the first discriminator to generate a first discrimination result; Optimizing the first generator and the first discriminator based on the first discrimination result; Inputting the pseudo-MRI image blocks into the second generator to obtain pseudo-CT image blocks; Inputting the pseudo-CT image blocks and the CT image blocks into the second discriminator to generate a second discrimination result; Optimizing the second generator and the second discriminator based on the second discrimination result; Based on the optimized first generator, second generator, first discriminator, and second discriminator, obtain the trained target cross-modal medical image generation model.

4. The cross-modal medical image generation method according to claim 3, wherein The inputting the CT image blocks into the first generator to obtain pseudo-MRI image blocks corresponding to the CT image data includes: Input the CT image block into the shallow feature extraction module of the first generator to extract shallow features at different levels of the CT image block, generate several first feature maps with different resolutions, extract the feature map with the smallest resolution among the several first feature maps, and output the remaining several first feature maps as the shallow feature maps of the tissue and organ; the shallow feature extraction module includes at least several downsampling convolutional blocks and several residual convolutional blocks; the shallow feature maps of the tissue and organ include at least the detailed features and the edge features of the tissue and organ; Input the feature map with the smallest resolution into the deep feature extraction module of the first generator to extract the deep features of the feature map with the smallest resolution and generate deep feature maps of the tissue and organ with different sizes; the deep feature extraction module includes at least several variable-size window attention blocks; the deep feature maps of the tissue and organ include at least the first local features and the first global features of the tissue and organ at different scales; Input the shallow feature maps of the tissue and organ and the deep feature maps of the tissue and organ into the upsampling module of the first generator for feature fusion and output a pseudo-MRI image block.

5. The cross-modal medical image generation method according to claim 3, wherein, The step of inputting the MRI image block and the pseudo-MRI image block into the first discriminator to generate a first discrimination result includes: Input the MRI image block and the pseudo-MRI image block into the first discriminator, and the feature extraction convolutional block of the first discriminator extracts the tissue and organ detail information of the MRI image block and the pseudo-MRI image block; the first discriminator is of a U-shaped structure; Input the tissue and organ detail information into the global perception channel focusing module of the first discriminator. The multi-size attention pooling layer of the global perception channel focusing module performs pooling operations on the tissue and organ detail information at different scales to obtain the second local features of the tissue and organ at different scales. The global average pooling layer of the global perception channel focusing module performs average pooling operations on each channel of the tissue and organ detail information to obtain the second global feature of the tissue and organ; Fuse the second global features and the second local features of the tissue and organ at different scales and perform adaptive attention weight calculation to obtain an attention weight map; Link the attention weight map with the MRI image block and the pseudo-MRI image block respectively and perform upsampling processing to obtain the true / false discrimination features of each pixel; Generate a first discrimination result based on the true / false discrimination features.

6. The cross-modal medical image generation method according to claim 3, wherein, The total objective function of the cross-modal medical image generation model includes a structural consistency loss function, a perceptual loss function, an adversarial loss function, and an identity loss function; The total objective function is expressed as: L total = L GAN + L identity + λ scycle · L scycle + λ perc · L perc Among them, L GAN represents the adversarial loss function term, L identity represents the identity loss function term, L scycle represents the structural consistency loss function term, L perc represents the perceptual loss function term, λ scycle 、λ perc represent weight parameters to balance the structural consistency loss function term and the perceptual loss function term; The structural consistency loss function is: Among them, L1 is the cycle consistency loss, which is represented by the pixel difference between the real image x and the pseudo-image y. λ is the weight parameter, μ x is the average value of x, μ y is the average value of y, δ x is the variance of x, δ y is the variance of y, δ xy is the covariance of x and y, and c1 and c2 are constants; the real image includes the CT image data and the MRI image data, and the pseudo-image includes the pseudo-MRI image block and the pseudo-CT image block; The perceptual loss function is: Among them, j is a specific layer in the cross-modal medical image generation model, is a feature extraction function, G x is a pseudo-image, x target is a real image; The adversarial loss function is: L GAN (G MRI ,G CT ,D MRI ,D CT ) = E[log(D MRI (I MRI ))] + E[1 - log(D MRI (G MRI (I CT )))] +E[log(D CT (I CT ))]+E[1 - log(D CT (G CT (G MRI (I CT ))))] Among them, G MRI represents the first generator, G CT represents the second generator, D MRI represents the first discriminator, D CT represents the second discriminator, I MRI represents an MRI image patch, I CT represents a CT image patch; The identity loss function is: Lidentity(G CT ,G MRI ) = E[||G CT (I MRI )||] + E[||G MRI (I CT ) - I CT ||].

7. The cross-modal medical image generation method according to claim 1, wherein The step of constructing a target MRI image dataset according to the output result of the target cross-modal medical image generation model includes: Perform similarity verification on the output result of the target cross-modal medical image generation model, and generate the target MRI image dataset based on the output result that passes the similarity verification.

8. A cross-modal medical image generation system implementing the cross-modal medical image generation method according to any one of claims 1-7, characterized in that, The system includes: a sample image data acquisition unit, a model construction unit, a model training unit, and a model application unit; The sample image data acquisition unit is configured to acquire sample image data, where the sample image data includes CT image data of a first modality and MRI image data of a second modality; The model construction unit: is configured to construct a cross-modal medical image generation model based on a cyclic generative adversarial network; The model training unit: is configured to train the cross-modal medical image generation model according to the sample image data to obtain a target cross-modal medical image generation model. During the training process, the first generator is configured to convert the first feature set of the extracted CT image data into pseudo-MRI image data, and the second generator is configured to convert the second feature set extracted from the pseudo-MRI image data into pseudo-CT image data; the adversarial network architecture in the cross-modal medical image generation model at least includes the first generator and the second generator; both the first feature set and the second feature set at least include: tissue and organ detail features, edge features, first tissue and organ local features of different scales, and first tissue and organ global features; The model application unit: during the actual image generation process, input each to-be-converted CT image collected into the target cross-modal medical image generation model, and construct a target MRI image dataset according to the output result of the target cross-modal medical image generation model.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a transceiver, which are connected through a bus; the memory is used to store a set of computer program instructions and data, and transmit the stored data to the processor, and the processor executes the program instructions stored in the memory to execute the cross-modal medical image generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored in the computer-readable storage medium, and when the computer program is run, the cross-modal medical image generation method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Medical image modal conversion method and device based on pre-training StyleGAN2

    CN121391591A