Synthetic CT image generation method, system and device based on multi-sequence MR images, medium and product

Feature extraction and fusion of multi-sequence MR images are performed through cross-attention generation adversarial network (CAGAN), and high-quality synthetic CT images are generated, solving the problem of insufficient cross-modal feature alignment and multi-scale details retention in single-sequence MR images, improving the diagnostic and treatment accuracy of images.

CN120411287APending Publication Date: 2025-08-01SHANDONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510555645.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing synthetic CT imaging method based on single-sequence MR images has shortcomings in cross-modal feature alignment and multi-scale detail retention, which is difficult to meet clinical needs.

Method used

Cross Attention Generation Adversarial Network (CAGAN) is used to extract features and fusion of complementary information of multi-sequence MR images, and high-quality synthetic CT images are generated through the cooperation of generators and discriminators.

Benefits of technology

The generated synthetic CT images optimize the anatomical structure and detailed information of the prior art, provide more accurate anatomical and tissue density information, and improve the accuracy of clinical diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411287A_ABST
    Figure CN120411287A_ABST
Patent Text Reader

Abstract

The invention discloses a synthetic CT image generation method, system and device based on a multi-sequence MR image, a medium and a product, and relates to the field of medical image processing, and the method comprises the steps: obtaining a to-be-detected multi-sequence MR image; constructing and training a cross attention generative adversarial network; and inputting a to-be-tested multi-sequence MR image into the trained generator of the cross attention generative adversarial network to generate a synthetic CT image. According to the invention, complementary information fusion of different sequence MR images can be realized, so that a high-quality sCT image is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical image processing, and particularly to a method, system, device, medium, and product for generating synthetic CT images based on multi-sequence MR images. Background Art

[0002] Radiation therapy is one of the main means for treating various cancers. With the development of radiology, radiotherapy methods have shifted from conventional external beam radiotherapy to three-dimensional conformal radiotherapy (3DCRT), image-guided radiotherapy (IGRT), and intensity-modulated radiotherapy (IMRT). Before implementing IMRT, it is necessary to obtain the patient's computed tomography (CT) images for radiotherapy dose planning. At the same time, magnetic resonance (MR) images of the corresponding spatial structure are also required to segment tumor tissues from surrounding normal tissues. MR provides high-contrast soft tissue imaging, while CT provides excellent bone imaging. CT imaging technology is based on the difference in the absorption degree of X-rays by tissues with different densities, which enables the clear display of the fine texture of bone structures. In addition, the pixel intensity values in CT images are closely related to the electron density data necessary for formulating radiotherapy plans. However, due to the small density differences among various soft tissues inside the human body, CT imaging faces challenges in accurately identifying these soft tissues. At the same time, the ionizing radiation associated with CT scans poses a potential risk to the health of patients, which may cause additional physical damage. To avoid the risk of this additional radiation, MR-only based radiotherapy plans have been applied during patient treatment, enabling more patients to avoid additional CT scans. By generating pseudo-CT images from MR, radiotherapy plans can obtain more comprehensive information in one image, enabling patients to receive more accurate diagnoses and treatment plans.

[0003] How to generate synthetic CT images (sCT) from MR images has become an important topic in medical image research. Existing methods are mainly based on single-sequence MR images, but the limitations of single-modal information lead to obvious deficiencies in cross-modal feature alignment, multi-scale detail retention, etc. of sCT images, making it difficult to meet clinical requirements. Summary of the Invention

[0004] The purpose of this application is to provide a method, system, device, medium and product for generating synthetic CT images based on multi-sequence MR images. By using a Cross-Attention Generative Adversarial Network (CAGAN), the complementary information of different sequence MR images is fused to generate high-quality sCT images, providing more accurate imaging basis for clinical diagnosis and treatment.

[0005] To achieve the above object, this application provides the following solutions:

[0006] In the first aspect, this application provides a method for generating synthetic CT images based on multi-sequence MR images, including:

[0007] Obtain multi-sequence MR images to be measured;

[0008] Construct and train a cross-attention generative adversarial network; the cross-attention generative adversarial network includes a generator and a discriminator; the generator includes a dual-branch feature extraction module, a cross-attention feature fusion module, and a decoder; the discriminator includes multiple PatchGAN modules; the generator is used to generate synthetic CT images according to multi-sequence MR images; the discriminator is used to distinguish the authenticity of the synthetic CT images output by the generator;

[0009] Input the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images.

[0010] In the second aspect, this application provides a system for generating synthetic CT images based on multi-sequence MR images, including:

[0011] A data acquisition module, configured to obtain multi-sequence MR images to be measured;

[0012] A model construction and training module, configured to construct and train a cross-attention generative adversarial network; the cross-attention generative adversarial network includes a generator and a discriminator; the generator includes a dual-branch feature extraction module, a cross-attention feature fusion module, and a decoder; the discriminator includes multiple PatchGAN modules; the generator is used to generate synthetic CT images according to multi-sequence MR images; the discriminator is used to distinguish the authenticity of the synthetic CT images output by the generator;

[0013] A synthetic CT image generation module, configured to input the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images.

[0014] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the above-described method for generating a synthetic CT image based on multi-sequence MR images.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-described method for generating a synthetic CT image based on multi-sequence MR images.

[0016] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-described method for generating a synthetic CT image based on multi-sequence MR images.

[0017] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0018] The present application provides a method, system, device, medium, and product for generating a synthetic CT image based on multi-sequence MR images. Through the dual-branch feature extraction module and cross-attention feature fusion module of the generator in the CAGAN network, feature extraction and complementary information fusion of multi-sequence MR images are achieved, thereby generating high-quality sCT images, realizing the conversion from MR images to sCT images, optimizing the anatomical structure and detail information of the generated images, and solving problems such as missing bone structures, blurred soft tissues, and anatomical alignment errors caused by single-sequence MR image input in the prior art. The present application is applicable to scenarios such as medical image synthesis, radiotherapy planning, and image-assisted diagnosis, and can provide more accurate anatomical and tissue density information, improving the accuracy of clinical diagnosis and treatment. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0020] Figure 1 It is an application environment diagram of a method for generating a synthetic CT image based on multi-sequence MR images in an embodiment of the present application;

[0021] Figure 2 It is a flowchart of a method for generating a synthetic CT image based on multi-sequence MR images provided by an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of the working process of a cross-attention generative adversarial network;

[0023] Figure 4 Schematic diagram of the generator architecture in the cross-attention generative adversarial network;

[0024] Figure 5 Schematic diagram of the structure of the cross-attention feature fusion module;

[0025] Figure 6 Schematic diagram of the structure of the differential feature processing sub-module;

[0026] Figure 7 Schematic diagram of the structure of the common feature processing sub-module;

[0027] Figure 8 Schematic diagram of the discriminator architecture in the cross-attention generative adversarial network;

[0028] Figure 9 sCT image intention generated from the input dual-sequence T1 / T2 MR images and single-sequence T1 MR images;

[0029] Figure 10 Box plot schematic diagram of the quantitative results of sCT generation by the cross-attention generative adversarial network and other generative models; among them, (a) is the box plot schematic diagram of PSNR comparison, (b) is the box plot schematic diagram of SSIM comparison, and (c) is the box plot schematic diagram of MAE comparison;

[0030] Figure 11 Schematic diagram of the anatomical structure alignment effect of sCT generation;

[0031] Figure 12 Box plot schematic diagram of the quantitative results comparison between the cross-attention generative adversarial network and the cross-attention module; among them, (a) is the box plot schematic diagram of PSNR comparison, (b) is the box plot schematic diagram of SSIM comparison, and (c) is the box plot schematic diagram of MAE comparison;

[0032] Figure 13 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0034] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0035] The synthetic CT image generation method based on multi-sequence MR images provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown in the figure. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the multi-sequence MR images to be measured to the server 104. After receiving the multi-sequence MR images to be measured, the server 104 constructs and trains a cross-attention generative adversarial network, and inputs the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images. The server 104 can feedback the obtained synthetic CT images to the terminal 102. In addition, in some embodiments, the synthetic CT image generation method based on multi-sequence MR images can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly generate synthetic CT images for the multi-sequence MR images to be measured, or the server 104 can obtain the multi-sequence MR images to be measured from the data storage system and generate synthetic CT images for the multi-sequence MR images to be measured.

[0036] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0037] In an exemplary embodiment, as Figure 2 shown in the figure, a synthetic CT image generation method based on multi-sequence MR images is provided. This method is executed by a computer device, and can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method applied to Figure 1 the server 104 in the figure as an example for illustration, it includes the following steps S1 to S3.

[0038] Among them:

[0039] S1: Obtain multi-sequence MR images to be measured.

[0040] S2: Construct and train a cross-attention generative adversarial network; the cross-attention generative adversarial network includes a generator and a discriminator; the generator includes a dual-branch feature extraction module, a cross-attention feature fusion module, and a decoder; the discriminator includes multiple PatchGAN modules (processing images of different resolutions respectively). As Figure 3 shown, the generator is used to generate a synthetic CT image based on multi-sequence MR images (T1MR + T2MR) (i.e., G(x): sCT); the discriminator is used to distinguish the authenticity of the synthetic CT image output by the generator.

[0041] S3: Input the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images.

[0042] In a specific embodiment, the architecture of the generator in the cross-attention generative adversarial network is as Figure 4 shown. The generator consists of two core modules: feature extraction and fusion. The former extracts multi-scale features of T1 / T2 sequence MR images through the dual-branch feature extraction module respectively, and the latter realizes the feature fusion of differences and commonalities through the cross-attention feature fusion module (i.e., the CA module), and finally decodes through the decoder to reconstruct and generate a synthetic CT image.

[0043] The dual-branch feature extraction module adopts a symmetric encoding-decoding structure. The encoder of each branch contains 5 cascaded convolutional modules, and each convolutional module includes two groups of convolutional layers with a kernel size of 3×3 (for extracting shallow local features), a batch normalization layer, and a ReLU activation function. Different-scale feature information is extracted through batch normalization (BN) and the ReLU activation function, and each channel uses shared parameters to ensure the consistency of feature fusion of different modalities. Except for the first layer, each module is configured with a downsampling operation with a stride of 2, and the output channel numbers increase sequentially to 32 / 64 / 128 / 256 / 512, effectively capturing multi-resolution feature representations. The corresponding decoder includes multiple transposed convolutional modules, which are 4 transposed convolutional modules in this embodiment. The transposed convolutional module uses bilinear interpolation upsampling to replace the downsampling layer, and the channel numbers decrease to 128 / 64 / 32 / 1, and feature reconstruction is achieved through skip connections. It should be noted that the dual-branch feature extraction module adopts a parallel training mechanism to ensure the alignment of modal features through parameter co-optimization.

[0044] The cross-attention feature fusion module innovatively adopts a three-level CA attention mechanism to implement cross-modal feature interaction at the 3rd - 5th levels of the encoder. The feature extraction and fusion process is as follows: Through channel-domain feature recalibration, dynamic weighted fusion of T1 / T2 sequence features is achieved. First, the global spatial statistics of each channel are calculated, then channel attention weights are generated through a fully connected layer, and finally, the weighted multi-modal features are input into the decoder for image reconstruction. This hierarchical fusion strategy not only retains the detailed texture information of the low levels but also integrates the semantic features of the high levels, enabling the finally generated sCT image to accurately reflect the contrast characteristics of the target modality while maintaining the continuity of the anatomical structure.

[0045] The cross-attention feature fusion module includes a differential feature processing sub-module and a common feature processing sub-module. The differential feature processing (DFP) sub-module is used to extract the differential features of multi-sequence MR images and perform weighted fusion through the channel attention mechanism. The common feature processing (CFP) sub-module is used to extract the common features of multi-sequence MR images and use a multi-layer perceptron (MLP) for information encoding to improve the effectiveness of feature fusion.

[0046] The goal of this application is to obtain high-quality sCT images with significant feature information while retaining rich texture details. Therefore, how to make full use of the differential and common feature information existing in multi-sequence MR images is a key factor in fusion performance. Based on the argument that the cross-attention mechanism can effectively extract the common features between images, this embodiment constructs a cross-attention feature fusion module including a DFP sub-module and a CFP sub-module, where there are two CFP sub-modules. The detailed structure of the CA module is as Figure 5 shown. The reconstructed T1 / T2 sequence MR images are used as the input and fed into the CA module to fuse and output the differential and common features of the two sequences of MR. The CA module realizes the information fusion between different modalities or different features through the interaction of the three basic elements of the Transformer: Query, Key, and Value.

[0047] Assume the input features are F T1 、F T2 , and the fused feature F CA is generated after passing through the CA module. The feature fusion process can be formulated as:

[0048] F CA =CA(F T1 ,F T2 )

[0049] Among them, CA(·) includes a DFP(DFP(·)) and a pair of CFP(CFP(·)), whose function is to extract global differential features and common features. It can be expressed by the formula as:

[0050] Z1 = DFP(F T1 , F T2 )

[0051] Z2 = CFP(F T2 , Z1)

[0052] Z3 = CFP(F T1 , Z2)

[0053] CA(F T1 , F T2 ) = Z3 + Z1

[0054] Among them, Z1, Z2, and Z3 are the outputs of the DFP sub-module, the first CFP sub-module, and the second CFP sub-module respectively.

[0055] In order to effectively obtain the differences between the T1MR and T2MR image features reconstructed in the previous stage, a DFP sub-module as shown in Figure 6 is constructed, which takes F T1 , F T2 as inputs and the difference information features between them as outputs. Specifically, in order to study the global feature relationship of the input T1 / T2 sequence MR, F T1 , F T2 are divided into s local feature segments, and the formula is as follows:

[0056] Q1,..., Q s = Partition(F T2 )

[0057] K1,..., K s = Partition(F T1 )

[0058] V1,..., V s = Partition(F T1 )

[0059] Among them and s = h × w. Then, a linear layer is used to convert them into query Q, key K, and value V. The linear mapping can be expressed as:

[0060] Q i = Linear Q (Q i )

[0061] K i = Linear K (K i )

[0062] V i = LinearV (V i )

[0063] where \(i = 1,\ldots,s\), and \(Linear(\cdot)\) is a linear mapping operator shared across different segments.

[0064] To explore the common information of T1MR and T2MR image features and consider the global feature relationship, a dot-generation attention layer is used to calculate the similarity matrix between \(Q\) i and \(K\) j (\(i\) and \(j\) belong to 1 to \(s\)), and then multiply by \(V\) to infer the common feature information between \(Q\) and \(V\). This process can be expressed as:

[0065]

[0066] where \(d\) k is a scaling factor, which can mitigate the softmax function from converging to the region of the minimum gradient when the dot product increases. Subsequently, the differential feature information between \(Q\) and \(V\) is obtained by removing the common feature information. This process can be expressed as:

[0067] DF QV = Linear(V - CF QV )

[0068] To obtain supplementary information from T1MR and T2MR images, the differential feature information is injected into \(Q\), which can be formulated as:

[0069] F add = DF QV + Q

[0070] F dfp = MLP(LN(F add )) + F add

[0071] where \(LN(\cdot)\) represents the normalization layer, \(MLP(\cdot)\) represents the multi-layer perceptron, and \(F\) dfp is the output of the DFP sub-module.

[0072] To obtain the common information between T1MR and T2MR image features to further enhance the fusion features, a CFP sub-module as shown in Figure 7 is constructed after the DFP sub-module to make the generated sCT image closer to the real CT. Assume that the segments of \(F\) dfp are used as \(Q\) 1,…,s , the segments of \(F\) T2 are used as \(K\) 1,…,s and \(V\) 1,…,s , and the common feature information between \(F\) dfp and \(F\) T2 can be expressed as:

[0073]

[0074] Then F dfp With F T2 Common feature information CF between T2 With F dfp Add. This process can be expressed as:

[0075] F add =Linear(CF T2 )·+Q

[0076]

[0077] in Indicates the output of the first CFP submodule.

[0078] Afterwards, and F T1 The common feature information between them is input into the second CFP submodule to enrich the fusion features. This process follows the same way as the first CFP submodule.

[0079] In a specific embodiment, the discriminator architecture in the cross-attention generative adversarial network is as follows: Figure 8As shown. In this embodiment, the discriminator is a multi-scale discriminator, which is used to distinguish the authenticity of the synthetic CT images output by the generator and provide feedback for optimizing the generator. The discriminator adopts a composite discrimination mechanism based on the pyramid structure. The discriminator constructs a multi-resolution image processing channel by gradually downsampling. Each scale corresponds to an independent PatchGAN module for feature discrimination. The discriminant results of each hierarchical PatchGAN module are independently output and weighted and fused to improve the detail fidelity of the sCT images and enhance the global consistency of the generated images. This hierarchical processing mechanism can effectively extract the feature representations of the image data in different spatial dimensions. The discriminator at the high-resolution level focuses on the analysis of pixel-level texture details, while the discriminator at the low-resolution level (downsampling scale) focuses on the semantic consistency verification of the overall image structure. Compared with the traditional single-scale discriminator, this hierarchical discrimination architecture forms complementary advantages in the feature perception range. More precisely, the local receptive field in the high-dimensional feature space can accurately capture microscopic texture features (such as texture details, edge sharpness, etc.), while the global receptive field in the low-dimensional feature space can effectively verify the rationality of the macroscopic layout. Through the weighted fusion of the multi-scale discriminant outputs, the model can provide more comprehensive gradient feedback for the generator during the adversarial training process, thereby guiding it to optimize both the local authenticity and global coordination of image generation simultaneously. The discriminator in this embodiment is composed of three PatchGAN modules (i.e., D1-D3) working in series. Each module adopts the same network structure but processes input data of different scales. This design enhances the feature decoupling ability of the discriminator through the scale difference in the feature space while ensuring the simplicity of the architecture. This composite discrimination mechanism can significantly improve the performance of the generated sCT images in the perceptual quality index, especially showing obvious advantages in the detail reconstruction of complex scenes and the cross-modal and cross-scale feature consistency.

[0080] In a specific embodiment, the training process of the cross-attention generative adversarial network is as follows:

[0081] (1) Obtain the original sample data (multi-sequence MR images and CT images). To ensure the consistency of the data and eliminate the deviation caused by differences in different imaging devices or imaging parameters, and to improve the image clarity and resolution, the acquired data is preprocessed by standardization.

[0082] 1) Perform standardization processing on the original sample data to obtain standardized data to ensure the consistency of the data and eliminate the deviation caused by differences in different imaging devices or imaging parameters.

[0083] 2) Use the N4 bias field correction method to remove the bias field caused by factors such as the magnetic field inhomogeneity of the magnetic resonance device and the sensitivity difference of the radio frequency coil in the multi-sequence MR images.

[0084] 3) Use image processing algorithms to eliminate noise and artifacts in the normalized data, obtaining intermediate data to improve the clarity and resolution of the images;

[0085] 4) Use a registration algorithm to align the intermediate data, obtaining preprocessed data to ensure the spatial consistency of the images at each time point and providing an accurate basis for subsequent analysis.

[0086] In adaptive radiotherapy, by spatially aligning images at different time points or of different modalities, the registration algorithm can help identify changes in tumors and surrounding normal tissues, providing an important basis for adjusting the treatment plan. The following is a detailed implementation process of a common image registration algorithm based on affine transformation and mutual information, applicable to the registration of multi-modal images such as CT and MRI, where:

[0087] Step 1: Data preparation. First, two sets of medical image data for registration need to be obtained: the reference image and the image to be registered. The reference image is the benchmark image for registration, usually the patient's initial image data (i.e., treatment history data, such as CT or MRI). The image to be registered is the real-time or stage image during the current treatment, usually obtained before or during each treatment, and may not be completely consistent with the reference image due to factors such as patient position and tumor growth.

[0088] Step 2: Preprocessing. Before registration, some preprocessing operations need to be performed on the images to improve the accuracy of registration:

[0089] Denoising: Remove noise in the images through filters (such as Gaussian filtering or median filtering, etc.).

[0090] Image normalization: Normalize the gray values of the images so that the intensity ranges of different modal images are consistent, reducing the gray deviation caused by imaging differences.

[0091] Image cropping: Only retain the key areas that need to be registered to reduce the computational amount.

[0092] Registration Model Selection. During the radiotherapy of actual patients, different registration models can be used to describe the geometric transformation between the reference image and the image to be registered. Commonly used registration models include Rigid Registration and Affine Registration. In Rigid Registration, it is assumed that the image to be registered only undergoes translation and rotation relative to the reference image, without considering scale changes or non-linear deformations. Affine Registration allows the image to be registered to undergo translation, rotation, scaling, and shearing relative to the reference image, and is applicable to cases with certain shape changes but not involving complex deformations. In order to handle relatively complex anatomical structure changes, the affine registration model is adopted in this embodiment.

[0093] Step 4: Similarity Metric Selection. The key to registration lies in defining a similarity metric to measure the degree of alignment between two images. For multimodal images (such as CT and MRI), the commonly used similarity metric is Mutual Information (MI). Mutual Information is a metric based on information entropy that can measure the statistical dependence of two images in terms of gray values. The better the alignment of two images, the greater their mutual information.

[0094] Step 5: Optimization Algorithm Selection. In order to find the optimal affine transformation parameters that maximize the mutual information, an optimization algorithm needs to be selected. Commonly used optimization algorithms include Gradient Descent and stochastic optimization algorithms (such as the Powell method). Among them, Gradient Descent calculates the gradient of the similarity metric and gradually adjusts the parameters in the affine transformation matrix to find the optimal solution. Stochastic optimization algorithms (such as the Powell method) search for different transformation parameters step by step and do not rely on gradient information, which is suitable for handling complex multimodal image registration problems. In this embodiment, the Powell optimization method is adopted to effectively avoid the local optimum problem and is suitable for multimodal image registration.

[0095] Step 6: Implement registration based on the selected optimization algorithm:

[0096] Step 1: Initialize the affine transformation matrix.

[0097] Set the initial affine transformation matrix as the identity matrix, indicating that the reference image and the image to be registered completely coincide in the initial state.

[0098] Step 2: Calculate the initial mutual information.

[0099] Use the current affine transformation matrix to project the image to be registered into the reference image space and calculate the initial mutual information value.

[0100] Step 3: Iterative optimization.

[0101] Using the Powell optimization algorithm, gradually adjust the parameters in the affine transformation matrix. After each adjustment, recalculate the mutual information value between the image to be registered and the reference image, and compare the change in mutual information before and after the adjustment. If the new mutual information value is greater than the mutual information value of the previous round, accept the new affine transformation matrix. Otherwise, continue to adjust the parameters until the mutual information value converges or reaches the maximum number of iterations.

[0102] Step 4: When the change amount of mutual information is lower than the set threshold (for example, less than 10-4), or when the optimization process reaches the maximum number of iterations, stop the optimization and output the final affine transformation matrix.

[0103] Step 5: Image resampling.

[0104] According to the final affine transformation matrix, resample the image to be registered so that it is spatially aligned with the reference image.

[0105] Step Seven: Post-processing and verification.

[0106] Step 1: Image overlay and verification.

[0107] Overlay and display the registered image to be registered and the reference image, and verify the registration effect through visualization means. If the two images are well aligned at key structures (such as tumor edges, organ contours, etc.), it indicates successful registration.

[0108] Step 2: Registration accuracy evaluation.

[0109] Use quantitative metrics such as target overlap (Dice coefficient), mean surface distance (MSD), etc. to evaluate the registration accuracy. Among them, the Dice coefficient is used to measure the similarity between two images, and the value closer to 1 indicates a higher overlap degree. MSD calculates the average distance of the key structure boundaries of the two registered images, and the smaller the distance, the better the registration effect.

[0110] Step Eight: Result feedback and use.

[0111] After registration is completed, the verified registered image will be transmitted to the downstream module for subsequent tumor target delineation, dose planning, and real-time treatment adjustment.

[0112] In summary, the registration algorithm can accurately align images of different modalities or different time points, providing a high-precision anatomical reference for the subsequent radiotherapy process. It is particularly suitable for the registration of multimodal images (such as CT and MR), which helps to track the dynamic changes of tumors and normal tissues in real time, thereby optimizing the treatment plan and reducing the radiation dose received by normal tissues during the treatment process.

[0113] The processed data is divided into a training set and a test set according to a ratio of 7:3, and they are respectively fed into the cross-attention generative adversarial network for training and testing. The generator and discriminator are optimized through the objective loss function to improve the quality of sCT images, making them highly matched with real CTs in terms of anatomical structure, edge sharpness, and tissue density.

[0114] In this application, the objective loss function of the CAGAN network consists of the loss function L G of the generator and the loss L D of the discriminator, which are two parts. Among them, the loss function L G includes four parts: adversarial loss L aG , pixel-level L1 loss L L1 , feature matching loss L F , and cross-attention regularization loss L d . The loss function L G is defined by the following formula:

[0115] L G = L aG + λ L1 ·L L1 + λ F ·L F + λ d ·L d

[0116] Among them, λ L1 , λ F , and λ d are important weight parameters for balancing relevant losses, and are respectively set to 50, 10, and 0.1 according to the experience in the experimental process. The following will introduce these loss terms in detail.

[0117] The adversarial loss L1 is used to promote the generator to generate realistic sCT images, and at the same time train the discriminator to distinguish real CTs from the generated sCTs. The generator hopes that the discriminator will judge the sCT images it generates as "true", then the adversarial loss L a of the generator is defined by the following formula:

[0118]

[0119] The discriminator needs to judge real CTs as "true" and generated CTs as "false", then the loss L D of the discriminator is defined by the following formula:

[0120] L D = L aD

[0121]

[0122] The generator forces the generated sCT to be pixel - level aligned with the real CT, preserving the consistency of the anatomical structure, and there is a pixel - level L1 loss L L1 It is defined as follows:

[0123]

[0124] The model further improves the quality of the generated sCT by constraining the feature distributions of the generated sCT and the real CT to be consistent in the intermediate layers of the discriminator. Assuming that the discriminator has K intermediate layers, the feature matching loss L FM can be defined as:

[0125]

[0126] where D k represents the feature map of the k - th layer of the discriminator.

[0127] The encoding parts of the generator that process different sequences need to be constrained for complementarity when outputting the reconstructed MR. Here, the cross - attention regularization loss is added as a loss term and is defined as follows:

[0128] L d =-λ div ·Corr(MR1(X),MR2(X))

[0129] where MR1 and MR2 are the reconstructed MRs output by the two encoders, and Corr calculates the correlation.

[0130] In this embodiment, distributed training is supported, and the model optimization is accelerated through a GPU cluster, shortening the training cycle to 30% of the original duration.

[0131] In a specific embodiment, step S3 specifically includes: inputting the multi - sequence MR images to be measured into the dual - branch feature extraction module for feature extraction and reconstruction to obtain the reconstructed multi - sequence MR images; inputting the reconstructed multi - sequence MR images into the cross - attention feature fusion module, extracting the differential features and common features of the reconstructed multi - sequence MR images, and performing weighted fusion on the differential features and common features to generate synthetic CT images.

[0132] In another exemplary embodiment, the sCT images generated from multi - sequence MR images in the pelvic region by the above - mentioned method are compared with the generation results of other advanced models. Specifically:

[0133] To verify the effectiveness and reliability of the CAGAN network, it was compared with five representative single - sequence generation models, namely Pix2pix, AttentionU - Net, DCNN, RegGAN, and CSAGAN.

[0134] The visualization results of the input dual-sequence T1 / T2 MR images and the sCT images generated from the single-sequence T1 MR images are presented in Figure 9 . Figure 9 the first row of []. From left to right in sequence are the input T1-sequence MR image, T2-sequence MR image, real CT image, sCT result generated by the Pix2pix model, sCT result generated by the Attention-UNet model, sCT result generated by the DCNN model, sCT result generated by the RegGAN model, sCT result generated by the CSAGAN model, and sCT result generated by the CAGAN model. The second row is the enlarged effect diagram of the ROI with significant differences. The third row is the feature fusion visualization difference diagram between the real CT and the generated sCT. The fourth row is the enlarged effect diagram of the ROI with significant differences in the feature fusion visualization difference diagram between the real CT and the generated sCT.

[0135] As Figure 9 can be seen, compared with other models, the CSAGAN model shows more excellent sCT generation effects. For the sCT images generated by the models with single-sequence MR images as the input source, there are varying degrees of bone structure deficiencies in areas such as the ischial tuberosity, femoral margin, and symphysis pubis. The reason for this may be that there are significant differences in the contrast of bones between the MR sequence images and the CT images, which makes the mapping relationship between the two complex, thereby increasing the difficulty of model generation. The CSAGAN model can combine and complement the features of the two MR sequence images, making the anatomical manifestations of the above-mentioned bone structures closer to the real CT images. This means that the CAGAN model can effectively complement the T1 MR sequence images with the feature information of the T2 MR sequence images.

[0136] Table 1 and Figure 10 (a)-(c) in [] present the quantitative comparative analysis of the results of different generation models. It can be clearly seen from this that the CAGAN model shows excellent performance in all evaluation indicators, and the values of all its evaluation indicators are higher than those of the single-sequence MR image generation models. When comparing the CAGAN model with the best-performing CSAGAN among the single-sequence generation models, the PSNR (peak signal-to-noise ratio) value of the CAGAN generation result has increased by 1.61 dB, the SSIM (structural similarity index) value has increased by 0.02, and the HU value error of the CT has decreased by 0.05. The MAE (mean absolute error) difference obtained by the CAGAN model is 42.54±4.57 HU. Compared with the worst-performing Pix2pix model, the MAE value of the CAGAN has decreased by approximately 17.3%. In terms of the SSIM evaluation indicator, the RegGAN is second only to the CAGAN model, and the SSIM value of the CAGAN is 0.01 higher than that of the RegGAN.

[0137] Through the above qualitative and quantitative analyses, it can be clearly found that compared with the sCT images generated from single-sequence MR images, the sCT images generated from dual-sequence MR images show obvious advantages in quality. This result strongly indicates that fusing the characteristic information of different-sequence images can effectively obtain high-quality sCT images. In the actual scenario of medical clinical applications, radiologists often need to comprehensively consider various modalities of medical images when diagnosing diseases and formulating treatment plans. By comparing the differences in anatomical structures, tissue densities, and lesion characteristics among different-modal images, doctors can grasp the condition of the disease more comprehensively and accurately. This way of multi-modal image comparison and analysis is highly consistent with the concept of fusing the characteristics of different-sequence MR images in this application to generate high-quality sCT images, further highlighting the potential application value of this application in clinical practice.

[0138] Table 1

[0139]

[0140]

[0141] In another exemplary embodiment, to verify the advancement and effectiveness of the CA module, ablation experiments were conducted on the proposed module for comparison. Specifically: to verify whether the CA module can effectively fuse the characteristics of the two images, the complete CAGAN was compared and analyzed with the network models without the CA module, without the DFP sub-module, and without the CFP sub-module respectively. It can be intuitively found from Figure 11 that for the sCT images generated by the model without the DFP sub-module, there is noise interference in areas such as the pubic symphysis, ischial tuberosity, and greater trochanter of the femur, resulting in discontinuous bone connections; for the model without the CFP sub-module, there are varying degrees of blurring and discontinuity at the inferior pubic ramus, pubic symphysis, and the soft tissue fascia between the obturator externus and pubic muscles; when the CA module is not used in the model, that is, when neither DFP nor CFP is adopted, the above situations are more prominent.

[0142] Table 2

[0143] Model PSNR (dB↑) SSIM (↑) MAE (HU↓) Without DFP sub-module 31.28±1.26 0.868±0.062 48.42±6.64 Without CFP sub-module 30.89±1.34 0.863±0.051 51.26±5.53 Without CA module 29.07±1.71 0.797±0.054 53.17±7.24 CAGAN 33.75±0.96 0.919±0.038 42.54±4.57

[0144] Compared with these three, the sCT generated by the complete model is closer to the characteristic distribution of the real CT. Further, from the perspective of quantitative analysis, the complete CAGAN model is significantly better than the models without the CA module, without the DFP sub-module, and without the CFP sub-module in all evaluation indicators. The specific data are shown in Table 2 and Figure 12As shown in (a)-(c) therein, compared with the network without the CA module, the sCT generated by CAGAN has a 10.63 HU reduction in the MAE value, a 4.68 dB increase in the PSNR value, and an approximately 15.3% increase in the SSIM value. For the model without the DFP sub-module and the CFP sub-module, there is little difference in the PSNR and SSIM values, but there is a 2.84 HU difference in the MAE value. The above ablation experiment results fully demonstrate that the CA module in this application can guide the model to focus on more key information during the multi-sequence feature fusion process, effectively suppress features irrelevant to the generation of the target modality image, and thus significantly enhance the synthesis ability of the model.

[0145] This application extracts features through a dual-branch feature extraction module and uses cross-attention mechanism, feature fusion, and multi-scale discrimination strategy to achieve high-precision sCT image synthesis. The CAGAN network proposed in this application adopts an improved Pix2Pix framework and combines the CA module to achieve feature fusion of multi-sequence MR images, generating high-quality synthetic CT (sCT) images, solving problems such as missing bone structures and blurred soft tissues when generating sCT from single-sequence MR, and significantly improving the anatomical structure accuracy and image quality of the generated images.

[0146] Based on the same inventive concept, the embodiment of this application also provides a synthetic CT image generation system based on multi-sequence MR images. The implementation solutions for solving problems provided by this system are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the synthetic CT image generation system based on multi-sequence MR images provided below can refer to the limitations on the synthetic CT image generation method based on multi-sequence MR images in the above text, and will not be elaborated here.

[0147] In an exemplary embodiment, a synthetic CT image generation system based on multi-sequence MR images is provided, including the following modules.

[0148] A data acquisition module for acquiring multi-sequence MR images to be measured.

[0149] A model construction and training module for constructing and training a cross-attention generative adversarial network; the cross-attention generative adversarial network includes a generator and a discriminator; the generator includes a dual-branch feature extraction module, a cross-attention feature fusion module, and a decoder; the discriminator includes multiple PatchGAN modules; the generator is used to generate synthetic CT images according to multi-sequence MR images; the discriminator is used to distinguish the authenticity of the synthetic CT images output by the generator.

[0150] A synthetic CT image generation module for inputting the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images.

[0151] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented. The computer device may be a server or a terminal, and its internal structure diagram may be as shown in Figure 13 the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data to be processed. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for generating a synthetic CT image based on multi-sequence MR images is implemented.

[0152] Those skilled in the art can understand that Figure 13 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0153] In an exemplary embodiment, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0154] In an exemplary embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0155] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0156] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memories (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0157] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0158] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these technical feature combinations do not conflict, they should all be considered as within the scope described in this specification.

[0159] In this text, specific examples are used to elaborate on the principles and implementation modes of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for generating synthetic CT images based on multi-sequence MR images, characterized in that Including: Obtain multi-sequence MR images to be measured; Construct and train a cross-attention generative adversarial network; the cross-attention generative adversarial network includes a generator and a discriminator; the generator includes a dual-branch feature extraction module, a cross-attention feature fusion module, and a decoder; the discriminator includes multiple PatchGAN modules; the generator is used to generate synthetic CT images according to the multi-sequence MR images; the discriminator is used to distinguish the authenticity of the synthetic CT images output by the generator; Input the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images.

2. The method for generating a synthetic CT image based on multi-sequence MR images according to claim 1, wherein Inputting the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images specifically includes: Input the multi-sequence MR images to be measured into the dual-branch feature extraction module for feature extraction and reconstruction to obtain the reconstructed multi-sequence MR images; Input the reconstructed multi-sequence MR images into the cross-attention feature fusion module, extract the differential features and common features of the reconstructed multi-sequence MR images, and perform weighted fusion on the differential features and common features to obtain fused features; Input the fused features into the decoder to generate synthetic CT images.

3. The method for generating a synthetic CT image based on multi-sequence MR images according to claim 1, wherein, The dual-branch feature extraction module adopts a symmetric encoding-decoding structure. Each encoder in each branch of the dual-branch feature extraction module contains 5 cascaded convolutional modules. Each convolutional module includes two groups of convolutional layers with a kernel size of 3×3, a batch normalization layer, and a ReLU activation function; each decoder in each branch of the dual-branch feature extraction module includes multiple transposed convolutional modules.

4. The method for generating a synthetic CT image based on multi-sequence MR images according to claim 1, wherein The cross-attention feature fusion module includes a differential feature processing sub-module and a common feature processing sub-module; The differential feature processing sub-module is used to extract the differential features of the multi-sequence MR images and perform weighted fusion through a channel attention mechanism; The common feature processing sub-module is used to extract the common features of the multi-sequence MR images and perform information encoding using a multi-layer perceptron.

5. The method for generating a synthetic CT image based on multi-sequence MR images according to claim 1, wherein The loss function of the generator in the cross-attention generative adversarial network during training includes adversarial loss, pixel-level L1 loss, feature matching loss, and cross-attention regularization loss.

6. The method for generating a synthetic CT image based on multi-sequence MR images according to claim 1, wherein The discriminator in the cross-attention generative adversarial network adopts a composite discrimination mechanism based on a pyramid structure.

7. A synthetic CT image generation system based on multi-sequence MR images, characterized in that, Including: A data acquisition module for obtaining multi-sequence MR images to be measured; A model construction and training module for constructing and training a cross-attention generative adversarial network; the cross-attention generative adversarial network includes a generator and a discriminator; the generator includes a dual-branch feature extraction module, a cross-attention feature fusion module, and a decoder; the discriminator includes multiple PatchGAN modules; the generator is used to generate synthetic CT images according to the multi-sequence MR images; the discriminator is used to distinguish the authenticity of the synthetic CT images output by the generator; A synthetic CT image generation module for inputting the multi-sequence MR images to be measured into the generator of the trained cross-attention generative adversarial network to generate synthetic CT images.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that the processor executes the computer program to implement the method for generating a synthetic CT image based on multi-sequence MR images according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating a synthetic CT image based on multi-sequence MR images according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for generating a synthetic CT image based on multi-sequence MR images according to any one of claims 1-6.