Method and system for generating CT image from mr image under multiple constraint conditions
Patent Information
- Application Number
- PCT/CN2026/086540
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure CN2026086540_01102026_PF_FP_ABST
Abstract
Description
Methods and systems for generating CT images from MR images under multiple constraints
[0001] Cross-reference to related applications
[0002] This invention claims priority to Chinese Patent Application No. 202510376490.3, filed with the China National Intellectual Property Administration on March 28, 2025, entitled "Method and System for Generating CT Images from MR Images under Multiple Constraints", the entire contents of which are incorporated herein by reference and constitute a part of this invention for all purposes. Technical Field
[0003] This invention belongs to the field of image processing technology, and in particular relates to a method and system for generating CT images from MR images under multiple constraints. Background Technology
[0004] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0005] CT (Computed Tomography) is a medical imaging technique. It uses X-ray beams to perform tomographic scanning of the human body and then uses a computer to process the data to produce detailed images of the body's internal structures. CT scans are widely used for the diagnosis of various diseases due to their fast and clear imaging capabilities.
[0006] MR (Magnetic Resonance) is a relatively new medical imaging technique, also known as Magnetic Resonance Imaging (MRI). It uses static magnetic fields and radio frequency magnetic fields to image human tissues. During the imaging process, it can obtain clear images with high contrast without the use of electron ionizing radiation or contrast agents.
[0007] Generating CT images from MR images can address the need for alternatives in certain clinical applications where CT is unavailable, contraindicated, or carries a high radiation risk. For example, dose calculation in radiotherapy planning requires precise electron density maps (usually provided by CT) for dose calculation. However, MR does not emit ionizing radiation and does not directly provide electron density information. Therefore, predicting "synthetic CT" from MR images for use in radiotherapy planning systems can effectively avoid duplicate CT scans; reduce radiation exposure for patients requiring multiple image follow-ups (such as children and cancer patients), avoiding duplicate CT scans can significantly reduce the cumulative dose of ionizing radiation; and fully utilize the characteristics and advantages of both MR and CT images, enabling automatic image registration and assisting physicians in simultaneously referencing information from both modalities on a single platform, improving diagnostic and treatment accuracy.
[0008] Existing technology provides a method for synthesizing CT images from MR images based on residual Transformer generative adversarial networks. By setting a deep feature extraction module in the generator and a feature enhancement module in the deep feature extraction module, a new generator is obtained. The deep feature extraction module in this generator can extract deeper features and can make up for the shortcomings of traditional generative adversarial networks in capturing global information. However, the quality of the synthesized CT images from MR images in this existing technology depends on the quality and quantity of training samples. This may lead to low training accuracy when there are insufficient samples, thus affecting the quality of the synthesized CT images. Summary of the Invention
[0009] To address the aforementioned technical problems, this invention provides a method and system for generating CT images from MR images under multiple constraints. This method considers referencing three-dimensional (3D) CT images of similar cases and does not rely entirely on the quality and quantity of training samples, thereby improving the quality of the synthesized CT images.
[0010] To achieve the above objectives, the present invention adopts the following technical solution:
[0011] One or more examples of the present invention provide a method for generating CT images from MR images under multiple constraints, comprising:
[0012] Acquire the current patient's MR images;
[0013] Obtain 3D CT images of similar cases based on the current patient's symptoms;
[0014] Using a feature encoder, the MR image of the current patient and the 3D CT image of a similar reference case are sampled and encoded to obtain two-dimensional (2D) MR image features and 3D CT image features, respectively.
[0015] By using a cascaded cross-domain transformer module, the 2D MR image features are extended one dimension, combined with batch dimension transformation, and then stitched with 3D CT features along the depth dimension to obtain the first stitched feature. The first stitched feature is then merged along the depth dimension and the batch dimension, and then merged along the width and height dimensions to obtain the merged feature. A set of trainable parameters is then stitched with the merged feature along the channel dimension to obtain the second stitched feature. The second stitched feature is then positionally encoded along the width and height dimensions, and then positionally encoded along the depth dimension. The depth-encoded feature is then passed through a depth-dimensional transformer module to obtain the first fused feature. The first fused feature is then passed through a height-width transformer module to obtain the second fused feature. A multi-head attention mechanism is used to filter the first and second fused features to obtain the filtered fused feature.
[0016] The filtered fusion features are decoded using a feature decoder, and the CT image corresponding to the current patient's MR image is obtained through decoding.
[0017] Based on the obtained CT images, the lesion site of the current patient is located, and the characteristics of the lesion tissue (such as tumor) in the lesion site are obtained, and the size and degree of the lesion tissue are identified.
[0018] The generator is configured to include:
[0019] Feature encoder, cascaded cross-domain transformer module, and feature decoder;
[0020] The feature encoder is configured to: in two branches, downsample and encode the MR image of the current patient and the 3D CT image of a reference similar case respectively to obtain 2D MR image features and 3D CT image features respectively;
[0021] The cascaded cross-domain transformer module is configured to: expand 2D MRI features to one dimension, combine them with batch dimension transformation, and then stitch them with 3D CT features along the depth dimension; merge the stitched features sequentially along the depth dimension and batch dimension, and then merge them along the width and height dimensions; then stitch a set of trainable parameters with the merged features along the channel dimension to obtain stitched features; sequentially perform position encoding on the width and height dimensions, and position encoding on the depth dimension; pass the depth dimension encoded features through the depth dimension Transformer module to obtain a first fused feature; then pass the first fused feature through the height-width Transformer module to obtain a second fused feature; and use a multi-head attention mechanism to filter the first and second fused features to obtain the filtered fused features.
[0022] The feature decoder is configured to decode the filtered fusion features and obtain the CT image corresponding to the current patient's MR image through decoding.
[0023] In one implementation, each of the two branches of the feature encoder contains four sampling modules; wherein, the last two sampling modules of the four sampling modules contain a degradation model, the degradation model including a blurring process and a noise addition process, the blurring process and the noise addition process being independent of each other.
[0024] As one implementation method, the expression for the degradation model is:
[0025] Among them, f L and f L+1 Represents the features of layer L and layer L+1, ↓ s This represents the downsampling process, where k is the fuzzy kernel, n is the added noise, and f is the input noise. tmp This represents the features obtained through a fuzzy process.
[0026] As one implementation, in the cascaded cross-domain transformer module, the formula for filtering the first and second fusion features using a multi-head attention mechanism is: f fusion =MLP(f out_D +f out_HW );
[0027] Among them, f fusion The filtered fusion features; f out_D As the first fusion feature, f out_HW For the second fusion feature; f out_D +f out_HW The system consists of height-width-depth global attention features; MLP is a multilayer perceptron function.
[0028] In one implementation, the feature decoder includes four layers, wherein the first two layers are upsampling layers with gating fusion mechanisms, and the last two layers are upsampling layers with self-attention mechanisms.
[0029] As one implementation, the expression for the upsampling layer of the gated fusion mechanism is: f Up_l+1 =σ(W z ·[f Up_l ,f Down_l ])*tanh(W Up ·f Up_l )+(1-σ(W z ·[f Up_l ,f Down_l ]))*tanh(W Down ·fDown_l );
[0030] Among them, W Up W Down W z It is the training parameter matrix, f Up_l f represents the input of layer l during the upsampling process. Down_l f represents the output of the l-layer degenerate model during downsampling. Up_l+1 σ represents the output of layer l during the upsampling process; σ is the sigmoid function, which produces values between 0 and 1; tanh is the tangent function.
[0031] In one or more examples of this disclosure, a system for generating CT images from MR images under multiple constraints is provided, the system comprising:
[0032] Image acquisition unit, which is used to acquire the current patient's MR image and reference similar case 3DCT images;
[0033] The CT image generation unit is used to generate the CT image corresponding to the current patient's MR image using a trained generator, under the constraints of the current patient's MR image and reference similar case 3D CT images.
[0034] In the CT image generation unit, the generator includes a feature encoder, a cascaded cross-domain transformer module, and a feature decoder; wherein,
[0035] The feature encoder comprises two branches, which sample and encode the current patient's MR image and the reference similar case's 3DCT image, respectively, to obtain 2D MR image features and 3D CT image features;
[0036] The cascaded cross-domain transformer module is used to expand 2D MR image features by one dimension, combine them with batch dimension transformation, and then stitch them with 3D CT image features along the depth dimension. The stitched features are then merged sequentially along the depth dimension and batch dimension, as well as the width and height dimensions. A set of trainable parameters is then stitched with the merged features along the channel dimension to obtain the stitched features. The stitched features are then sequentially encoded in the width and height dimensions, as well as in the depth dimension. The depth-encoded features are then processed by the depth dimension transformer module to obtain the first fused feature. The first fused feature is then processed by the height-width transformer module to obtain the second fused feature. A multi-head attention mechanism is used to filter the first and second fused features to obtain the filtered fused features.
[0037] The feature decoder is used to decode the filtered fusion features to obtain the CT image corresponding to the current patient's MR image.
[0038] In one implementation, in the CT image generation unit, each branch of the feature encoder contains four sampling modules, wherein the last two sampling modules contain a degradation model, which includes two independent processes: a blurring process and a noise addition process.
[0039] In one or more examples of this disclosure, a computer-readable storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the steps in the method for generating CT images from MR images under multiple constraints as described above.
[0040] In one or more examples of this disclosure, an electronic device is provided, the electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method for generating CT images from MR images under multiple constraints as described above.
[0041] Compared with the prior art, the beneficial effects of the present invention are:
[0042] (1) The cascaded cross-domain transformer module designed in this invention uses the attention mechanism of 2D features and 3D features in the depth dimension and the attention mechanism of 2D features and 3D features in the height and width dimensions to achieve feature screening and heterogeneous features based on the attention mechanism. It takes into account the reference similar case 3DCT images, does not completely depend on the quality and quantity of training samples, and improves the quality of synthesized CT images.
[0043] (2) The generator of the present invention adopts an encoder-decoder structure, introduces a degradation model in the decoder stage, and introduces a gating fusion mechanism and a self-attention mechanism in the decoder stage, which improves the generator effect and ultimately improves the quality of the synthesized CT image. Attached Figure Description
[0044] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same reference numerals denote the same structures, wherein: FIG1 is a schematic flowchart illustrating a method for generating CT images from MR images under multiple constraints according to an example of the present invention;
[0045] Figure 2 is a schematic diagram of a generator structure shown in an example of the present invention;
[0046] Figure 3 is a cascaded cross-domain transformer structure diagram illustrating an example of the present invention;
[0047] Figure 4 is a gated fusion structure diagram illustrating an example of the present invention;
[0048] Figure 5 is a comparison diagram of the input MR image, the true value of the CT image corresponding to the input MR image, and the synthesized CT image shown in a verification example of the present invention.
[0049] Figure 6 is a schematic diagram of the system structure for generating CT images from MR images under multiple constraints, as shown in an example of the present invention.
[0050] Figure 7 is a schematic diagram of an electronic device according to an example of the present invention. Detailed Implementation
[0051] The present invention will be further described below with reference to the accompanying drawings and examples.
[0052] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0053] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0054] This disclosure provides a method 100 for generating CT images from MR images under multiple constraints, exemplified by one or more embodiments of this disclosure. Figure 1 is a flowchart illustrating a method for generating CT images from MR images under multiple constraints, exemplified by one or more embodiments of this disclosure. As shown in Figure 1, the method 100 may include:
[0055] S101, acquire the MR image of the current patient, and acquire 3D CT images of similar cases based on the symptoms of the current patient;
[0056] S102, using a feature encoder, sample and encode the current patient's MR image and the reference similar case's 3D CT image respectively to obtain two-dimensional (2D) MR image features and 3D CT image features respectively;
[0057] S103: Through a cascaded cross-domain transformer module, the 2D MR image features are extended by one dimension, combined with batch dimension transformation, and then stitched with 3D CT features along the depth dimension to obtain the first stitched feature. The first stitched feature is merged along the depth dimension and the batch dimension, and then merged along the width and height dimensions to obtain the merged feature. A set of trainable parameters is then stitched with the merged feature along the channel dimension to obtain the second stitched feature. The second stitched feature is positionally encoded in the width and height dimensions, and then positionally encoded in the depth dimension. The depth dimension encoded feature is passed through a depth dimension transformer module (D-Transformer) to obtain the first fused feature. The first fused feature is then passed through a height-width transformer module (HW-Transformer) to obtain the second fused feature. A multi-head attention mechanism is used to filter the first and second fused features to obtain the filtered fused feature.
[0058] S104, using a feature decoder to decode the filtered fused features to obtain the CT image corresponding to the current patient's MR image; and,
[0059] S105, based on the obtained CT images, locates the lesion site of the current patient, obtains the characteristics of the lesion tissue (such as soft tissue, bone, tumor, etc.) in the lesion site, and identifies the size and degree of the lesion tissue.
[0060] In some examples, to maximize the clarity and completeness of the final synthesized CT image, multiple constraints can be established based on image information from the current patient's MR image and 3D CT images of similar cases, combined with the desired diagnostic result. These constraints, along with the image information, are then fed into the generator for data processing and CT image generation. Those skilled in the art can use any known, conventional means to set these constraints. This feature is not within the scope of this disclosure and will not be elaborated further.
[0061] In some examples, the feature encoder used to perform step S102 comprises two branches, each of which contains four sampling modules to downsample the input image information. The last two sampling modules further contain a degradation model, which includes a blurring process and a noise addition process, which are independent of each other.
[0062] It should be noted that in other examples, the number of sampling modules and the number of degradation models contained in each branch of the feature encoder can be set according to the actual situation.
[0063] Specifically, in these examples, the first branch of the feature encoder takes 3D CT images of similar cases as input and passes through four sampling modules to downsample the input 3D CT image information and output 3D CT image features.
[0064] In some examples, the structure of the first branch is shown in Table 1 below, which contains four sampling modules. Table 1 shows the kernel size, number of kernels, and stride size for each sampling module. The 3D CT image features obtained after downsampling processing by the four sampling modules are represented as: f CT ∈H×W×D×C.
[0065] Table 1. First Branch Structure and 3D CT Image Feature Downsampling Table
[0066] In other examples, the second branch of the feature encoder takes the current patient's MR image as input and also uses four sampling modules to downsample the input MR image information. The structures of the first two sampling modules are shown in Table 2. Table 2 shows the kernel size, number of kernels, and stride size of the first two sampling modules.
[0067] Table 2. Second Branch Structure and 2D MR Image Feature Downsampling Table
[0068] In these examples, the latter two sampling modules of the second branch employ a degradation model, which comprises two processes: a blurring process and a noise addition process. The blurring process and the noise addition process are described by formula (1):
[0069] Among them, f L and f L+1 Represents the features of layer L and layer L+1, ↓ s This represents the downsampling process, where k is the fuzzy kernel, n is the added noise, and f is the input noise. tmp This represents the features obtained through a fuzzy process.
[0070] In some examples, for the fuzzy kernel k, a first prior random variable z following a multidimensional normal distribution is first defined. k Then, the fuzzy kernel k is learned through a set of convolutional networks:
[0071] in, k=3 indicates the magnitude of the convolution weights. This suggests that the learned blur kernel k is spatially dependent, and h=w=1 indicates that k can be approximated by a space-invariant kernel. kThis is the dimension of the normal distribution. The third sampling module has 256 kernels, and the fourth sampling module has 512 kernels. The last layer adds a Softmax function to ensure that the sum of each row of the blur kernel k is 1.
[0072] As mentioned above, n is the added noise, and the fuzzing process and the noise process are independent of each other.
[0073] Unlike previous examples, in one or more examples disclosed herein, the added noise n is feature-dependent.
[0074] In these examples, for the added noise n, a second prior random variable z that follows a multidimensional normal distribution is first defined. n Then the feature f output by the fuzzing process tmp and the second prior random variable z n The input is fed into a set of convolutional networks to learn the input noise n:
[0075] in, The third sampling module has h = w = 64, c = 256, and f n The dimension of the normal distribution is set to 64; the fourth sampling module has h = w = 32, c = 512, and f n This is the dimension of the normal distribution, set to 64.
[0076] In summary, the degradation model formula (3) described above can be further modified as follows:
[0077] Among them, f L and f L+1 Represents the features of layer L and layer L+1, ↓ s This represents the downsampling process, where k is the fuzzy kernel; z k It is the first prior random variable that follows a multidimensional normal distribution; z n It is the second prior random variable that follows a multidimensional normal distribution.
[0078] After two layers of degradation modeling, the output 2D MR image features are: f MR ∈H×W×C.
[0079] In one or more examples of this disclosure, the cascaded cross-domain transformer model architecture used to perform step S103 is shown in Figure 3. The cascaded cross-domain transformer model architecture contains two cascaded transformer models: a depth-dimensional transformer model (D-Transformer) and a height-width-dimensional transformer model (HW-Transformer). These two cascaded transformer models, according to their hierarchical cascading relationship, sequentially perform heterogeneous feature fusion and feature filtering on 3D CT image features and 2D MR image features.
[0080] In these examples, specifically, the 2D MR image features f are first analyzed. MR After extending ∈H×W×C by one dimension, the batch dimension is introduced, resulting in f′. MR ∈B×H×W×1×C; f′ MR Features of 3D CT images f CT By stitching along the depth dimension, we obtain f. D+1 ∈B×H×W×(D+1)×C, which is the first concatenation feature; f D+1 The depth dimension and batch dimension are merged, and then the width dimension and width dimension are merged to obtain f′. D+1 ∈(BD+B)×C×(HW), i.e., merged features; set a set of trainable parameters as cls_token, and merged features f′ D+1 By concatenating along the channel dimension, we obtain f. C+1 ∈(BD+B)×(C+1)×(HW), which is the second splicing feature.
[0081] Specifically, setting a set of trainable parameters as the cls_token has the following three effects:
[0082] Aspect 1: The set of trainable parameters is randomly initialized and continuously updated as the network is trained, which can encode the statistical characteristics of the entire dataset;
[0083] Aspect 2: It is possible to deal with f MR and f CT The feature information is aggregated, and this global feature aggregation effect is achieved through cls_token and the transformer model architecture. Furthermore, since it is not based on image content, it avoids relying on f... MR and f CT It produces bias;
[0084] Aspect 3: cls_token is placed in f MR and f CTUsing a fixed positional encoding at the very beginning of the depth dimension can prevent the output from being affected by positional encoding.
[0085] Furthermore, firstly regarding f C+1 Position encoding is performed in the width and height dimensions. This position encoding is still a set of trainable parameters PE_HW, which is related to f. C+1 Perform tensor addition as shown in formula (5): f HWP =f C+1 +PE_HW; (5)
[0086] After that, for f HWP Perform positional encoding in the depth dimension. This positional encoding is still a set of trainable parameters PE_D, and this set of trainable parameters PE_D is related to f. HWP Perform tensor addition as shown in formula (6): f HWDP =f HWP +PE_D. (6)
[0087] The encoded feature f HWDP The encoded f is fed into the D-Transformer. HWDP This corresponds to Q (Query), K (Key), and V (Value) in the multi-head attention mechanism of D-Transformer; the multi-head attention mechanism is represented as follows:
[0088] Multi-head attention mechanism Attention(Q,K,V) (i.e., f) HWDP After normalization, and f HWDP Perform the addition to obtain f′ out Then let f′ out After passing through the feedforward neural network, normalization is performed again, and then compared with f′. out Perform addition to obtain f out_D Its formula is expressed as follows: f′ out =LN(f HWDP )+f HWDP (8)
[0089] f out_D =LN(FFN(f′) out ))+f′ out (9)
[0090] After that, f out_D First, it passes through a fully connected layer (FC Layer) to obtain f.out_D_C Then it is fed into the HW-Transformer, which has the same structure as the D-Transformer, and performs the following functions:
[0091] For the input f out_D_C After normalization, and f out_D_C Performing addition yields f″ out Then let f″ out After passing through the feedforward neural network, normalization and union are performed again, and then combined with f″. out Perform addition to obtain f out_HW Its formula is expressed as follows: f″ out =LN(f out_D_C )+f out_D_C ; f out_HW =LN(FFN(f″) out ))+f″ out .
[0092] The obtained f out_D and f out_HW The data is fed into a multilayer perceptron (MLP) for feature selection, as shown below: f fusion =MLP(f out_D +f out_HW (10)
[0093] Among them, f fusion For the filtered fusion features, f fusion ∈B×H×W×C;f out_D As the first fusion feature, f out_HW For the second fusion feature; f out_D +f out_HW The height-width-depth global attention feature is used; MLP is a multilayer perceptron function.
[0094] Features with high attention scores are filtered out using MLP to remove redundant features.
[0095] In one or more examples of this disclosure, the feature decoder for performing step S104 includes four layers, wherein the first two layers are upsampling layers containing a gated fusion mechanism, and the last two layers are upsampling layers containing a self-attention mechanism. The filtered fused feature f fusion The image is fed into the feature decoder for decoding, ultimately generating the CT slice image of the current patient.
[0096] As shown in Figure 4, in one example, the expression for the upsampling layer employing the gating fusion mechanism is: f Up_l+1 =σ(Wz ·[f Up_l ,f Down_l ])*tanh(W Up ·f Up_l )+(1-σ(W z ·[f Up_l ,f Down_l ]))*tanh(W Down ·f Down_l );
[0097] Among them, W Up W Down W z It is the training parameter matrix, f Up_l Representative of f fusion During the upsampling process, the input of the l-th layer, f Down_l Representative of f fusion The output of the degenerate model of the l-th layer during downsampling, f Up_l+1 Representative of f fusion During the upsampling process, the output of the l-th layer is σ, which is the sigmoid function that produces values between 0 and 1; tanh is the tangent function.
[0098] Those skilled in the art should know that the gating fusion mechanism is a method that controls the fusion process of different data sources or features by introducing gating units. Its core lies in dynamically adjusting the contribution of different modal or hierarchical features, thereby more accurately capturing complex relationships and contextual dependencies.
[0099] In these examples, the last two upsampling layers of the feature decoder both contain a self-attention mechanism, which is applied to the output f of the l-th layer features after upsampling. Up_l+1 This can be decoded using an attention mechanism similar to the width and height shown in the example in Figure 3. Unlike the D-Transformer or HW-Transformer mentioned above, this is for f... Up_l+1 Without positional encoding, the Q, K, and V features in the self-attention mechanism of the last two upsampling layers directly use the features f output by the first two upsampling layers. Up_l+1 The processing mechanism described above—multi-head attention → normalization and addition → feedforward neural network → normalization and addition—is used to process the output f of the l-th layer. Up_l+1 Decode to obtain
[0100] Those skilled in the art should know that the self-attention mechanism is an attention mechanism that associates different positions of a single sequence to compute the representation of the same sequence, and its core components include three weight matrices: query, key, and value.
[0101] In the example described above, the feature encoder, the cascaded cross-domain transformer model architecture, and the feature decoder together constitute the generator.
[0102] The generator needs to be pre-trained before use. For the pre-training of the generator, conventional training methods for transformer model architectures in this field can be used, and will not be elaborated upon in this disclosure.
[0103] In some examples of generators, the total loss function of the generator is... It includes perceptual loss, contextual loss, and structural loss, and its formula is as follows:
[0104] in, It is perceptual loss, which is achieved by estimating... And the real I CT The data is fed into a pre-trained 19-layer convolutional neural network (i.e., Visual Geometry Group Network, VGG-19) for feature extraction. Then, the high-level features of the two are compared, as shown in the following formula:
[0105] Where Φ represents the operation of extracting high-level features.
[0106] This is the context loss, used to ensure the consistency of local texture features. Its formula is as follows:
[0107] Among them, w l It is a balance coefficient used to balance the continuity effect of multi-scale hierarchical features; CX indicates the selection of multiple hierarchical features.
[0108] The structural loss is calculated using the Learned Perceptual Image Patch Similarity (LPIPS) method, and the results are as follows:
[0109] As described above, in several examples disclosed herein, a cascaded cross-domain transformer module is designed, which has two cascaded transformer models, wherein the core of the first transformer model is an attention mechanism between 2D features and 3D features in the depth dimension, and the core of the second transformer model is an attention mechanism between 2D features and 3D features in the height and width dimensions.
[0110] In summary, the cascaded cross-domain transformer module in the multiple examples disclosed in this publication is the first proposed attention mechanism module for feature selection and heterogeneous feature fusion of 3D and 2D features.
[0111] In some examples of this disclosure, the generators employ an encoder-decoder structure. To improve the generator's performance, a degradation model is introduced in the encoder stage, and a gating fusion mechanism and a self-attention mechanism are introduced in the decoder stage.
[0112] The accuracy of the generator was verified in several validation examples of this disclosure, as shown in Figure 5. Figure 5 illustrates a comparison between the input MR image, the ground truth CT image corresponding to the MR image, and the synthesized CT image. The first column represents the input MR image, the second column represents the ground truth CT image corresponding to the MR image, and the third column represents the synthesized CT image based on the MR image. As can be seen from Figure 5, the synthesized CT images obtained using the generator designed in several examples of this disclosure achieve 85-90% consistency with the ground truth, which can be used as a clinical reference.
[0113] In some other examples of this disclosure, a system 600 for generating CT images from MR images under multiple constraints is also provided, as shown in FIG6. FIG6 illustrates a schematic diagram of a system structure for generating CT images from MR images under multiple constraints, which corresponds to a method for generating CT images from MR images under multiple constraints as shown in FIG1. As shown in FIG6, the system 600 may include:
[0114] Image acquisition unit 601 is used to acquire the current patient's MR image and reference similar case 3D CT images;
[0115] The CT image generation unit 602 is used to obtain the CT image corresponding to the current patient's MR image using a trained generator 603 under the constraints of the current patient's MR image and reference similar case 3D CT images.
[0116] In the CT image generation unit, the generator 603 includes a feature encoder 604, a cascaded cross-domain transformer module 605, and a feature decoder 606. The feature encoder 604 has two branches, which sample and encode the current patient's MR image and the reference similar case's 3D CT image, respectively, to obtain 2D MR image features and 3D CT features.
[0117] The cascaded cross-domain transformer module 605 is used to expand 2D MRI features to one dimension, combine them with batch dimension transformation, and then stitch them together with 3D CT features along the depth dimension. The stitched features are then merged sequentially along the depth dimension and batch dimension, as well as along the width and height dimensions. A set of trainable parameters is then stitched together with the merged features along the channel dimension to obtain stitched features. The stitched features are then sequentially encoded in the width and height dimensions, as well as in the depth dimension. The depth dimension encoded features are then processed by the depth dimension transformer module to obtain the first fused feature. The first fused feature is then processed by the height-width transformer module to obtain the second fused feature. A multi-head attention mechanism is used to filter the first and second fused features to obtain the filtered fused features.
[0118] The feature decoder 606 is used to decode the filtered fusion features to obtain the corresponding CT image.
[0119] In the CT image generation unit 602, each branch of the feature encoder 604 contains four sampling modules, and the last two sampling modules contain a degradation model, which includes two independent processes: a blurring process and a noise addition process.
[0120] It should be noted that each module in the system 600 for generating CT images from MR images under multiple constraints in Figure 6 corresponds one-to-one with each step in the method 100 for generating CT images from MR images under multiple constraints in Figure 1. Their specific implementation processes are the same and will not be repeated here.
[0121] Referring to Figure 7, a schematic diagram of an electronic device is given. It should be noted that the electronic device 700 shown in Figure 7 is merely an example and should not impose any limitations on the functionality and scope of use of this invention.
[0122] As shown in Figure 7, the electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 702 or programs loaded from storage section 708 into random access memory (RAM) 703. The RAM 703 also stores various programs and data required for system operation. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0123] The following components are connected to I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a local area network (LAN) card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 710 as needed so that computer programs read from it can be installed into storage section 708 as needed.
[0124] When the central processing unit 701 in the electronic device of this example executes the program, it implements the steps in the method for generating CT images from MR images under multiple constraints as shown in FIG1.
[0125] In particular, according to the examples of this disclosure, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, examples of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in FIG1. In such an example, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by central processing unit 701, it performs the various functions defined in the apparatus of this application.
[0126] The computer program instructions corresponding to the method shown in Figure 1 may also be stored in a computer-readable storage medium that can guide a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0127] Those skilled in the art will understand that all or part of the processes in the above-described example methods can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the examples described above. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0128] The above description is merely a preferred example of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating CT images from MR images under multiple constraints, characterized in that, include: Acquire the current patient's MR image and 3D CT images of similar reference cases; Based on the constraints of the current patient's MR image and reference similar case 3D CT images, the trained generator is used to obtain the CT image corresponding to the current patient's MR image. The generator includes a feature encoder, a cascaded cross-domain transformer module, and a feature decoder. The feature encoder has two branches, which sample and encode the current patient's MR image and the reference similar case's 3D CT image, respectively, to obtain 2D MRI features and 3D CT features. The cascaded cross-domain transformer module is used to expand 2D MRI features to one dimension, combine them with batch dimension transformation, and then stitch them with 3D CT features along the depth dimension. The stitched features are then merged sequentially along the depth dimension and batch dimension, as well as the width and height dimensions. A set of trainable parameters is then stitched with the merged features along the channel dimension to obtain the stitched features. The stitched features are then sequentially encoded in the width and height dimensions, as well as in the depth dimension. The depth-encoded features are then processed by the depth dimension Transformer module to obtain the first fused feature. The first fused feature is then processed by the height-width Transformer module to obtain the second fused feature. A multi-head attention mechanism is used to filter the first and second fused features to obtain the filtered fused features. The feature decoder is used to decode the filtered fused features to obtain the corresponding CT image.
2. The method for generating CT images from MR images under multiple constraints as described in claim 1, characterized in that, Each branch of the feature encoder contains four sampling modules, and the last two sampling modules contain a degradation model, which includes two independent processes: a blurring process and a noise addition process.
3. The method for generating CT images from MR images under multiple constraints as described in claim 2, characterized in that, The expression for the degradation model is: Among them, f L and f L+1 Represents the features of layer L and layer L+1, ↓ s This represents the downsampling process, where k is the fuzzy kernel, n is the added noise, and f is the input noise. tmp This represents the features obtained through a fuzzy process.
4. The method for generating CT images from MR images under multiple constraints as described in claim 1, characterized in that, In the cascaded cross-domain transformer module, the formula for filtering the initial fused features using a multi-head attention mechanism is: f fusion =MLP(f out_D +f out_HW ) Among them, f fusion The filtered fusion features; f out_D As the first fusion feature, f out_HW For the second fusion feature; f out_D +f out_HW The height-width-depth global attention feature is used; MLP is a multilayer perceptron function.
5. The method for generating CT images from MR images under multiple constraints as described in claim 1, characterized in that, The feature decoder includes a first two-layer gated fusion upsampling layer and a last two-layer self-attention upsampling layer.
6. The method for generating CT images from MR images under multiple constraints as described in claim 5, characterized in that, The expression for the gated fusion upsampling layer is: f Up_l+1 =σ(W z ·[f Up_l ,f Down_l ])*tanh(W Up ·f Up_l )+(1-σ(W z ·[f Up_l ,f Down_l ]))*tanh(W Down ·f Down_l ); Among them, W Up W Down W z It is the training parameter matrix, f Up_l f represents the input of layer l during the upsampling process. Down_l f represents the output of the l-layer degenerate model during downsampling. Up_l+1 σ represents the output of layer l during the upsampling process; σ is the sigmoid function, which produces values between 0 and 1; tanh is the tangent function.
7. A system for generating CT images from MR images under multiple constraints, characterized in that, include: Image acquisition unit, which is used to acquire the current patient's MR image and reference similar case 3D CT images; The CT image generation unit is used to generate the CT image corresponding to the current patient's MR image using a trained generator, under the constraints of the current patient's MR image and reference similar case 3D CT images. In the CT image generation unit, the generator includes a feature encoder, a cascaded cross-domain transformer module, and a feature decoder. The feature encoder has two branches, which sample and encode the current patient's MR image and a reference similar case's 3D CT image, respectively, to obtain 2D MRI features and 3D CT features. The cascaded cross-domain transformer module is used to expand 2D MRI features to one dimension, combine them with batch dimension transformation, and then stitch them with 3D CT features along the depth dimension. The stitched features are then merged sequentially along the depth dimension and batch dimension, as well as the width and height dimensions. A set of trainable parameters is then stitched with the merged features along the channel dimension to obtain the stitched features. The stitched features are then sequentially encoded in the width and height dimensions, as well as in the depth dimension. The depth-encoded features are then processed by the depth dimension Transformer module to obtain the first fused feature. The first fused feature is then processed by the height-width Transformer module to obtain the second fused feature. A multi-head attention mechanism is used to filter the first and second fused features to obtain the filtered fused features. The feature decoder is used to decode the filtered fused features to obtain the corresponding CT image.
8. The system for generating CT images from MR images under multiple constraints as described in claim 7, characterized in that, In the CT image generation unit, each branch of the feature encoder contains four sampling modules, and the last two sampling modules contain a degradation model, which includes two independent processes: a blurring process and a noise addition process.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the method for generating CT images from MR images under multiple constraints as described in claim 1.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for generating CT images from MR images under multiple constraints as described in claim 1.