Method and device for generating image and method and device for training image generation model

By acquiring two orthogonal projection views and performing feature expansion and fusion processing, the problem of high time cost of three-dimensional image reconstruction in the prior art is solved, and efficient three-dimensional image reconstruction is achieved.

CN120182469APending Publication Date: 2025-06-20SHANGHAI UNITED IMAGING HEALTHCARE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311745379.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art requires multiple data acquisitions and a large number of computing resources when generating high-precision three-dimensional image reconstruction, resulting in high time costs.

Method used

By acquiring two orthogonal projected views, feature expansion and fusion processing are performed, the amount of parameters of the image generation model is reduced, and the reconstructed image is generated through the decoding reconstruction process.

Benefits of technology

It realizes that three-dimensional image reconstruction can be completed using only two orthogonal projection views, reducing the number of data acquisition times and computing resource requirements, and reducing time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182469A_ABST
    Figure CN120182469A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating an image, a method and a device for training an image generation model, computer equipment and a storage medium. The method comprises the following steps: acquiring a first projection view and a second projection view of an object in an orthogonal relationship in a projection direction; performing feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view; performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map; and performing decoding reconstruction processing on the fused feature map to obtain a reconstructed image corresponding to the object. By adopting the method, the purpose of completing three-dimensional image reconstruction only by using two orthogonal projection views can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to a method for generating an image, a method for training an image generation model, an apparatus, a computer device, a storage medium, and a computer program product. Background Art

[0002] The three-dimensional images collected by computer devices can represent the three-dimensional structural information inside the objects being collected, bringing great convenience and value to fields such as clinical medical diagnosis and item security inspection. For example, the computed tomography (CT) imaging collected by CT devices can reveal the detailed three-dimensional structural information inside the scanned object. Therefore, this technology has been widely applied, and accordingly, three-dimensional image reconstruction technology has emerged.

[0003] In related technologies, in order to generate high-precision reconstructed images, a large number of multi-angle projection views are required, and at the same time, powerful computing capabilities and a large amount of storage space are also needed. However, in practical applications, in order to obtain sufficient projection views, multiple data collection processes are required, resulting in a relatively high time cost. Summary of the Invention

[0004] Based on this, in order to solve the technical problem of the relatively long time consumed by the above method, it is necessary to provide a method for generating an image, a method for training an image generation model, an apparatus, a computer device, a computer-readable storage medium, and a computer program product.

[0005] In a first aspect, this application provides a method for generating an image. The method includes:

[0006] Obtain a first projection view and a second projection view of an object; the first projection view and the second projection view are orthogonal in the projection direction;

[0007] Perform feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view;

[0008] Perform fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fusion feature map; and

[0009] Perform decoding and reconstruction processing on the fusion feature map to obtain a reconstructed image corresponding to the object.

[0010] In one embodiment, after performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fusion feature map, it further includes:

[0011] Perform feature enhancement processing on the fusion feature map to obtain an enhanced feature map;

[0012] Perform a decoding and reconstruction process on the enhanced feature map to obtain a reconstructed image corresponding to the object.

[0013] In one embodiment, the performing a feature expansion process on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view includes:

[0014] Perform an encoding process on the first projection view and the second projection view to obtain a first encoded image corresponding to the first projection view and a second encoded image corresponding to the second projection view; and

[0015] Perform a feature expansion process on the first encoded image and the second encoded image to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0016] In one embodiment, after performing a decoding and reconstruction process on the fused feature map to obtain a reconstructed image corresponding to the object, it further includes:

[0017] Perform a segmentation process on the reconstructed image to obtain a segmented image corresponding to the reconstructed image.

[0018] In a second aspect, the present application provides a method for training an image generation model. The method includes:

[0019] Obtain a sample data set, where the sample data set includes a sample image of a sample object, a first projection view and a second projection view corresponding to the sample image, and an annotation image corresponding to the sample image, the annotation image annotates information of one or more parts of the sample object, and the first projection view and the second projection view are orthogonal in the projection direction;

[0020] Through a two-dimensional encoding module in the image generation model, perform a feature expansion process on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view;

[0021] Through a three-dimensional encoder in the image generation model, perform a fusion process on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map;

[0022] Through a three-dimensional decoder in the image generation model, perform a decoding and reconstruction process on the fused feature map to obtain a predicted image of the sample object; and

[0023] Train the image generation model based on one or more of the predicted image, the sample image, and the labeled image.

[0024] In one embodiment, after obtaining the predicted image of the sample object, it further includes:

[0025] Perform segmentation processing on the predicted image through the segmentation module in the image generation model to obtain a segmentation image corresponding to the predicted image;

[0026] Input the predicted image into a discriminator to obtain a discrimination result; and

[0027] Train the image generation model based on one or more of the segmentation image, the discrimination result, the predicted image, the sample image, and the labeled image.

[0028] In one embodiment, the training of the image generation model based on one or more of the segmentation image, the discrimination result, the predicted image, the sample image, and the labeled image includes:

[0029] Obtain a reconstruction loss based on the difference information between the predicted image and the sample image;

[0030] Obtain a segmentation loss based on the difference information between the labeled image and the segmentation image;

[0031] Respectively obtain the projection views of the predicted image and the sample image in the same at least one direction, and obtain a projection loss based on the at least one projection view corresponding to the predicted image and the at least one projection view corresponding to the sample image;

[0032] Obtain an adversarial loss between the discriminator and the image generation model based on the discrimination result and the predicted image; and

[0033] Train the image generation model according to one or more of the reconstruction loss, the segmentation loss, the projection loss, and the adversarial loss.

[0034] In a third aspect, the present application also provides an apparatus for generating an image. The apparatus includes:

[0035] A view acquisition unit for acquiring a first projection view and a second projection view of an object; the first projection view and the second projection view are orthogonal to each other in the projection direction;

[0036] A feature expansion unit for performing feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view;

[0037] A feature fusion unit for performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map; and

[0038] A decoding and reconstruction unit for performing decoding and reconstruction processing on the fused feature map to obtain a reconstructed image corresponding to the object.

[0039] Fourthly, the present application also provides an apparatus for training an image generation model. The apparatus includes:

[0040] A sample acquisition unit for fetching a sample data set, the sample data set including a sample image of a sample object, a first projection view and a second projection view corresponding to the sample image, and an annotation image corresponding to the sample image, the annotation image annotating information of one or more parts of the sample object, and the first projection view and the second projection view being orthogonal in the projection direction;

[0041] A feature expansion unit for performing feature expansion processing on the first projection view and the second projection view through a two-dimensional encoding module in the image generation model to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view;

[0042] A feature fusion unit for performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map through a three-dimensional encoder in the image generation model to obtain a fused feature map;

[0043] A decoding and reconstruction unit for performing decoding and reconstruction processing on the fused feature map through a three-dimensional decoder in the image generation model to obtain a predicted image of the sample object; and

[0044] A model training unit for training the image generation model based on one or more of the predicted image, the sample image, and the annotation image to obtain a trained image generation model.

[0045] Fifthly, the present application also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the method for generating an image or the method for training an image generation model described in any one of the above embodiments.

[0046] Sixth aspect, the present application also provides a computer-readable storage medium. On the computer-readable storage medium, a computer program is stored, and when the computer program is executed by a processor, the method for generating an image or the method for training an image generation model described in any of the foregoing embodiments is implemented.

[0047] Seventh aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the method for generating an image or the method for training an image generation model described in any of the foregoing embodiments is implemented.

[0048] For the above method for generating an image, method for training an image generation model, device, computer device, storage medium, and computer program product, after obtaining the first projection view and the second projection view of the object to be subjected to image reconstruction, feature expansion is performed on the first projection view and the second projection view, and the first three-dimensional feature map and the second three-dimensional feature map obtained by expansion are fused, thereby realizing feature expansion and fusion of the two projection views in the encoding stage, and thus the number of parameters of the image generation model can be reduced. Finally, decoding and reconstruction processing is performed on the fused feature map to obtain the reconstructed image of the object, thereby realizing the purpose of completing three-dimensional image reconstruction only by using two orthogonal projection views, and can overcome the defect in the prior art that more projection views are required and multiple data acquisition processes are required, resulting in a relatively high time cost. Description of the Drawings

[0049] Figure 1 It is a schematic flowchart of the method for generating an image in an embodiment;

[0050] Figure 2 It is a schematic diagram of the network structure of the image generation model in an embodiment;

[0051] Figure 3 It is a schematic flowchart of the method for generating an image in another embodiment;

[0052] Figure 4 It is a schematic flowchart of the method for training an image generation model in an embodiment;

[0053] Figure 5 It is a schematic diagram of the training process of the image generation model in an embodiment;

[0054] Figure 6 It is a block diagram of the structure of the device for generating an image in an embodiment;

[0055] Figure 7 It is a block diagram of the structure of the device for training an image generation model in an embodiment;

[0056] Figure 8 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0059] In one embodiment, as Figure 1 shown, a method for generating an image is provided. In this embodiment, an example is given where the method is applied to a terminal. It can be understood that the method can also be applied to a server and can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:

[0060] Step S110, obtaining a first projection view and a second projection view of an object; the first projection view and the second projection view are orthogonal to each other in the projection direction.

[0061] Among them, the object is an object for which a three-dimensional image needs to be reconstructed, and specifically can be an image of a living / non-living object or a certain part of a living / non-living object. For example, the object can be the lungs or heart of a human body or a phantom for calibration, etc.

[0062] Among them, the first projection view and the second projection view are two projection views of the object in orthogonal directions, and both the first projection view and the second projection view are two-dimensional views.

[0063] Specifically, the first projection view and the second projection view can be directly acquired from two orthogonal projection directions, or can be obtained by projecting a three-dimensional image of an object. For example, by performing projection processing on the three-dimensional image through a projection algorithm, the first projection view and the second projection view are obtained. For illustration, when the three-dimensional image is a CT image, the projection view corresponding to the CT image is a two-dimensional planar X-ray image. Then, based on the imaging principle of X-rays, the CT image can be combined with the projection algorithm to generate the projection view to simulate a traditional X-ray film. Among them, the specific process of generating the projection view through the projection algorithm includes: calculating projection geometric parameters, simulating the interaction between rays and voxels, calculating projection pixel values, and generating the projection view, etc.

[0064] Step S120: Perform feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0065] Specifically, methods such as converting the feature shape, using convolution operations, and using attention mechanisms can be used to expand the two-dimensional projection view into a three-dimensional feature map. Among them, converting the feature shape is to convert the two-dimensional feature map into a three-dimensional feature map by adjusting the shape of the two-dimensional feature map. Using convolution operations is to expand the two-dimensional feature map into a three-dimensional feature map by using convolution operations. In the convolution operation, different-sized convolution kernels can be used to perform convolution operations on the two-dimensional feature map to obtain a three-dimensional feature map. Using attention mechanisms is to introduce an attention mechanism to weight the two-dimensional feature map to obtain a three-dimensional feature map. The attention mechanism can give different attentions to different parts of the input feature map to obtain a richer feature representation.

[0066] The two-dimensional projection view can also be feature-expanded through the two-dimensional encoding module of the image generation model to obtain a three-dimensional feature map.

[0067] Among them, the image generation model is a model that takes two orthogonal projection views as input variables and a three-dimensional reconstructed image as output.

[0068] In specific implementation, the image generation model can be pre-trained through a sample data set. The sample data set can include sample images of sample objects, the first projection view and the second projection view corresponding to the sample images, and the annotated images corresponding to the sample images. More specifically, the image generation model can be trained with the first projection view and the second projection view corresponding to the sample image as input variables, the annotated image corresponding to the sample image as supervision information, and the predicted image of the sample object as output variables.

[0069] After completing the training of the image generation model, when performing 3D image reconstruction for any object, the orthogonal first projection view and second projection view collected for the object can be input into the image generation model, and the image generation model performs image reconstruction.

[0070] More specifically, the image generation model includes a 2D encoder, a 3D encoder, and a 3D decoder, and the 2D encoder, 3D encoder, and 3D decoder are connected in sequence, that is, the input of the 3D encoder is the output of the 2D encoder, and the input of the 3D decoder is the output of the 3D encoder. After inputting the first projection view and the second projection view into the image generation model, the first projection view and the second projection view first enter the 2D encoder, and the 2D encoder performs feature expansion on the two projection views respectively to obtain a first 3D feature map corresponding to the first projection view and a second 3D feature map corresponding to the second projection view.

[0071] Step S130: Perform a fusion process on the first 3D feature map and the second 3D feature map to obtain a fused feature map.

[0072] Specifically, fusion algorithms such as element-wise addition, element-wise multiplication, concatenation, and attention mechanism can be used for the fusion process. Among them, element-wise addition is to add the corresponding elements of the two 3D feature maps to obtain a new 3D feature map. Element-wise multiplication is to multiply the corresponding elements of the two 3D feature maps to obtain a new 3D feature map. Concatenation is to concatenate the two 3D feature maps along the channel direction and then perform a convolution operation to obtain a new 3D feature map. The attention mechanism is to perform weighted fusion on the two 3D feature maps to obtain a new 3D feature map. The attention mechanism can give different attentions to different parts of the input feature map, so as to achieve more flexible feature fusion.

[0073] It is also possible to use the 3D encoder of the image generation model to fuse the 3D feature maps corresponding to the two projection views.

[0074] In specific implementation, after the first projection view and the second projection view pass through the 2D encoder and output the first 3D feature map and the second 3D feature map, the first 3D feature map and the second 3D feature map enter the 3D encoder, and the 3D encoder performs a fusion process on the first 3D feature map and the second 3D feature map to obtain a fused feature map.

[0075] Step S140: Perform a decoding and reconstruction process on the fused feature map to obtain a reconstructed image corresponding to the object.

[0076] Among them, the reconstructed image is a 3D image generated for the object. For example, the reconstructed image can be a 3D magnetic resonance image (MR image), a computed tomography image (CT image), etc.

[0077] Specifically, methods such as 3D convolutional neural network decoder, 3D generative adversarial network, 3D interpolation and reconstruction can be used to decode and reconstruct the fused feature map. Among them, the 3D convolutional neural network decoder uses a 3D CNN decoder to reverse the process of the encoder, and restores the encoded feature map to the original three-dimensional feature map through upsampling or deconvolution operations. The generator in the 3D generative adversarial network is used to map the random vector in the latent space of the fused feature map to the three-dimensional feature map space, while the discriminator is responsible for distinguishing the generated predicted image and the real sample image. 3D interpolation and reconstruction is to use interpolation methods to reconstruct the lost or compressed representation of the three-dimensional feature map.

[0078] The fused feature map can also be decoded and reconstructed through a three-dimensional decoder. Specifically, the fused feature map obtained by fusing the first three-dimensional feature map and the second three-dimensional feature map through a three-dimensional encoder can be input into the three-dimensional decoder, and the three-dimensional decoder decodes and reconstructs the fused feature map to obtain a reconstructed image.

[0079] In the above method for generating an image, after obtaining the first projection view and the second projection view of the object to be image-reconstructed, feature expansion is performed on the first projection view and the second projection view, and the first three-dimensional feature map and the second three-dimensional feature map obtained by the expansion are fused, realizing feature expansion and fusion of the two projection views in the encoding stage, thereby reducing the number of parameters of the image generation model. Finally, the fused feature map is decoded and reconstructed to obtain the reconstructed image of the object, thus achieving the purpose of completing three-dimensional image reconstruction only by using two orthogonal projection views, and overcoming the defect that in the prior art, more projection views are required and multiple data acquisition processes are required, resulting in a relatively high time cost.

[0080] In an exemplary embodiment, after the step S130 of fusing the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map, it further includes: performing feature enhancement processing on the fused feature map to obtain an enhanced feature map; performing decoding and reconstruction processing on the enhanced feature map to obtain a reconstructed image corresponding to the object.

[0081] In a specific implementation, in order to further improve the quality of the reconstructed image, after obtaining the fused feature map and before decoding and reconstructing the fused feature map, the fused feature map can be subjected to feature enhancement processing. Specifically, the fused feature map can be enhanced through 3D convolution operations, 3D pooling operations, and 3D attention mechanisms, etc. Among them, by using 3D convolution operations, different convolutional kernels can be applied to extract features of different scales and directions, thereby enhancing the expressive ability of the feature map. Using 3D pooling operations can perform downsampling to reduce the size of the feature map while retaining important features. The 3D attention mechanism can give different attentions to different parts of the feature map, thereby enhancing the expressive ability of important regions in the feature map and improving the performance of the model.

[0082] It is also possible to set a conversion module between the 3D encoder and the 3D decoder. The conversion module performs feature enhancement processing on the fused feature map output by the 3D encoder. After the enhancement processing, the obtained enhanced feature map is input into the 3D decoder for decoding and reconstruction, thereby improving the quality of the obtained reconstructed image. Among them, the conversion module is used to enhance the features of the input image to improve the quality of the subsequently generated reconstructed image.

[0083] More specifically, the conversion module can adopt a Transformer model based on the attention mechanism to avoid the problems of information loss and insufficient information fusion in general generation networks when processing complex data, enhance the global perception ability of the image generation model for the input fused feature map, and thus optimize the quality of the subsequently generated reconstructed image. Among them, multiple conversion modules can be set, such as 12. The conversion module adopting the Transformer model receives the feature input, outputs after passing through the multi-head attention mechanism and the feed-forward neural network, and the feature shape does not change. Before performing the decoding and reconstruction processing on the fused feature map, the fused feature map can be processed by the Transformer model.

[0084] In this embodiment, by performing feature enhancement processing on the fused feature map, the quality of the image to be reconstructed is better, thereby improving the quality and generation effect of the reconstructed image decoded and reconstructed based on the enhanced feature map, making the reconstructed image closer to the real image.

[0085] In an exemplary embodiment, in the above step S120, when performing feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view, it specifically includes: performing encoding processing on the first projection view and the second projection view to obtain a first encoded image corresponding to the first projection view and a second encoded image corresponding to the second projection view; performing feature expansion processing on the first encoded image and the second encoded image to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0086] In specific implementation, the feature expansion of the projection view includes two steps: encoding the projection view to obtain an encoded image, and then performing feature expansion on the encoded image to obtain a three-dimensional feature map. Therefore, when performing feature expansion processing on the first projection view and the second projection view through a two-dimensional encoding module, the two-dimensional encoding module includes a two-dimensional encoder and a feature expansion sub-module. Through the two-dimensional encoding module, the specific process of feature expansion processing on the first projection view and the second projection view is as follows: first, through the two-dimensional encoder in the two-dimensional encoding module, perform encoding processing on the first projection view and the second projection view to obtain a first encoded image corresponding to the first projection view and a second encoded image corresponding to the second projection view. Then, through the feature expansion sub-module in the two-dimensional encoding module, perform feature expansion processing on the first encoded image and the second encoded image to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0087] In this embodiment, by performing encoding processing on the first projection view and the second projection view, the gray value of each pixel point in the projection view can be converted into a binary code with a fixed length, and the two-dimensional projection view can be converted into a two-dimensional binary encoded image to compress the image information, reduce the storage space and the amount of calculation. Then, perform three-dimensional feature expansion on the two encoded images to obtain a three-dimensional feature map, and the three-dimensional feature map can contain more information, so as to ensure the accuracy of the image decoded and reconstructed based on the three-dimensional feature map.

[0088] In an exemplary embodiment, after the above step S140 performs decoding and reconstruction processing on the fusion feature map to obtain a reconstructed image corresponding to the object, it further includes: performing segmentation processing on the reconstructed image to obtain a segmentation image corresponding to the reconstructed image.

[0089] Among them, the information of each part of the object can be marked in the segmentation image. For example, if the object is the lung, the information of parts such as the lung, vertebra, and rib can be marked in the segmentation image.

[0090] Specifically, image segmentation is the process of dividing an image into regions or objects with semantics. The reconstructed image can be segmented by a threshold-based segmentation method, an edge-based segmentation method, or a deep learning-based segmentation method. Among them, threshold-based segmentation compares the pixel values of the image with a predefined threshold and assigns the pixels to different regions according to the comparison results. Edge-based segmentation uses edge detection algorithms (such as Sobel, Canny, etc.) to detect the edges in the image, and then segments the image into different regions according to the edge information. Deep learning-based segmentation uses a convolutional neural network for end-to-end image segmentation. For example, by setting up a segmentation module, the reconstructed image is segmented to obtain a segmented image.

[0091] In the image generation model of the present application, in addition to a two-dimensional encoding module, a three-dimensional encoder, and a three-dimensional decoder for performing three-dimensional image reconstruction, a segmentation module is also provided. When training the image generation model, the segmentation module is co-trained so that the generated reconstructed image has richer semantic information, and for key structures, three-dimensional semantic segmentation results can be obtained.

[0092] In one embodiment, to facilitate those skilled in the art to understand the method for generating an image provided in the above embodiment, the following is specifically described with reference to the Figure 2 example.

[0093] Refer to Figure 2 , which is a schematic diagram of the network structure of the image generation model provided by the present application. As Figure 2 shown, the image generation model includes a two-dimensional encoder, a feature expansion sub-module, a three-dimensional encoder, a transformation module (Transformer), and a three-dimensional decoder. As Figure 3 shown, the process of image reconstruction based on the Figure 2 image generation model specifically includes the following steps:

[0094] Step S310, obtaining a first projection view and a second projection view of the object; the first projection view and the second projection view are orthogonal in the projection direction.

[0095] Step S320, encoding the first projection view and the second projection view through the two-dimensional encoder in the two-dimensional encoding module to obtain a first encoded image corresponding to the first projection view and a second encoded image corresponding to the second projection view.

[0096] Specifically, there are two two-dimensional encoders, and the first projection view and the second projection view can be encoded through the two two-dimensional encoders respectively to obtain a first encoded image corresponding to the first projection view and a second encoded image corresponding to the second projection view.

[0097] Step S330: Through the feature expansion sub-module in the two-dimensional encoding module, perform feature expansion processing on the first encoded image and the second encoded image to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0098] Specifically, there are also two feature expansion sub-modules. The first encoded image and the second encoded image can be respectively subjected to feature expansion processing through the two feature expansion sub-modules to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0099] Step S340: Through the three-dimensional encoder in the image generation model, perform fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map.

[0100] Step S350: Through the conversion module in the image generation model, perform feature enhancement processing on the fused feature map to obtain an enhanced feature map.

[0101] Step S360: Through the three-dimensional decoder in the image generation model, perform decoding and reconstruction processing on the enhanced feature map to obtain a reconstructed image corresponding to the object.

[0102] Step S370: Through the segmentation module in the image generation model, perform segmentation processing on the reconstructed image to obtain a segmentation image corresponding to the reconstructed image.

[0103] The method provided in this embodiment can synthesize a three-dimensional reconstructed image with two orthogonally oriented projection views as input, thus avoiding the need to collect a large number of projection views and reducing time waste. By introducing a conversion module based on the attention mechanism to perform feature enhancement on the fused feature map, the problems of information loss and insufficient information fusion in general generation networks when processing complex data can be effectively avoided, improving the global perception ability of the image generation model for images, and thus optimizing the quality of the generated reconstructed image. Through the two-dimensional encoding module and the three-dimensional encoder, feature expansion and fusion of the first projection view and the second projection view are achieved, thereby realizing feature expansion and fusion at the encoding stage of the two projection views, and thus reducing the number of parameters of the image generation model.

[0104] In one embodiment, as Figure 4 shown, a method for training an image generation model is also provided. In this embodiment, the method includes the following steps:

[0105] Step S410: Obtain a sample data set, which includes sample images of sample objects, a first projection view and a second projection view corresponding to the sample images, and an annotation image corresponding to the sample images. The annotation image annotates information of one or more parts of the sample object, and the first projection view and the second projection view are orthogonal to each other in the projection direction.

[0106] Among them, the sample image is a three-dimensional image.

[0107] Among them, there can be multiple sample objects, but they need to belong to the same type of object. For example, they are all lungs, or they are all hearts.

[0108] Among them, the annotation image is an image with the same size as the sample image, indicating the categories of different voxel points on the sample image. Specifically, the annotation image can be obtained by using different voxel values to annotate different parts. For example, for a lung image as the sample image, the parts it contains may include the lungs, vertebrae, and ribs. Then, for the annotation image of the lung image, it can be: using voxel value 0 to represent voxel points belonging to the background; using voxel value 1 to represent voxel points belonging to the lungs; using voxel value 2 to represent voxel points belonging to the vertebrae and rib parts.

[0109] In specific implementation, for the sample image of the sample object, it can be acquired by a computer device. For example, when the sample image is a CT image, the CT image of the sample object can be acquired by a CT device. For the first projection view and the second projection view, the sample image of the sample object can be obtained first, and then the sample image is projected by a projection algorithm to obtain the first projection view and the second projection view corresponding to the sample image. It can also be directly acquired from two orthogonal projection directions. When obtained by projection algorithm processing, the process includes: calculating projection geometric parameters, simulating the interaction between rays and voxels, calculating projection pixel values, and generating projection views, etc.

[0110] Step S420: Through the two-dimensional encoding module in the image generation model, perform feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0111] Specifically, the image generation model includes a two-dimensional encoder, a three-dimensional encoder, and a three-dimensional decoder, and the two-dimensional encoder, the three-dimensional encoder, and the three-dimensional decoder are connected in sequence, that is, the input of the three-dimensional encoder is the output of the two-dimensional encoder, and the input of the three-dimensional decoder is the output of the three-dimensional encoder. After inputting the first projection view and the second projection view into the image generation model, the first projection view and the second projection view first enter the two-dimensional encoder, and the two-dimensional encoder performs feature expansion on the two projection views respectively to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0112] Step S430: Use the 3D encoder in the image generation model to fuse the first 3D feature map and the second 3D feature map to obtain a fused feature map.

[0113] Specifically, after the first projection view and the second projection view pass through the 2D encoder and output the first 3D feature map and the second 3D feature map, the first 3D feature map and the second 3D feature map enter the 3D encoder, and the 3D encoder fuses the first 3D feature map and the second 3D feature map to obtain a fused feature map.

[0114] Step S440: Use the 3D decoder in the image generation model to perform decoding and reconstruction processing on the fused feature map to obtain a predicted image of the sample object.

[0115] Specifically, after the fusion of the first 3D feature map and the second 3D feature map by the 3D encoder, the obtained fused feature map will enter the 3D decoder, and the 3D decoder performs decoding and reconstruction on the fused feature map to obtain a predicted image.

[0116] Step S450: Train the image generation model based on one or more of the predicted image, the sample image, and the annotated image.

[0117] Specifically, the training of the image generation model includes the training of the 2D encoding module, the 3D encoder, and the 3D decoder. Specifically, the training loss can be determined by one or more of the predicted image, the sample image, and the annotated image, and the image generation model is trained based on the training loss value.

[0118] In the above method for training the image generation model, after obtaining the first projection view and the second projection view of the sample object to be reconstructed, the 2D encoding module is used to perform feature expansion on the first projection view and the second projection view, and the 3D encoder is used to fuse the first 3D feature map and the second 3D feature map obtained by the expansion, realizing the feature expansion and fusion of the two projection views in the encoding stage, thereby reducing the number of parameters of the image generation model. Finally, the 3D decoder performs decoding and reconstruction processing on the fused feature map to obtain the reconstructed image of the object, thereby achieving the purpose of completing 3D image reconstruction only using two orthogonal projection views, and can overcome the defect in the prior art that more projection views are required and multiple data acquisition processes are required, resulting in a relatively high time cost.

[0119] In one embodiment, after obtaining the sample data set, the sample data set can also be preprocessed, and the image generation model is trained using the preprocessed sample data set. More specifically, the preprocessing of the sample data set can include the following methods:

[0120] Method 1: For the sample images and labeled images in the sample dataset, perform the first normalization process and voxel normalization process successively to obtain the processed images and processed labeled images.

[0121] Among them, the first normalization process is used to normalize the voxel values and sizes of the sample images and labeled images.

[0122] Specifically, both the sample images and labeled images are three-dimensional images. The normalization process for them includes resampling to standardize the voxels and size normalization. Specifically, through bilinear interpolation or nearest neighbor interpolation algorithms, the sample images and labeled images can be resampled to the standard voxel size, such as (1, 1, 1), and then the resampled sample images and labeled images are cropped to adjust their sizes to a unified size, such as 128×128×128, to achieve the normalization process for the sample images and labeled images. Then, the normalized sample images and normalized labeled images are normalized according to the maximum value MAX1 and minimum value MIN1 of the voxels. The normalization method can be expressed by the formula:

[0123] Normalized voxel value = (Voxel value before normalization - MIN1) / (MAX1 - MIN1)

[0124] Among them, before normalizing the normalized sample images and normalized labeled images, the voxel values of the normalized sample images and normalized labeled images are truncated in the range [0, 2500].

[0125] Method 2: For the first projection view and the second projection view in the sample dataset, perform the second normalization process and pixel normalization process successively to obtain the first processed projection view and the second processed projection view.

[0126] Among them, the second normalization process is used to normalize the sizes of the first projection view and the second projection view.

[0127] Specifically, both the first projection view and the second projection view are two-dimensional images. The normalization process for them includes size normalization. The specific size normalization is to crop the first projection view and the second projection view to adjust their sizes to a unified size, such as 128×128, thereby achieving the normalization process for the first projection view and the second projection view. Then, the normalized first projection view and the normalized second projection view labeled images are normalized according to the maximum value MAX2 and minimum value MIN2 of the pixels. The normalization method can be expressed by the formula:

[0128] Normalized pixel value = (Pixel value before normalization - MIN2) / (MAX2 - MIN2)

[0129] Finally, the obtained processed image, processed annotation image, first processed projection view, and second processed projection view are combined to form a preprocessed sample data set.

[0130] In this embodiment, the sample data set is standardized and normalized to ensure the consistency of the images in the sample data set, so as to improve the training efficiency and training accuracy of the image generation model.

[0131] In an exemplary embodiment, after obtaining the predicted image of the sample object in step S440 above, it further includes:

[0132] Step S441, perform segmentation processing on the predicted image through the segmentation module in the image generation model to obtain a segmentation image corresponding to the predicted image;

[0133] Step S442, input the predicted image into the discriminator to obtain a discrimination result;

[0134] Step S443, train the image generation model based on one or more of the segmentation image, discrimination result, predicted image, sample image, and annotation image.

[0135] Specifically, the image generation model of the present application further includes a segmentation module. After obtaining the predicted image decoded and reconstructed by the three-dimensional decoder, the predicted image is input into the segmentation module, and the segmentation module performs segmentation processing on the predicted image to obtain a segmentation image. Further, in order to discriminate the generation result of the predicted image, the predicted image can be input into the discriminator for judgment to obtain a true or false discrimination result. Then, the image generation model is trained based on one or more of the segmentation image, discrimination result, predicted image, sample image, and annotation image.

[0136] Among them, the two-dimensional encoding module, three-dimensional encoder, and three-dimensional decoder are used to generate the predicted image and can be regarded as the generation network. The segmentation module is used to segment the image and can be regarded as the segmentation network. Thus, the image generation model can be regarded as a model including a generation network and a segmentation network. Among them, the algorithm model adopted by the generation network is the Generative Adversarial Nets (GAN). The principle of image generation for two orthogonal projection views is: the two-dimensional encoder performs independent downsampling encoding on the two projection views respectively, and then the feature expansion module converts the two-dimensional features of the two projection views into three-dimensional features. Finally, in the upsampling decoding stage, the three-dimensional encoder first fuses the two three-dimensional features, and the three-dimensional decoder decodes and generates a reconstructed image based on the fused features.

[0137] In this application, the generative adversarial network can accept two orthogonally oriented projection views as inputs to synthesize a three-dimensional reconstructed image. The introduction of the generative adversarial network can ensure that the generated reconstructed image is visually closer to the real image, so as to achieve the purpose of completing image reconstruction to a certain extent using only two orthogonally oriented projection views. Moreover, by using the strategy of co-training the generative network and the segmentation network, the quality of the semantic information of the subsequently generated reconstructed image can be further improved, and a three-dimensional segmentation result of a specific structure can be obtained from two orthogonally oriented projection views, with wider applicability.

[0138] In one exemplary embodiment, in step S443, based on one or more of the segmentation image, the discrimination result, the prediction image, the sample image, and the annotation image, training is performed on the image generation model, which specifically includes:

[0139] Step S443a, based on the difference information between the prediction image and the sample image, a reconstruction loss is obtained.

[0140] Specifically, the comparison between the prediction image and the sample image is to determine the accuracy of the generated image. Therefore, in order to ensure that the generated prediction image is closer to the real sample image, the mean square error loss can be used to calculate the reconstruction loss between the prediction image and the sample image. If the sample image is denoted as y, the first projection view and the second projection view of the sample image are denoted as x, and the prediction image generated by the image generation model is denoted as G(x), then the reconstruction loss l rec can be expressed by the formula: .

[0141] Step S443b, based on the difference information between the annotation image and the segmentation image, a segmentation loss is obtained.

[0142] Specifically, the comparison between the annotation image and the segmentation image is to determine the accuracy of the segmentation result of the segmentation module. In order to improve the semantic information of the generated image, the dice loss function (Dice Loss) and the cross-entropy loss function (Cross-Entropy Loss) can be used to calculate the segmentation loss.

[0143] The segmentation loss l seg can be expressed by the formula:

[0144] l seg = w 1 [ dice S y ,GT +ce S y ,GT ]+ w 2 [ dice S G(x) ,GT +ce S G(x) ,GT ]+ w 3 [ dice S G x ,S(y) +ce S G x ,S(y) ]

[0145] Among them, dice and ce respectively represent the dice loss and the cross-entropy loss. , , respectively represent the weights of the segmentation loss for segmenting each part of the sample image. G is the generation module, D is the discriminator, and S is the segmentation network. The sample image is y, and the first projection view and the second projection view of the sample image are x.

[0146] In step S443c, obtain the projection views of the predicted image and the sample image in at least one same direction respectively, and based on at least one projection view corresponding to the predicted image and at least one projection view corresponding to the sample image, obtain the projection loss.

[0147] Specifically, after obtaining the predicted image, the predicted image and the sample image can be projected respectively to obtain at least one projection view corresponding to the predicted image and at least one projection view corresponding to the sample image. Among them, the projection directions of the predicted image and the sample image for projection must be the same, and the obtained projection views are paired. For example, project the predicted image and the sample image in the cross-sectional direction respectively to obtain the cross-sectional projection view of the predicted image and the cross-sectional projection view of the sample image. Another example is to project the predicted image and the sample image in the cross-sectional direction and the sagittal direction respectively to obtain the cross-sectional projection view and the sagittal projection view of the predicted image, and the cross-sectional projection view and the sagittal projection view of the sample image. Then, based on the difference information between each group of projection views of the predicted image and the sample image, obtain the projection loss.

[0148] More specifically, the L1 loss function can be used to calculate the projection loss on the two-dimensional plane, and when there are multiple projection directions, the projection losses in each direction can be calculated separately first, and then based on the projection losses in each direction, determine the total projection loss.

[0149] For example, taking the projection of the predicted image and the sample image in three directions, namely the cross-sectional direction, the sagittal direction, and the coronal direction, as an example, after obtaining the three projection views of the predicted image and the three projection views of the sample image, the loss can be calculated separately on the cross-section, the sagittal plane, and the coronal plane first, and then calculate the average value of the corresponding losses of the three planes, and take the obtained average value as the projection loss.

[0150] The projection loss l proj The calculation formula of can be expressed as:

[0151]

[0152] Among them, , , respectively represent orthogonal projection operations on the cross-sectional plane, sagittal plane, and coronal plane. The sample image is y, the first projection view and the second projection view of the sample image are x, and the predicted image generated by the image generation model is G(x).

[0153] Step S443d, based on the discrimination result and the predicted image, obtain the adversarial loss between the discriminator and the image generation model.

[0154] Specifically, if G is still set as the generation module, D is the discriminator, the sample image is y, the first projection view and the second projection view of the sample image are x, and the predicted image generated by the image generation model is G(x), then the adversarial loss can be expressed by the formula:[[]] .

[0155] The loss of the discriminator can be expressed by the formula:[[]] .

[0156] Step S443e, train the image generation model according to one or more of the reconstruction loss, segmentation loss, projection loss, and adversarial loss.

[0157] Specifically, weights can be assigned to the reconstruction loss, segmentation loss, projection loss, and adversarial loss, and then weighted summation is performed according to the weights to obtain the total loss. For the purpose of reducing the total loss, the image generation model is trained until the loss converges or reaches the preset number of training times, and the training is ended to obtain the trained image generation model.

[0158] In this embodiment, the image generation model is trained based on one or more of the reconstruction loss, segmentation loss, projection loss, and adversarial loss to ensure the accuracy of the trained image generation model.

[0159] In one embodiment, for the convenience of those skilled in the art to understand the embodiments of the present application, hereinafter, taking the image to be reconstructed as a lung CT image as an example, the principle of the present application will be specifically described in combination with the specific examples of the attached Figure 5 drawings.

[0160] Refer to Figure 5 , which shows a schematic diagram of the process of training an image generation model. As Figure 5 shown, the image generation model can be divided into a generation network and a segmentation network. Among them, the generation network is an adversarial generation network, including a generator and a discriminator. The generation network includes a 2D encoder, a feature expansion sub-module, a 3D encoder, and a 3D decoder. When training the image generation model, the following steps are included:[[]]

[0161] (1)Input the first and second projection views at orthogonal 0° and 90°, and through the 2D encoder, feature expansion sub-module, 3D encoder, and 3D decoder in the generator, output the reconstructed predicted CT image.

[0162] Among them, the two input projection views in orthogonal directions can also be replaced with plain films.

[0163] (2)Input the predicted image into the segmentation network to obtain the predicted segmented image. Also, input the predicted image into the discriminator to obtain the discrimination result on the authenticity of the predicted image.

[0164] (3)Based on the comparison between the predicted image and the real sample image, obtain the reconstruction loss. Based on the discrimination result and the predicted image, obtain the adversarial loss. Orthogonally project the predicted image in three directions: cross-section, sagittal plane, and coronal plane to obtain three projection views. Similarly, orthogonally project the corresponding sample image in the cross-section, sagittal plane, and coronal plane to obtain three projection views. Compare the projection views corresponding to the predicted image and the sample image to obtain the projection loss. Compare the predicted segmented image with the labeled image of the pre-labeled sample image to obtain the segmentation loss.

[0165] (4)Based on the reconstruction loss, adversarial loss, two-dimensional projection loss, and segmentation loss, perform weighted summation as the total loss function L of the model. Continuously optimize the parameters of the image generation model through backpropagation in the image generation model, and train the entire model to minimize L. Continuously repeat the above training process, so that the predicted image generated by the generator of the generation network is continuously close to the real CT image, and the discrimination ability of the discriminator of the generation network for true and false CT also improves accordingly, and the two reach a balance in continuous training.

[0166] The CT image reconstruction method based on deep learning proposed in this embodiment has the following beneficial effects: (1) Compared with traditional methods, this method only needs to use two projection views in orthogonal directions to reconstruct a three-dimensional CT image with moderate accuracy. It can significantly reduce the need to obtain a large number of projection views, directly reduce the time and cost of data acquisition, and at the same time can also reduce the radiation exposure risk caused by repeated CT scans of the acquisition object. (2) It realizes the reduction of the number of model parameters and the enhancement of the semantic information of the generated image, thereby improving the image generation efficiency and image generation effect. (3) This method reduces the requirements for scanning and imaging equipment, making computed tomography technology more available and expandable. Especially in the case of limited resource environment, this method can obtain the internal three-dimensional structure information of the scanned object at a lower cost, expand the application range of computed tomography technology, and maximize the value.

[0167] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0168] Based on the same inventive concept, the embodiments of the present application further provide an apparatus for generating an image for implementing the method for generating an image involved above, and an apparatus for training an image generation model for implementing the method for training an image generation model involved above. The implementation solutions for solving problems provided by this apparatus are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more of the following apparatus embodiments can refer to the limitations on the corresponding method in the above text, and will not be repeated here.

[0169] In one embodiment, as Figure 6 shown, an apparatus for generating an image is provided, including: a view acquisition unit 610, a feature expansion unit 620, a feature fusion unit 630, and a decoding and reconstruction unit 640, where:

[0170] The view acquisition unit 610 is configured to acquire a first projection view and a second projection view of an object; the first projection view and the second projection view are in an orthogonal relationship in the projection direction;

[0171] The feature expansion unit 620 is configured to perform feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view;

[0172] The feature fusion unit 630 is configured to perform fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fusion feature map; and

[0173] The decoding and reconstruction unit 640 is configured to perform decoding and reconstruction processing on the fusion feature map to obtain a reconstructed image corresponding to the object to be reconstructed.

[0174] In one embodiment, the above apparatus further includes a feature enhancement module configured to perform feature enhancement processing on the fusion feature map to obtain an enhanced feature map;

[0175] The decoding and reconstruction unit 640 is further configured to perform decoding and reconstruction processing on the enhanced feature map to obtain a reconstructed image corresponding to the object.

[0176] In one embodiment, the feature expansion unit 620 is further configured to perform encoding processing on the first projection view and the second projection view to obtain a first encoded image corresponding to the first projection view and a second encoded image corresponding to the second projection view; perform feature expansion processing on the first encoded image and the second encoded image to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

[0177] In one embodiment, the above device further includes a segmentation unit configured to perform segmentation processing on the reconstructed image to obtain a segmented image corresponding to the reconstructed image.

[0178] In one embodiment, as Figure 7 shown, a device for training an image generation model is provided, including: a sample acquisition unit 710, a feature expansion unit 720, a feature fusion unit 730, a decoding and reconstruction unit 740, and a model training unit 750, wherein:

[0179] The sample acquisition unit 710 is configured to acquire a sample data set; the sample data set includes a sample image of a sample object, a first projection view and a second projection view corresponding to the sample image, and an annotation image corresponding to the sample image; the annotation image annotates information of each part of the sample object, and the first projection view and the second projection view are orthogonal in the projection direction;

[0180] The feature expansion unit 720 is configured to perform feature expansion processing on the first projection view and the second projection view through a two-dimensional encoding module in the image generation model to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view;

[0181] The feature fusion unit 730 is configured to perform fusion processing on the first three-dimensional feature map and the second three-dimensional feature map through a three-dimensional encoder in the image generation model to obtain a fused feature map;

[0182] The decoding and reconstruction unit 740 is configured to perform decoding and reconstruction processing on the fused feature map through a three-dimensional decoder in the image generation model to obtain a predicted image of the sample object; and

[0183] The model training unit 750 is configured to train the image generation model based on one or more of the predicted image, the sample image, and the annotation image to obtain a trained image generation model.

[0184] In one embodiment, the above device further includes a segmentation unit, which is configured to perform segmentation processing on the predicted image through the segmentation unit in the image generation model to obtain a segmentation image corresponding to the predicted image; input the predicted image into a discriminator to obtain a discrimination result; and

[0185] The model training unit 750 is further configured to train the image generation model based on one or more of the segmentation image, the discrimination result, the predicted image, the sample image, and the labeled image.

[0186] In one embodiment, the above model training unit 750 is further configured to obtain a reconstruction loss based on the difference information between the predicted image and the sample image; obtain a segmentation loss based on the difference information between the labeled image and the segmentation image; respectively obtain the projection views of the predicted image and the sample image in the same at least one direction, and obtain a projection loss based on at least one projection view corresponding to the predicted image and at least one projection view corresponding to the sample image; the projection directions of the three projection views of the predicted image and the three projection views of the sample image are the same; obtain an adversarial loss between the discriminator and the image generation model based on the discrimination result and the predicted image; and train the image generation model according to one or more of the reconstruction loss, the segmentation loss, the projection loss, and the adversarial loss.

[0187] Each unit in the above device for generating an image can be implemented in whole or in part by software, hardware, and their combination. The above units can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective units.

[0188] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a method for generating images. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball, or touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0189] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0190] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0191] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0192] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0193] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0194] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0195] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0196] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for generating an image, characterized in that, The method includes: Obtaining a first projection view and a second projection view of an object; the first projection view and the second projection view are orthogonal in the projection direction; Performing feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view; Performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map; and Performing decoding and reconstruction processing on the fused feature map to obtain a reconstructed image corresponding to the object.

2. The method according to claim 1, characterized in that, After performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map, it further includes: Performing feature enhancement processing on the fused feature map to obtain an enhanced feature map; Performing decoding and reconstruction processing on the enhanced feature map to obtain a reconstructed image corresponding to the object.

3. The method according to claim 1, characterized in that, The performing feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view includes: Performing encoding processing on the first projection view and the second projection view to obtain a first encoded image corresponding to the first projection view and a second encoded image corresponding to the second projection view; and Performing feature expansion processing on the first encoded image and the second encoded image to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view.

4. The method according to claim 1, characterized in that, After performing decoding and reconstruction processing on the fused feature map to obtain a reconstructed image corresponding to the object, it further includes: Performing segmentation processing on the reconstructed image to obtain a segmented image corresponding to the reconstructed image.

5. A method for training an image generation model, characterized in that, The method includes: Obtaining a sample data set, the sample data set including a sample image of a sample object, a first projection view and a second projection view corresponding to the sample image, and an annotation image corresponding to the sample image, the annotation image annotating information of one or more parts of the sample object, the first projection view and the second projection view being orthogonal in the projection direction; Performing feature expansion processing on the first projection view and the second projection view through a two-dimensional encoding module in the image generation model to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view; Performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map through a three-dimensional encoder in the image generation model to obtain a fused feature map; Performing decoding and reconstruction processing on the fused feature map through a three-dimensional decoder in the image generation model to obtain a predicted image of the sample object; and Training the image generation model based on one or more of the predicted image, the sample image, and the annotation image.

6. The method according to claim 5, characterized in that, After obtaining the predicted image of the sample object, it further includes: Through the segmentation module in the image generation model, perform segmentation processing on the predicted image to obtain a segmentation image corresponding to the predicted image; Input the predicted image into a discriminator to obtain a discrimination result; and Train the image generation model based on one or more of the segmentation image, the discrimination result, the predicted image, the sample image, and the annotation image.

7. The method according to claim 6, characterized in that, The training of the image generation model based on one or more of the segmentation image, the discrimination result, the predicted image, the sample image, and the annotation image includes: Obtain a reconstruction loss based on the difference information between the predicted image and the sample image; Obtain a segmentation loss based on the difference information between the annotation image and the segmentation image; Respectively obtain the projection views of the predicted image and the sample image in at least one same direction, and obtain a projection loss based on at least one projection view corresponding to the predicted image and at least one projection view corresponding to the sample image; Obtain an adversarial loss between the discriminator and the image generation model based on the discrimination result and the predicted image; and Train the image generation model according to one or more of the reconstruction loss, the segmentation loss, the projection loss, and the adversarial loss.

8. An apparatus for generating an image, characterized in that, The device includes: A view acquisition unit for acquiring a first projection view and a second projection view of an object; the first projection view and the second projection view are orthogonal to each other in the projection direction; A feature expansion unit for performing feature expansion processing on the first projection view and the second projection view to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view; A feature fusion unit for performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map to obtain a fused feature map; and A decoding and reconstruction unit for performing decoding and reconstruction processing on the fused feature map to obtain a reconstructed image corresponding to the object.

9. An apparatus for training an image generation model, characterized in that, The device includes: A sample acquisition unit for acquiring a sample data set, the sample data set including a sample image of a sample object, a first projection view and a second projection view corresponding to the sample image, and an annotation image corresponding to the sample image, the annotation image annotating information of one or more parts of the sample object, the first projection view and the second projection view being orthogonal to each other in the projection direction; A feature expansion unit for performing feature expansion processing on the first projection view and the second projection view through a two-dimensional encoding module in the image generation model to obtain a first three-dimensional feature map corresponding to the first projection view and a second three-dimensional feature map corresponding to the second projection view; A feature fusion unit for performing fusion processing on the first three-dimensional feature map and the second three-dimensional feature map through a three-dimensional encoder in the image generation model to obtain a fused feature map; A decoding and reconstruction unit, configured to perform decoding and reconstruction processing on the fused feature map through a three-dimensional decoder in the image generation model to obtain a predicted image of the sample object; and A model training unit, configured to train the image generation model based on one or more of the predicted image, the sample image, and the annotated image to obtain a trained image generation model.

10. A computer-readable storage medium having instructions stored thereon, characterized in that, The instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Tunnel karst advanced geological forecast automatic interpretation method, device and medium

    CN120852694A