X-ray image three-dimensional reconstruction method and system and electronic equipment

By generating an adversarial network and conditional segmentation guidance module, combined with anatomical prior knowledge, the Seg-XCT network was constructed, solving the problem of poor three-dimensional reconstruction of complex internal tissues and organs in the existing technology, and achieving a more efficient and accurate three-dimensional reconstruction effect.

CN120014175APending Publication Date: 2025-05-16HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510174489.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing X-ray three-dimensional reconstruction methods based on deep learning are poor in processing complex internal tissues and organs, and it is difficult to effectively integrate prior knowledge of human anatomy and spatial relationships.

Method used

The generative adversarial network is used to construct the X-ray image three-dimensional reconstruction network Seg-XCT, including generator, discriminator and projection space transformer. Through multi-angle projection loss and conditional segmentation guidance module, combined with anatomical prior knowledge, the reconstruction effect is improved.

Benefits of technology

It significantly improves the accuracy and robustness of X-ray image three-dimensional reconstruction, and can more effectively restore complex structures and organ relationships inside the human body.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014175A_ABST
    Figure CN120014175A_ABST
Patent Text Reader

Abstract

The invention discloses an X-ray image three-dimensional reconstruction method and system and electronic equipment, and the method comprises the steps: constructing and training an X-ray image three-dimensional reconstruction network Seg-XCT based on a generative adversarial network, the Seg-XCT comprising a generator, a discriminator and a projection space transformer; the generator outputs a reconstructed CT according to the input double-view-angle X-ray image; the projection space transformer generates a multi-angle reconstruction CT projection image and a multi-angle real CT projection image according to the input reconstruction CT and the real CT, and the projection loss of the reconstruction CT projection image and the projection loss of the real CT projection image are calculated from multiple angles; the discriminator takes the double-view-angle X-ray image as a priori condition, and the reconstruction loss between the reconstructed CT and the real patient CT is calculated according to the input reconstructed CT and the real CT; and utilizing the trained X-ray image three-dimensional reconstruction network Seg-XCT to realize X-ray image three-dimensional reconstruction. According to the method, the authenticity and accuracy of the reconstructed image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular relates to a method, system and electronic equipment for three-dimensional reconstruction of X-ray images. Background Art

[0002] In modern clinical diagnosis, medical imaging technologies, such as computed tomography (CT), magnetic resonance imaging (MRI), and X-ray imaging, are widely used to evaluate the patient's internal conditions. These technologies have their own advantages, but they also have certain limitations. CT imaging can provide detailed three-dimensional images of internal structures and is an important tool for diagnosing lesions, tumors and other diseases. However, the high-dose radiation used in the CT imaging process poses a risk to the patient's health, especially for patients who need multiple examinations. MRI, as a radiation-free imaging method, has better safety and is particularly suitable for soft tissue imaging. However, MRI equipment is expensive, the examination time is long, and it is not suitable for patients with metal implants, which limits its popularization and application. At the same time, CT and MRI examinations usually require patients to remain in a lying position, which is not completely consistent with clinical diagnostic needs, and thus affects the diagnosis of diseases related to certain body positions.

[0003] Compared with CT and MRI, X-ray imaging has the advantages of low cost, low radiation dose and suitability for standing patients, making it one of the most commonly used imaging examination methods in clinical practice. Doctors obtain frontal and side X-ray images to make preliminary diagnoses of common fractures, lung diseases, etc. However, X-ray imaging only provides two-dimensional projections of three-dimensional objects and cannot present comprehensive structural information inside the human body. This limitation results in X-ray images being unable to effectively display complex organ and tissue structures. Especially when dealing with complex fractures or visceral injuries, simple two-dimensional images may not be able to accurately assess the nature and extent of the injury, which in turn affects clinical decision-making and the choice of treatment options.

[0004] In order to solve the limitations of 3D reconstruction of X-ray images, in recent years, many researchers have proposed more efficient 3D reconstruction methods by combining deep learning technology. Deep learning, especially convolutional neural networks (CNN), has made significant breakthroughs in the field of image processing, and is particularly good at automatically extracting complex features from a large number of images. Compared with traditional methods that rely on manually extracted features and prior knowledge, deep learning can achieve efficient 3D reconstruction through training with a large amount of data. Deep learning models can not only automatically identify subtle differences and complex patterns in images, but also process large amounts of image data in a shorter time, thereby improving the accuracy and efficiency of 3D reconstruction of X-ray images. For example, when processing X-ray images, the CNN-based reconstruction method not only improves the spatial resolution of the image, but also can accurately capture the subtle differences between organs, significantly improving the quality of 3D reconstruction.

[0005] However, most of the current deep learning-based 3D reconstruction methods rely on convolutional neural networks (CNNs) for modeling. These methods show good performance in most scenarios, but are often limited to pure data-driven. Therefore, when dealing with complex internal tissues and organs, the reconstruction effects of these methods are often poor. Since the anatomical structure and spatial relationship of internal organs and tissues of the human body are highly similar, it is difficult to achieve good results by simply relying on codecs for reconstruction. Summary of the invention

[0006] In view of the above problems, the present invention provides a method, system and electronic equipment for three-dimensional reconstruction of X-ray images, aiming to improve the three-dimensional reconstruction effect of X-ray images.

[0007] According to a first aspect of an embodiment of the present disclosure, a method for three-dimensional reconstruction of an X-ray image is provided, the method comprising the following steps:

[0008] Based on the generative adversarial network, an X-ray image 3D reconstruction network Seg-XCT is constructed and trained. Seg-XCT includes a generator, a discriminator, and a projection space transformer.

[0009] The generator outputs a reconstructed CT according to the input dual-view X-ray image;

[0010] The projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and real CT, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles;

[0011] The discriminator uses the dual-view X-ray image as a priori condition and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT;

[0012] The trained X-ray image 3D reconstruction network Seg-XCT is used to realize 3D reconstruction of X-ray images.

[0013] A further technical solution of the present invention is: the projection space transformer projects the reconstructed CT from multiple angles and compares it with the projection of the real CT, and performs supervision by calculating the multi-angle projection loss to improve the authenticity and accuracy of the reconstructed image.

[0014] A further technical solution of the present invention is: the generator includes a dual encoder-decoder network corresponding to the front view X-ray image and the side view X-ray image, wherein the encoder corresponding to the front view X-ray image also includes a conditional segmentation guidance module for segmenting the front view X-ray image, and the segmentation map output by the conditional segmentation guidance module is used as prior knowledge to guide three-dimensional reconstruction.

[0015] A further technical solution of the present invention is: the generator also includes a multi-scale fusion module, and the multi-scale fusion module includes a decoding main branch for integrating information from the dual encoder-decoder and outputting the generated multi-scale feature map.

[0016] A further technical solution of the present invention is: a densely connected module is used in a dual encoder to generate feature representation, and the densely connected module includes a downsampling module, a densely connected convolution module, and a compression module that reduces the output channel by half.

[0017] A further technical solution of the present invention is: the conditional segmentation guidance module uses a pre-trained segmentation network UNet++ to segment the input front view X-ray image, identify key features, and fuse the key features with the front view X-ray image to guide three-dimensional reconstruction.

[0018] A further technical solution of the present invention is: using the segmentation map output by the conditional segmentation guidance module as prior knowledge to guide 3D reconstruction, specifically including:

[0019] Generate condition information using the segmentation map, and generate two adjustment parameters based on the condition information through a mapping function;

[0020] The input feature map is scaled and translated under the action of two adjustment parameters, which is implemented through element-level addition and multiplication;

[0021] The conditional information is passed to each 2D encoder as a shared intermediate variable so that the encoders can share the same adjustment parameters, thereby effectively adjusting the feature map.

[0022] A further technical solution of the present invention is: the multi-scale fusion module fuses and connects the features extracted from the two perspectives by the dual encoder-decoder network, calculates the dual-perspective average feature map, and transmits the average feature map back to the two decoder branches; the average feature map is connected to the feature map of the main branch of the previous round of decoding, and then the fused final feature map is obtained through the upsampling module; the final feature map obtained in each layer is multi-scale normalized and superimposed to achieve multi-scale rendering from coarse-grained to fine-textured.

[0023] A further technical solution of the present invention is: the discriminator is constructed based on a full convolution form, which is used to convert the input image into a matrix, and each value in the matrix represents the probability that the corresponding area in the original image is a real image.

[0024] According to a second aspect of an embodiment of the present disclosure, a system for three-dimensional reconstruction of an X-ray image is provided, the system comprising:

[0025] Generator, Discriminator, and Projection Space Transformer;

[0026] The generator outputs a reconstructed CT according to the input dual-view X-ray image;

[0027] The projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and real CT, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles;

[0028] The discriminator takes the dual-view X-ray image as a priori condition and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT.

[0029] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned X-ray image three-dimensional reconstruction method when executing the program.

[0030] The disclosed embodiments provide a method, system and electronic device for three-dimensional reconstruction of X-ray images, which are mainly used in the field of medical image processing, especially in clinical image diagnosis. The present invention constructs and trains a three-dimensional reconstruction network Seg-XCT of X-ray images based on a generative adversarial network. Seg-XCT includes a generator, a discriminator and a projection space transformer. The generator outputs a reconstructed CT according to the input dual-view X-ray image. The projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and real CT, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles. The beneficial effects of the present invention include:

[0031] The discriminator uses the dual-view X-ray image as a priori condition and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT. The generator part incorporates a multi-scale fusion module to enhance its ability to process complex image data, and uses a 3D coordinate convolution layer to replace the traditional convolution layer.

[0032] In order to effectively supervise the learning process of the model, a multi-angle projection loss function based on ProjectiveSpatial Transformer (PST) is introduced. This method not only helps to combine prior knowledge in the field of anatomy, but also enables the model to use large training data sets to accurately learn the mapping relationship from X-ray images to CT images.

[0033] The Projection Space Transformer uses projection space transformation technology to project the reconstructed CT from multiple angles and compare it with the projection of the real CT. It supervises by calculating the multi-angle projection loss, thereby further improving the authenticity and accuracy of the reconstructed image.

[0034] The anteroposterior and lateral chest X-ray images are also provided as priors to the discriminator, which enhances the information available to the discriminator and improves its performance. This in turn facilitates the adversarial training process of the generator, thereby improving the image quality of the reconstructed CT.

[0035] The conditional segmentation guidance module of the present invention aims to introduce prior knowledge constraints and optimize the reconstruction process. In addition to relying on the data itself, it also combines a deep understanding of human anatomy and uses pre-segmented key tissue and organ information to provide strong guidance for the 3D reconstruction process. This prior knowledge not only includes the precise location and morphology of tissues and organs, but also reflects the complex spatial relationship between them, thereby significantly improving the accuracy and robustness of the reconstruction algorithm.

[0036] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the description, serve to explain the principles of the present invention.

[0038] Figure 1 2 is a schematic diagram of the structure of the Seg-XCT network for three-dimensional reconstruction of X-ray images in an embodiment of the present invention;

[0039] Figure 2 is a schematic diagram of a generator network structure in an embodiment of the present invention;

[0040] Figure 3 is a schematic diagram of a conditional segmentation guidance module method according to an embodiment of the present invention;

[0041] Figure 4 is a schematic diagram of the structure of a conditional segmentation guidance module in an embodiment of the present invention;

[0042] Figure 5 is a schematic diagram of the structure of a multi-scale fusion module in an embodiment of the present invention;

[0043] Figure 6 is a schematic diagram of densely connected modules in an embodiment of the present invention;

[0044] Figure 7 is a schematic diagram of the principle of a 3D decoder in an embodiment of the present invention;

[0045] Figure 8 It is a schematic diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0047] It should be mentioned before discussing the exemplary embodiments in more detail that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0048] The embodiments of the present invention are mainly applied to the field of medical image processing, especially in clinical image diagnosis, especially in three-dimensional reconstruction technology based on X-ray images. Specific applications include but are not limited to the following aspects:

[0049] 1. 3D CT reconstruction: By using only two X-ray images (such as front and side X-rays), the present invention can efficiently reconstruct 3D CT images, thereby reducing radiation exposure in traditional CT scans. This technology is particularly suitable for clinical diagnostic scenarios that require fast, low-radiation imaging, such as fracture detection, joint lesion assessment, and preliminary screening of lung diseases.

[0050] 2. Auxiliary diagnostic tool: In medical imaging, the present invention can be used as a tool to assist doctors in decision-making, especially in emergency treatment or resource-limited environments, to provide faster three-dimensional image reconstruction support and improve diagnostic efficiency and accuracy.

[0051] 3. Telemedicine and mobile devices: The present invention can be widely used in telemedicine and mobile medical equipment, especially in remote areas or places where high-end CT equipment is not available. Accurate three-dimensional reconstruction can be achieved through simple X-ray images, thereby improving the diagnostic capabilities of grassroots hospitals.

[0052] 4. Innovation of medical imaging equipment: The present invention can also be integrated into existing X-ray or portable medical imaging equipment as a new diagnostic function extension to help doctors achieve more comprehensive image analysis.

[0053] 5. Potential applications: In addition to medical imaging, the present invention can also be extended to other fields, such as internal defect detection in material science, industrial non-destructive testing, and structural analysis in security monitoring.

[0054] In summary, the technology of the present invention not only has broad application prospects in the field of medical imaging, but can also provide innovative solutions for many other industries and has good market potential.

[0055] The present invention provides a method, system and electronic device for three-dimensional reconstruction of X-ray images, and provides the following embodiments:

[0056] A three-dimensional reconstruction method for X-ray images comprises the following steps: constructing and training an X-ray image three-dimensional reconstruction network Seg-XCT based on a generative adversarial network, wherein the Seg-XCT comprises a generator, a discriminator and a projection space transformer; wherein the generator outputs a reconstructed CT according to an input dual-view X-ray image; the projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and real CT, respectively, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles; the discriminator uses the dual-view X-ray image as a priori condition, and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT; and the trained X-ray image three-dimensional reconstruction network Seg-XCT is used to realize the three-dimensional reconstruction of the X-ray image.

[0057] Specifically, the conditional segmentation-based X-ray image 3D reconstruction network architecture constructed in the embodiment is named Seg-XCT, which is used to reconstruct 3D CT images from dual-view orthogonal X-ray films. The overall network architecture of Seg-XCT is as follows: Figure 1 As shown in the figure, the framework of generative adversarial network is adopted, in which the generator part integrates a multi-scale fusion module to enhance its ability to process complex image data, and a 3D coordinate convolution layer is used to replace the traditional convolution layer. In order to effectively supervise the learning process of the model, a multi-angle projection loss function based on Projective Spatial Transformer (PST) is introduced.

[0058] The invention not only helps to incorporate prior knowledge in the field of anatomy, but also enables the model to use large training datasets to accurately learn the mapping relationship from X-ray images to CT images. Extensive experimental results on large-scale datasets show that Seg-XCT exhibits excellent performance.

[0059] In addition, the present invention designs a chest segmentation framework called Seg-Block, which combines traditional image processing and deep learning models to segment the lungs and spine, guiding the generator to extract features and reconstruct three-dimensional CT.

[0060] The Projection Space Transformer projects the reconstructed CT from multiple angles and compares it with the projection of the real CT. Supervision is performed by calculating the multi-angle projection loss to improve the authenticity and accuracy of the reconstructed image.

[0061] like Figure 2 As shown, the generator includes a dual encoder-decoder network corresponding to the front view X-ray image and the side view X-ray image, wherein the encoder corresponding to the front view X-ray image also includes a conditional segmentation guidance module for segmenting the front view X-ray image, and the segmentation map output by the conditional segmentation guidance module is used as prior knowledge to guide three-dimensional reconstruction.

[0062] Seg-XCT follows the adversarial network paradigm with an additional digitally reconstructed radiography (DRR) branch. The dual-view X-ray images are simultaneously input into the Seg-XCT generator as prior conditions to obtain the reconstructed CT, and then the reconstruction loss between the reconstructed CT and the real patient CT is calculated. At the same time, the reconstructed CT is projected from multiple angles using the projection space transformation technique and compared with the projection of the real CT. Supervision is performed by calculating the multi-angle projection loss, thereby further improving the authenticity and accuracy of the reconstructed image. In addition, the frontal and lateral chest X-ray images are also provided to the discriminator as prior conditions, which enhances the information available to the discriminator and improves the performance of the discriminator. This in turn promotes the adversarial training process of the generator, thereby improving the image quality of the reconstructed CT.

[0063] like Figure 2 As shown, the generator also includes a multi-scale fusion module, which includes a decoding main branch for integrating information from the dual encoder-decoder and outputting a generated multi-scale feature map.

[0064] like Figure 2 As shown, a densely connected module is used in the dual encoder to generate feature representations, and the densely connected module includes a downsampling module, a densely connected convolution module, and a compression module that reduces the output channels by half.

[0065] At present, most deep learning-based 3D reconstruction methods rely on convolutional neural networks (CNNs) for modeling. These methods show good performance in most scenarios, but are often limited to pure data-driven, lack a deep understanding of the human anatomical structure and spatial relationships, and are difficult to effectively incorporate prior knowledge. Therefore, when dealing with complex internal tissues and organs, the reconstruction effects of these methods are often poor. Since the anatomical structure and spatial relationships of internal organs and tissues in the human body are highly similar, it is difficult to achieve good results by simply relying on codecs for reconstruction.

[0066] To solve this problem, the present invention proposes a conditional segmentation guidance module, which aims to constrain and optimize the reconstruction process by introducing prior knowledge. The core idea of ​​this module is that, in addition to relying on the data itself, it also combines a deep understanding of human anatomy and uses pre-segmented key tissue and organ information to provide strong guidance for the 3D reconstruction process. This prior knowledge not only includes the precise location and morphology of tissues and organs, but also reflects the complex spatial relationships between them, thereby significantly improving the accuracy and robustness of the reconstruction algorithm.

[0067] The conditional segmentation guidance module uses the pre-trained segmentation network UNet++ to segment the input front view X-ray image, identify key features, and fuse the key features with the front view X-ray image to guide 3D reconstruction.

[0068] like Figure 3 As shown in the figure, the conditional segmentation guidance module first uses the pre-trained segmentation network UNet++ to finely segment the input medical image and identify key tissues and organs such as the lungs and spine. Subsequently, the segmentation results are introduced into the deep learning model as conditional information and combined with the original image data through a specific fusion strategy to jointly guide the 3D reconstruction process.

[0069] The segmentation map output by the conditional segmentation guidance module is used as prior knowledge to guide 3D reconstruction, including:

[0070] Generate condition information using the segmentation map, and generate two adjustment parameters based on the condition information through a mapping function;

[0071] The input feature map is scaled and translated under the action of two adjustment parameters, which is implemented through element-level addition and multiplication;

[0072] The conditional information is passed to each 2D encoder as a shared intermediate variable so that the encoders can share the same adjustment parameters, thereby effectively adjusting the feature map.

[0073] Specifically, Figure 2 As shown in the figure, the generator architecture in Seg-XCT includes two identical encoder-decoder components, which process image data acquired from different perspectives respectively. The front view includes a conditional segmentation guidance module (Seg-Block), and a decoding main branch is added between the decoding branches of the two views to generate the final reconstructed CT. Specifically, the dual encoder-decoder processes the front and side of the chest X-ray film in different paths. Since a single front chest X-ray image cannot capture the side information of the object, and only the side chest X-ray image will lose the front information, Seg-XCT adopts a dual orthogonal view input strategy, that is, inputting the front and side chest X-ray images at the same time to overcome this limitation.

[0074] The dual encoder-decoder network extracts features from the images of these two perspectives and then passes these features to the subsequent fusion and decoding modules. Human 3D CT has strong prior knowledge. Each organ has its own distinct anatomical structure and texture information. Different tissues and organs should be processed in different ways. A conditional segmentation module is built on the main view encoder network to guide the encoder to better extract features. The multi-scale fusion module consists of a decoding main branch that integrates information from dual orthogonal viewpoints and an output branch that generates feature maps covering multiple spatial levels.

[0075] Seg-XCT uses densely connected modules as its basic unit in the encoder. These densely connected modules are composed of three parts: a downsampling module with a stride of 2, a densely connected convolution block, and a compression block that reduces the output channels by half, in order to generate feature representations that ensure that key information is retained at different spatial levels from the input image.

[0076] like Figure 4 As shown in Figure 2, the conditional segmentation guidance module generates two adjustment parameter pairs (α, β) through a mapping function to adjust the input feature map. Specifically, the input feature map (F n ) performs scaling and translation operations under the action of these two parameters, and the transformation process is as follows: (α, β) = M(e), where M(e) is a mapping function, the input is conditional information (e), and the output is a pair of adjustment parameters (α, β) used to control the transformation of the feature map.

[0077] Next, the obtained parameters α and β are applied to the transformation of the feature map, which is implemented by element-level addition and element-level multiplication, as shown below: F = F n α+β, where F n is the input feature map, and the transformed feature map is F. The conditional information (e) comes from the input segmentation map and is passed to each 2D encoder as a shared intermediate variable. These conditional information enable the model to share the same adjustment parameters between multiple encoders, thereby effectively adjusting the feature map. Specifically, the conditional information is generated based on the input segmentation map and shared among the 2D encoders of the model. In this way, the conditional guidance module can obtain a small number of adjustment parameters and optimize the mapping function M(e) through end-to-end training to generate suitable adjustment parameters and improve model performance. In this process, the model not only learns the feature representation of the image data, but is also constrained by the segmentation results, which enables the reconstructed three-dimensional image to not only maintain the consistency of the overall structure, but also finely restore the subtle structure of tissues and organs.

[0078] The multi-scale fusion module fuses and connects the features extracted from two perspectives by the dual encoder-decoder network, calculates the average feature map of the two perspectives, and transmits the average feature map back to the two decoder branches; the average feature map is connected with the feature map of the main branch of the previous round of decoding, and then the final fused feature map is obtained through the upsampling module; the final feature map obtained in each layer is normalized and superimposed at multiple scales to achieve multi-scale rendering from coarse-grained to fine-textured.

[0079] like Figure 5 As shown in the figure, the feature transformation of the dual perspectives, the dimensional order of the three-dimensional feature maps generated by different perspectives is inconsistent. For example, the front of the feature map generated by the front perspective is the actual front, while the front of the feature map generated by the side perspective is the actual side. Therefore, the embodiment designs a multi-scale fusion module (MFusion) to standardize the consistency of the dimensions. The features extracted from the two perspectives by the dual encoder-decoder network are then fused and connected in the decoding module. Figure 5 As shown in the figure, after calculating the average feature map of the two views, it is passed back to the two decoder branches to promote the fusion of information from the dual orthogonal views. The average feature map is then connected with the feature map of the main branch of the previous round of decoding, and then the fused final feature map is obtained through the upsampling module. At the same time, the final feature map obtained at each layer is multi-scale normalized and superimposed, thereby achieving multi-scale rendering from coarse-grained to fine-textured.

[0080] Existing methods still rely on traditional two-dimensional convolutional layers to extract image features, but this approach often performs poorly in complex tasks. Dense connections not only effectively alleviate the gradient vanishing problem in deep neural networks, but also significantly promote the reuse of multi-level features and information transfer, thereby improving the network's expressive power and computational efficiency. Figure 6 As shown in the figure, in a dense block, each layer not only receives the feature output of the previous layer, but also concatenates the feature maps of all previous layers to form the input of the current layer. The mathematical expression is:

[0081] X l =H([X0,X1,…,X2])

[0082] Among them, [X0,X1,…,X2] represents the concatenation of feature maps from layer 0 to layer l-1, and H(·) is a composite operation composed of batch normalization, ReLU activation function and 3x3 convolution. Through this design, the network not only improves the computational efficiency, but also enhances the expressiveness of features and the robustness of the network.

[0083] In order to better extract key information from two-dimensional X-ray images, the embodiment introduces a dense connection module in the encoding stage of the generator. Figure 2As shown in the figure, each module contains a downsampling block (downsampling by convolution with a stride of 2), a dense block, and a channel compression block (reducing the computational complexity by reducing the number of channels). These cascaded densely connected components can perform multi-level feature encoding on the input data, allowing the network to extract and fuse multiple feature information at different levels. By cascading these modules, the model can achieve more efficient feature encoding. Ultimately, the encoded information is not only passed to the subsequent processing modules of the network, but also directly passed to the corresponding level of the decoder through jump connections, realizing multi-level utilization and sharing of features. This design plays a key role in improving image reconstruction accuracy and maintaining image structural consistency.

[0084] The discriminator is built on a fully convolutional form, which is used to convert the input image into a matrix, where each value in the matrix represents the probability that the corresponding area in the original image is a real image.

[0085] Unlike traditional generative adversarial networks (GANs), Seg-XCT uses a fully convolutional form to construct a discriminator. In traditional GANs, the discriminator outputs a single value indicating the possibility that the input sample is a real sample. However, in the fully convolutional discriminator network of this embodiment, the input image is converted into a matrix. Each value in the matrix represents the probability that the corresponding area in the original image is a real image. This scheme has been proven to be effective, and the method has shown strong generalization capabilities in the field of high-resolution and high-definition image generation. In addition, the translation invariance provided by traditional convolution helps to learn more robust features, especially in tasks such as image classification. However, in the case of relatively static human body structures, the translation invariance of traditional convolution may limit its ability to capture position information. To solve this problem, coordinate convolution (Coordinate Convolution, CoordConv) is introduced. CoordConv integrates coordinate information as part of the feature map, allowing the network to maintain a certain degree of translation dependence during the learning process of the task. In the context of CT reconstruction, the position of each organ in the chest image is relatively fixed. Allowing the network to retain the position information of each organ can improve the performance of the model.

[0086] Seg-XCT expands the coordinate convolution layer into 3D format and replaces the traditional convolution layer in the original discriminator network. In addition, the discriminator of Seg-XCT is also built in 3D format. First, feature extraction and position information fusion are performed through the 3D coordinate convolution layer. Then, the feature map is passed through three cascaded convolution downsampling modules. Each module contains a 3D convolution layer with a stride of 2, a normalization layer, and a rectified linear unit. Finally, the feature map is compressed through a complex network structure composed of multiple up-sampling convolution blocks (Up-Conv Block), fully connected layers (Dense), and 2D-3D up-sampling convolution layers (2D-3D Up-Conv) to obtain the final output matrix.

[0087] like Figure 7 As shown, in the decoder stage, the present invention uses a 3D deconvolution module to reconstruct three-dimensional features. The 3D deconvolution module restores high-dimensional features to a three-dimensional image by gradually enlarging the feature map and performing a three-dimensional convolution operation.

[0088] Another embodiment is used to illustrate a system for three-dimensional reconstruction of X-ray images. Figure 1 As shown, the system includes: a generator, a discriminator and a projection space transformer; the generator outputs a reconstructed CT according to the input dual-view X-ray image; the projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and the real CT, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles; the discriminator takes the dual-view X-ray image as a priori condition, and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT.

[0089] In addition to the above modules, an X-ray image three-dimensional reconstruction system may also include other components. However, since these components are irrelevant to the content of the embodiments of the present disclosure, their illustration and description are omitted here.

[0090] The other specific working processes of an X-ray image three-dimensional reconstruction system refer to the description of the above-mentioned basic X-ray image three-dimensional reconstruction method embodiment and will not be repeated here.

[0091] Another embodiment is used to illustrate that the system of the present invention can also be used with the help of Figure 8 The architecture of the computing device shown is implemented. Figure 8 The architecture of the computing device is shown. Figure 8 As shown, a computer system 810, a system bus 830, one or more CPUs 840, an input / output 820, a memory 850, etc. The memory 850 can store various data or files used for computer processing and / or communication and program instructions including the embodiment X-ray image three-dimensional reconstruction method executed by the CPU. Figure 8 The architecture shown is only exemplary and can be adjusted according to actual needs when implementing different devices. Figure 8 One or more components in the memory 850. The memory 850, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the X-ray image three-dimensional reconstruction method in the embodiment of the present invention (for example, the generator, discriminator and projection space transformer in the X-ray image three-dimensional reconstruction system). One or more CPUs 840 execute various functional applications and data processing of the system of the present invention by running the software programs, instructions and modules stored in the memory 850, that is, to implement the above-mentioned X-ray image three-dimensional reconstruction method, which includes the following steps:

[0092] Based on the generative adversarial network, an X-ray image 3D reconstruction network Seg-XCT is constructed and trained. Seg-XCT includes a generator, a discriminator, and a projection space transformer.

[0093] The generator outputs a reconstructed CT according to the input dual-view X-ray image;

[0094] The projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and real CT, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles;

[0095] The discriminator uses the dual-view X-ray image as a priori condition and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT;

[0096] The trained X-ray image 3D reconstruction network Seg-XCT is used to realize 3D reconstruction of X-ray images.

[0097] Of course, the processor of the server provided in the embodiment of the present invention is not limited to executing the method operations described above, but can also execute related operations in the X-ray image three-dimensional reconstruction method provided in any embodiment of the present invention.

[0098] The memory 850 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 850 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 850 may further include a memory remotely arranged relative to one or more CPUs 840, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0099] The input / output 820 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The input / output 820 may also include a display device such as a display screen.

[0100] In this document, the terms "comprises," "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, such that a step or method that includes a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such step or method.

[0101] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A method for three-dimensional reconstruction of X-ray images, characterized in that: The method comprises the following steps: Based on the generative adversarial network, an X-ray image 3D reconstruction network Seg-XCT is constructed and trained. Seg-XCT includes a generator, a discriminator, and a projection space transformer. The generator outputs a reconstructed CT according to the input dual-view X-ray image; The projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and real CT, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles; The discriminator uses the dual-view X-ray image as a priori condition and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT; The trained X-ray image 3D reconstruction network Seg-XCT is used to realize 3D reconstruction of X-ray images.

2. The X-ray image three-dimensional reconstruction method according to claim 1, characterized in that: The generator includes a dual encoder-decoder network corresponding to the front view X-ray image and the side view X-ray image, wherein the encoder corresponding to the front view X-ray image also includes a conditional segmentation guidance module for segmenting the front view X-ray image, and the segmentation map output by the conditional segmentation guidance module is used as prior knowledge to guide three-dimensional reconstruction.

3. The X-ray image three-dimensional reconstruction method according to claim 1, characterized in that: The generator also includes a multi-scale fusion module, which includes a decoding main branch for integrating information from the dual encoder-decoder and outputting a generated multi-scale feature map.

4. The X-ray image three-dimensional reconstruction method according to claim 2, characterized in that: The dual encoder uses a densely connected module to generate feature representations, and the densely connected module includes a downsampling module, a densely connected convolution module, and a compression module that reduces the output channels by half.

5. The X-ray image three-dimensional reconstruction method according to claim 2, characterized in that: The conditional segmentation guidance module uses the pre-trained segmentation network UNet++ to segment the input front view X-ray image, identify key features, and fuse the key features with the front view X-ray image to guide three-dimensional reconstruction.

6. The X-ray image three-dimensional reconstruction method according to claim 2, characterized in that: The segmentation map output by the conditional segmentation guidance module is used as prior knowledge to guide 3D reconstruction, including: Generate condition information using the segmentation map, and generate two adjustment parameters based on the condition information through a mapping function; The input feature map is scaled and translated under the action of two adjustment parameters, which is implemented through element-level addition and multiplication; The conditional information is passed to each 2D encoder as a shared intermediate variable so that the encoders can share the same adjustment parameters, thereby effectively adjusting the feature map.

7. The X-ray image three-dimensional reconstruction method according to claim 3, characterized in that: The multi-scale fusion module fuses and connects the features extracted from the two perspectives by the dual encoder-decoder network, calculates the average feature map of the two perspectives, and transmits the average feature map back to the two decoder branches; The average feature map is connected with the feature map of the main branch of the previous decoding round, and then the fused final feature map is obtained through the upsampling module; The final feature map obtained at each layer is normalized and superimposed at multiple scales to achieve multi-scale rendering from coarse-grained to fine-textured.

8. The X-ray image three-dimensional reconstruction method according to claim 1, characterized in that: The discriminator is constructed based on a full convolutional form and is used to convert the input image into a matrix, in which each value represents the probability that the corresponding area in the original image is a real image.

9. A three-dimensional reconstruction system for X-ray images, characterized in that: The system comprises: Generator, Discriminator, and Projection Space Transformer; The generator outputs a reconstructed CT according to the input dual-view X-ray image; The projection space transformer generates multi-angle reconstructed CT projection imaging and multi-angle real CT projection imaging according to the input reconstructed CT and real CT, and calculates the projection loss of the reconstructed CT projection imaging and the real CT projection imaging from multiple angles; The discriminator takes the dual-view X-ray image as a priori condition and calculates the reconstruction loss between the reconstructed CT and the real patient CT according to the input reconstructed CT and the real CT.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the X-ray image three-dimensional reconstruction method as claimed in any one of claims 1 to 8 are implemented.