Method for reconstructing three-dimensional image from x-ray images by using artificial intelligence model, computing device therefor, and recording medium therefor

An AI-based diffusion model for X-ray image processing addresses the challenges of 3D reconstruction by generating high-quality 3D images from multiple angles, enhancing accuracy and efficiency, suitable for real-time applications in diverse fields.

WO2026019042A1PCT designated stage Publication Date: 2026-01-22KOREA INST OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/006861
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-17
Filing Date
2025-05-21
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing 3D shape reconstruction from X-ray data is time-consuming, costly, prone to artifacts, and limited by complex data processing, reducing its real-time applications and accuracy.

Method used

An artificial intelligence model based on a diffusion model is used to generate X-ray images from various angles by preprocessing, encoding, transforming latent vectors, and generating 3D images, with fine-tuning to minimize latent vector differences and adapt to noise levels.

Benefits of technology

This approach enables efficient, high-quality 3D image reconstruction with improved accuracy and robustness, overcoming limitations of conventional methods and enabling real-time applications in fields like medicine, security, and industrial inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025006861_22012026_PF_FP_ABST
    Figure KR2025006861_22012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for reconstructing a three-dimensional image from X-ray images, according to an embodiment, comprises, by using an artificial neural network model implemented on the basis of a diffusion model: a first step of preprocessing an image to extract an object of interest from an input X-ray image; a second step of encoding the image, preprocessed in the first step, into a latent vector; a third step of transforming, on the basis of transformation information, the latent vector encoded in the second step, to reconstruct same into a plurality of transformed X-ray images; and a fourth step of generating a 3D image on the basis of the plurality of transformed X-ray images reconstructed in the third step, wherein the artificial neural network model is trained in the direction of minimizing differences between latent vectors of the reconstructed images.
Need to check novelty before this filing date? Find Prior Art

Description

Method for restoring a three-dimensional image from an X-RAY image using an artificial intelligence model, a computing device therefor, and a recording medium thereof

[0001] The present invention relates to a technology for restoring a consistent three-dimensional image from an X-Ray image using an artificial intelligence model.

[0002]

[0003] The fog removal technology is based on the Atmospheric Scattering Model, and the present invention relates to a three-dimensional shape restoration technology using two-dimensional X-ray images. This restoration technology is a technology area whose importance is increasing in various fields such as the military, modern medicine, security, and industrial inspection.

[0004] The recent rapid advancement of deep learning technology, particularly generative models based on diffusion models, is opening up new possibilities in image generation and processing. Diffusion models can identify and reproduce the subtle features and distributions of data, enabling the generation of high-quality images. Furthermore, these models offer great flexibility for application to a variety of data types and problems, making them highly applicable to X-ray image processing.

[0005] However, 3D shape reconstruction technology using X-ray data still faces several challenges. The biggest problem is that acquiring X-ray data from all angles is extremely time-consuming and cost-intensive. Furthermore, generating high-quality 3D information from limited data requires sophisticated algorithms and complex calculations.

[0006] Previously proposed methods are prone to artifacts (distortion issues) during the reconstruction process and suffer from reduced image resolution. Furthermore, they involve complex data processing and interpretation processes, limiting their real-time applications. Furthermore, consistent 3D shape reconstruction from multiple angles is difficult, limiting accurate 3D structural identification.

[0007]

[0008] The purpose of the present invention is to propose a technology that uses an artificial intelligence model to generate X-ray images from various angles based on a limited set of X-ray images, and thereby identify and restore three-dimensional information. Unlike existing technologies, this technology focuses on reproducing three-dimensional structures, suggesting potential applications in a wider range of fields (e.g., medical, security, industrial inspection, military, etc.).

[0009]

[0010] In order to solve the above technical problem, a method for restoring a 3D image from an X-Ray image of an embodiment includes a first step of preprocessing an image to extract an object of interest from an input X-Ray image in an artificial neural network model implemented based on a Diffusion model, a second step of encoding an image preprocessed in the first step into a latent vector, a third step of transforming the latent vector encoded in the second step based on deformation information to restore a plurality of deformed X-Ray images, and a fourth step of generating a 3D image based on the plurality of deformed X-Ray images restored in the third step, wherein the artificial neural network model is trained in a direction of minimizing a difference in latent vectors between restored images.

[0011] The above artificial neural network model can be pre-trained by repeating the process in which a latent encoder receives an X-ray image with added noise and generates a latent vector with added noise, an image restoration module receives the latent vector with added noise and the transformation information, removes noise from the latent vector with added noise, generates a latent vector according to the transformation information, and a latent decoder generates a final image from the latent vector generated by the image restoration module.

[0012] As the above pre-learning progresses, the level of the above noise may gradually increase.

[0013] In the above artificial neural network model, the image restoration module (Gx) can be fine-tuned to minimize the pixel difference between a first image restored by recursively repeating image restoration based on deformation information and a second image restored by calculating the deformation information by the number of recursive iterations.

[0014] The above transformation information includes at least one of rotation, position transformation, and size transformation of the image.

[0015] Another embodiment of the present invention also includes a computing device and a recording medium for implementing the above-described method.

[0016]

[0017] In this embodiment, the system utilizes cutting-edge AI techniques, such as the Diffusion model, to generate consistent 3D images, while an additional learning process using specific loss functions enables more accurate 3D information extraction. This innovative approach overcomes the limitations of conventional techniques and opens up new possibilities in the field of X-ray-based 3D information extraction and visualization.

[0018]

[0019] Figure 1 is a diagram illustrating a general diffusion model.

[0020] Figure 2 schematically shows the overall network configuration.

[0021] Figure 3 schematically shows the learning process with added noise.

[0022] Figure 4 schematically shows the fine tuning process of the model.

[0023] Figure 5 is a flowchart showing a method for restoring a three-dimensional image from an X-ray image of an embodiment.

[0024] FIG. 6 is a block diagram illustrating a computational device that executes a method for restoring a three-dimensional image from an X-ray image of an embodiment.

[0025]

[0026] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, detailed descriptions of well-known functions or components that may obscure the gist of the present invention will be omitted in the following description and the attached drawings. Additionally, throughout the specification, the term "including" a component does not exclude other components, unless specifically stated otherwise, but rather implies the inclusion of other components.

[0027] Additionally, while terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms may be used to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.

[0028] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0029] Unless specifically defined otherwise, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning within the context of the relevant technology, and shall not be construed in an idealized or overly formal sense unless explicitly defined herein.

[0030]

[0031] The artificial intelligence model used in the embodiment can be implemented based on a diffusion model, and Fig. 1 is a diagram illustrating a general diffusion model.

[0032] The diffusion model is a generative model that learns a forward process that gradually adds noise to data and a reverse process that removes it.

[0033] This diffusion model is based on the U-Net architecture and can be configured to include an encoder and decoder. The terms shown in Figure 1 are explained in Table 1.

[0034] Terminology Input Data (X) X: Original pixel image data. Used as input to the model and located in the pixel space. Encoder (E) E: An encoder that converts the original image into a latent vector. This encoder converts the input image into a compressed latent space. Latent Space (Z) Z: A latent vector generated by the encoder. This vector contains image information in a compressed form. Diffusion Process The diffusion process refers to the process of gradually restoring an image to which noise has been added. In the embodiment, a latent vector to which noise has been added is used. Denoising UNET ( ): : It plays a role in restoring the original image by removing noise during the diffusion process. This U-Net structure receives a latent vector with added noise as input and removes noise step by step. Query, Key, Value (Q, K, V): Q (Query), K (Key), V (Value): Elements of the attention mechanism used in the diffusion process. Q: Query vector, representing information of the current step. K: Key vector, containing information of all steps. V: Value vector, representing the value associated with the key. Through the attention mechanism, the denoising U-Net helps produce the desired form of result by focusing on important information and removing noise.

[0035] This diffusion model adds Gaussian noise to the original data during the diffusion process, creating a completely noise-free state through multiple stages. The de-diffusion process gradually removes noise, starting from the noise state. At each stage of the de-diffusion process, a small amount of noise is removed, approaching the original data distribution. The diffusion model learns which noise to remove at each noise removal step through a learning process, enabling more robust image restoration.

[0036] This diffusion model operates in latent space, which improves computational efficiency and has strengths specialized for text-conditional image generation.

[0037] In the embodiment, a diffusion process is applied to a latent representation of an X-ray image through such a diffusion model, and X-ray images of various angles are generated.

[0038] In the embodiment, an artificial intelligence model is implemented based on a diffusion model, so that high-quality images rich in details can be generated, diversity and consistency can be achieved simultaneously, and conditional generation (e.g., generation of images at a specific angle) is possible.

[0039]

[0040] Hereinafter, an artificial intelligence model of an embodiment implemented based on the diffusion model is described through Figs. 2 and 3. Fig. 2 schematically shows the overall network configuration, and Fig. 3 schematically shows the learning process with added noise.

[0041] In explaining the artificial intelligence model of the example, the necessary terms are summarized in Table 2.

[0042] Term Description Input Data ( )X": X-ray image data with added noise. Used as input to the model. Latent vector encoding ( ) : This is the process of converting an image with added noise into a latent vector. In this process, is the latent vector Encoded with image restoration module ( ) : A module that receives an encoded latent vector as input and restores the image by reflecting the object's deformation information (R, T, S). The restored image ( ) : Encoded latent vector , and generate the first restored image X(1). Iterative image restoration( ): Recursively encode and restore the first restored image X(1) to generate the image X(n). This verifies and improves the consistency of the model. Object transformation information (R, T, S) R: Rotation information. Reflects the rotation information of the image during model learning. T: Translation information. Reflects the position transformation information of the image during model learning. S: Scaling information. Reflects the size transformation information of the image during model learning.

[0043] In an embodiment, the encoder (Ex) converts the input X-ray image (X) into a latent vector (latent vector) in the latent space ( ) to compress high-dimensional image data into low-dimensional vectors, thereby improving computational efficiency. The latent vector encoded in this way ( ) is given as an input to the image restoration module (Gx). The image restoration module (Gx) receives a latent vector (Z) and transformation information (R, T, S) as input, and transforms the input latent vector according to at least one of the transformation information, namely rotation, translation, and scaling, according to pre-learning, and outputs the transformed latent vector (Z). This latent vector (Z) is given as an input to the latent decoder (Dx).

[0044] The latent decoder (Dx) receives the transformed latent vector (Z) as input, converts it into an image, and outputs an X-ray image (X') with a new angle (S) and a new scale (S).

[0045] This process, in which the image restoration module (Gx) outputs a transformed latent vector (Z) based on the latent vector (Z) and transformation information (R, T, S), and the latent decoder (Dx) receives the transformed latent vector (Z) as input, converts it into an image, and outputs an X-ray image (X') at a new angle (S) and a new scale (S), is repeated to generate X-ray images at various angles (or scales).

[0046] These multi-angle X-ray images generated in this way can be easily restored into 3D using various known methods.

[0047] This process can be implemented through a visualization module, and an example of the process is as follows.

[0048] The visualization module receives multi-angle X-Ray images generated from the image generation module (a network such as that shown in Fig. 2) and sorts the generated images in order according to angle.

[0049] If necessary, the visualization module can interpolate intermediate images between the generated images to represent the successive rotations, and can render the aligned images successively to create a 3D rotation effect.

[0050] In an embodiment, through this process, 3D consistent X-Ray images (X') from various angles can be generated from a single X-Ray image (X), and these can be visualized as continuous 3D rotation images. This is a powerful tool that can be utilized in various fields such as medicine, security, industrial inspection, and military.

[0051]

[0052] The pre-training process of an artificial intelligence model that operates in this way is explained as follows.

[0053] In the pre-learning process, the training data set used can be configured to include an X-ray image dataset taken at various angles and positions and information on the shooting angle (R), scale (S), and position (T) for each image, and since the model is configured based on the Stable Diffusion architecture, it can be initialized with pre-learned Stable Diffusion weights.

[0054] During the learning process, the latent vector of the original X-ray image, target angle (R) or target scale (S) / position (T) information are given as input to the model, and the image restoration module (Gx) transforms the input latent vector to fit the target angle / position and outputs the transformed latent vector.

[0055] At this time, in the embodiment, a loss function defined as mathematical expression 1 is applied, and the potential loss ( ) is pre-trained to reduce the number of errors.

[0056]

[0057] In mathematical expression 1, = i-th element of the latent vector, = jth element of the latent vector, N = number of restored images, and this latent loss aims to minimize the difference in latent vectors between restored images.

[0058]

[0059] Additionally, in embodiments, the model can be trained using data with added noise during the pre-training process. This process is exemplified in Figure 3.

[0060] The purpose of adding noise is to improve the robustness of the model, prevent overfitting, and improve the ability to generate diverse images.

[0061] Random noise is added to the original X-ray image (X) to generate a noisy image (X''), and the level of this noise can be gradually adjusted during the learning process.

[0062] During the learning process of the model, X'' is input to the latent encoder (Ex) to generate a latent vector Ex(X″) with added noise, which is then given as input to the image restoration module (Gx).

[0063] The image restoration module (Gx) removes noise from the input noise-added latent vector Ex(X″) and deformation information (R, T, S), generates a latent vector (Gx(Ex(X''))) corresponding to the image at the target angle / scale / position, and generates the final image through the latent decoder (Dx).

[0064] The model learns to generate clean target images from noisy inputs, and in the process, the model learns to understand the essential characteristics and 3D structure of X-ray images.

[0065] In one embodiment, the noise level can start at a low level at the beginning of training and gradually increase as training progresses. This allows the model to adapt to a variety of noise levels.

[0066] As training progresses with images with added noise, it becomes more tolerant to noise that can occur in actual X-ray imaging environments, enabling the generation of more diverse and richer images.

[0067] Moreover, through this noise-adding process, the model gains the ability to understand and reconstruct the essential structure and characteristics of the image rather than simply copying the input image, ultimately enabling higher quality 3D X-ray image reconstruction.

[0068]

[0069] Figure 4 schematically shows the process of securing 3D consistency for a pre-learned model and fine-tuning it to suit the characteristics of an X-ray image. For convenience of explanation, the encoder and decoder are omitted.

[0070] In the fine tuning process, the image restoration module (Gx) is fine-tuned in a direction that minimizes the pixel difference between the first image (Gx(Gx(Ex(X''))) restored by recursively repeating image restoration based on deformation information, and the second image (2Gx(Ex(X''))) restored by calculating the deformation information for the number of recursive iterations.

[0071] Here is a concrete explanation using an example where the deformation information is an angle.

[0072] The first image (Gx(Gx(Ex(X''))) is generated through the following process.

[0073] In the fine tuning process, X-ray image data with added noise is used, but for convenience of explanation, it is referred to as X-ray image data and is denoted by X''.

[0074] X-ray image data (X'') is converted into a latent vector (Ex(X'')) through an encoder and then input to an image restoration module (Gx), and an image (Gx(Ex(X'')) rotated by 30 degrees is generated based on the deformation information (30, 0). Here, the deformation information (30, 0) indicates that the rotation angle is 30 degrees and that there is no position transformation (0).

[0075] The 30-degree rotated image (Gx(Ex(X'')) is encoded into a latent vector again through an encoder and input to the image restoration module (Gx), and based on the transformation information (30, 0), it is further rotated by 30 degrees to generate the first image (Gx(Gx(Ex(X'')))) which is a total 60-degree rotated image.

[0076] In this way, the first image is an image generated by recursively applying the rotation angle.

[0077] On the other hand, the second image (2(Gx(Ex(X''))) is an image generated by applying the total rotation angle of the first image once. Since the first image was generated by recursively applying 30 degrees twice, the second image is an image generated by applying 30 degrees*2=60 degrees once.

[0078] In the fine tuning of the embodiment, the model is fine-tuned by updating the weights in a direction that minimizes the pixel difference between the first image and the second image generated in this way.

[0079] The following mathematical expression 2 is the loss function (L recon ) is shown.

[0080]

[0081] In mathematical expression 2, =ith element of the restored image, =jth element of the restored image, N=number of restored images.

[0082]

[0083] Hereinafter, each step of a method for restoring a three-dimensional image from an X-ray image of an embodiment using an artificial neural network model is described. Fig. 5 is a flowchart showing a method for restoring a three-dimensional image from an X-ray image of an embodiment.

[0084] A method for restoring a 3D image from an X-Ray image of an embodiment comprises a first step (S10) of preprocessing an image to extract an object of interest from an input X-Ray image, a second step (S20) of encoding an image preprocessed in the first step into a latent vector, a third step (S30) of transforming the latent vector encoded in the second step based on deformation information to restore a plurality of deformed X-Ray images, and a fourth step (S40) of generating a 3D image based on the plurality of deformed X-Ray images restored in the third step.

[0085]

[0086] S10 stage

[0087] In this step, the Segment Anything model is applied to the input X-ray image to isolate the object of interest, as an example. The Segment Anything model is an image segmentation model developed by Meta AI. In addition to this model, various other known models can be used to segment objects of interest at this stage.

[0088]

[0089] S20 stage

[0090] This step is the process of encoding the preprocessed image into a latent vector.

[0091] The segmented X-ray image is input to a latent encoder (Ex), which transforms the input image into a low-dimensional latent vector. This latent vector represents the key features of the image in a compressed form.

[0092]

[0093] S30 stage

[0094] This step is a process of transforming the encoded latent vector based on deformation information and restoring it into multiple transformed X-Ray images.

[0095] In this step, the latent vector is transformed using the image restoration module (Gx).

[0096] The input of the image restoration module can be a set of latent vectors, a target angle (R), a target scale (S), and a matching position (T). This image restoration module (Gx) is a Stable Diffusion-based model that transforms the input latent vector to match the target angle / position.

[0097] The image restoration module (Gx) outputs a transformed latent vector, and inputs this value back into the latent decoder (Dx) to restore an X-ray image from a new angle.

[0098] This process is repeated for multiple angles (e.g., angles selected from four equal sections of the 0°-360° range) to generate multi-angle X-ray images.

[0099]

[0100] S40 stage

[0101] This step is a process of creating a 3D image based on the restored multiple transformed X-Ray images.

[0102] This step uses multi-angle X-ray images generated by S30 as input. At this step, various known 3D reconstruction algorithms can be applied to convert 2D images into 3D volume data.

[0103] The generated 3D volume data can be rendered to produce a final 3D X-ray image, which can be viewed from various angles and presented in a form that allows exploration of internal structures.

[0104]

[0105] Fig. 6 is a block diagram illustrating a computing device (800) that executes the method for restoring a three-dimensional image from the above-described X-Ray image, and is a reconstruction of a series of processing steps according to the above-described restoration method from the perspective of hardware configuration. Therefore, in order to avoid duplication of explanation, only an outline of the functions and operations of each configuration will be briefly described here.

[0106] The computing device (800) is configured to include a memory (810) that stores an artificial intelligence model (830) trained to generate multi-angle X-Ray images from an input X-Ray image, and a processor (820) that performs a series of operations for generating a 3D image based on the artificial intelligence model, wherein the artificial intelligence model encodes a preprocessed image into a latent vector, transforms the encoded latent vector based on deformation information to restore a plurality of deformed X-Ray images, and generates a 3D image based on the plurality of restored deformed X-Ray images, and is implemented based on a diffusion model and is trained in a direction that minimizes the difference in latent vectors between the restored images.

[0107] In addition, the artificial neural network model can be pre-trained by repeating the process in which a latent encoder receives an X-ray image with added noise and generates a latent vector with added noise, an image restoration module receives the latent vector with added noise and the transformation information, removes noise from the latent vector with added noise, generates a latent vector according to the transformation information, and a latent decoder generates a final image from the latent vector generated by the image restoration module.

[0108] Additionally, as the above pre-learning progresses, the level of the noise may gradually increase.

[0109] Additionally, in the artificial neural network model, the image restoration module (Gx) can be fine-tuned to minimize the pixel difference between a first image restored by recursively repeating image restoration based on deformation information and a second image restored by calculating the deformation information by the number of recursive iterations.

[0110]

[0111] Meanwhile, the method for restoring a three-dimensional image from an X-ray image of the above-described embodiment can be implemented as a computer-readable code on a computer-readable recording medium. The computer-readable recording medium includes all types of recording devices that store data that can be read by a computer system.

[0112] Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Furthermore, computer-readable recording media can be distributed across network-connected computer systems, allowing computer-readable code to be stored and executed in a distributed manner. Furthermore, functional programs, codes, and code segments for implementing the present invention can be readily inferred by programmers in the technical field to which the present invention pertains.

[0113] The present invention has been described above, focusing on various embodiments thereof. Those skilled in the art will appreciate that the present invention can be implemented in modified forms without departing from its essential characteristics. Therefore, the disclosed embodiments should be considered illustrative rather than limiting. The scope of the present invention is set forth in the claims, not the foregoing description, and all differences within the scope equivalent thereto should be construed as being encompassed by the present invention.

Claims

1. In an artificial neural network model implemented based on the diffusion model, The first step is image preprocessing to extract objects of interest from the input X-Ray image; A second step of encoding the image preprocessed in the first step into a latent vector; A third step of restoring the latent vector encoded in the second step into a plurality of transformed X-Ray images by transforming the latent vector based on the transformation information; A fourth step of generating a 3D image based on the plurality of deformed X-Ray images restored in the third step; Including, A method for restoring a three-dimensional image from an X-ray image, wherein the artificial neural network model is trained in a direction that minimizes the difference in latent vectors between restored images.

2. In paragraph 1, The above artificial neural network model is, The latent encoder receives an X-ray image with added noise as input and generates a latent vector with added noise. The image restoration module receives the latent vector with added noise and the deformation information, removes noise from the latent vector with added noise, and generates a latent vector according to the deformation information. The latent decoder generates a final image from the latent vector generated by the image restoration module. A method for restoring a three-dimensional image from a pre-trained X-Ray image by repeating the process.

3. In paragraph 2, A method for restoring a three-dimensional image from an X-Ray image, wherein the level of the noise gradually increases as the above pre-learning progresses.

4. In paragraph 1, In the above artificial neural network model, A method for restoring a three-dimensional image from an X-Ray image, wherein the image restoration module is fine-tuned to minimize the pixel difference between a first image restored by recursively repeating image restoration based on deformation information and a second image restored by calculating the deformation information for the number of recursive iterations.

5. In paragraph 1, A method for restoring a three-dimensional image from an X-Ray image, wherein the above transformation information includes at least one of rotation, position transformation, and size transformation of the image.

6. A memory storing an artificial intelligence model that implements a method for restoring a three-dimensional image from an X-Ray image as described in any one of paragraphs 1 to 5; and A processor that drives the above artificial intelligence model; A computing device including a .

7. A recording medium having recorded thereon a computer-readable coded program for restoring a three-dimensional image from an X-ray image described in any one of paragraphs 1 to 5.

Citation Information

Patent Citations

  • Method for converting back photo into spine X-ray image based on stable diffusion model

    CN117689533A

  • Training data generation device and training data generation method through voice recognition artificial intelligence model

    KR1020260018255A

  • Method for generating a virtual image using noise in latent space and computing device for the same method

    KR102663123B1