3D volume data conversion method and device using multi-slice-trained 2d diffusion model, and computer program

A multi-slice learned diffusion model addresses inconsistencies in 3D medical image conversion by learning and combining cross-sectional images, improving diagnostic accuracy and reliability through consistent style and structure.

WO2025264018A1PCT designated stage Publication Date: 2025-12-26NEWCURE M INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008522
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-13
Filing Date
2025-06-19
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing deep learning-based methods for converting 3D medical images into 2D slices and back to 3D images often result in stylistic and structural inconsistencies, leading to reduced diagnostic accuracy and reliability due to the lack of modeling inter-slice correlations.

Method used

A multi-slice learned diffusion model that extracts cross-sectional images from 3D volume data, learns correlations between them, and generates consistent 3D volume data by combining these images using a pre-learned artificial intelligence model, specifically a diffusion model trained to minimize a preset loss function.

Benefits of technology

The method ensures consistent style and structure across cross-sectional images, enhancing diagnostic accuracy and reliability by maintaining inter-slice connectivity and reducing artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008522_26122025_PF_FP_ABST
    Figure KR2025008522_26122025_PF_FP_ABST
Patent Text Reader

Abstract

This method by which a computing device uses a multi-slice-trained diffusion model so as to convert 3D volume data comprises the steps of: extracting a plurality of first cross-sectional images from first 3D volume data; generating, through a pre-trained artificial intelligence model, a plurality of second cross-sectional images according to the conversion of the plurality of extracted first cross-sectional images; and combining the plurality of generated second cross-sectional images so as to generate second 3D volume data, wherein the pre-trained artificial intelligence model is a diffusion model trained with the correlation between two or more cross-sectional images according to training, as a group, with two or more cross-sectional images associated with each other from among the plurality of cross-sectional images extracted from the 3D volume data.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device and computer program for converting 3D volume data using a multi-slice learned 2D diffusion model

[0001] The present disclosure relates to a method, device and computer program for converting 3D volume data using a multi-slice learned diffusion model.

[0002] Image transformation technology has recently attracted significant attention in the medical field. This is because medical image analysis and processing play a crucial role in medical diagnosis and treatment planning. In particular, medical imaging techniques such as computed tomography (CT) and magnetic resonance imaging (MRI) provide precise diagnostic information, but accurate analysis can be difficult if the acquired data is complex or contains noise. To address these issues, technologies that transform or correct medical images are essential.

[0003] Advances in image transformation technology have been made possible primarily by the introduction of deep learning.

[0004] Deep learning offers powerful tools for automatically performing various transformation tasks by learning image patterns. For example, deep learning-based generative models, such as generative adversarial networks (GANs) and diffusion models, are effective at transforming images from one domain to another. In medical imaging, these technologies are being utilized for tasks such as noise removal, resolution enhancement, and conversion between various image modalities. However, existing deep learning-based technologies often fail to adequately reflect the unique characteristics of medical imaging, limiting their application in real-world clinical settings.

[0005] This is particularly problematic when dealing with 3D medical images. To process 3D images, the image is typically divided into 2D slices, processed, and then reassembled. While this approach improves computational efficiency and reduces the complexity of deep learning models, it can also lead to stylistic or structural inconsistencies during the conversion process.

[0006] For example, because each 2D slice is processed independently, the final reconstructed 3D image may lose its original structural consistency or be visually distorted. Furthermore, given the critical importance of continuity between slices in medical imaging, these inconsistencies can reduce diagnostic accuracy and reliability.

[0007] The background technology described above is something that the inventor possessed or acquired in the process of deriving the disclosure of the present application, and cannot necessarily be said to be a publicly known technology disclosed to the general public prior to the present application.

[0008] One object of the present disclosure is to provide a method, device and computer program for converting 3D volume data using a multi-slice learned diffusion model.

[0009] An object of the present disclosure is to provide a method, apparatus and computer program for converting 3D volume data using a multi-slice learned diffusion model, which can generate cross-sectional images having consistent style and structure by extracting a plurality of first cross-sectional images from first 3D volume data, converting the plurality of first cross-sectional images into a bundle of interrelated cross-sectional images using a multi-slice learned diffusion model, thereby generating and combining a plurality of second cross-sectional images to generate second 3D volume data, thereby deriving more complete result data.

[0010] A method for converting 3D volume data using a multi-slice learned diffusion model according to one embodiment is a method performed by a computing device, the method comprising: extracting a plurality of first cross-sectional images from first 3D volume data; generating a plurality of second cross-sectional images by converting the extracted plurality of first cross-sectional images through a pre-learned artificial intelligence model; and generating second 3D volume data by combining the generated plurality of second cross-sectional images, wherein the pre-learned artificial intelligence model may be a diffusion model in which a correlation between two or more cross-sectional images is learned by learning two or more cross-sectional images that are mutually related among a plurality of cross-sectional images extracted from 3D volume data as a bundle.

[0011] In one embodiment, the step of generating the plurality of second cross-sectional images may include the step of deriving a plurality of prediction results for a specific first cross-sectional image among the plurality of extracted first cross-sectional images using the pre-learned artificial intelligence model, the step of calculating an average value of the derived plurality of prediction results, and the step of generating a specific second cross-sectional image corresponding to the specific first cross-sectional image based on the calculated average value.

[0012] In one embodiment, the diffusion model may be a model that generates learning data using a plurality of cross-sectional images extracted from mutually adjacent locations from 3D volume data, and is trained to minimize a preset loss function using the generated learning data.

[0013] In one embodiment, the preset loss function may be expressed as in the following mathematical expression 1.

[0014] <Mathematical Formula 1>

[0015]

[0016] Here, the above is the expected value, above is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the actual noise and the above may be noise predicted by the diffusion model.

[0017] In one embodiment, the preset loss function may be expressed as in the following mathematical expression 2 when the diffusion model is a Brownian Bridge Diffusion Model (BBDM).

[0018] <Mathematical Formula 2>

[0019]

[0020] Here, the above is the expected value, above is the second 3D volume data, which is the target data, is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the weight, above is the interpolation coefficient, above is the time-step noise coefficient, is the actual noise and the above may be noise predicted by the diffusion model.

[0021] In one embodiment, the step of generating the second 3D volume data may include the step of correcting the generated plurality of second cross-sectional images so that manifolds of the generated plurality of second cross-sectional images match, and the step of combining the corrected plurality of second cross-sectional images to generate the second 3D volume data.

[0022] In one embodiment, the step of correcting the plurality of second cross-sectional images generated may include the step of correcting the plurality of second cross-sectional images generated based on the following mathematical expression 3.

[0023] <Mathematical Formula 3>

[0024]

[0025] Here, the above is the mth corrected data at step t, is the m-1th corrected data at step t, is the correction intensity factor, is the time-step noise coefficient, is the dimension of the data and the above can be a score function.

[0026] In one embodiment, the score function may be expressed as in mathematical expression 4 below.

[0027] <Mathematical Formula 4>

[0028]

[0029] Here, the above is the score function, above is a specific point in time Intermediate state data in the above The prediction results derived by analyzing the first cross-sectional image through the above diffusion model and the above can be a time-step noise coefficient.

[0030] In one embodiment, the score function may be expressed as in the following mathematical expression 5 when the diffusion model is a Brownian bridge diffusion model.

[0031] <Mathematical Formula 5>

[0032]

[0033] Here, the above is the score function, above is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the interpolation coefficient, above The prediction results derived by analyzing the first cross-sectional image through the above diffusion model and the above can be a time-step noise coefficient.

[0034] A 3D volume data conversion device using a multi-slice learned diffusion model according to one embodiment includes a processor, a network interface, a memory, and a computer program loaded into the memory and executed by the processor, wherein the computer program may include an instruction for extracting a plurality of first cross-sectional images from first 3D volume data, an instruction for generating a plurality of second cross-sectional images by converting the extracted plurality of first cross-sectional images through a pre-learned artificial intelligence model, wherein the pre-learned artificial intelligence model is a diffusion model in which a relationship between two or more cross-sectional images is learned by learning two or more interrelated cross-sectional images among a plurality of cross-sectional images extracted from 3D volume data as a bundle, and an instruction for generating second 3D volume data by combining the generated plurality of second cross-sectional images.

[0035] In one embodiment, a computer-readable non-transitory recording medium having recorded thereon instructions for executing a method for converting 3D volume data using a multi-slice learned diffusion model, the method comprising: extracting a plurality of first cross-sectional images from first 3D volume data; generating a plurality of second cross-sectional images by converting the extracted plurality of first cross-sectional images using a pre-learned artificial intelligence model, wherein the pre-learned artificial intelligence model is a diffusion model in which a relationship between two or more cross-sectional images is learned by grouping two or more interrelated cross-sectional images among the plurality of cross-sectional images extracted from the 3D volume data into a single bundle; and generating second 3D volume data by combining the generated plurality of second cross-sectional images.

[0036] A method, device, and computer program for converting 3D volume data using a multi-slice learned diffusion model according to one embodiment extracts a plurality of first cross-sectional images from first 3D volume data, converts the plurality of first cross-sectional images using a multi-slice learned diffusion model by grouping the interrelated cross-sectional images into a single bundle, thereby generating and combining a plurality of second cross-sectional images to generate second 3D volume data, thereby generating cross-sectional images having consistent styles and structures, thereby deriving more complete result data.

[0037] A data transformation method using a multi-slice learned diffusion model according to one embodiment is not limited to transforming 3D volume data using a 2D diffusion model, and the multi-slice learned diffusion model can perform transformation by reflecting correlations between data of a dimension one higher than the learned dimension.

[0038] The effects of the 3D volume data conversion method, device and computer program using a multi-slice learned diffusion model according to the embodiment are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0039] The following drawings attached to this specification illustrate preferred embodiments of the present disclosure and, together with the detailed description of the invention, serve to further understand the technical idea of ​​the present disclosure, and therefore, the present disclosure should not be interpreted as being limited to matters described in such drawings.

[0040] Figure 1 is a drawing showing an MRI image and a target MRI image converted according to a conventional technique.

[0041] FIG. 2 is a flowchart of a method for converting 3D volume data using a multi-slice learned diffusion model according to one embodiment of the present disclosure.

[0042] FIG. 3 is a diagram illustrating a 3D volume data conversion process using a multi-slice learned diffusion model according to one embodiment of the present disclosure.

[0043] FIG. 4 is a flowchart of a second cross-sectional image generation method based on co-prediction according to one embodiment of the present disclosure.

[0044] FIG. 5 is a flowchart of a second 3D volume data generation method based on second cross-sectional image correction according to one embodiment of the present disclosure.

[0045] FIG. 6 is a diagram illustrating a second cross-sectional image generation process and a second cross-sectional image correction process based on multi-prediction according to one embodiment of the present disclosure.

[0046] FIG. 7 is a diagram illustrating a hardware configuration of a 3D volume data conversion device using a multi-slice learned diffusion model according to one embodiment of the present disclosure.

[0047] The various embodiments described in this specification are exemplified for the purpose of clearly explaining the technical concept of the present disclosure and are not intended to be limited to specific embodiments. The technical concept of the present disclosure includes various modifications, equivalents, alternatives, and embodiments selectively combined from all or part of the embodiments described herein. Furthermore, the scope of the technical concept of the present disclosure is not limited to the various embodiments presented below or the specific descriptions thereof.

[0048] Terms used herein, including technical or scientific terms, unless otherwise defined, may have the meaning commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0049] As used herein, expressions such as "includes," "may include," "comprises," "may have," "have," and "may have" indicate the presence of a target feature (e.g., a function, operation, or component), but do not exclude the presence of other additional features. In other words, such expressions should be understood as open-ended terms that imply the possibility of including a second embodiment.

[0050] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.

[0051] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0052] According to one embodiment of the present disclosure, a 'module' or 'unit' may be implemented as a processor and a memory. 'Processor' should be broadly construed to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, and the like. In some circumstances, a 'processor' may also refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), and the like. A 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors in conjunction with a DSP core, or any other such combination of configurations. In addition, 'memory' should be broadly construed to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.

[0053] As used herein, the expressions “first,” “second,” or “first,” “second,” etc., unless the context indicates otherwise, are used to refer to multiple similar objects and to distinguish one object from another, and do not limit the order or importance among the objects.

[0054] As used herein, the expressions "A, B, and C," "A, B, or C," "A, B, and / or C," or "at least one of A, B, and C," "at least one of A, B, or C," "at least one of A, B, and / or C," "at least one selected from A, B, and C," "at least one selected from A, B, or C," "at least one selected from A, B, and / or C," and the like can mean each listed item or all possible combinations of the listed items. For example, "at least one selected from A and B" can refer to (1) A, (2) at least one of A, (3) B, (4) at least one of B, (5) at least one of A and at least one of B, (6) at least one of A and B, (7) at least one of B and A, and (8) both A and B.

[0055] The expression "based on" as used herein is used to describe one or more factors that influence the decision, act of judgment, or action described in the phrase or sentence containing the expression, and this expression does not exclude additional factors that influence the decision, act of judgment, or action.

[0056] As used herein, the expression that a component (e.g., a first component) is “connected” or “connected” to another component (e.g., a second component) may mean that the component is directly connected or connected to the other component, as well as connected or connected via a new other component (e.g., a third component).

[0057] The expression "configured to" used herein may have the meanings of "set to", "having the ability to", "modified to", "made to", "capable of", etc., depending on the context. The expression is not limited to the meaning of "specifically designed in hardware", and for example, a processor configured to perform a specific operation may mean a generic-purpose processor that can perform the specific operation by executing software.

[0058] Hereinafter, various embodiments of the present disclosure will be described with reference to the attached drawings. In the attached drawings and the description of the drawings, identical or substantially equivalent components may be assigned the same reference numerals. Furthermore, in the description of various embodiments below, duplicate descriptions of identical or corresponding components may be omitted, but this does not mean that the corresponding components are not included in the embodiments.

[0059] Deep learning-based image transformation technology is a fundamental and widely applicable technology that can be applied to any problem that requires predicting another image from a specific one. In particular, in medical imaging, various problems such as image modality transformation (e.g., between MRI modalities, such as T1 MRI and T2FLAIR MRI, and between CT and MRI), denoising in PET / CT / MRI, artifact removal (e.g., motion artifacts, beam hardening in CT, magnetic susceptibility artifacts in MRI), super-resolution, and image restoration (inpainting) can all be interpreted as image transformation problems.

[0060] Many medical images, such as PET, CT, and MRI, often contain 3D data, which is volumetric and created by stacking 2D images. This data can be trained and transformed slice-by-slice using 2D encoder-decoder models (e.g., deep learning networks). For example, axial, sagittal, and coronal sections can be processed separately.

[0061] However, these 2D-based approaches (hereinafter referred to as "prior art") can lead to slice inconsistencies. This causes slice-level inferences to fail to maintain inter-slice connectivity. In particular, structural continuity is crucial for 3D data at the volume level, and such discontinuities can manifest as image quality degradation or artifacts.

[0062] Figure 1 is a drawing showing an MRI image and a target MRI image converted according to a conventional technique.

[0063] Referring to Fig. 1, when comparing an MRI image generated using a conventional technique with a target MRI image, it can be seen that the MRI image converted according to the conventional technique exhibits style inconsistencies (style SI) such as brightness and contrast. In addition, it can be seen that the MRI image converted according to the conventional technique exhibits structure inconsistencies (shape SI), which appear as a type of artifact due to inconsistencies in shape and structure between slices.

[0064] Meanwhile, 3D models have the advantage of resolving discontinuities between slices, but there are some technical limitations.

[0065] First, high memory usage is required. When comparing 3D models with the same batch size and model capacity, memory usage is significantly higher than that of 2D models, necessitating a high-performance GPU for training 3D models.

[0066] Second, due to this high memory usage, 3D models are trained with much smaller batch sizes than 2D models, which leads to reduced learning speed and stability.

[0067] Third, 3D models require more data than 2D models. Because the data dimensionality encoded by 3D models is relatively much larger, they require more data to learn the high-dimensional spatial manifold.

[0068] Accordingly, according to the 3D volume data conversion method using a multi-slice learned diffusion model according to various embodiments of the present disclosure, the problems of the prior art, namely, the style mismatch and structural mismatch that occur during image conversion, can be resolved. This will be described in more detail below with reference to FIGS. 2 to 6.

[0069] Here, the methods described with reference to FIGS. 2 to 6 may be performed through a 3D volume data conversion device (100) using a multi-slice learned diffusion model illustrated in FIG. 7 (hereinafter referred to as 'computing device (100)'), but are not limited thereto.

[0070] FIG. 2 is a flowchart of a method for converting 3D volume data using a multi-slice learned diffusion model according to an embodiment of the present disclosure, and FIG. 3 is a diagram illustrating a process for converting 3D volume data using a multi-slice learned diffusion model according to an embodiment of the present disclosure.

[0071] As mentioned earlier, the most fundamental reason for the style mismatch (style SI) issue between cross-sectional images is that each cross-sectional image is individually inferred and transformed, so the correlation between interrelated cross-sectional images (e.g., cross-sectional images extracted from adjacent locations) is not modeled.

[0072] Of course, if you are lucky enough to have all the data in the data set have the same style, the result will also reflect the same style, so there will be no problem with style mismatch (style SI). However, shape mismatch (shape SI), where the structure and boundaries between consecutive cross-sectional images are subtly misaligned, is bound to occur.

[0073] Taking these points into consideration, a computing device (100) according to one embodiment of the present disclosure can perform a 3D volume data conversion method using a multi-slice learned diffusion model for the purpose of simultaneously resolving not only style mismatch but also structural mismatch problems between cross-sectional images.

[0074] In addition, the data transformation method using a multi-slice learned diffusion model according to an embodiment of the present disclosure is not limited to transforming 3D volume data using a 2D diffusion model. In one embodiment, the multi-slice learned diffusion model can perform transformation by reflecting the correlation between data in a dimension one higher than the learned dimension. For example, in the case of 4D medical image data such as functional magnetic resonance imaging (fMRI) including a time axis, 4D volume data can be generated using only the 3D diffusion model of the present invention.

[0075] Referring to FIGS. 2 and 3, in step S110, the computing device (100) can extract a plurality of first cross-sectional images from the first 3D volume data.

[0076] Here, the first 3D volume data is 3D CT data, and the plurality of first cross-sectional images may be, but are not limited to, a plurality of cross-sectional CT images extracted from the 3D CT data.

[0077] At step S120, the computing device (100) can generate a plurality of second cross-sectional images by converting each of the plurality of first cross-sectional images extracted through step S110.

[0078] In one embodiment, the computing device (100) can generate a plurality of second cross-sectional images by transforming a plurality of first cross-sectional images through a pre-learned artificial intelligence model.

[0079] Here, the artificial intelligence model may be a model that learns the correlation between two or more cross-sectional images by learning by grouping two or more cross-sectional images that are mutually related among multiple cross-sectional images extracted from 3D volume data into one bundle.

[0080] An artificial intelligence model (e.g., a neural network) consists of one or more network functions, which may be comprised of a set of interconnected computational units, generally referred to as "nodes." These "nodes" may also be referred to as "neurons." One or more network functions comprise at least one node. The nodes (or neurons) comprising one or more network functions may be interconnected by one or more "links."

[0081] Within an AI model, one or more nodes connected via links can form a relationship between input nodes and output nodes. The concepts of input nodes and output nodes are relative, meaning that any node in an output node relationship with one node can also be in an input node relationship with another node, and vice versa. As described above, input node-to-output node relationships can be created around links. One input node can be connected to one or more output nodes via links, and vice versa.

[0082] In a relationship between input nodes and output nodes connected through a single link, the value of the output node can be determined based on the data input to the input node. Here, the node interconnecting the input nodes and output nodes can have a weight. The weight can be variable and can be varied by a user or an algorithm so that the artificial intelligence model can perform a desired function. For example, when one or more input nodes are interconnected to one output node through each link, the output node can determine the output node value based on the values ​​input to the input nodes connected to the output node and the weight set for the link corresponding to each input node.

[0083] As described above, an AI model consists of one or more nodes interconnected through one or more links, forming input and output node relationships within the AI ​​model. The characteristics of an AI model can be determined based on the number of nodes and links, the relationships between nodes and links, and the weights assigned to each link. For example, if two AI models exist with the same number of nodes and links but different weight values ​​between the links, the two AI models may be perceived as different from each other.

[0084] Some of the nodes that make up an AI model can form a layer based on their distances from the initial input node. For example, a set of nodes that are n distances from the initial input node can form n layers. The distance from the initial input node can be defined by the minimum number of links that must be passed to reach the node from the initial input node. However, this definition of a layer is arbitrary for illustrative purposes, and the order of layers within an AI model can be defined in a different way than described above. For example, the layer of nodes can also be defined by their distance from the final output node.

[0085] The initial input node may refer to one or more nodes in an AI model into which data is directly input without going through a link in its relationship with other nodes. Alternatively, in an AI model network, in terms of the relationship between nodes based on links, it may refer to nodes that do not have other input nodes connected by links. Similarly, the final output node may refer to one or more nodes in an AI model that do not have an output node in its relationship with other nodes. In addition, a hidden node may refer to nodes that constitute an AI model other than the initial input node and the final output node. An AI model according to one embodiment of the present disclosure may be an AI model in which the number of nodes in the input layer may be greater than that in the hidden layer closer to the output layer, and the number of nodes decreases as it progresses from the input layer to the hidden layer.

[0086] An AI model may include one or more hidden layers. Hidden nodes in a hidden layer can receive the output of the previous layer and the output of surrounding hidden nodes as input. The number of hidden nodes in each hidden layer may be the same or different. The number of nodes in the input layer may be determined based on the number of data fields in the input data and may be the same as or different from the number of hidden nodes. Input data entered into the input layer may be computed by the hidden nodes in the hidden layer and output by the output layer, a fully connected layer (FCL).

[0087] In various embodiments, the artificial intelligence model may be a deep learning model.

[0088] A deep learning model (e.g., a deep neural network (DNN)) can refer to an artificial intelligence model that includes multiple hidden layers in addition to input and output layers. Using a deep neural network, it is possible to identify latent structures in data. That is, it is possible to identify latent structures in photos, text, videos, voices, and music (e.g., what objects are in the photo, what the content and emotion of the text are, what the content and emotion of the voice are, etc.).

[0089] Deep neural networks may include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, generative adversarial networks (GANs), restricted boltzmann machines (RBMs), deep belief networks (DBNs), Q networks, U networks, Siamese networks, etc.

[0090] In various embodiments, the network function may include an autoencoder, which may be a type of artificial neural network that outputs output data similar to the input data.

[0091] An autoencoder may include at least one hidden layer, and an odd number of hidden layers may be positioned between the input and output layers. The number of nodes in each layer may be reduced from the number of nodes in the input layer to an intermediate layer called the bottleneck layer (encoding), and then expanded symmetrically from the bottleneck layer to the output layer (symmetrical to the input layer). The nodes of the dimensionality reduction layer and the dimensionality restoration layer may or may not be symmetrical. In addition, the autoencoder may perform nonlinear dimensionality reduction. The number of input and output layers may correspond to the number of sensors remaining after preprocessing the input data. In the autoencoder structure, the number of nodes in the hidden layer included in the encoder may have a structure that decreases as it moves away from the input layer. The number of nodes in the bottleneck layer (the layer with the fewest nodes located between the encoder and the decoder) may be maintained above a certain number (e.g., more than half of the number of nodes in the input layer) because too small a number may not convey sufficient information.

[0092] In various embodiments, the artificial intelligence model may be a diffusion model in which an association between two or more cross-sectional images is learned.

[0093] For example, a diffusion model may be a model that learns the correlation between cross-sectional images at mutually adjacent locations by generating learning data using multiple cross-sectional images extracted at mutually adjacent locations from 3D volume data and learning to minimize a preset loss function using the generated learning data.

[0094] At this time, the preset loss function for the diffusion model can be expressed as in the following mathematical expression 1.

[0095] <Mathematical Formula 1>

[0096]

[0097] Here, the above is the expected value, above is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the actual noise and the above may be noise predicted by the diffusion model.

[0098] Meanwhile, the diffusion model may be a Brownian bridge diffusion model, and when the diffusion model is a Brownian bridge diffusion model, the loss function set for the diffusion model may be expressed as in the following mathematical expression 2.

[0099] <Mathematical Formula 2>

[0100]

[0101] Here, the above is the expected value, above is the second 3D volume data, which is the target data, is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the weight, above is the interpolation coefficient, above is the time-step noise coefficient, is the actual noise and the above may be noise predicted by the diffusion model.

[0102] In one embodiment, the computing device (100) can generate a second cross-sectional image based on multiple predictions (e.g., the co-prediction process of FIG. 6) using a pre-trained artificial intelligence model. This will be described below with reference to FIG. 4.

[0103] FIG. 4 is a flowchart of a second cross-sectional image generation method based on co-prediction according to one embodiment of the present disclosure.

[0104] Referring to FIG. 4, in step S210, the computing device (100) can derive multiple prediction results for the first cross-sectional image using an artificial intelligence model.

[0105] In one embodiment, the computing device (100) can derive multiple prediction results by inputting a specific first cross-sectional image into the artificial intelligence model multiple times.

[0106] In step S220, the computing device (100) can calculate an average value of multiple prediction results derived through step S210.

[0107] In step S230, the computing device (100) can generate a second cross-sectional image corresponding to the first cross-sectional image using the average value calculated through step S220.

[0108] Again, referring to FIGS. 2 and 3, at step S130, the computing device (100) can generate second 3D volume data by combining a plurality of second cross-sectional images generated through step S120.

[0109] Here, the plurality of second cross-sectional images are a plurality of MRI cross-sectional images generated by converting a plurality of CT cross-sectional images, and the second 3D volume data may be 3D MRI data generated by combining a plurality of MRI cross-sectional images. However, the present invention is not limited thereto.

[0110] In one embodiment, the computing device (100) can correct a plurality of second cross-sectional images (e.g., the Correction process of FIG. 6) and generate second 3D volume data using the corrected plurality of second cross-sectional images.

[0111] Here, the operation of correcting multiple second cross-sectional images may be omitted for reduced computational complexity, but is not limited thereto. A more detailed description will be provided below with reference to FIG. 5.

[0112] FIG. 5 is a flowchart of a second 3D volume data generation method based on second cross-sectional image correction according to one embodiment of the present disclosure.

[0113] Referring to FIG. 5, in step S310, the computing device (100) can correct the generated plurality of second cross-sectional images so that the manifolds of the plurality of second cross-sectional images match.

[0114] In one embodiment, the computing device (100) can correct a plurality of second cross-sectional images based on the following mathematical expression 3.

[0115] <Mathematical Formula 3>

[0116]

[0117] Here, the above is the mth corrected data at step t, is the m-1th corrected data at step t, is the correction intensity factor, is the time-step noise coefficient, is the dimension of the data and the above can be a score function.

[0118] At this time, the score function can be expressed as in mathematical equation 4 below.

[0119] <Mathematical Formula 4>

[0120]

[0121] Here, the above is the score function, above is a specific point in time Intermediate state data in the above The prediction results derived by analyzing the first cross-sectional image through the above diffusion model and the above can be a time-step noise coefficient.

[0122] Meanwhile, when the diffusion model is a Brownian bridge diffusion model, the score function can be expressed as in mathematical expression 5 below.

[0123] <Mathematical Formula 5>

[0124]

[0125] Here, the above is the score function, above is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the interpolation coefficient, above The prediction results derived by analyzing the first cross-sectional image through the above diffusion model and the above can be a time-step noise coefficient.

[0126] A method for converting 3D volume data using a multi-slice learned diffusion model according to an embodiment of the present disclosure may apply a style key (Style Key, C_SKC) as a technique for adjusting a style within the converted 3D volume data. The style key is a value indicating the style of a medical image to be generated, and may be an index that numerically expresses image features such as brightness, contrast, and color change of the image. For example, the style of a medical image may refer to features such as the overall brightness, contrast, and intensity distribution by tissue of the medical image, and the style key may be a value defined by quantifying these image features.

[0127] In one embodiment, to ensure that a medical image to be generated has a specific style, a computing device (100) according to an embodiment of the present disclosure may apply a diffusion model including a style key as an input condition. The style key may be utilized as a condition to reflect the style of the target image when performing a specific medical image transformation, and may be used as an input condition in the process of optimizing the score function and loss function of the diffusion model.

[0128] In step S320, the computing device (100) can generate second 3D volume data by combining a plurality of second cross-sectional images corrected through step S310. Hereinafter, with reference to FIG. 7, a hardware configuration of the computing device (100) that performs a 3D volume data conversion method using a multi-slice learned diffusion model according to various embodiments of the present disclosure will be described.

[0129] FIG. 7 is a diagram illustrating a hardware configuration of a 3D volume data conversion device using a multi-slice learned diffusion model according to one embodiment of the present disclosure.

[0130] Referring to FIG. 7, a computing device (100) according to one embodiment of the present disclosure may include one or more processors (110), a memory (120) for loading a computer program (151) executed by the processor (110), a bus (130), a communication interface (140), and a storage (150) for storing the computer program (151). Here, only components related to the embodiment of the present disclosure are illustrated in FIG. 7. Therefore, a person skilled in the art to which the present disclosure pertains may understand that other general components may be included in addition to the components illustrated in FIG. 7.

[0131] The processor (110) controls the overall operation of each component of the computing device (100). The processor (110) may be configured to include a CPU (Central Processing Unit), an MPU (Micro Processor Unit), an MCU (Micro Controller Unit), a GPU (Graphics Processing Unit), or any other type of processor well known in the art of the present disclosure.

[0132] Additionally, the processor (110) may perform operations for at least one application or program for executing a method according to embodiments of the present disclosure, and the computing device (100) may have one or more processors.

[0133] In various embodiments, the processor (110) may further include a Random Access Memory (RAM, not shown) and a Read-Only Memory (ROM, not shown) that temporarily and / or permanently store signals (or data) processed within the processor (110). In addition, the processor (110) may be implemented in the form of a system on chip (SoC) that includes at least one of a graphics processing unit, RAM, and ROM.

[0134] The memory (120) stores various data, commands, and / or information. The memory (120) can load a computer program (151) from the storage (150) to execute methods / operations according to various embodiments of the present disclosure. When the computer program (151) is loaded into the memory (120), the processor (110) can perform the method / operation by executing one or more instructions constituting the computer program (151). The memory (120) may be implemented as a volatile memory such as RAM, but the technical scope of the present disclosure is not limited thereto.

[0135] The bus (130) provides a communication function between components of the computing device (100). The bus (130) may be implemented as various types of buses, such as an address bus, a data bus, and a control bus.

[0136] The communication interface (140) supports wired and wireless Internet communication of the computing device (100). Furthermore, the communication interface (140) may support various communication methods other than Internet communication. To this end, the communication interface (140) may be configured to include a communication module well known in the technical field of the present disclosure. In some embodiments, the communication interface (140) may be omitted.

[0137] The storage (150) can non-temporarily store a computer program (151). When performing a medical image conversion process using a style condition-based diffusion model and a 3D volume data conversion process using a multi-slice learned diffusion model through a computing device (100), the storage (150) can store various information necessary to provide the medical image conversion process using a style condition-based diffusion model and the 3D volume data conversion process using a multi-slice learned diffusion model.

[0138] Storage (150) may be configured to include non-volatile memory such as ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), flash memory, a hard disk, a removable disk, or any form of computer-readable recording medium well known in the art to which the present disclosure pertains.

[0139] The computer program (151) may include one or more instructions that, when loaded into the memory (120), cause the processor (110) to perform a method / operation according to various embodiments of the present disclosure. That is, the processor (110) may perform the method / operation according to various embodiments of the present disclosure by executing the one or more instructions.

[0140] In one embodiment, the computer program (151) may include one or more instructions for performing a 3D volume data conversion method using a multi-slice learned diffusion model, including the steps of extracting a plurality of first cross-sectional images from first 3D volume data, generating a plurality of second cross-sectional images by converting the extracted plurality of first cross-sectional images using a learned artificial intelligence model, and generating second 3D volume data by combining the generated plurality of second cross-sectional images.

[0141] The steps of a method or algorithm described in connection with the embodiments of the present disclosure may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in a random access memory (RAM), a read only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable recording medium well known in the art to which the present disclosure pertains.

[0142] The components of the present disclosure may be implemented as programs (or applications) and stored on a medium to be executed in conjunction with a computer, which is hardware. The components of the present disclosure may be implemented as software programs or software elements. Similarly, the embodiments may be implemented in a programming or scripting language such as C, C++, Java, or an assembler, including various algorithms implemented as a combination of data structures, processes, routines, or other programming components. Functional aspects may be implemented as algorithms that are executed on one or more processors.

[0143] As described above, those skilled in the art will appreciate that the present disclosure can be implemented in other specific forms without altering the technical spirit or essential characteristics thereof. Therefore, the above-described embodiments should be understood as illustrative in all respects and not restrictive. The scope of the present disclosure is defined by the following claims rather than the detailed description, and all changes or modifications derived from the meaning and scope of the claims and equivalent concepts should be construed as being included within the scope of the present disclosure.

[0144] The features and advantages described in this specification are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art upon review of the drawings, specification, and claims. Furthermore, it should be noted that the language used in this specification has been primarily selected for readability and instructional purposes, and may not be intended to delineate or circumscribe the subject matter of the present disclosure.

[0145] The above description of the embodiments of the present disclosure has been presented for illustrative purposes. It is not intended to limit the present disclosure to the precise form disclosed, nor is it intended to be exhaustive. Those skilled in the art will appreciate that numerous modifications and variations are possible in light of the above disclosure.

[0146] Therefore, the scope of this disclosure is not limited by the detailed description, but is defined by any claims of the application based on this description. Accordingly, the disclosure of embodiments of this disclosure is illustrative and does not limit the scope of this disclosure, which is set forth in the following claims.

Claims

1. In a method performed by a computing device, A step of extracting a plurality of first cross-sectional images from first 3D volume data; A step of generating a plurality of second cross-sectional images by converting the plurality of extracted first cross-sectional images through a learned artificial intelligence model; and A step of combining the plurality of second cross-sectional images generated above to generate second 3D volume data, The above-mentioned artificial intelligence model is, A diffusion model in which the correlation between two or more cross-sectional images is learned by grouping two or more cross-sectional images that are related to each other among multiple cross-sectional images extracted from 3D volume data into one bundle. A method for transforming 3D volume data using a multi-slice learned diffusion model.

2. In paragraph 1, The step of generating the plurality of second cross-sectional images is: A step of deriving multiple prediction results for a specific first cross-sectional image among the plurality of extracted first cross-sectional images using the above-mentioned learned artificial intelligence model; A step of calculating an average value of the above-described multiple prediction results; and A step of generating a specific second cross-sectional image corresponding to the specific first cross-sectional image based on the calculated average value, A method for transforming 3D volume data using a multi-slice learned diffusion model.

3. In paragraph 1, The above diffusion model is, A model that generates learning data using multiple cross-sectional images extracted from mutually adjacent locations from 3D volume data, and learns to minimize a preset loss function using the generated learning data. A method for transforming 3D volume data using a multi-slice learned diffusion model.

4. In paragraph 3, The above preset loss function is, It is expressed as the following mathematical formula 1, A method for transforming 3D volume data using a multi-slice learned diffusion model. <Mathematical Formula 1> Here, the above is the expected value, above is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the actual noise and the above is the noise predicted by the diffusion model 5. In paragraph 3, The above preset loss function is, When the above diffusion model is a Brownian Bridge Diffusion Model (BBDM), it is expressed as in the following mathematical equation 2. A method for transforming 3D volume data using a multi-slice learned diffusion model. <Mathematical Formula 2> Here, the above is the expected value, above is the second 3D volume data, which is the target data, is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the weight, above is the interpolation coefficient, above is the time-step noise coefficient, is the actual noise and the above is the noise predicted by the diffusion model 6. In paragraph 1, The step of generating the second 3D volume data is: A step of correcting the plurality of second cross-sectional images so that the manifolds of the plurality of second cross-sectional images are identical; and A step of combining the plurality of corrected second cross-sectional images to generate second 3D volume data, A method for transforming 3D volume data using a multi-slice learned diffusion model.

7. In paragraph 6, The step of correcting the plurality of second cross-sectional images generated above is: A step of correcting the plurality of second cross-sectional images generated based on the following mathematical expression 3, A method for transforming 3D volume data using a multi-slice learned diffusion model. <Mathematical Formula 3> Here, the above is the mth corrected data at step t, is the m-1th corrected data at step t, is the correction intensity factor, is the time-step noise coefficient, is the dimension of the data and the above is a score function 8. In paragraph 7, The above scoring function is, It is expressed as the following mathematical formula 4, A method for transforming 3D volume data using a multi-slice learned diffusion model. <Mathematical Formula 4> Here, the above is the score function, above is a specific point in time Intermediate state data in the above The prediction results derived by analyzing the first cross-sectional image through the above diffusion model and the above is the noise coefficient for each time step 9. In paragraph 7, The above scoring function is, When the above diffusion model is a Brownian bridge diffusion model, it is expressed as in the following mathematical expression 5. A method for transforming 3D volume data using a multi-slice learned diffusion model. <Mathematical Formula 5> Here, the above is the score function, above is a specific point in time Intermediate state data in the above is the first 3D volume data, which is the initial data, is the interpolation coefficient, above The prediction results derived by analyzing the first cross-sectional image through the above diffusion model and the above is the noise coefficient for each time step 10. Processor; network interface; memory; and A computer program loaded into the above memory and executed by the above processor, The above computer program, An instruction for extracting a plurality of first cross-sectional images from first 3D volume data; An instruction for generating a plurality of second cross-sectional images by transforming the plurality of extracted first cross-sectional images through a pre-learned artificial intelligence model, wherein the pre-learned artificial intelligence model is a diffusion model in which a relationship between two or more cross-sectional images is learned by grouping two or more interrelated cross-sectional images extracted from 3D volume data into a single bundle; and An instruction for generating second 3D volume data by combining the plurality of second cross-sectional images generated above, A 3D volume data conversion device using a multi-slice learned diffusion model.

11. A step of extracting a plurality of first cross-sectional images from the first 3D volume data; A step of generating a plurality of second cross-sectional images by converting the plurality of extracted first cross-sectional images through a pre-learned artificial intelligence model - wherein the pre-learned artificial intelligence model is a diffusion model in which a relationship between two or more cross-sectional images is learned by grouping two or more interrelated cross-sectional images among a plurality of cross-sectional images extracted from 3D volume data into a single bundle; and A computer-readable non-transitory recording medium having recorded thereon commands for executing a method for converting 3D volume data using a multi-slice learned diffusion model, the method including the step of combining the plurality of second cross-sectional images generated above to generate second 3D volume data.

Citation Information

Patent Citations

  • Method for training a neural network model for semiconductor design

    KR1020230124461A

  • Dash grommet for vehicle

    KR1020250064116A

  • Ship operating system and method using sea fog removal device

    KR102321675B1

  • Music box for exhibition place

    KR102558114B1

  • Apparatus and method for power transmission of distribution line

    KR102909320B1