Method and apparatus for analyzing computed tomography image using artificial intelligence model
Patent Information
- Application Number
- PCT/KR2025/020461
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-24
- Filing Date
- 2025-12-02
- Publication Date
- 2026-10-01
Smart Images

Figure KR2025020461_01102026_PF_FP_ABST
Abstract
Description
Method and apparatus for analyzing computed tomography images using an artificial intelligence model
[0001] The present disclosure relates to a method and apparatus for analyzing computed tomography images. Specifically, the present disclosure relates to a method and apparatus for analyzing computed tomography images using an artificial intelligence model.
[0002] Clinicians must identify anatomical structures or pathological signs appearing in medical images in 2D or 3D form. In this process, segmentation is performed to assign labels to each pixel, which is the unit of a 2D image, or each voxel, which is the unit of a 3D image.
[0003] Segmenting specific structures in computed tomography (CT) scans significantly impacts decision-making in both diagnosis and treatment. Performing this task manually can lead to problems. For instance, clinicians may spend a considerable amount of time distinguishing complex structures or microtissues when segmenting specific structures in CT scans. Furthermore, errors or omissions by the clinician may occur during the aforementioned process.
[0004] Artificial intelligence models (e.g., deep learning algorithms) can be used to solve the aforementioned problems. Using artificial intelligence models, the region of a pancreatic tumor can be effectively segmented in computed tomography scan images. However, even with the use of artificial intelligence models, there are still limitations in simultaneously achieving accurate segmentation of the pancreatic tumor region in computed tomography scan images and segmentation of the pancreatic tumor region under various imaging conditions.
[0005] When applying a 3D CNN structure to an AI model, the processing burden on the AI model may increase in terms of data sets and computational complexity. Conversely, when applying a 2D-based structure to an AI model, it may be difficult to sufficiently reflect 3D context. Therefore, segmenting lesions with faint boundaries or small tumors in computed tomography scan images may not be perfect.
[0006] The present application is an invention related to the Artificial Intelligence Convergence Innovation Talent Development (Ewha Womans University) (Project No.: 2710033947), which was carried out under the Artificial Intelligence Convergence Innovation Talent Development (R&D) project of the Ministry of Science and ICT. The project was carried out at Ewha Womans University, the project implementing agency, between 2025-01-01 and 2025-12-31, and was managed by the Information and Communications Planning and Evaluation Institute, the project management (specialized) agency.
[0007] The problem that the present disclosure aims to solve is to provide a method and apparatus for analyzing computed tomography images using an artificial intelligence model to more precisely segment regions of small and difficult-to-identify pancreatic tumors.
[0008] The problems that the present disclosure aims to solve are not limited to those mentioned above, and other problems and advantages of the present disclosure not mentioned can be understood from the following description and will be more clearly understood from the embodiments of the present disclosure. Furthermore, it will be seen that the problems and advantages that the present disclosure aims to solve can be realized by the means and combinations thereof set forth in the claims.
[0009] As a technical means for achieving the aforementioned technical problem, a method for analyzing computed tomography images using an artificial intelligence model according to one embodiment of the present disclosure comprises: a step of acquiring a plurality of computed tomography images; and a step of outputting a first mask as the plurality of computed tomography images are input into a first artificial intelligence model that has been previously trained; wherein the first artificial intelligence model may include a second artificial intelligence model that has been previously trained based on a MAR model.
[0010] In addition, the first artificial intelligence model further includes a decoder, and the decoder may include a Masked-attention Mask Transformer.
[0011] In addition, the second artificial intelligence model can be trained based on tokens generated as a single computed tomography image is input into the second artificial intelligence model.
[0012] Additionally, the method may further include the step of generating a second mask and a third mask as a plurality of computed tomography images are input to the first artificial intelligence model and the third model; the step of updating a second parameter of the third model based on the second mask and the third mask; and the step of training the first artificial intelligence model by updating a first parameter of the first artificial intelligence model based on the result of updating the second parameter.
[0013] Additionally, the step of training the first artificial intelligence model may include the step of training the first artificial intelligence model by updating the first parameter as the updated second parameter is transmitted to the first artificial intelligence model in an exponential moving average manner.
[0014] Additionally, the step of updating the second parameter may include the step of updating the second parameter based on the sum of the first value and the second value.
[0015] Additionally, the first value may include a value calculated by comparing the third mask and the fourth mask, and the second value may include a value calculated by comparing the second mask and the third mask.
[0016] Additionally, the step of generating the second mask and the third mask may include the step of generating the second mask and the third mask as the augmented images of the plurality of computed tomography images are input into the first artificial intelligence model and the third model.
[0017] An apparatus for segmenting computed tomography images using an artificial intelligence model according to one embodiment of the present disclosure comprises: a memory storing at least one program; and a processor that operates by executing the at least one program, wherein the processor acquires a plurality of computed tomography images and outputs a first mask as the plurality of computed tomography images are input into a first artificial intelligence model that has been previously trained, and the first artificial intelligence model may include a second artificial intelligence model that has been previously trained based on a MAR model.
[0018] A computer-readable recording medium according to one embodiment of the present disclosure may include a computer-readable recording medium having a program for executing the above-described method on a computer.
[0019] Other aspects, features, and advantages other than those described above will become clear from the following drawings, claims, and detailed description of the invention.
[0020] According to one embodiment of the present disclosure, as a plurality of computed tomography images are input into an artificial intelligence model, it is possible to provide high accuracy by reinforcing three-dimensional information without an additional 3D network.
[0021] In addition, the Mean-Teacher Framework makes it possible to stably and consistently segment pancreatic tumor regions in computed tomography scans, even in noisy environments.
[0022] Furthermore, as the AI model clearly identifies even small, low-contrast areas of pancreatic tumors, it is possible to precisely segment the tumor area in computed tomography scans.
[0023] In addition, the Masked-attention Mask Transformer makes it possible to effectively segment pancreatic tumor regions of various locations and sizes in computed tomography scan images.
[0024] The effects of the embodiments of the present disclosure are not limited to the effects mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description in this specification.
[0025] FIG. 1 is an example of a method for analyzing computed tomography images using an artificial intelligence model according to one embodiment.
[0026] FIG. 2 is an exemplary diagram illustrating the process of learning a second artificial intelligence model according to one embodiment.
[0027] FIG. 3 is an exemplary diagram illustrating the learning step of a first artificial intelligence model according to one embodiment.
[0028] FIG. 4 is an exemplary drawing for explaining a first artificial intelligence model according to one embodiment.
[0029] FIG. 5 is a block diagram of a device for analyzing computed tomography images using an artificial intelligence model according to one embodiment.
[0030] A method for analyzing computed tomography images using an artificial intelligence model according to one embodiment of the present disclosure comprises: a step of acquiring a plurality of computed tomography images; and a step of outputting a first mask as the plurality of computed tomography images are input into a first artificial intelligence model that has been previously trained; wherein the first artificial intelligence model may include a second artificial intelligence model that has been previously trained based on a MAR model.
[0031] The advantages and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described in detail together with the accompanying drawings. However, the present disclosure is not limited to the embodiments presented below, but can be implemented in various different forms and should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. The embodiments presented below are provided to make the disclosure complete and to fully inform those skilled in the art of the scope of the invention. In describing the present disclosure, detailed descriptions of related prior art are omitted if it is determined that such detailed descriptions may obscure the essence of the present disclosure.
[0032] The terms used herein are used merely to describe specific embodiments and are not intended to limit the disclosure. Unless otherwise defined, all terms used herein have the same meaning as generally understood by those skilled in the art to which this disclosure pertains.
[0033] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0034] Additionally, terms including ordinal numbers, such as "first" or "second" as used herein, may be used to describe various components, but they are used solely for the purpose of distinguishing one component from another. The order and / or importance of the components are not limited by terms including ordinal numbers. For example, a component named as the first component in parts of this specification may be named as the second component without departing from the scope of the technology disclosed herein.
[0035] Phrases such as "in one embodiment," "according to one embodiment," "related to one embodiment," or "according to an implementation of one embodiment" in this specification do not necessarily refer to the same embodiment. Furthermore, throughout this specification, "examples" are arbitrary distinctions to facilitate the description of the present disclosure, and each embodiment does not need to be mutually exclusive. For example, configurations mentioned for the description of one embodiment may be applied and / or implemented in other embodiments, and may be modified and applied and / or implemented to the extent that they do not depart from the scope of the present disclosure.
[0036] Some embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific function.
[0037] For example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Additionally, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing, etc. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations. Furthermore, terms such as "-part," "-module," etc. refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or as a combination of hardware and software.
[0038] Furthermore, the connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or added functional connections, physical connections, or circuit connections.
[0039] The present disclosure will be described in detail below with reference to the attached drawings.
[0040] FIG. 1 is an example of a method for analyzing computed tomography images using an artificial intelligence model according to one embodiment.
[0041] Referring to FIG. 1, a method for analyzing computed tomography images using an artificial intelligence model according to one embodiment may consist of a learning step (110) of the artificial intelligence model and a utilization step (120) of the artificial intelligence model.
[0042] In one embodiment, the learning step (110) of the artificial intelligence model may include a step in which a processor trains a first artificial intelligence model based on the Mean-Teacher Framework. Additionally, the learning step (110) of the artificial intelligence model may include a process in which a second artificial intelligence model is trained based on the MAR model. At this time, the processor may include the processor (520) of FIG. 5. A detailed description of the learning step of the artificial intelligence model will be provided later with reference to FIG. 2 and FIG. 3.
[0043] In one embodiment, the step of utilizing the artificial intelligence model (120) may be a step in which a first mask is output as a plurality of computed tomography images are input into a first artificial intelligence model that has been trained. At this time, the first artificial intelligence model may include a second artificial intelligence model that has been trained based on a MAR model. And, the processor may acquire a plurality of computed tomography images. And, the first mask may include a mask in which the region of the pancreas and the region of the pancreatic tumor are segmented in the computed tomography scan image. A detailed description of the step of utilizing the artificial intelligence model will be described later with reference to FIG. 4.
[0044] In machine learning technology and cognitive science, an artificial intelligence model may refer to a statistical learning algorithm implemented based on the structure of a biological neural network, or a structure that executes such an algorithm.
[0045] For example, an artificial intelligence model can represent a model with problem-solving capabilities by having nodes, which are artificial neurons forming a network through the combination of synapses as in biological neural networks, learn to reduce the error between the correct output corresponding to a specific input and the inferred output by repeatedly adjusting the weights of the synapses. For example, an artificial intelligence model may include arbitrary probability models and artificial intelligence models used in artificial intelligence learning methods such as machine learning and deep learning, as well as CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network).
[0046] For example, an artificial intelligence model can be implemented as a multilayer perceptron (MLP) composed of multiple layers of nodes and connections between them. An artificial intelligence model according to the present embodiment can be implemented using one of various artificial neural network model structures including an MLP. For example, an artificial intelligence model may be composed of an input layer that receives input signals or data from the outside, an output layer that outputs output signals or data corresponding to the input data, and at least one hidden layer located between the input layer and the output layer, which receives a signal from the input layer, extracts a feature, and transmits it to the output layer. The output layer receives signals or data from the hidden layer and outputs them to the outside.
[0047] FIG. 2 is an exemplary diagram illustrating the process of learning a second artificial intelligence model according to one embodiment.
[0048] The processor can input a single computed tomography image (201) into the second artificial intelligence model (200). Additionally, the second artificial intelligence model (200) may include a token generator (210), a Masked Autoencoder (hereinafter MAE) encoder (220), an MAE decoder (230), and a noise remover (240).
[0049] For example, MAE can refer to a method that reconstructs input data by observing only a portion of the input data. Since only a part of the image is input into the MAE, the entire image input into the MAE can be reconstructed. In this case, the MAE may include an encoder and a decoder.
[0050] In the present disclosure, masking may mean obscuring a part of image information.
[0051] For example, a patch of an image may be masked arbitrarily. In this case, the patch may refer to a square block created as the image is divided into square blocks (e.g., 16x16 pixels). Additionally, the patch may be set to an arbitrary size. As the unmasked patch is input into an encoder included in the MAE, an encoded patch may be generated. Then, as the encoded patch and the masked patch are input into a decoder included in the MAE, a restored patch may be generated. And, based on the restored patch, the image input into the MAE may be restored.
[0052] For example, the MAE encoder (220) may be a component of the MAE. Also, the MAE decoder (230) may be a component of the MAE.
[0053] The processor can input a single computed tomography image (201) into a token generator (210). The token generator (210) can convert the single computed tomography image (201) into a continuous vector. Additionally, the token generator (210) can divide the continuous vector into multiple patches. Furthermore, the token generator (210) can tokenize the multiple patches to generate a token (211).
[0054] The processor can mask a portion of the token (211). Accordingly, a masked token (213) and an unmasked token (212) may be generated.
[0055] The processor can input an unmasked token (212) into the MAE encoder (220).
[0056] For example, the MAE encoder (220) can generate a latent vector (221) based on information from an unmasked token (212). In this case, the latent vector (221) may include a vector representing specific information of the unmasked token (212). Here, the specific information may refer to information used for the reconstruction of a single computed tomography image (201). Additionally, the specific information may refer to information containing features that distinguish a single computed tomography image (201) from other images.
[0057] The processor can input input data (222) into the MAE decoder (230). At this time, the input data (222) may include a latent vector (221) and a masked token (213).
[0058] For example, the MAE decoder (230) can generate a recovery value (231) for a token (211) based on input data (222). At this time, the recovery value (231) may include a token recovered based on the token (211) generated by the token generator (210).
[0059] The processor can generate a noise-containing token (214) by adding noise to the token (211). Additionally, the processor can input the noise-containing token (214) into a noise remover (240). At this time, the noise remover (240) may include multiple perceptrons. Furthermore, the processor can input a recovery value (231) into the noise remover (240). At this time, the recovery value (231) may be used as a condition for removing noise from the noise-containing token (214).
[0060] For example, the noise remover (240) can generate a noise-removed token (241) based on a noise-containing token (214). During the process of generating the noise-removed token (241), the processor can calculate the Diffusion Loss. Additionally, the processor can update the parameters of the second artificial intelligence model (200) based on the Diffusion Loss.
[0061] More specifically, the processor can calculate the difference between the first noise predicted by the noise remover (240) and the second noise added to the noise-containing token (214). The processor can calculate the Diffusion Loss based on the difference in noise. Then, the processor can update the parameters of the second artificial intelligence model (200) so that the noise is minimized. At this time, the Diffusion Loss may include a loss function. For example, the loss function may induce the second artificial intelligence model to be trained in a way that restores the input data while minimizing the noise.
[0062] The second artificial intelligence model (200) can be trained so that the MAE encoder (220) accurately extracts the latent vector (221) based on the updated parameters.
[0063] In conclusion, the second artificial intelligence model (200) can be trained based on tokens (241) generated as a single computed tomography image (201) is input into the second artificial intelligence model (200).
[0064] For example, the MAR model may include a token generator (210), an MAE encoder (220), an MAE decoder (230), and a noise remover (240). Accordingly, referring to the description above, the second artificial intelligence model (200) can be trained based on the MAR model.
[0065] FIG. 3 is an exemplary diagram illustrating the learning step of a first artificial intelligence model according to one embodiment.
[0066] Referring to FIG. 3, the first artificial intelligence model (310) may include the Teacher Model in the Mean-Teacher Framework. Additionally, the third model (320) may include the Student Model in the Mean-Teacher Framework.
[0067] For example, the Mean-Teacher Framework may include a structure consisting of a Teacher Model and a Student Model. In this case, the Teacher Model can be optimized as the difference between the prediction results of the Teacher Model and the Student Model decreases.
[0068] For example, the parameters of the Teacher Model can be updated based on the exponential moving average method as the parameters of the Student Model are passed to the Teacher Model.
[0069] For example, the exponential moving average method may include a method that assigns higher weight to recently entered data.
[0070] For example, in the Mean-Teacher Framework, the exponential moving average method may include a method in which the weights of the Teacher Model are updated based on the weights of the exponential moving average and the weights of the Student Model. As the weights of the Teacher Model are updated, the Teacher Model can be updated. In this case, the update speed of the Teacher Model can be controlled as the weights of the exponential moving average are adjusted.
[0071] The processor can generate a second mask (330) and a third mask (340) as a plurality of computed tomography images (300) are input into a first artificial intelligence model (310) and a third model (320). At this time, the second mask (330) may include a mask generated from the first artificial intelligence model (310). Additionally, the third mask (340) may include a mask generated from the third model (320). Furthermore, the second mask (330) and the third mask (340) may include masks for the region of the pancreas and the region of the pancreatic tumor.
[0072] The processor can input augmented images of a plurality of computed tomography images (300) into the first artificial intelligence model (310) and the third model (320). At this time, the augmented images may include images augmented by a jitter augmentation method. Additionally, there may be multiple augmented images. Accordingly, the processor can generate a second mask (330) and a third mask (340) as the augmented images of a plurality of computed tomography images (300) are input into the first artificial intelligence model (310) and the third model (320).
[0073] For example, a jitter enhancement method may include a method in which an image is enhanced as the characteristics of the image change. More specifically, a jitter enhancement method may include a method in which an image is enhanced as the brightness, contrast, and saturation of the image change.
[0074] The processor can update the second parameter of the third model (320) based on the second mask (330) and the third mask (340). At this time, the processor can update the second parameter based on the sum of the first value (360) and the second value (370). More specifically, the processor can update the second parameter so that an appropriate latent vector is generated in the second artificial intelligence model included in the third model (320).
[0075] For example, the first value (360) may include Cross-Entropy Loss. Additionally, the first value (360) may include a value calculated by comparing the third mask (340) and the fourth mask (350). In this case, the fourth mask (350) may include a mask marked manually by a clinician for the pancreas region and the pancreatic tumor region.
[0076] For example, the Cross-Entropy Loss may include a loss function that uses the predicted probability of the third model (320) and the correct probability. Then, the difference between the probability value predicted by the third model (320) and the actual correct value can be measured through the Cross-Entropy Loss. Using the measured difference, the second parameter of the third model (320) can be updated to predict a higher probability for the correct answer.
[0077] For example, the second value (370) may include a Consistency Loss. Additionally, the second value (370) may include a value calculated by comparing the second mask (330) and the third mask (340).
[0078] For example, Consistency Loss may include a loss function that induces the third model (320) to make consistent predictions for the same input data. By using Consistency Loss, the prediction of the third model (320) can be maintained even if the input data is modified. Additionally, by using Consistency Loss, the difference between the prediction value of the first artificial intelligence model (310) and the prediction value of the third model (320) can be minimized.
[0079] The processor can train the first artificial intelligence model (310) by updating the first parameter of the first artificial intelligence model (310) based on the result of updating the second parameter. At this time, the processor can train the first artificial intelligence model (310) by updating the first parameter as the updated second parameter (380) is transmitted to the first artificial intelligence model (310) in an exponential moving average manner.
[0080] More specifically, the first parameter may be updated so that a latent vector expressing specific information is generated in the second artificial intelligence model included in the first artificial intelligence model (310). Here, the specific information may refer to information used for the reconstruction of a single input computed tomography image. Additionally, the specific information may refer to information containing features that distinguish the single input computed tomography image from other images.
[0081] FIG. 4 is an exemplary drawing for explaining a first artificial intelligence model according to one embodiment.
[0082] For example, the first artificial intelligence model (400) may be configured with the same model structure as the third model. Therefore, the third model can also be understood by referring to FIG. 4. In addition, the following description regarding the first artificial intelligence model (400) may be applied equally to the third model.
[0083] For example, the first artificial intelligence model (400) may include a second artificial intelligence model (410). In other words, the first artificial intelligence model (400) may include a second artificial intelligence model (410) that has been previously trained based on a MAR model.
[0084] For example, the first artificial intelligence model (400) may include a decoder (440). The decoder (440) may include a Masked-attention Mask Transformer.
[0085] For example, a Masked-attention Mask Transformer can refer to a transformer-based model for general-purpose image segmentation. A Masked-attention Mask Transformer may include a Backbone, a Pixel Decoder, and a Transformer Decoder.
[0086] For example, Backbone can extract low-resolution features from images.
[0087] For example, a pixel decoder can convert low-resolution features extracted from the backbone into high-resolution features.
[0088] For example, the transducer decoder can generate a mask based on the Masked Attention technique. In this case, the Masked Attention technique may include a method that is trained limited to the mask region predicted by the query. Additionally, the transducer decoder can perform Cross-Attention and Self-Attention.
[0089] For example, Attention may refer to a mechanism in which weights are assigned to specific parts of input data. In this case, the specific parts may be determined by a learned first artificial intelligence model (400).
[0090] For example, Cross-Attention can refer to a mechanism in which one data is learned by referencing a specific part of another data. In this case, the specific part can be determined by the learned first artificial intelligence model (400).
[0091] For example, Self-Attention can identify the relationships between tokens included in the input data. Also, Self-Attention can identify the associations between tokens included in the input data. Furthermore, Self-Attention can refer to a mechanism that efficiently conveys information by focusing on important parts included in the relationships and associations between tokens. At this time, specific parts can be determined by the learned first artificial intelligence model (400).
[0092] The processor can acquire a plurality of computed tomography images (401). The processor can input the plurality of computed tomography images (401) into a first artificial intelligence model (400). More specifically, the plurality of computed tomography images (401) can be input into a second artificial intelligence model (410) included in the first artificial intelligence model (400). The second artificial intelligence model (410) may include a token generator (420) and an MAE encoder (430). At this time, the token generator (420) and the MAE encoder (430) may be initialized by parameters of the second artificial intelligence model (410) that have been previously trained.
[0093] For example, a token generator (420) can generate a token (421). Then, a processor can input the token (421) into an MAE encoder (430). Then, the MAE encoder (430) can generate a latent vector (431). Then, a processor can input the latent vector (431) into a decoder (440). Then, the decoder (440) can generate a first mask (450). At this time, the decoder (440) can be trained based on the process of training the first artificial intelligence model (400). At this time, the first mask (450) may include a mask for the region of the pancreas and the region of the pancreatic tumor.
[0094] Accordingly, the processor can output a first mask (450) as a plurality of computed tomography images (401) are input into the first artificial intelligence model (400).
[0095] FIG. 5 is a block diagram of a device for analyzing computed tomography images using an artificial intelligence model according to one embodiment.
[0096] Referring to FIG. 5, the device (500) may include a memory (510) and a processor (520). Only the components related to the embodiment are shown in the device (500) of FIG. 5. Therefore, a person skilled in the art will understand that the device (500) may include other general-purpose components in addition to the components shown in FIG. 5.
[0097] The memory (510) is hardware that stores various data processed within the device (500) and can store programs for various operations, processing, and control of the processor (520).
[0098] The memory (510) may include RAM (Random Access Memory), such as DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), CD-ROM, Blu-ray or other optical disc storage, HDD (hard disk drive), SSD (solid state drive), or flash memory.
[0099] The processor (520) controls the overall operation of the device (500). For example, the processor (520) can control the memory (510), an input unit (not shown) and / or an output unit (not shown), etc., by executing programs stored in the memory (510). The processor (520) can control the operation of the device (500) by executing at least one program stored in the memory (510).
[0100] The processor (520) can control at least some of the operations of the device (500) described above with reference to FIGS. 1 to 4. For example, the processor (520) can acquire a plurality of computed tomography images and output a first mask as the plurality of computed tomography images are input into a first artificial intelligence model that has been trained.
[0101] Meanwhile, a specific example of the operation of the processor (520) is the same as described above with reference to FIGS. 1 to 4. Therefore, a specific description of the operation of the processor (520) is omitted below.
[0102] The processor (520) may be implemented using at least one of ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), controllers, microcontrollers, microprocessors, and other electrical units for performing functions.
[0103] In one embodiment, the device (500) may be a server device. The server device may be implemented as a computer device or a plurality of computer devices that communicate through a network to provide commands, code, files, content, services, etc.
[0104] In another embodiment, the device (500) may be a mobile electronic device. For example, the device (500) may be implemented as a smartphone, tablet PC, PC, smart TV, laptop, a device equipped with a camera, and other mobile electronic devices.
[0105] In another embodiment, the process performed in the device (500) may be performed by at least some of the server device and the electronic device having mobility.
[0106] Meanwhile, embodiments according to the present disclosure may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. In this case, the medium may include, but is not limited to, magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROM, RAM, and flash memory.
[0107] Meanwhile, the above computer program may be one specifically designed and configured for the present disclosure or one known and available to those skilled in the art of computer software. Examples of computer programs may include machine code, such as that produced by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0108] According to one embodiment, the method according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0109] Unless explicitly stated otherwise, the steps constituting the method according to the present disclosure may be performed in a suitable order. The present disclosure is not necessarily limited by the order in which the steps are described. The use of all examples or exemplary terms (e.g., etc.) in the present disclosure is merely for the purpose of describing the present disclosure in detail and, unless limited by the claims, the scope of the present disclosure is not limited by such examples or exemplary terms. Furthermore, those skilled in the art will understand that various modifications, combinations, and changes may be made according to design conditions and factors within the scope of the claims or equivalents to which they are added.
[0110] Accordingly, the scope of the present disclosure should not be limited to the embodiments described above, and all scopes equivalent to or equivalently modified from the claims set forth below, as well as the claims set forth below, shall be considered to fall within the scope of the scope of the present disclosure.
Claims
1. In a method for analyzing computed tomography images using an artificial intelligence model, A step of acquiring a plurality of computed tomography images; and A step of outputting a first mask as a plurality of computed tomography images are input into a first artificial intelligence model that has been previously trained; Includes, The above-mentioned first artificial intelligence model is, A method comprising a second artificial intelligence model that has been trained based on a MAR model.
2. In Paragraph 1, The above-mentioned first artificial intelligence model is, Includes a decoder, A method in which the above decoder includes a Masked-attention Mask Transformer.
3. In Paragraph 1, The above second artificial intelligence model is, A method of learning based on tokens generated as a single computed tomography image is input into the second artificial intelligence model.
4. In Paragraph 1, The above method is, A step of generating a second mask and a third mask as a plurality of computed tomography images are input into the first artificial intelligence model and the third model; A step of updating the second parameter of the third model based on the second mask and the third mask; and A step of training the first artificial intelligence model by updating the first parameter of the first artificial intelligence model based on the update result of the second parameter; A method that further includes.
5. In Paragraph 4, The step of training the above-mentioned first artificial intelligence model is, A step of training the first artificial intelligence model by updating the first parameter as the updated second parameter is transmitted to the first artificial intelligence model in an exponential moving average manner; A method including 6. In Paragraph 4, The step of updating the second parameter above is, A step of updating the second parameter based on the sum of the first value and the second value; A method including 7. In Paragraph 6, The above first value is, It includes a value calculated by comparing the third mask and the fourth mask, and The above second value is, A method comprising a value calculated by comparing the second mask and the third mask.
8. In Paragraph 4, The step of generating the second mask and the third mask is, A step of generating a second mask and a third mask as the augmented images of the plurality of computed tomography images are input into the first artificial intelligence model and the third model; A method including 9. Memory in which at least one program is stored; and A processor that operates by executing at least one of the above programs; Includes, The above processor is, Acquiring multiple computed tomography images, and outputting a first mask as the multiple computed tomography images are input into a pre-trained first artificial intelligence model, and The above-mentioned first artificial intelligence model is, A device for segmenting computed tomography images using an artificial intelligence model, comprising a second artificial intelligence model that has been trained based on a MAR model.
10. A computer-readable recording medium storing a program for executing the method according to paragraph 1 on a computer.