Method and device for training artificial intelligence model for image compatibility of different oct equipment
The AI model enhances OCT device compatibility by converting images with different resolutions and algorithms, improving diagnostic precision for glaucoma by accurately measuring retinal nerve fiber layer thickness.
Patent Information
- Application Number
- PCT/KR2025/000336
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-01-02
- Filing Date
- 2025-01-07
- Publication Date
- 2025-07-17
AI Technical Summary
Existing OCT devices provide inconsistent diagnostic results due to differences in image resolution and thickness measurement algorithms, hindering compatibility between different devices and complicating disease diagnosis and monitoring.
A method and device for training an artificial intelligence model to convert images between different OCT devices using a loss function targeting the retinal nerve fiber layer boundary, employing generators and discriminators to enhance compatibility and accuracy.
Improves compatibility and accuracy of image conversion between OCT devices, enabling precise identification of retinal nerve fiber layer thickness for better glaucoma diagnosis.
Smart Images

Figure KR2025000336_17072025_PF_FP_ABST
Abstract
Description
Method and device for training an artificial intelligence model for image compatibility of different OCT devices
[0001] The present invention relates to a method and device for training an artificial intelligence model for image compatibility of different OCT devices.
[0002] This study was conducted in connection with a research project (No. 14-2022-0024) supported by Seoul National University Bundang Hospital with funding from Seoul National University Bundang Hospital in 2022.
[0003] For reference, this application claims priority to Korean Patent Application No. 10-2025-0000305, filed January 2, 2025. The entire contents of that application, which serves as the basis for this priority claim, are incorporated herein by reference.
[0004] OCT (optical coherence tomography) equipment is a device that obtains cross-sectional images of tissues by utilizing the interference phenomenon of light, and is widely used to observe structural abnormalities in the optic nerve and retina.
[0005] In particular, quantitative measurement of the thickness of the retinal nerve fiber layer around the optic nerve head using OCT is an essential test for the diagnosis and follow-up of glaucoma.
[0006] However, there is a limitation in that the image resolution and algorithm for measuring the retinal nerve fiber layer thickness are different for each OCT device, so there are frequent cases where different diagnostic results are shown for each OCT device.
[0007] These equipment compatibility issues make communication between hospitals with different equipment difficult, which acts as a factor in the power system failure and has a serious impact not only on the diagnosis of the disease but also on the monitoring of its progress.
[0008] Accordingly, there is a need to develop technologies that can improve compatibility between different devices.
[0009] (Prior art literature)
[0010] (Patent Document 0001) Korean Patent Publication No. 10-2023-0152647 (November 3, 2023)
[0011] The problem to be solved by the present invention is to increase compatibility between equipment having different image resolutions and thickness measurement algorithms for the retinal nerve fiber layer.
[0012] In addition, the problem to be solved by the present invention is to train an artificial intelligence model to be able to generate an eye image with layer dividing lines by receiving an eye image without layer dividing lines as input.
[0013] However, the problems to be solved by the present invention are not limited to those mentioned above, and other problems to be solved that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from the description below.
[0014] A method for training an artificial intelligence model for image compatibility of different OCT devices according to one embodiment of the present invention comprises the steps of: obtaining a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device; and training the artificial intelligence model to enable mutual conversion between the first eye images and the second eye images, wherein, in the learning process of the artificial intelligence model, a loss function designed to target a boundary of a retinal nerve fiber layer included in each eye image may be used.
[0015] Here, the learning dataset may include each of the plurality of second eye images having a corresponding relationship with each of the plurality of first eye images, and the first eye image and the second eye image having the corresponding relationship may be characterized in that they were captured for the same eye.
[0016] In addition, the step of obtaining the learning dataset may further include the step of obtaining a first segmentation mask label corresponding to the retinal nerve fiber layer from each of the first eye images using a pre-learned segmentation model; and the step of obtaining a second segmentation mask label corresponding to the retinal nerve fiber layer from each of the second eye images using the pre-learned segmentation model.
[0017] Meanwhile, the artificial intelligence model may include a first generator for generating a first-second eye image corresponding to a second domain for the second eye image from the first eye image; a second generator for generating a second-first eye image corresponding to a first domain for the first eye image from the second eye image; a first discriminator for identifying a difference between the first eye image and the second-first eye image; and a second discriminator for identifying a difference between the second eye image and the first-second eye image.
[0018] Here, the step of training the artificial intelligence model may include a step of first training the artificial intelligence model using the training data set, and the step of first training the artificial intelligence model may include a step of inputting the second eye image to the second generator to output the 2-1 eye image; a step of training the first discriminator using a first loss function determined based on a difference between the first eye image and the 2-1 eye image; and a step of training the second generator using a third loss function determined based on a difference between the 2-1-2 eye image and the second eye image output by inputting the 2-1 eye image to the first generator.
[0019] In addition, the step of first training the artificial intelligence model may include a step of inputting the first eye image to the first generator and outputting the 1-2 eye image; a step of training the second discriminator using a second loss function determined based on the difference between the second eye image and the 1-2 eye image; and a step of training the first generator using the third loss function determined based on the difference between the 1-2-1 eye image output by inputting the 1-2 eye image to the second generator and the first eye image.
[0020] Meanwhile, the step of training the artificial intelligence model may include a step of training the artificial intelligence model a second time using the fourth loss function determined based on the learning dataset and the thickness of the retinal nerve fiber layer.
[0021] Here, the step of training the artificial intelligence model for the second time may include a step of training the first generator to input the first eye image and output a first segmentation mask with reference to the first segmentation mask label; and a step of training the first generator to input the first segmentation mask and output a second segmentation mask and the second eye image with reference to the second segmentation mask label.
[0022] In addition, the step of training the artificial intelligence model for the second time may include a step of training the second generator to input the second eye image and output the second segmentation mask with reference to the second segmentation mask label; and a step of training the second generator to input the second segmentation mask and output the first segmentation mask and the first eye image with reference to the first segmentation mask label.
[0023] According to another embodiment of the present invention, a device for training an artificial intelligence model for image compatibility of different OCT devices includes: a memory in which an artificial intelligence model training program is stored; and a processor for loading the artificial intelligence model training program from the memory and executing the artificial intelligence model training program, wherein the processor obtains a training dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device, and trains the artificial intelligence model to enable mutual conversion between the first eye images and the second eye images, wherein, in the training process of the artificial intelligence model, a loss function designed to target a boundary of a retinal nerve fiber layer included in each eye image may be used.
[0024] Here, the learning dataset may include each of the plurality of second eye images having a corresponding relationship with each of the plurality of first eye images, and the first eye image and the second eye image having the corresponding relationship may be characterized in that they were captured for the same eye.
[0025] In addition, the processor can obtain a first segmentation mask label corresponding to the retinal nerve fiber layer from each of the first eye images using the pre-learned segmentation model, and can obtain a second segmentation mask label corresponding to the retinal nerve fiber layer from each of the second eye images using the pre-learned segmentation model.
[0026] Meanwhile, the artificial intelligence model may include a first generator for generating a first-second eye image corresponding to a second domain for the second eye image from the first eye image; a second generator for generating a second-first eye image corresponding to a first domain for the first eye image from the second eye image; a first discriminator for identifying a difference between the first eye image and the second-first eye image; and a second discriminator for identifying a difference between the second eye image and the first-second eye image.
[0027] Here, the processor may train the artificial intelligence model for the first time using the learning data set, wherein the processor may input the second eye image to the second generator to output the 2-1 eye image, train the first discriminator using a first loss function determined based on the difference between the first eye image and the 2-1 eye image, and train the second generator using a third loss function determined based on the difference between the 2-1-2 eye image and the second eye image output by inputting the 2-1 eye image to the first generator.
[0028] In addition, the processor may input the first eye image to the first generator to output the 1-2 eye image, train the second discriminator using the second loss function determined based on the difference between the second eye image and the 1-2 eye image, and train the first generator using the third loss function determined based on the difference between the 1-2-1 eye image and the first eye image output by inputting the 1-2 eye image to the second generator.
[0029] Meanwhile, the processor can perform secondary training on the artificial intelligence model using the fourth loss function determined based on the learning dataset and the thickness of the retinal nerve fiber layer.
[0030] Here, the processor may train the first generator to input the first eye image and output the first segmentation mask with reference to the first segmentation mask label, and may train the first generator to input the first segmentation mask and output the second segmentation mask and the second eye image with reference to the second segmentation mask label.
[0031] In addition, the processor may train the second generator to input the second eye image and output the second segmentation mask with reference to the second segmentation mask label, and may train the second generator to input the second segmentation mask and output the first segmentation mask and the first eye image with reference to the first segmentation mask label.
[0032] According to another embodiment of the present invention, a device for converting an OCT image includes: a memory in which an OCT image conversion program is stored; and a processor for loading the OCT image conversion program from the memory and executing the OCT image conversion program, wherein the processor obtains a first eye image captured by a first OCT device, and, when provided with the image captured by the first OCT device, provides the obtained first eye image to an artificial intelligence model that has been trained to output an image of a quality similar to a predetermined level to an image captured by a second OCT device having a different quality of the capture result from the first OCT device, and obtains a second eye image corresponding to the first eye image from the artificial intelligence model, wherein, in the learning process of the artificial intelligence model, a loss function designed to target a boundary of a retinal nerve fiber layer included in the eye image may be used.
[0033] A non-transitory computer-readable recording medium storing a computer program according to another embodiment of the present invention comprises the steps of: obtaining a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device; and training an artificial intelligence model to enable mutual conversion between the first eye images and the second eye images, wherein, in the learning process of the artificial intelligence model, the processor may include instructions for performing an artificial intelligence model learning method in which a loss function designed to target a boundary of a retinal nerve fiber layer included in each eye image is used.
[0034] According to an embodiment of the present invention, by training an artificial intelligence model using a loss function designed to target the boundary of the retinal nerve fiber layer included in an eye image, compatibility between different first OCT devices and second OCT devices and conversion accuracy between the first eye image and the second eye image can be improved.
[0035] In addition, according to an embodiment of the present invention, by converting a first eye image into a second eye image using a pre-learned artificial intelligence model, the degree of loss of the thickness of the retinal nerve fiber layer around the optic nerve head can be accurately determined, and a doctor (or user) can diagnose glaucomatous damage more precisely.
[0036] FIG. 1 is a block diagram showing a device for training an artificial intelligence model for image compatibility of different OCT devices according to an embodiment of the present invention.
[0037] Figure 2 is a block diagram conceptually illustrating the function of an artificial intelligence model learning program according to an embodiment of the present invention.
[0038] FIG. 3 is a flowchart illustrating a method for training an artificial intelligence model for image compatibility of different OCT devices according to one embodiment of the present invention.
[0039] FIG. 4 is a diagram illustrating the first learning of an artificial intelligence model according to one embodiment of the present invention.
[0040] FIG. 5 is a diagram exemplarily showing secondary learning of the artificial intelligence model according to one embodiment of the present invention.
[0041] FIG. 6 is a diagram exemplarily showing the results of performing OCT image conversion using an artificial intelligence model learned according to one embodiment of the present invention.
[0042] FIG. 7 is a diagram illustrating an example of comparing the results of performing OCT image conversion of a first-learned artificial intelligence model and a second-learned artificial intelligence model according to one embodiment of the present invention.
[0043] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined solely by the scope of the claims.
[0044] When describing embodiments of the present invention, detailed descriptions of known functions or configurations will be omitted if they are deemed to unnecessarily obscure the gist of the invention. Furthermore, the terms described below are defined in light of their functions in the embodiments of the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the overall content of this specification.
[0045] The terms 'unit', 'unit', etc. used below mean a unit that processes at least one function or operation, and this can be implemented by hardware, software, or a combination of hardware and software.
[0046] FIG. 1 is a block diagram showing a device for training an artificial intelligence model for image compatibility of different OCT devices according to an embodiment of the present invention.
[0047] Referring to FIG. 1, the device (100) may include a processor (110), an input / output device (120), and a memory (130).
[0048] The processor (110) can control the overall operation of the device (100).
[0049] The processor (110) can receive eye images of multiple patients captured by different OCT devices using the input / output device (120). Here, the patient's eye images are used to diagnose glaucomatous damage and can be captured to include the retinal nerve fiber layer, which is the uppermost layer around the optic nerve head.
[0050] In the present invention, it has been described that eye images of multiple patients captured by different OCT devices are input through the input / output device (120), but this is not limited thereto. That is, according to an embodiment, the device (100) may include a transceiver (not shown), and the device (100) may receive eye images of multiple patients captured by different OCT devices using the transceiver (not shown), and the eye images of multiple patients captured by different OCT devices may be generated within the device (100).
[0051] The processor (110) can obtain a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device, and train the artificial intelligence model to enable mutual conversion between the first eye images and the second eye images.
[0052] The input / output device (120) may include one or more input devices and / or one or more output devices. For example, the input devices may include a microphone, a keyboard, a mouse, a touch screen, etc., and the output devices may include a display, a speaker, etc.
[0053] The memory (130) can store an artificial intelligence model learning program (200) and information required for executing the artificial intelligence model learning program (200).
[0054] In this specification, an artificial intelligence model learning program (200) may mean software that includes commands for acquiring a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device, and for learning the artificial intelligence model to enable mutual conversion between the first eye images and the second eye images.
[0055] The processor (110) can load the artificial intelligence model learning program (200) and information necessary for executing the artificial intelligence model learning program (200) from the memory (130) to execute the artificial intelligence model learning program (200).
[0056] The processor (110) can execute an artificial intelligence model learning program (200) to input a first eye image captured by a first OCT device and a second eye image captured by a second OCT device, and train an artificial intelligence model to enable mutual conversion between the first eye image and the second eye image.
[0057] In the present invention, the first OCT device according to one embodiment may mean a Cirrus SD-OCT device from Zeiss, and the first eye image may mean an image without a dividing line for a layer around the optic nerve head.
[0058] Additionally, in the present invention, the second OCT device according to one embodiment may mean a Spectralis SD-OCT device from Heidelberg, and the second eye image may mean an image having a dividing line for a layer around the optic nerve head.
[0059] Meanwhile, the first OCT device, the second OCT device, the first eye image, and the second eye image are merely examples and may be variously changed within a range that can achieve the purpose of the present invention to improve compatibility between different OCT devices.
[0060]
[0061] The functions and / or operations of the artificial intelligence model learning program (200) will be examined in detail with reference to Fig. 2.
[0062] Figure 2 is a block diagram conceptually illustrating the function of an artificial intelligence model learning program according to an embodiment of the present invention.
[0063] Referring to FIG. 2, the artificial intelligence model learning program (200) may include a dataset acquisition unit (210) and a model learning unit (220).
[0064] The dataset acquisition unit (210) and model learning unit (220) illustrated in FIG. 2 conceptually divide the functions of the artificial intelligence model learning program (200) to easily explain the functions of the artificial intelligence model learning program (200), but are not limited thereto. According to embodiments, the functions of the dataset acquisition unit (210) and model learning unit (220) can be merged / separated, and can also be implemented as a series of commands included in a single program.
[0065] First, the dataset acquisition unit (210) can acquire a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device.
[0066] For example, the training dataset may include each of a plurality of first eye images (e.g., 244 images captured using an OCT device from Zeiss) and each of a plurality of second eye images (e.g., 570 images captured using an OCT device from Heidelberg) that have a corresponding relationship.
[0067] Here, according to one embodiment, the first eye image and the second eye image having a corresponding relationship may be characterized as having been captured for the same eye. However, the corresponding relationship is not limited to a pair-wise relationship.
[0068] Meanwhile, the dataset acquisition unit (210) can obtain a first segmentation mask label corresponding to the retinal nerve fiber layer from each of the first eye images using a pre-learned segmentation model (e.g., U-net, Y-net, GCU-net, etc.).
[0069] In addition, the dataset acquisition unit (210) can obtain a second segmentation mask label corresponding to the retinal nerve fiber layer from each of the second eye images using a pre-learned segmentation model (e.g., U-net, Y-net, GCU-net, etc.).
[0070] According to one embodiment, the first segmentation mask label and the second segmentation mask label may mean the boundary of the retinal nerve fiber layer detected using the learned segmentation model, and may mean a pseudo label rather than a real label.
[0071] Meanwhile, the first segmentation mask label and the second segmentation mask label can be used for secondary learning of the artificial intelligence model, and the secondary learning process of the artificial intelligence model will be described later.
[0072] Next, the model learning unit (220) can train an artificial intelligence model to enable mutual conversion between the first eye image and the second eye image. Here, the artificial intelligence model is an artificial intelligence model that enables domain conversion between the first eye image and the second eye image, and may refer to a generative model (e.g., cycleGAN).
[0073] However, the above cycleGAN is only an example, and the artificial intelligence model can be changed in various ways within the scope that can achieve the purpose of the present invention.
[0074] Specifically, the artificial intelligence model may include a first generator for generating first-second eye images corresponding to a second domain for a second eye image from a first eye image. Here, according to one embodiment, the second domain may refer to a domain (e.g., resolution, quality, presence of layer dividing lines, etc.) corresponding to an image captured using a second OCT device.
[0075] Additionally, the AI model may include a second generator for generating a second-first eye image corresponding to a first domain for the first eye image from the second eye image. Here, according to one embodiment, the first domain may refer to a domain (e.g., resolution, quality, presence of layer dividing lines, etc.) corresponding to an image captured using the first OCT device.
[0076] Additionally, the artificial intelligence model may include a first discriminator for identifying differences between the first eye image and the second-first eye image.
[0077] Additionally, the artificial intelligence model may include a second discriminator for identifying differences between the second eye image and the first-second eye images.
[0078] Meanwhile, the model learning unit (220) can perform initial training of the artificial intelligence model using the learning dataset.
[0079] Specifically, the model learning unit (220) can input the second eye image into the second generator and output the second-1 eye image.
[0080] Additionally, the model learning unit (220) can learn the first discriminator using the first loss function determined based on the difference between the first eye image and the second-first eye image.
[0081] Here, the first loss function according to one embodiment can be expressed as in the following mathematical expression 1.
[0082]
[0083]
[0084]
[0085] Here, may mean a second constructor, may denote a first discriminator, z may denote a first eye image, and s may denote a second eye image.
[0086] Meanwhile, the model learning unit (220) can input the first eye image into the first generator and output the first-second eye image.
[0087] Additionally, the model learning unit (220) can learn the second discriminator using a second loss function determined based on the difference between the second eye image and the first-second eye images.
[0088] Here, the second loss function according to one embodiment can be expressed as in the following mathematical expression 2.
[0089]
[0090]
[0091]
[0092] Here, can mean the first constructor, may mean the first discriminator.
[0093] Meanwhile, the model learning unit (220) can train the second generator using the third loss function determined based on the difference between the second eye image and the second eye image output by inputting the second eye image to the first generator.
[0094] Additionally, the model learning unit (220) can train the first generator using a third loss function determined based on the difference between the first eye image and the first eye image output by inputting the first-second eye image to the second generator.
[0095] Here, the third loss function according to one embodiment can be expressed as in the following mathematical expression 3.
[0096]
[0097]
[0098]
[0099] According to one embodiment, the model learning unit (220) can determine a first generator and a second generator that enable mutual transformation between the first eye image and the second eye image by updating the parameters of the artificial intelligence model through backpropagation so as to maximize or minimize the first loss function determined by weighting the first loss function, the second loss function, and the third loss function.
[0100] Here, the first loss function according to one embodiment can be expressed as in the following mathematical expression 4.
[0101]
[0102]
[0103]
[0104] Here, may mean a hyperparameter representing a weight for the third loss function.
[0105] This allows for improved compatibility between different first and second OCT devices, even if the first and second eye images included in the learning dataset are not exactly paired.
[0106] Next, the model learning unit (220) can perform secondary learning of the artificial intelligence model using the fourth loss function determined based on the learning dataset and the thickness of the retinal nerve fiber layer.
[0107] Here, the fourth loss function may mean a loss function designed to target the boundary of the retinal nerve fiber layer included in each eye image.
[0108] For example, the fourth loss function may mean a curve-shaped transformation similarity determined based on the thickness and boundary of the retinal nerve fiber layer included in the first eye image and the second eye image.
[0109] Specifically, the fourth loss function can be determined by weighting the correct loss determined based on the difference between the segmentation mask label including the curve for layer division of the retinal nerve fiber layer and the segmentation mask output by the generator of the artificial intelligence model, the smooth loss for layer division in the form of a curve and smoothly connecting the curves, and the adversary loss for learning the discriminator.
[0110] The fourth loss function according to one embodiment can be expressed as in the following mathematical expression 5.
[0111]
[0112]
[0113]
[0114] Here, R may mean a pre-trained segmentation model, G may mean a generator of an artificial intelligence model, D may mean a discriminator of an artificial intelligence model, and x may mean an eye image. may mean a split mask label.
[0115] In this way, the model learning unit (220) trains the artificial intelligence model a second time to minimize or maximize the fourth loss function, thereby improving the layer distinction accuracy for the retinal nerve fiber layer, thereby enabling appropriate mutual conversion between the first eye image and the second eye image.
[0116] Meanwhile, there was a limitation in that it was difficult to accurately detect a segmentation mask regarding the thickness and boundary of the retinal nerve fiber layer using a basic boundary detection algorithm such as the existing canny filter. However, in the present invention, in order to overcome this limitation, the first segmentation mask label and the second segmentation mask label, which are pseudo-labels detected through a pre-learned segmentation model, can be used.
[0117] Specifically, the model learning unit (220) can train the first generator to input a first eye image and output a first segmentation mask by referring to the first segmentation mask label.
[0118] Additionally, the model learning unit (220) can train the first generator to input the first segmentation mask and output the second segmentation mask or the second eye image by referring to the second segmentation mask label.
[0119] Here, the first segmentation mask and the second segmentation mask according to one embodiment may include a layer dividing line for the retinal nerve fiber layer.
[0120] Additionally, the model learning unit (220) can train the second generator to input a second eye image and output a second segmentation mask by referring to the second segmentation mask label.
[0121] Additionally, the model learning unit (220) can train the second generator to input the second segmentation mask by referring to the first segmentation mask label and output the first segmentation mask or the first eye image.
[0122] Through this, a unique effect can be achieved that improves compatibility between different first OCT devices and second OCT devices, and the accuracy of conversion between first eye images and second eye images.
[0123]
[0124] Meanwhile, according to another embodiment of the present invention, a processor (not shown) can control the overall operation of an OCT conversion device (not shown).
[0125] A processor (not shown) obtains a first eye image captured by a first OCT device, and when provided with the image captured by the first OCT device, provides the acquired first eye image to an artificial intelligence model that has been trained to output an image of a quality similar to that captured by a second OCT device whose quality of the capture result is different from that of the first OCT device, and obtains a second eye image corresponding to the first eye image from the artificial intelligence model.
[0126] The memory (not shown) can store an OCT image conversion program (not shown) and information required for executing the OCT image conversion program (not shown).
[0127] An OCT image conversion program (not shown) may mean software including commands for obtaining a first eye image captured by a first OCT device, providing the acquired first eye image to an artificial intelligence model that has been trained to output an image of a quality similar to that captured by a second OCT device whose quality of the captured result is different from that of the first OCT device, and obtaining a second eye image corresponding to the first eye image from the artificial intelligence model.
[0128] A processor (not shown) can load an OCT image conversion program (not shown) and information necessary for executing the OCT image conversion program (not shown) from a memory (not shown) to execute the OCT image conversion program (not shown).
[0129] A processor (not shown) executes an OCT image conversion program (not shown) to obtain a first eye image captured by a first OCT device, and when provided with an image captured by the first OCT device, provides the obtained first eye image to an artificial intelligence model that has been trained to output an image of a quality similar to that captured by a second OCT device having a different quality of capture result from the first OCT device, and obtains a second eye image corresponding to the first eye image from the artificial intelligence model, thereby converting the first eye image into a second eye image.
[0130] In addition, it may be characterized in that, in the learning process of the artificial intelligence model, a loss function designed to target the boundary of the retinal nerve fiber layer included in the eye image is used.
[0131] Here, the first eye image according to one embodiment may mean an image that does not include a layer dividing line for the retinal nerve fiber layer, and the second eye image may mean an image that includes a layer dividing line for the retinal nerve fiber layer.
[0132] In this way, by converting the first eye image into a second eye image using the above-mentioned learned artificial intelligence model, the degree of loss of the thickness of the retinal nerve fiber layer around the optic nerve head can be accurately determined, and a unique effect can be achieved in which the doctor (or user) can precisely diagnose glaucomatous damage.
[0133]
[0134] FIG. 3 is a flowchart illustrating a method for training an artificial intelligence model for image compatibility of different OCT devices according to one embodiment of the present invention.
[0135] Referring to FIG. 3, the dataset acquisition unit (210) can acquire a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device (S310).
[0136] Next, the model learning unit (220) can train an artificial intelligence model to enable mutual conversion between the first eye image and the second eye image (S320).
[0137] Here, in the learning process of the artificial intelligence model, a loss function designed to target the boundary of the retinal nerve fiber layer included in each eye image can be used.
[0138]
[0139] FIG. 4 is a diagram illustrating the first learning of an artificial intelligence model according to one embodiment of the present invention.
[0140] Referring to FIG. 4, the model learning unit (220) generates a first eye image (z) using a first generator ( ) to input the 1st-2nd eye image ( ) can be printed.
[0141] Next, the model learning unit (220) learns the second eye image (s) and the first-second eye image ( ) to identify the difference between the two. ) can be taught.
[0142] Next, the model learning unit (220) learns the 1st and 2nd eye images ( ) as the second constructor ( ) and output the 1-2-1 eye image ( ) and the first eye image (z), so as to minimize the loss function determined based on the difference between the first generator ( ) can be taught.
[0143] Likewise, the model learning unit (220) trains the second eye image (s) on the second generator ( ) and enter the 2-1 eye image ( ) can be printed.
[0144] Next, the model learning unit (220) learns the first eye image (z) and the second-first eye image ( ) to identify the difference between the first discriminator ( ) can be taught.
[0145] Next, the model learning unit (220) learns the 2-1 eye image ( ) as the first constructor( ) and output the 2nd-1-2 eye image ( ) and the second eye image(s) to minimize the loss function determined based on the difference between the second generator( ) can be taught.
[0146]
[0147] FIG. 5 is a diagram exemplarily showing secondary learning of the artificial intelligence model according to one embodiment of the present invention.
[0148] Referring to FIG. 5, the model learning unit (220) can train a generator (i.e., corresponding to the first generator) to receive a first segmentation mask as input and output a second segmentation mask.
[0149] Specifically, the model learning unit (220) can input the first segmentation mask into a feature extractor to extract features related to the first segmentation mask, and can train a second segmentation mask label classifier to minimize the difference between the features and the second segmentation mask label. In addition, the model learning unit (220) can train a generator to output the second segmentation mask by passing the features related to the first segmentation mask and the features extracted through the second segmentation mask classifier through a plurality of UNet blocks and upscaling them.
[0150] Meanwhile, based on the second segmentation mask output through the above-mentioned learned generator, the first eye image can be transformed to have a similar quality to the second eye image.
[0151]
[0152] FIG. 6 is a diagram exemplarily showing the results of performing OCT image conversion using an artificial intelligence model learned according to one embodiment of the present invention.
[0153] FIG. 6 illustrates a first eye image (601) with low resolution and no layer separation lines, a second eye image (602) with high resolution and layer separation lines, and a first eye image (611) converted using a pre-trained artificial intelligence model.
[0154] Here, the artificial intelligence model is learned based on a loss function designed to target the boundary of the retinal nerve fiber layer included in each eye image.
[0155] Looking at the converted first eye image (611) of FIG. 6, it can be seen that the noise of the first eye image (601) has been reduced, the overall resolution has been improved, and the curve that separates the retinal nerve fiber layer has been well created.
[0156]
[0157] FIG. 7 is a diagram illustrating an example of comparing the results of performing OCT image conversion of a first-learned artificial intelligence model and a second-learned artificial intelligence model according to one embodiment of the present invention.
[0158] FIG. 7 shows a first eye image (701) and its enlarged portion (711) converted using a first-learned artificial intelligence model, and a first eye image (702) and its enlarged portion (712) converted using a second-learned artificial intelligence model.
[0159] At this time, the first eye images (701, 702) are superimposed to show the results of each artificial intelligence model learning five times.
[0160] Referring to the enlarged parts (711, 712) in Fig. 7, it can be confirmed that when the artificial intelligence model is trained a second time using a loss function designed to target the boundary of the retinal nerve fiber layer included in the eye image, the error between the layer dividing lines for the retinal nerve fiber layer is small.
[0161] In this way, as the layer distinction accuracy for the retinal nerve fiber layer is improved, mutual conversion between the first eye image and the second eye image can be performed appropriately.
[0162]
[0163] The combination of each block of the block diagram and each step of the flowchart attached to the present invention may be performed by computer program instructions. These computer program instructions may be installed in an encoding processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, so that the instructions executed by the encoding processor of the computer or other programmable data processing equipment create a means for performing the functions described in each block of the block diagram or each step of the flowchart. These computer program instructions may also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to implement the functions in a specific manner, so that the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes an instruction means for performing the functions described in each block of the block diagram or each step of the flowchart. Since the computer program instructions can also be installed on a computer or other programmable data processing device, a series of operational steps are performed on the computer or other programmable data processing device to create a computer-executable process, and the instructions that cause the computer or other programmable data processing device to perform the steps for executing the functions described in each block of the block diagram and each step of the flowchart can also provide steps for executing the functions described in each block of the block diagram and each step of the flowchart.
[0164] Additionally, each block or step may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specific logical function(s). It should also be noted that in some alternative embodiments, the functions mentioned in the blocks or steps may occur out of order. For example, two blocks or steps depicted in succession may actually be performed substantially concurrently, or the blocks or steps may sometimes be performed in reverse order, depending on the functionality they perform.
[0165] The above description is merely an illustrative illustration of the technical idea of the present invention, and those skilled in the art will appreciate that various modifications and variations can be made without departing from the essential quality of the present invention. Therefore, the embodiments disclosed in the present invention are intended to illustrate, rather than limit, the technical idea of the present invention, and the scope of the technical idea of the present invention is not limited by these embodiments. The scope of protection of the present invention should be interpreted by the following claims, and all technical ideas within a scope equivalent thereto should be interpreted as being included in the scope of the rights of the present invention.
Claims
1. A method for training an artificial intelligence model for image compatibility of different OCT devices. A step of obtaining a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device; and Including a step of training the artificial intelligence model to enable mutual conversion between the first eye image and the second eye image, In the learning process of the above artificial intelligence model, A loss function designed to target the boundary of the retinal nerve fiber layer included in each eye image is used. How to train an artificial intelligence model.
2. In paragraph 1, The above learning dataset is, Each of the plurality of second eye images having a corresponding relationship with each of the plurality of first eye images, The first eye image and the second eye image having the above correspondence relationship are characterized in that they were captured for the same eye. How to train an artificial intelligence model.
3. In paragraph 2, The steps for obtaining the above learning dataset are: A step of obtaining a first segmentation mask label corresponding to the retinal nerve fiber layer from each of the first eye images using the learned segmentation model; and A step of obtaining a second segmentation mask label corresponding to the retinal nerve fiber layer from each of the second eye images using the above-mentioned learned segmentation model is further included. How to train an artificial intelligence model.
4. In paragraph 3, The above artificial intelligence model is, A first generator for generating a first-second eye image corresponding to a second domain for the second eye image from the first eye image; A second generator for generating a second-first eye image corresponding to a first domain for the first eye image from the second eye image; a first discriminator for identifying the difference between the first eye image and the second-first eye image; and a second discriminator for identifying the difference between the second eye image and the first-second eye images; How to train an artificial intelligence model.
5. In paragraph 4, The step of training the above artificial intelligence model is: A step of first training the artificial intelligence model using the above learning dataset is included, The first step of training the above artificial intelligence model is: A step of inputting the second eye image into the second generator and outputting the 2-1 eye image; A step of training the first discriminator using a first loss function determined based on the difference between the first eye image and the second-first eye image; and A step of training a second generator by using a third loss function determined based on the difference between the second eye image and the second eye image output by inputting the second-1 eye image to the first generator. How to train an artificial intelligence model.
6. In paragraph 5, The first step of training the above artificial intelligence model is: A step of inputting the first eye image into the first generator and outputting the first-second eye image; A step of training the second discriminator using a second loss function determined based on the difference between the second eye image and the first-second eye images; and A step of training the first generator by using the third loss function determined based on the difference between the first-2-1 eye image output by inputting the first-2 eye image to the second generator and the first eye image. How to train an artificial intelligence model.
7. In paragraph 6, The step of training the above artificial intelligence model is: A step of performing secondary training on the artificial intelligence model using the fourth loss function determined based on the above learning data set and the thickness of the retinal nerve fiber layer. How to train an artificial intelligence model.
8. In paragraph 7, The second step of training the above artificial intelligence model is: A step of training the first generator to input the first eye image and output the first segmentation mask by referring to the first segmentation mask label; and A step of training the first generator to input the first segmentation mask and output the second segmentation mask and the second eye image by referring to the second segmentation mask label. How to train an artificial intelligence model.
9. In paragraph 8, The second step of training the above artificial intelligence model is: A step of training the second generator to input the second eye image and output the second segmentation mask by referring to the second segmentation mask label; and A step of training the second generator to input the second segmentation mask and output the first segmentation mask and the first eye image by referring to the first segmentation mask label. How to train an artificial intelligence model.
10. A device that trains an artificial intelligence model for image compatibility of different OCT devices. Memory in which the artificial intelligence model learning program is stored; and A processor for loading the artificial intelligence model learning program from the memory and executing the artificial intelligence model learning program, The above processor, Obtain a learning dataset including multiple first eye images captured by a first OCT device and multiple second eye images captured by a second OCT device, The artificial intelligence model is trained to enable mutual conversion between the first eye image and the second eye image. In the learning process of the above artificial intelligence model, A loss function designed to target the boundary of the retinal nerve fiber layer included in each eye image is used. device.
11. In clause 10, The above learning dataset is, Each of the plurality of second eye images having a corresponding relationship with each of the plurality of first eye images, The first eye image and the second eye image having the above correspondence relationship are characterized in that they were captured for the same eye. device.
12. In paragraph 11, The above processor, Using the learned segmentation model, a first segmentation mask label corresponding to the retinal nerve fiber layer is obtained from each of the first eye images, Using the above-mentioned learned segmentation model, a second segmentation mask label corresponding to the retinal nerve fiber layer is obtained from each of the second eye images. device.
13. In paragraph 12, The above artificial intelligence model is, A first generator for generating a first-second eye image corresponding to a second domain for the second eye image from the first eye image; A second generator for generating a second-first eye image corresponding to a first domain for the first eye image from the second eye image; a first discriminator for identifying the difference between the first eye image and the second-first eye image; and a second discriminator for identifying the difference between the second eye image and the first-second eye images; device.
14. In paragraph 13, The above processor, The artificial intelligence model is trained for the first time using the above learning dataset. The above processor, Input the second eye image into the second generator and output the second-first eye image, The first discriminator is trained using the first loss function determined based on the difference between the first eye image and the second-first eye image, The second generator is trained using the third loss function determined based on the difference between the second eye image and the second eye image output by inputting the second eye image to the first generator. device.
15. In paragraph 14, The above processor, Inputting the first eye image into the first generator and outputting the first-second eye image, The second discriminator is trained using a second loss function determined based on the difference between the second eye image and the first-second eye images, The first generator is trained using the third loss function determined based on the difference between the first-2-1 eye image output by inputting the first-2 eye image to the second generator and the first eye image. device.
16. In paragraph 15, The above processor, The artificial intelligence model is trained for the second time using the fourth loss function determined based on the above learning data set and the thickness of the retinal nerve fiber layer. device.
17. In paragraph 16, The above processor, By referring to the first segmentation mask label, the first generator is trained to input the first eye image and output the first segmentation mask, With reference to the second segmentation mask label, the first generator is trained to input the first segmentation mask and output the second segmentation mask and the second eye image. device.
18. In paragraph 17, The above processor, By referring to the second segmentation mask label, the second generator is trained to input the second eye image and output the second segmentation mask, With reference to the first segmentation mask label, the second generator is trained to input the second segmentation mask and output the first segmentation mask and the first eye image. device.
19. As a device for converting OCT images, Memory where the OCT image conversion program is stored; and A processor for loading the OCT image conversion program from the memory and executing the OCT image conversion program, The above processor, Obtain the first eye image captured by the first OCT device, When an image captured by the first OCT device is provided, the acquired first eye image is provided to an artificial intelligence model that has been trained to output an image of a similar quality to that captured by a second OCT device whose quality of capture result is different from that of the first OCT device. From the above artificial intelligence model, a second eye image corresponding to the first eye image is obtained, In the learning process of the above artificial intelligence model, A loss function designed to target the boundary of the retinal nerve fiber layer included in the eye image is used. Device for converting OCT images.
20. A non-transitory computer-readable recording medium storing a computer program, The above computer program, when executed by a processor, A step of obtaining a learning dataset including a plurality of first eye images captured by a first OCT device and a plurality of second eye images captured by a second OCT device; and A step of training an artificial intelligence model to enable mutual conversion between the first eye image and the second eye image, In the learning process of the above artificial intelligence model, A loss function designed to target the boundary of the retinal nerve fiber layer included in each eye image is used. A method for learning an artificial intelligence model comprising instructions for causing the processor to perform the method. A non-transitory computer-readable recording medium.
Citation Information
Patent Citations
Condition setting support apparatus
JP2014193193A
system
JP2019208602A
Correction of flow projection artifacts in OCTA volumes using neural networks
JP2023520001A
Medical image processing device, medical image processing method, computer-readable medium, and learning completion model
KR102543875B1
Image diagnostic apparatus, image diagnostic method, medical image server and medical image storage method
US20120278359A1