Inference device, inference model creation method, and training data creation method

The inference device generates substitute images using prompts to improve training data quality, addressing issues with unexpected or degraded images, thereby enhancing inference model performance.

WO2026013858A1PCT designated stage Publication Date: 2026-01-15OLYMPUS MEDICAL SYST CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/025135
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing image classification systems face issues due to unexpected images obtained under incorrect shooting conditions or degraded quality, leading to unusable image features for comparison and reduced inference performance.

Method used

An inference device and method that generates alternative images using prompts based on historical information and annotation, creating substitute images to enhance training data and improve inference model reliability.

Benefits of technology

Enhances inference model performance by utilizing substitute images to address issues with personal information protection and image quality, ensuring accurate image feature determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024025135_15012026_PF_FP_ABST
    Figure JP2024025135_15012026_PF_FP_ABST
Patent Text Reader

Abstract

An inference device according to the present invention includes an acquisition unit that acquires a series of a plurality of images obtained in advance and serving as candidates for training data, a first inference unit that is created by learning first training data obtained by adding first annotation information to a first image frame from among the series of the plurality of images acquired by the acquisition unit, a prompt generation unit that generates a prompt relating to a second image frame that is different from the first image frame from among the series of the plurality of images acquired by the acquisition unit, on the basis of history information imparted to the series of the plurality of images, an alternative image generation unit that generates an alternative image that is similar to the second image frame on the basis of the prompt, and a second inference unit that is created by learning second training data obtained by adding second annotation information, based on the prompt, to the alternative image.
Need to check novelty before this filing date? Find Prior Art

Description

Inference device, inference model creation method, and teacher data creation method

[0001] The present invention relates to an inference device, an inference model creation method, and a training data creation method that enable the construction of an effective inference model while protecting personal information.

[0002] An endoscope is a device that is inserted into the body and enables observation of diseased areas that cannot be seen from the outside. An endoscope has an insertion section that is inserted into the body, and an imaging device is provided, for example, at the tip of the insertion section. During an examination using an endoscope, a doctor sequentially displays images acquired by the imaging device at the tip of the insertion section inserted into the body, and adjusts the position of the tip of the insertion section while checking the displayed images to diagnose the patient's health or disease state.

[0003] Image information obtained from the process of inserting an endoscope into the body until it is removed is recorded. By effectively utilizing the rich image information recorded, it is possible to develop technologies that enable the location of affected areas inside the body, guidance of the insertion part into lumens, which has been difficult until now, and assistance in diagnosis and reporting using endoscopic images. For example, by using images captured by an endoscope as training data for learning, it is possible to build an inference model that verifies the meaning of captured images and judges each scene in medical procedures.

[0004] Furthermore, as a technology for improving learning effectiveness, Japanese Patent Application Laid-Open Publication No. 2020-87310 discloses a technology for learning object recognition functions by generating new learning data based on the gradient of the error for each pixel of a composite image calculated from a composite image including a CG model of the object and a teacher signal of the object.

[0005] Japanese Patent Application Laid-Open No. 2020-87310

[0006] However, when attempting to classify images obtained during an examination or treatment, not limited to endoscopy, using image features, an unexpected image may be obtained as a classification result due to incorrect shooting conditions or processing of the image that differs from the expected. Such unexpected images pose a problem in that the similarity of image features, etc., becomes unusable when comparing them with other images of expected image quality or processing. To solve this problem, the present invention aims to provide an inference device, an inference model creation method, and a training data creation method that can create an expected image as an auxiliary image and perform image comparison (assuming similar image determination or image feature determination by inference), thereby enabling image feature determination similar to that of the expected image.

[0007] An inference device according to one aspect of the present invention comprises an acquisition unit that acquires a series of images that are candidates for previously obtained teacher data; a first inference unit created by learning first teacher data obtained by adding first annotation information to a first image frame of the series of images acquired by the acquisition unit; a prompt generation unit that generates a prompt related to a second image frame that is different from the first image frame of the series of images acquired by the acquisition unit based on history information assigned to the series of images; an alternative image generation unit that generates an alternative image similar to the second image frame based on the prompt; and a second inference unit created by learning second teacher data obtained by adding second annotation information based on the prompt to the alternative image.

[0008] An inference model creation method according to one aspect of the present invention involves acquiring a series of multiple images, creating a first inference unit by learning first teacher data obtained by adding first annotation information to a first image frame among the acquired series of multiple images, and if the reliability value of the inference result of the first inference unit when a specific image is input to the first inference unit is lower than a predetermined threshold, generating a prompt related to a second image frame different from the first image frame among the acquired series of multiple images based on history information assigned to the series of multiple images, generating an alternative image similar to the second image frame based on the generated prompt, and creating a second inference unit by learning second teacher data obtained by adding second annotation information based on the prompt to the generated alternative image.

[0009] An inference device according to another aspect of the present invention comprises: a first inference unit created by learning first teacher data obtained by inputting, via an acquisition unit, a video obtained by previously capturing a scene similar to a specific examination or treatment scene in which a series of multiple images are captured and input via an image input unit, and adding first classification scene text information as first annotation information to image frames in the input video that have specific image features; and a second inference unit created by learning second teacher data obtained by providing each image frame of the images input via the image input unit to the created first inference unit and adding the second classification scene text information as second annotation information to a generated substitute image generated using second classification scene text information other than the first classification scene text information for image frames whose reliability of the inference result is below a predetermined threshold when an inference result is obtained from the first inference unit.

[0010] A method for creating teacher data according to one aspect of the present invention includes the steps of: obtaining first teacher data by adding first classification scene text information as first annotation information to image frames having specific image features in multiple videos that have been previously filmed showing specific examination and treatment scenes; and obtaining second teacher data by adding second classification scene text information as second annotation information to generated substitute images generated using second classification scene text information other than the first classification scene text information assigned to the multiple videos.

[0011] Another aspect of the present invention provides an inference model creation method that acquires multiple video files, determines the image features of each frame that constitutes each video, creates a first inference unit by learning first teacher data obtained by adding first annotation information to a group of images that have a first image feature among the frames that constitute each video, generates a prompt based on the attached information attached to the video file for a group of images that have a second image feature different from the first image feature among the frames that constitute each video, generates an alternative image similar to the second image frame based on the generated prompt, and creates a second inference unit by learning second teacher data obtained by adding second annotation information based on the prompt to the generated alternative image. Another aspect of the present invention provides a method for creating teacher data, which includes acquiring multiple video files, determining the characteristics of each frame that constitutes each video, adding first annotation information to a group of images that have a first image feature among the frames that constitute each video to create first teacher data, generating a prompt based on the annotation information added to the video file for a group of images that have a second image feature different from the first image feature among the frames that constitute each video, generating an alternative image similar to the second image frame based on the generated prompt, and adding second annotation information based on the prompt to the generated alternative image to create second teacher data.

[0012] Another aspect of the present invention provides a method for creating an inference model, which involves acquiring multiple medical videos that capture examination and treatment processes performed in accordance with a specific affected area and surgical procedure, creating a first inference unit by learning first training data obtained by adding first classification scene text information for classifying image frames having a first image feature from the acquired videos as scenes from a specific examination and treatment process, and creating a second inference unit by learning second training data obtained as training data from generated images obtained using prompt information that uses second classification scene text information that is different from the first classification scene text information for classifying image frames having a second image feature that is different from the first image feature from the acquired videos as scenes from a specific examination and treatment process.

[0013] According to the present invention, when a series of images including some images that cannot be used as training data is input, a prompt is generated to regenerate images related to some of the images, and alternative images based on the generated prompt are generated and used as training data, thereby enabling the construction of an inference model with high inference performance.

[0014] Fig. 1 is a block diagram showing an inference device according to a first embodiment of the present invention; Fig. 2 is a block diagram showing a specific example of a prompt generation unit in Fig. 1; Fig. 3 is a flowchart for explaining the operation of the first embodiment; Fig. 4 is an explanatory diagram for explaining the operation of the first embodiment; Fig. 5 is a flowchart for explaining an operation flow adopted in a second embodiment;

[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0016] For example, when building an inference model for classifying medical procedures, some of the images in a series of images acquired during the use of an endoscope may not be usable as training data for building the inference model due to concerns about protecting personal information. Furthermore, images with significantly degraded image quality may also not be usable as training data for building the inference model. In such cases, building an inference model using images other than the selected selected images results in a problem of degraded inference performance for images related to the selected selected images.

[0017] In this embodiment, we will explain how to respond when a series of images including some images that cannot be used as training data is input when determining image features using an inference model. For example, we will explain an example of providing an inference device, an inference model creation method, and a training data creation method that can build an inference model with high inference performance by generating prompts to regenerate images related to some images that are unsuitable for training data, and generating substitute images based on the generated prompts and using them as training data.

[0018] In other words, in this embodiment, when a series of images including some images that cannot be used as training data is input, a prompt is generated to regenerate images related to some of the images, and alternative images based on the generated prompt are generated and used as training data, thereby enabling the construction of an inference model with high inference performance.

[0019] (First embodiment) Fig. 1 is a block diagram showing an inference device according to a first embodiment of the present invention, and Fig. 2 is a block diagram showing a specific example of the prompt generation unit in Fig. 1.

[0020] In AI (artificial intelligence) development, it is generally considered preferable to aggregate and learn from various data. Although distributed learning methods have been proposed, they have not been able to surpass centralized learning methods in terms of performance. Therefore, when creating inference models for diagnosis, etc., it is necessary to aggregate useful data and build a learning environment while respecting the wishes of patients and others who provide images for creating training data.

[0021] For example, an inference model may be constructed by classifying each scene in endoscopic images of a series of medical procedures, such as an endoscopic procedure, and obtaining classification results, such as which part of the body each surgical procedure targets and what the procedure entails. Such an inference model can be created by learning using images acquired by an endoscope as training data. However, from the perspective of protecting the personal information of patients and others, some images acquired during endoscopic examinations and diagnoses may not be usable as training data. For example, during a procedure using a laparoscope, the laparoscope may be temporarily removed from the body for cleaning. In this case, the removed laparoscope may capture images of the operating room, such as images of the patient's or medical staff's faces. Since such images are undesirable from the perspective of protecting personal information, they are processed, for example, to black images, before generating training data. That is, when generating training data, some images in the series of endoscopic images are replaced with, for example, black images, and therefore cannot be used as training data. Furthermore, images with significantly degraded image quality, such as blurred images, cannot be used as training data. In an inference model constructed using images other than these few images as training data, even if an image obtained by taking a picture of the operating room after a laparoscope is removed from the body is input, the scene cannot be identified through inference.

[0022] Therefore, in this embodiment, when creating training data, if an image different from the original image due to image processing or an image with significantly deteriorated image quality is given, a substitute image similar to the original image is generated using a prompt for generating a substitute image, and an inference model is constructed by learning using this substitute image, thereby making it possible to construct a highly reliable inference model. Note that the following description will explain an example of an inference device that performs inference on endoscopic images from an endoscope, but the target images are not limited to endoscopic images.

[0023] 1, the inference device includes a control unit 10, a hospital management information recording unit 20, a teacher data creation unit 30, a prompt generation unit 40, a substitute image generation unit 50, an inference model unit 60, and an information output unit 70. Each unit constituting the inference device in FIG. 1 may be configured by a processor using a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), an NPU (Neural Processing Unit), or the like, may operate according to a program stored in a memory (not shown) to control each unit, or may realize some or all of its functions using hardware electronic circuits.

[0024] The control unit 10 performs overall control of the entire inference device of Fig. 1. The information input unit 11 is composed of input devices such as a keyboard (not shown) and various interface devices that receive information from external devices, and accepts user operations to supply operation signals to the control unit 10, or receives information from external devices (not shown), and supplies the received information to the control unit 10. The control unit 10 controls each part of the inference device based on the information and operation signals from the information input unit 11. For example, the information input unit 11 also provides the control unit 10 with information regarding the specifications of the inference model.

[0025] The hospital management information recording unit 20 is installed, for example, in a hospital or the like and is composed of a recording device capable of recording confidential information. The hospital management information recording unit 20 records all information related to the contents of patient examinations, etc. For example, the patient name, doctor name, disease name, type of examination or procedure, and images 21a obtained by, for example, an endoscopic examination showing the contents of the procedure are recorded. The images 21a are still images or moving images. Each image 21a recorded in the hospital management information recording unit 20 is accompanied by various information, such as the patient information, disease information, and surgical procedure information, as metadata 21b. The metadata 21b also indicates that the image 21a was captured during, for example, laparoscopic surgery. In other words, the metadata 21b is information associated with the shooting scene, shooting environment, and image (file) when it was recorded, and is history information including information on the recording environment, recording folder, etc. From the images organized in this way, many useful images are selected as training data. Images captured during surgery are videos, but they are organized and recorded as files in specific folders. This information (or related information) can be used to collect similar videos and create training data. Videos consist of many frames, so collecting similar frames from each video can yield a large amount of training data. The videos recorded in the hospital management information recording unit 20 (many medical institutions have similar recording units, and these may also be used) can be acquired and used by the acquisition unit described below. When a medical institution provides videos to an external party, the images are processed at the medical institution or other facility to remove personal information. The acquisition unit acquires the processed images. The text information used to create annotations and prompts may include information such as the file names of the videos and still images, the folder names in which they are recorded, the source of their recording, the sender, and the conditions at the time of reception. Text information contained in the videos and still images may also be used by referencing contract information at the time of data transfer, or encoded information may be converted to text and used.

[0026] The processing unit 22 reads out images for creating teacher data from the hospital management information recording unit 20. For images read out from the hospital management information recording unit 20 that are not desirable to be made public from the viewpoint of protecting personal information, the processing unit 22 processes the images into images that make it impossible to infer the original image, such as black images, and then outputs the images to the teacher data creation unit 30. The processing unit 22 also outputs metadata 21b to the teacher data creation unit 30.

[0027] The teacher data creation unit 30 includes an annotation unit 31 and a condition unmet information unit 32. The teacher data creation unit 30, which serves as an acquisition unit, acquires a series of images from the processing unit 22. The teacher data creation unit 30 creates teacher data based on the images from the processing unit 22. The annotation unit 31 of the teacher data creation unit 30 annotates each image from the processing unit 22. For example, the annotation unit 31 adds annotation information to the input images, indicating the scene (site, procedure, etc.) of each image, such as each surgical step, to create annotated images P1, P2, ... and test images. At this time, the scene the image corresponds to may be expressed in text information, or coded track information or the like may be added as annotation information. Furthermore, the annotation information may include textual representations of features that can be read from each image, textual representations of tissues or objects of examination or treatment, and coordinates representing their positions.

[0028] The condition non-satisfaction information unit 32, under the control of the control unit 10, determines whether an input image satisfies the conditions for adding annotation information (including whether adding annotation information is meaningful) based on the specifications of the inference model. If the conditions are not met, the input image is determined to be a condition non-satisfaction image. The condition non-satisfaction information unit 32 determines an image supplied from the processing unit 22 that cannot be annotated as a condition non-satisfaction image. For example, the condition non-satisfaction information unit 32 can determine the black image described above as a black image that cannot be annotated by determining the pixel level. It can also determine whether not only black images, but also white images, partially erased images, and fully or partially blurred images satisfy the conditions based on the characteristics of the image signal. The condition non-satisfaction information unit 32 may also determine images with poor image quality, such as out of focus, that are unsuitable for annotation, or images of unknown treatment tools, as condition non-satisfaction images. The condition non-satisfaction information section 32 may determine the condition non-satisfying images by using information included in the metadata 21b.

[0029] For example, if the inference model to be created is one that infers scenes from each step of laparoscopic surgery and is specified to use endoscopic images from laparoscopic surgery as training data, and if the information in the metadata 21b indicates that the image in question was acquired during laparoscopic surgery, the condition non-satisfaction information unit 32 may determine that a black image is an input image that does not satisfy the condition. The condition non-satisfaction information unit 32 may also determine that an image that does not satisfy a predetermined standard, such as an out-of-focus image, is an image that does not satisfy the condition. If a condition non-satisfaction image is found, the condition non-satisfaction information unit 32 outputs information on the determination result that the condition is not met and the metadata 21b to the prompt generation unit 40.

[0030] Images P1, P2, ... that have been annotated with text information or the like are supplied to a first inference model 61 of an inference model unit 60 as a first teacher data group. The inference model unit 60 includes a first inference model 61 and a second inference model 62. The first inference model 61 and the second inference model 62 are each composed of a network N1 or N2. A network design is determined for the network N1 or N2 by performing learning, for example, deep learning, using a large amount of teacher data so that an output corresponding to each input can be obtained.

[0031] Annotated images P1, P2, ... are provided to the network N1 as training data. As described above, the images P1, P2, ... are annotated with information indicating, for example, which scene the image represents. By performing deep learning on the network N1 using the annotated images P1, P2, ... as training data, the network design of the network N1 is determined and a first inference model 61 is constructed. A test image is input to the first inference model 61 to determine whether an inference result with the desired reliability can be obtained. If an inference result with the desired reliability cannot be obtained, learning using the annotated images is repeated. This constructs a first inference model 61 with the desired inference performance.

[0032] Deep learning is a multilayered version of the machine learning process using neural networks. A typical example is a forward propagation neural network, which sends information from front to back and makes a judgment. In its simplest form, it requires three layers: an input layer consisting of m1 neurons, a hidden layer consisting of m2 neurons determined by parameters, and an output layer consisting of m3 neurons corresponding to the number of classes to be discriminated. The neurons in the input and hidden layers, and those in the hidden and output layers, are connected by connection weights, and a bias value is added between the hidden and output layers, making it easy to form logic gates. While three layers are sufficient for simple discrimination, increasing the number of hidden layers makes it possible to learn how to combine multiple features during the machine learning process. In recent years, neural networks with 9 to 152 layers have become practical due to their training time, judgment accuracy, and energy consumption.

[0033] Various known networks may be used as the networks N1 and N2 used for machine learning. For example, R-CNN (Regions with CNN features) or FCN (Fully Convolutional Networks) using CNN (Convolution Neural Network) may be used. This involves a process called "convolution" that compresses image features, operates with minimal processing, and is strong in pattern recognition. Furthermore, a "recurrent neural network" (fully connected recurrent neural network) that can handle more complex information and allows information analysis in which the meaning changes depending on the order or sequence of information may be used, allowing information to flow bidirectionally.

[0034] To realize these technologies, conventional general-purpose arithmetic processing circuits such as CPUs and FPGAs can be used, but because much of the processing in neural networks involves matrix multiplication, GPUs and Tensor Processing Units (TPUs), which are specialized for matrix calculations, may also be used.In recent years, such dedicated artificial intelligence (AI) hardware, called "neural network processing units (NPUs)," have been designed to be integrated and embeddable with CPUs and other circuits, and may even become part of the processing circuit.

[0035] By providing an endoscopic image Ps from the image input unit 63 to the inference model (network N1) 61 constructed in this manner, an inference result is obtained that indicates which scene in a surgical procedure or other process the endoscopic image Ps represents. This inference result also includes reliability (degree of confidence). The inference result and reliability information from the first inference model 61 are supplied to the information output unit 70. The information output unit 70 outputs the inference result and reliability information from the first inference model 61 to an external device. For example, the information output unit 70 provides the inference result and reliability to a display device (not shown) for display.

[0036] Now, let us assume that an image obtained by capturing images of the operating room using a laparoscope temporarily removed from the body during surgery is input to the training data creation unit 30 as a condition-unsatisfied image Pm, such as a black image, from the viewpoint of protecting personal information. In this case, it is conceivable that the first inference model 61 cannot make correct inferences for an endoscopic image Ps that resembles the original image Po (hereinafter referred to as the unprocessed image) before processing the condition-unsatisfied image Pm. In this case, the reliability of the inference results from the first inference model 61 will be extremely low.

[0037] In this embodiment, it is assumed that the reason why the first inference model 61 does not obtain a highly reliable inference result is that the condition-failed image Pm is not used in learning, and so to obtain a highly reliable result, the second inference model 62 is learned using a substitute image of the condition-failed image Pm. That is, in this embodiment, multiple substitute images similar to the unprocessed image Po of the condition-failed image Pm are generated, and after annotating the generated substitute images, they are provided as training data to the network N2 that constitutes the second inference model 62 for learning.

[0038] In addition, since an inference model can increase the number of image features that can be inferred and improve accuracy and reliability by using a large amount of training data, it is better to use image frames that constitute a group of images for each video that are organized and recorded as multiple video files. The first image frame used as training data for the first inference model is a similar image included in the multiple video files.

[0039] By providing the endoscopic image Ps to an inference model (network N2) obtained by learning using substitute images, an inference result indicating which scene the endoscopic image Ps is from is obtained. As will be described later, learning is performed to ensure that this inference result is sufficiently reliable.

[0040] As described above, in this embodiment, whether or not to train the second inference model 62 is determined based on the reliability of the inference result of the first inference model 61. For this determination, information on the inference result and reliability from the first inference model 61 is also supplied to the reliability determination unit 71. When information on reliability lower than a predetermined threshold is input from the first inference model 61, the reliability determination unit 71 outputs this information to the training data creation unit 30. The output of the reliability determination unit 71 indicates that an image other than an image that can be inferred with a reliability equal to or higher than the predetermined threshold has been input to the first inference model 61. Based on the information from the reliability determination unit 71, the condition non-satisfaction information unit 32 of the training data creation unit 30 estimates that the first inference model 61 cannot infer with sufficient accuracy an image similar to the unprocessed image Po of the condition non-satisfaction image Pm. In this case, the condition non-satisfaction information unit 32 instructs the prompt generation unit 40 to create a substitute image for the original unprocessed image Po of the condition non-satisfaction image Pm.

[0041] Here, an example has been introduced in which an image unsuitable for inference using the first inference model 61 is determined using the first inference model 61. However, images unsuitable for inference using the first inference model 61 may be determined by examining the numerical characteristics of the image signal according to the image's characteristics, as described above. Alternatively, the image may be determined by comparing it with an image suitable for inference using the first inference model 61. For example, if the similarity differs significantly, it may be better to perform inference using a model other than the first inference model 61. The second inference model 62 may be created by recording the image characteristics corresponding to the first inference model 61, inputting images that satisfy the conditions into the first inference model 61, and inferring those that do not. A second image frame that is different from the first image frame described above can be determined by differences in image data characteristics (such as color, brightness, contrast, or processing) or by a lower similarity (below a predetermined value) when compared with the first image frame.

[0042] The prompt generation unit 40 generates a prompt for creating a substitute image based on information from the condition failure information unit 32. The information from the condition failure information unit 32 includes information on the reliability of the inference result from the first inference unit 61, historical information, and information on the specifications of the first and second inference units 61 and 62. The prompt generation unit 40 generates the prompt based on at least one of these pieces of information. The simplest example is a prompt that differs from the text information used for annotation when the first inference model 61 was trained. For example, if the annotation is "is ____," the prompt can be text such as "is not ____." Various information can be added to the prompt. For example, while the annotation "inside the colon" generates text such as "other than inside the colon," by using text information such as "surgery" or "hospital" obtained during image acquisition or recording (which may be file names, folder names, or metadata associated with the image, file, or folder), a prompt such as "colon surgery image, other than inside the colon, at the hospital" can be generated. Based on this prompt, the image generation AI can generate images, such as images from an operating room. Furthermore, various prompts can be created by modifying the text with synonyms or modifiers of these word texts. Various images can be generated using such prompts and used as training data candidates. Then, by annotating and training the corresponding text (prompts, or portions or summaries thereof) for each image, various inference models can be obtained through training, in this example, for inferring "other than the inside of the colon." With these various models, even if an image other than "the inside of the colon" is input, it is possible to infer and output text that can explain the scene. In this way, it is possible to create training data characterized by the first annotation information used to train the first inference unit 61 being text information, and the prompts generated to train the second inference unit 62 being information different from the text information for the first annotation, i.e., information that the first inference unit 61 does not intend to infer. It is possible to create training data and provide an inference device using this training data.The second annotation information based on the prompt and added to the substitute image is characterized in that it is a summary or part of the prompt.

[0043] For example, the prompt generation unit 40 may receive metadata 21b from the condition non-satisfaction information unit 32 and generate a prompt based on the metadata 21b. For example, if the metadata 21b indicates that the unedited image Po was taken during laparoscopic surgery and the information from the condition non-satisfaction information unit 32 indicates that the condition non-satisfaction image Pm was a black image, the prompt generation unit 40 may determine that the unedited image was an image captured in an operating room and generate a prompt such as "operating room." Note that if the editing unit 22 provides information to the condition non-satisfaction information unit 32 indicating that the image has been edited and the type of image the unedited image was, the condition non-satisfaction information unit 32 can easily determine that the image is a condition non-satisfaction image Pm, and the prompt generation unit 40 can easily create a prompt based on the information from the condition non-satisfaction information unit 32. Note that even if the prompt generation unit 40 is not provided with information regarding editing by the editing unit 22, the prompt generation unit 40 can estimate the type of image the unedited image Po is based on the metadata 21b from the hospital management information recording unit 20 and generate a prompt.

[0044] The prompt generation unit 40 includes a prompt amplification unit 41. The prompt amplification unit 41 increases the number of prompts by generating synonyms of the generated prompts, modifying them, or synthesizing them into sentences. For example, the prompt amplification unit 41 may refer to a thesaurus to search for and use words such as "operating room," "operation room," "surgery room," "surgical room," and "operating table" for the generated "operating room," or may search for words related to these words that appear in examples using a dictionary containing those words. In other words, various prompts such as "surgical device," "medical equipment," "monitor," "bed," "doctor," and "nurse" may be generated using these methods.

[0045] Words can be classified using tools like Word2Vec, which is commonly used in natural language processing to vectorize word meaning. Since the words in this example are from a medical context, vectorizing their meaning can be done by defining the elements of the vector as "medical care," "examination / treatment," and "affected area." For example, the three words "operating room," "medical equipment," and "reception / accounting" can be represented as three-dimensional vectors (expressed here using three numerical examples with values ​​less than or equal to 1) as follows: "operating room" = [0.9, 0.6, 0.; "medical equipment" = [0.9, 0.7, 0.; "reception / accounting" = [0.7, 0.2, 0.]). For example, the vectors for "operating room" and "endoscope" are similar, while the vectors for "operating room" and "accounting" are dissimilar. Representing words as vectors in this way makes it possible to numerically determine whether words are similar. Vectorizing word meanings in this way makes it possible to determine the similarity and dissimilarity of words. These values ​​can be obtained using a model trained to acquire and process large amounts of text and achieve high reliability.

[0046] In this way, the semantic similarity and dissimilarity of words can be determined using multiple evaluation values ​​(vectors). Therefore, for each timing of a medical video capturing an examination / treatment process performed in response to a specific affected area or surgical procedure, appropriate words (text) used to distinguish a scene as a first scene, a second scene, etc. can be obtained, with some evaluations indicating similarity and other evaluations indicating dissimilarity. In other words, appropriate scene classification is selected as first classification scene text information for classifying image frames having a first image feature in the video as scenes of a specific examination / treatment process, and second classification scene text information different from the first classification scene text information for classifying image frames having a second image feature different from the first image feature as scenes of a specific examination / treatment process. For example, if the first classification scene text information is a word such as "duodenum," the evaluations for "medical care," "examination / treatment," and "affected area" all have high values, and similar words such as "operating room" and "medical equipment" can be detected. On the other hand, words that are completely unrelated to medical videos and are inappropriate for this purpose (classifying medical video scenes), such as "car" or "dining table," will have low values ​​when quantified based on evaluation criteria and perspectives such as "medical department," "examination / treatment," and "equipment," and will not be used.

[0047] Of course, dictionary examples can be used as prompts. Examples include "The operating room was clean, with various medical equipment neatly arranged" and "The first thing I noticed when I entered the operating room was the large bed." Instead of using a dictionary, an internet search for "operating room" could yield a hit, such as "Operating room nurses and perioperative management team nurses are on-site." When converting words into sentences to generate prompts, example sentences using each word can be created using generative AI, or they can be searched in dictionaries or online. The words used here can be expressed as classification scene text information or classification scene word information. In other words, the first classification scene text information and second classification scene text information for classifying video scenes can be words that appear in synonyms or example sentences found in dictionary or internet text searches, or words whose meanings are expressed using multiple evaluation values ​​and vectorized to determine the similarity or difference between the words' meanings. Words with a similarity greater than or equal to a meaningful numerical difference (vector value) for scene classification are used.

[0048] FIG. 2 is a block diagram showing an example of a specific configuration of the prompt generating unit 40. As shown in FIG.

[0049] The prompt generation unit 40 includes a situation information recording unit 42, an incomplete / missing information determination unit 43, a word generation unit 44, a text generation unit 45, an output unit 46, a key word list 47, and an input unit 48. The input unit 48 receives various information from the condition unmet information unit 32. The situation information recording unit 42 records various information related to the unprocessed image Po. For example, the situation information recording unit 42 may record surgical procedure information, medical institution information, and examination information related to the unprocessed image Po. The surgical procedure information includes information on the surgical method and steps, the instruments and devices used in the surgical procedure, etc. The medical institution information is information related to the medical institution. The examination information is information on the types of various examinations performed on the patient, the examination results, etc. The situation information recording unit 42 may collect this information based on, for example, information in the metadata 21b input via the input unit 48.

[0050] The incomplete / missing information determination unit 43 determines at least one of the reasons for the occurrence of an image lacking essential information and the reason for the occurrence of an incomplete image with low image quality that does not correspond to the annotation information, based on the information from the condition non-satisfaction information unit 32, the surgical procedure information, the examination information, etc. For example, when the laparoscope is withdrawn from the body during laparoscopic surgery, images taken during the period when the laparoscope is withdrawn from the body are converted to black images to protect personal information. Therefore, the absence reason determination unit 43a of the incomplete / missing information determination unit 43 may determine that the occurrence of such an image lacking information is due to the capture of an operating room. Furthermore, for example, the image quality of a captured image is significantly degraded when water is pumped into the body. Therefore, the incomplete reason determination unit 43b of the incomplete / missing information determination unit 43 may determine that the occurrence of such an incomplete image that does not correspond to the annotation information by the annotation unit 31 is due to the performance of a water pumping process.

[0051] The word generation unit 44 includes a related word extraction unit 44a and a search unit 44b. The word generation unit 44 generates words using, for example, an important word list 47. It is possible to determine what kind of scene in each step of a medical procedure the image input to the teacher data creation unit 30 represents from the information recorded in the situation information recording unit 42. For example, it is possible to determine each step of a medical procedure on a patient from surgical procedure information, examination information, etc., and also to determine what kind of image is acquired in each step. For example, it is possible to create an important word list 47 corresponding to the surgical procedure information and examination information. For example, the important word list 47 may be created based on a procedure manual, a paper, etc. corresponding to the surgical procedure information, examination information, etc.

[0052] The search unit 44b of the word generation unit 44 searches for words (text) corresponding to missing or incomplete images by searching papers, important word lists, etc., based on the information recorded in the situation information recording unit 42 for images determined by the incomplete / missing information determination unit 43. The related word extraction unit 44a extracts related words from the words searched for by the search unit 44b. This makes it possible to extract related words such as "operating room" for an image that is a black image and determined by the condition unsatisfaction information unit 32 to be a condition unsatisfying image Pm, for example.

[0053] The multiple texts generated by the word generation unit 44 are provided to the text generation unit 45. When multiple words are extracted by the related word extraction unit 44a, the word combination unit 45a of the text generation unit 45 generates new words by modifying, sentence-forming, combining, and generating synonyms from these words. This increases the number of words generated by the word generation unit 44. The word combination unit 45a in FIG. 2 constitutes the prompt amplification unit 41 in FIG. 1. The text generated by the text generation unit 45 is supplied as a prompt to the substitute image generation unit 50 by the output unit 46. That is, the prompt generation unit 40 can generate a text prompt related to any of the scenes of the medical procedure based on surgical procedure information indicating the surgical procedure of the medical procedure included in the history information.

[0054] The substitute image generation unit 50 generates a substitute image PA using, for example, a generation AI. The substitute image generation unit 50 includes an image search unit 51 and a prompt image creation unit 52. The image search unit 51 searches the Internet or the like for images similar to the original image before processing of the condition-unsatisfied image Pm. A prompt is given to the image search unit 51 from the prompt generation unit 40. The image search unit 51 searches the Internet or the like using the prompt (text) from the prompt generation unit 40 to search for one or more similar images that are similar to the unprocessed image Po.

[0055] The prompt image creation unit 52 includes an image input unit 53 that inputs images as real data, a training unit 54 that trains the generator G1 that constitutes the generation AI, and a generation unit 55 that generates images using the generator G1. The generation AI is an AI that automatically generates content based on learned data. For example, the generation AI improves the judgment accuracy of the recognition AI by having the recognition AI, a classifier D1, distinguish between the "real data" that it has created and collected and the "fake data" that it has created, while gradually making the fake data of the generation AI more realistic, thereby generating fake data that resembles real data.

[0056] The image input unit 53 takes in similar images searched by the image search unit 51 and provides them to the training unit 54. The training unit 54 includes a generator G1 and a classifier D1. Noise is provided to the generator G1 as input information. The generator G1 generates fake data (image Pg). Meanwhile, images (similar images) that are real data taken in by the image input unit 53 are supplied to the classifier D1. The classifier D1 compares the image Pg from the generator G1 with the image from the image input unit 53 and provides the comparison result to the generator G1. This type of training is repeated, and the generator G1 comes to generate fake data (image Pg) that is similar to the real data.

[0057] The generator 55 uses a generator G1. Noise is provided to the generator G1 as input information. The generator G1 is trained in the training unit 54, and when input information is provided, the generator G1 generates and outputs an alternative image PA similar to the similar image input by the image input unit 53. Note that although an example has been shown in which the alternative image PA is generated based on a similar image obtained by a search by the image search unit 51, the similar image searched by the image search unit 51 may be used as the alternative image PA as is.

[0058] The training unit 54 and the generation unit 55 may be provided with language processors L1 and L2. When the language processor L1 supplies a similar image from the image input unit 53 to the classifier D1, it supplies the classifier D1 with a prompt from the prompt generation unit 40 regarding the characteristics of the similar image. As a result, the classifier D1 is supplied with a similar image having characteristics corresponding to the prompt from the prompt generation unit 40 as real data. Furthermore, when the language processor L2 supplies input information to the generator G1, it supplies the generator G1 with a prompt from the prompt generation unit 40 regarding the characteristics of the substitute image PA to be generated. As a result, the generator G1 can generate a substitute image PA and a test image having characteristics corresponding to the prompt from the prompt generation unit 40.

[0059] In addition, the substitute image generation unit 50 may be configured to generate multiple substitute images corresponding to the prompt generated by the prompt generation unit 40, or may be configured to generate multiple substitute images based on multiple similar images obtained by searching for one prompt using the image search unit 51.

[0060] The substitute image generated by the substitute image generation unit 50 is annotated by the annotation unit 58. In this embodiment, the annotation unit 58 may add the prompt generated by the prompt generation unit 40 to the substitute image. The teacher data PA1, PA2, ... annotated by the annotation unit 58 is provided to the network N2 for learning. A second inference model 62 is constructed through this learning.

[0061] Although an example has been described in which annotation is performed using prompts as second-class scene text information, annotation may also be performed using text or symbols that describe scenes different from the first-class scenes.

[0062] When training the second inference model 62, a test is performed using a test image. Information about the inference result and reliability of the second inference model 62 for the test image is supplied to the control unit 10 via the information output unit 70. The control unit 10 adopts the second inference model 62 if the received reliability is equal to or greater than a predetermined threshold. If the reliability from the second inference model 62 is lower than the predetermined threshold, the control unit 10 causes the substitute image generation unit 50 to generate a substitute image PA and performs training using new training data. In this manner, the second inference model 62 with sufficient inference performance is constructed. When regenerating the substitute image PA, the control unit 10 may control the prompt generation unit 40 to generate a prompt different from the previous prompt and supply it to the substitute image generation unit 50. For example, when an image capturing an operating room during laparoscopic surgery is input, the second inference model 62 can output an inference result such as "a scene in which the laparoscope is withdrawn from the body."

[0063] Next, the operation of the embodiment configured as above will be described with reference to Figures 3 to 5. Figures 3 and 4 are flow charts for explaining the operation of the first embodiment. Figure 5 is an explanatory diagram for explaining the operation of the first embodiment.

[0064] Figure 3 shows the operation when creating the first and second inference models 61, 62. In S1 of Figure 3, information regarding the specifications of the inference model to be created is input to the control unit 10 by the information input unit 11. For example, information such as creating an inference model that infers scenes of each process in laparoscopic surgery and uses endoscopic images from laparoscopic surgery as training data is supplied to the control unit 10. In S2, an image is input from the processing unit 22. The input image may be displayed on a display device (not shown). The annotation unit 31 of the training data creation unit 30 annotates the image input from the processing unit 22. The annotated images P1, P2, ... are supplied to the network N1 for learning. This creates the first inference model 61.

[0065] The images used here are assumed to be, for example, videos acquired from various hospitals and other medical institutions, showing numerous tests, surgeries, and other procedures performed by different patients and treating doctors, and are organized by affected area, surgical procedure, and patient profile such as gender and age.The assumed case is to use videos of the same category selected as needed from such images, and annotate them with classified scene text information that allows the type of scene in each frame to be separated by track.

[0066] For example, by adding first classification scene text information as first annotation information to image frames having specific image features in multiple videos previously filmed of specific examination and treatment scenes involving different patients and treating doctors, first training data for training the first inference model 61 can be obtained.

[0067] The control unit 10 provides a test image to the network N1 and performs inference using the first inference model 61. The reliability determination unit 71 determines whether the reliability (reliability value) of the first inference result from the first inference model 61 is equal to or greater than a predetermined threshold, and if a determination result equal to or greater than the predetermined threshold is not obtained, the control unit 10 repeats the learning of the first inference model 61 by selecting and rejecting teacher data or having the teacher data creation unit 30 create teacher data again. When the reliability determination unit 71 indicates that a determination result equal to or greater than the predetermined threshold has been obtained, the control unit 10 terminates the learning of the first inference model 61. In this way, the first inference model 61 is constructed (S3).

[0068] Here, the judgment and branching process of inputting test data into the inference model obtained through learning and redoing the learning according to its reliability is not shown, but if necessary, the learning process may not end with "end" and may be redone. This is omitted in Figure 3 to avoid complicating the explanation.

[0069] In S4 of Fig. 3, the control unit 10 provides a specific image to the first inference model 61 and performs inference using the first inference model 61. The reliability determination unit 71 determines whether the reliability of the inference result of the first inference model 61 is equal to or greater than a predetermined threshold (S5). If a reliability equal to or greater than the threshold is obtained, it is determined in S6 that the inference process has ended (S6). If not, the process returns to S4 and inference is performed again. The processes of S4 to S6 are repeated until the inference process ends.

[0070] The upper part of Figure 5 shows inference by the first inference model 61. The first inference model 61 is trained using training data from the training data creation unit 30. In the example of Figure 5, the first inference model 61 is capable of detecting A and B for an input image. Three input images have features A, B, and C, respectively. The first inference model 61 outputs A as the inference result for an image input having feature A. The first inference model 61 also outputs B as the inference result for an image input having feature B. However, the first inference model 61 cannot detect feature C as the inference result for an image input having feature C. For an image having feature C, the reliability determination unit 71 determines that the reliability of the inference result of the first inference model 61 is lower than a predetermined threshold and outputs the determination result to the training data creation unit 30.

[0071] When the reliability determination unit 71 provides a judgment result indicating low reliability to the condition-unsatisfied image Pm, the condition-unsatisfied information unit 32 of the teacher data creation unit 30 infers that an image similar to the unedited image Po corresponding to the condition-unsatisfied image Pm was input to the first inference model 61, and provides information about the unedited image Po to the prompt generation unit 40 so that learning is performed using an alternative image similar to the unedited image Po as teacher data. The condition-unsatisfied information unit 32 also provides the prompt generation unit 40 with the judgment result of the reliability determination unit 71 and specification information. The prompt generation unit 40 converts words and other elements that will become prompts into text in accordance with the judgment result, specification, and other elements of the reliability determination unit 71 (S7). In this case, various texts related to the judgment result, specification, and other elements of the reliability determination unit 71 are also collected. The prompt amplification unit 41 modifies and sentences the collected text to increase the number of texts (text amplification) (S8).

[0072] The prompt generation unit 40 provides the text prompt to the substitute image generation unit 50. The substitute image generation unit 50 generates a group of images by searching using the prompt or processing by a generation AI (S9). Note that this group of images includes images for training data and images for test data.

[0073] The annotation unit 58 creates the training data by adding, as annotation information, text or the like generated as a prompt by the prompt generation unit 40 to the image for training data from the substitute image generation unit 50. The control unit 10 provides the annotated training data to the network N2 for learning (S10).

[0074] The control unit 10 provides a test image to the network N2 and performs inference using the second inference model 62. The control unit 10 determines whether the reliability of the second inference result from the second inference model 62 is equal to or greater than a predetermined threshold (inference reliability OK) (S11). If a determination result equal to or greater than the predetermined threshold is not obtained, the control unit 10 repeats the learning of the second inference model 62 by selecting and rejecting training data, regenerating alternative images, redoing annotations using the annotation unit 58, etc. (S12). If a determination result equal to or greater than the predetermined threshold is obtained, the control unit 10 terminates the learning of the second inference model 62. In this way, the second inference model 62 is completed (S13).

[0075] The bottom part of Figure 5 shows inference by the second inference model 62. The second inference model 62 is trained using substitute images from the substitute image generation unit 50. The substitute image generation unit 50 generates substitute images based on a prompt based on feature C. As a result, the second inference model 62 can detect feature C from an image that contains feature C.

[0076] For example, the first and second inference models 61 and 62 are constructed to infer scenes of each step in laparoscopic surgery, using endoscopic images taken during laparoscopic surgery as training data. In this case, for example, features A and B indicate each step in which an internal body part is imaged using a laparoscope, and feature C indicates that the operating room is imaged using a laparoscope removed from the body. The first inference model 61 is trained using endoscopic images annotated with the contents of features A and B by the annotation unit 31, and can detect features A and B from the images. However, the image of feature C is converted to a black image, for example, to protect personal information, and is not trained by the first inference model 61. As a result, the first inference model 61 cannot detect feature C from the image of feature C, as shown in the upper part of Figure 5.

[0077] In this embodiment, the prompt generation unit 40 creates a prompt based on feature C. Based on this prompt, the substitute image generation unit 50 generates a substitute image. This substitute image is annotated using the prompt from the prompt generation unit 40, and the second inference model 62 is trained. As a result, the second inference model 62 can detect feature C from the image of feature C, as shown in the lower part of Figure 5.

[0078] FIG. 4 illustrates the classification process of an endoscopic image using the first and second inference models 61 and 62.

[0079] In S20 of FIG. 4 , an endoscopic image to be inferred is input from the image input unit 63. The image input unit 63 may be configured to provide the input endoscopic image to a display device (not shown) for display (S21). The image input unit 63 provides the input image to the first and second inference models 61, 62 (S22). Note that the first and second inference models 61, 62, for example, classify each scene of a medical procedure (area and procedure content) and output the classification results as inference results. The first and second inference models 61, 62 each perform inference on the input image. The first and second inference models 61, 62 output the inference results together with reliability information to the information output unit 70.

[0080] The control unit 10 determines whether inference results having reliability equal to or greater than a predetermined threshold are obtained from the first and second inference models 61 and 62 (S23). For example, in the example of FIG. 5, features A and B are detected with a relatively high reliability by the first inference model 61, and feature C is detected with a relatively high reliability by the second inference model 62. The control unit 10 controls the information output unit 70 to output and display the inference results having reliability equal to or greater than a predetermined threshold on the display device (S24). The control unit 10 also associates the inference results with the endoscopic images (S25). The endoscopic images associated with the inference results may be recorded in a recording device (not shown) and displayed. The control unit 10 determines whether the process is complete in S26. If the process is not complete, the control unit 10 returns to S20.

[0081] If the control unit 10 does not obtain an inference result with a reliability equal to or higher than a predetermined threshold value from the first and second inference models 61, 62 in S23, the control unit 10 proceeds to S27. In S27, the control unit 10 creates training data by referring to user opinions, comments, etc., and retrains the first and second inference models 61, 62, and proceeds to S26.

[0082] For example, it may be possible that the prompt generation unit 40 is unable to initially generate the prompt required to generate a substitute image similar to the unedited image Po. Even in this case, it may be possible to generate a substitute image similar to the unedited image Po by repeatedly regenerating the prompt in S27. This makes it possible to construct first and second inference models 61, 62 with high inference performance.

[0083] That is, in this embodiment, assuming that a series of multiple images constituting each frame of a video capturing a specific examination or treatment scene, for example, in which different patients or treating physicians are involved, is input to the image input unit 63, the first inference unit 61 is created by learning from first teacher data obtained by adding first classification scene text information as first annotation information to image frames having specific image features in a video captured in advance of scenes similar to the images input from the image input unit 63. When each frame of images from the image input unit 63 is provided to the created first inference unit 61 and an inference result is obtained from the first inference unit 61, a generated alternative image is generated for image frames whose reliability of the inference result is below a predetermined threshold using second classification scene text information other than the first classification scene text information, and the second classification scene text information is added to the generated generated alternative image as second annotation information, and the second inference unit is created by learning from second teacher data obtained. An inference device using these first and second inference units can reliably infer which scene each image frame belongs to.

[0084] In this way, in this embodiment, when an image different from the original image is given due to image processing during the creation of training data, a substitute image similar to the original image is generated using a prompt, and an inference model is constructed by learning using this substitute image, thereby enabling the construction of a highly reliable inference model.

[0085] Second Embodiment Fig. 6 is a flowchart showing an operation flow employed in a second embodiment. In Fig. 6, the same steps as those in Fig. 3 are assigned the same reference numerals, and a description thereof will be omitted. The hardware configuration in this embodiment is the same as that in the embodiment in Fig. 1.

[0086] In the first embodiment, whether or not to create the second inference model 62 was determined based on the judgment result of the reliability of the inference result output from the first inference model 61. In contrast, in this embodiment, the second inference model 62 is created in accordance with the judgment result of the condition unmet information section 32.

[0087] In S1 of Figure 6, information regarding the specifications of the first and second inference models 61, 62 is input. In S2, an image is input from the processing unit 22, and the input image is displayed on the display device. In this embodiment, the condition non-fulfillment information unit 32 of the teacher data creation unit 30 determines whether the input image is a condition non-fulfillment image Pm. If the input image is not a condition non-fulfillment image Pm (NO in S31), the annotation unit 31 annotates the image input from the processing unit 22 to create teacher data for the first inference model 61 (S32).

[0088] If the condition non-satisfaction information unit 32 determines that the input image is a condition non-satisfaction image Pm, it provides the prompt generation unit 40 with historical information about the unprocessed image Po of the condition non-satisfaction image Pm and information about the specifications of the inference model to create a prompt. The prompt generation unit 40 converts words and other elements that will become prompts into text according to the historical information and specifications (S33). In this case, various texts related to the historical information and specification information are also collected. The prompt amplification unit 41 modifies and organizes the collected text to increase the amount of text (text amplification) (S8).

[0089] The prompt generation unit 40 provides the text prompt to the substitute image generation unit 50. The substitute image generation unit 50 generates a group of images by searching using the prompt or processing by the generation AI (S9). Note that this group of images includes images for training data and test data.

[0090] The annotation unit 58 annotates the image for the training data from the substitute image generation unit 50 with the text generated as a prompt by the prompt generation unit 40 to create training data (S10).

[0091] The control unit 10 provides the annotated training data from the annotation unit 31 to the network N1 for learning, and also provides the annotated training data from the annotation unit 58 to the network N2 for learning (S34). The control unit 10 provides test data to the first inference model 61 and the second inference model 62 to perform inference using the first and second inference models 61 and 62. The control unit 10 determines whether the reliability of the inference results from the first and second inference models 61 and 62 is equal to or greater than a predetermined threshold (inference reliability OK) (S11). If a determination result equal to or greater than the predetermined threshold is not obtained, the control unit 10 repeats the learning process by selecting and rejecting training data, regenerating substitute images, redoing annotations, etc. (S12). If a determination result equal to or greater than the predetermined threshold is obtained, the control unit 10 terminates the learning of the first and second inference models 61 and 62. In this way, the first and second inference models 61 and 62 are completed (S36).

[0092] Other functions are the same as those of the first embodiment.

[0093] In this way, the same effects as those of the first embodiment can be obtained in this embodiment.

[0094] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied without departing from the spirit of the invention. Furthermore, the present invention can be applied to technologies such as improving inference models using hard-to-obtain images, and does not need to be limited to the medical field. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some of the components shown in the embodiments may be deleted. Furthermore, components from different embodiments may be appropriately combined.

[0095] Furthermore, among the technologies described herein, many of the controls and functions, mainly those described in the flowcharts, can be set by a program, and the above-mentioned controls and functions can be realized by a computer reading and executing the program. The program can be recorded or stored, in whole or in part, as a computer program product on portable media such as non-volatile memory, including flexible disks and CD-ROMs, or on storage media such as hard disks and volatile memory, and can be distributed or provided at the time of product shipment, via portable media, or via communication lines. A user can easily realize the inference device and inference model creation method of this embodiment by downloading the program via a communication network and installing it on a computer, or by installing it on a computer from a recording medium.

Claims

1. An inference device comprising: an acquisition unit that acquires a series of multiple images that are candidates for pre-obtained teacher data; a first inference unit created by learning first teacher data obtained by adding first annotation information to a first image frame of the series of multiple images acquired by the acquisition unit; a prompt generation unit that generates a prompt related to a second image frame different from the first image frame of the series of multiple images acquired by the acquisition unit based on history information assigned to the series of multiple images; a substitute image generation unit that generates a substitute image similar to the second image frame based on the prompt; and a second inference unit created by learning second teacher data obtained by adding second annotation information based on the prompt to the substitute image.

2. The inference device described in claim 1, wherein the first annotation information used when training the first inference unit is text information, and the prompt generated when training the second inference unit is different from the text information of the first annotation information.

3. The inference device of claim 1, wherein the second annotation information based on the prompt that is added to the alternative image is a summary or part of the prompt.

4. The inference device according to claim 1, wherein the second image frame is determined to be different from the first image frame by having different image data characteristics or by having a similarity lower than a predetermined value when compared with the first image frame.

5. The inference device described in claim 1, wherein the series of multiple images acquired by the acquisition unit includes a group of images for each video organized and recorded as multiple video files, and the first image frame is a similar image included in the multiple video files.

6. The inference device described in claim 1, wherein the history information assigned to the image may be at least one of the file name of the video or still image, the folder name, the recording source of the video or still image, the sender, and the conditions at the time of reception, or may be information obtained by using text information contained in the video or still image by referring to contract information at the time of data transfer, or may be information obtained by converting encoded information into text.

7. The inference device described in claim 1, wherein the second image frame is acquired by the acquisition unit after being converted into an image that does not correspond to the first annotation information by image processing, or is an image of image quality that does not correspond to the first annotation information.

8. The inference device according to claim 1, wherein the prompt generation unit generates the substitute image based on the prompt using a generation AI.

9. The inference device according to claim 1, wherein the second inference unit is created when the reliability of the inference result of the first inference unit when a specific image is input to the first inference unit is lower than a predetermined threshold value.

10. The inference device described in claim 4, wherein the prompt generation unit generates the prompt based on at least one of information on the reliability of the inference result of the first inference unit, the history information, and information on the specifications of the first and second inference units.

11. The inference device described in claim 1, wherein the first annotation information includes information classifying scenes of medical procedures, and the prompt generation unit generates the prompt in the form of text related to one of the scenes of the medical procedures based on surgical procedure information indicating the surgical procedure of the medical procedures included in the history information.

12. The inference device of claim 1, wherein the prompt generator generates the prompt using at least one of text, qualification, sentence formation, occurrence of synonyms, and word combinations related to any of the medical procedure scenes.

13. The inference device described in claim 1, wherein the prompt generation unit determines at least one of the reasons why the second image frame is missing and why the second image frame is an incomplete image that does not correspond to the first annotation information, and generates the prompt based on the determination result.

14. A method for creating an inference model, comprising: acquiring a series of multiple images; creating a first inference unit by learning first teacher data obtained by adding first annotation information to a first image frame of the acquired series of multiple images; if the reliability value of the inference result of the first inference unit when a specific image is input to the first inference unit is lower than a predetermined threshold, generating a prompt related to a second image frame different from the first image frame of the acquired series of multiple images based on history information assigned to the series of multiple images; generating an alternative image similar to the second image frame based on the generated prompt; and creating a second inference unit by learning second teacher data obtained by adding second annotation information based on the prompt to the generated alternative image.

15. The inference model creation method described in claim 14, wherein the second image frame is an image that has been converted by image processing into an image that does not correspond to the first annotation information, or an image of image quality that does not correspond to the first annotation information.

16. An inference device comprising: a first inference unit created by learning from first teacher data obtained by inputting, via an acquisition unit, a video obtained by previously capturing a scene similar to a specific examination or treatment scene in which a series of multiple images are input via an image input unit, and adding first classification scene text information as first annotation information to image frames in the input video that have specific image features; and a second inference unit created by learning from second teacher data obtained by providing each image frame of the images input via the image input unit to the created first inference unit and obtaining an inference result from the first inference unit, and adding the second classification scene text information as second annotation information to a generated substitute image generated using second classification scene text information other than the first classification scene text information for image frames whose reliability of the inference result is below a predetermined threshold.

17. A method for creating teacher data comprising the steps of: obtaining first teacher data by adding first classification scene text information as first annotation information to image frames having specific image features in multiple videos previously filmed of specific examination and treatment scenes; and obtaining second teacher data by adding second classification scene text information as second annotation information to generated substitute images generated using second classification scene text information other than the first classification scene text information assigned to the multiple videos.

18. A method for creating an inference model, comprising: acquiring a plurality of video files; determining the image features of each frame constituting each video; creating a first inference unit by learning first teacher data obtained by adding first annotation information to a group of images among the frames constituting each of the videos that have a first image feature; generating a prompt based on the annotation information added to the video file for a group of images among the frames constituting each of the videos that have a second image feature different from the first image feature; generating a substitute image similar to the second image frame based on the generated prompt; and creating a second inference unit by learning second teacher data obtained by adding second annotation information based on the prompt to the generated substitute image.

19. A method for creating teacher data, comprising: acquiring a plurality of video files; determining the characteristics of each frame constituting each video; creating first teacher data by adding first annotation information to a group of images among the frames constituting each of the videos that have a first image characteristic; generating a prompt based on the annotation information added to the video file for a group of images among the frames constituting each of the videos that have a second image characteristic different from the first image characteristic; generating a substitute image similar to the second image frame based on the generated prompt; and creating second teacher data by adding second annotation information based on the prompt to the generated substitute image.

20. A method for creating an inference model, comprising: acquiring a plurality of medical videos capturing examination and treatment processes performed in accordance with a specific affected area and surgical procedure; creating a first inference unit by learning first training data obtained by adding first classification scene text information for classifying image frames having a first image feature from the acquired videos as scenes from the specific examination and treatment process; and creating a second inference unit by learning second training data obtained as training data from generated images obtained using prompt information that utilizes second classification scene text information different from the first classification scene text information for classifying image frames having a second image feature different from the first image feature from the acquired videos as scenes from the specific examination and treatment process.

21. The inference model creation method described in claim 20, wherein the second classification scene text information, which is different from the first classification scene text information, uses words obtained by determining the similarity and difference of the meanings of words by representing the meanings of words or words searched as synonyms or sentence examples through a text search using a dictionary or the Internet using multiple evaluation values ​​and vectorizing them.

Citation Information

Patent Citations

  • Image file generating apparatus, image file generating method, image management apparatus, and image management method

    JP2020123174A

  • Detection device and detection method

    JP2022040912A