Information processing method and information processing device
The method addresses the challenge of sensor diversity and unknown situations by selecting training data based on similarity with historical data, enhancing the efficiency and accuracy of inference model generation.
Patent Information
- Application Number
- JP2025086428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-13
AI Technical Summary
Existing signal analysis systems fail to account for differences in sensor characteristics and unknown situations when generating machine learning models, leading to inconsistent annotations and difficulty in handling diverse data inputs.
An information processing method that records and compares endoscopic images with historical data to select candidate training data based on similarity, utilizing the history of training data and image groups to efficiently generate inference models for sensors with different characteristics or unknown situations.
Enables efficient generation of inference models by selecting relevant training data, reducing unnecessary effort and improving model accuracy for diverse sensor inputs.
Smart Images

Figure 2025119024000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing method and an information processing device that can efficiently and effectively utilize the acquired image data when annotating images acquired by an inspection device such as an endoscope, creating training data, and generating an inference model using this training data. [Background technology]
[0002] A signal analysis system and method have been proposed that first collects raw data to generate an inference model (AI model), then extracts and interprets data features from the raw data for generating an inference model (see Patent Document 1). That is, in this signal analysis system, raw signal data is input, the origin of the features is traced back to the signal data, and the features are mapped to domain / application knowledge. Then, features are extracted using a deep learning network, a machine learning model is implemented for sensor data analysis, and causal analysis is performed for fault prediction. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-87221 Summary of the Invention [Problem to be solved by the invention]
[0004] The signal analysis system described in the aforementioned Patent Document 1 continuously performs classification according to instructions from experts, enabling effective use of data. However, the sensors used to generate raw data have different characteristics, and therefore, when performing deep learning to generate a machine learning model (inference model), differences in characteristics between models (e.g., sensors attached to devices) must be taken into consideration. However, Patent Document 1 does not take this into consideration at all. For this reason, with Patent Document 1, it is difficult to ensure that annotations for different models are of the same quality.
[0005] Furthermore, there may be cases where a machine learning model (inference model) is generated by inputting data from a sensor or the like having unknown characteristics, but this point is also not taken into consideration in Patent Document 1. For this reason, it is difficult for Patent Document 1 to deal with unknown situations.
[0006] The present invention has been made in consideration of the above circumstances, and aims to provide an information processing method and information processing device that can efficiently generate an inference model when data from sensors with different characteristics is input, or when data that can be considered an unknown situation is input. [Means for solving the problem]
[0007] In order to achieve the above-mentioned object, the information processing method of the first invention records, when creating a first inference model, information obtained when an endoscope is inserted and during examination, along with information indicating the history of the adopted training data, and compares the newly acquired group of endoscopic images with a first group of images at a time when an image of the affected area is included and a second group of images at a second other time, which are included in the historical video of the training data, and selects a group of images from the newly acquired group of endoscopic images to be candidates for training data for the inference model based on the similarity.
[0008] An information processing method according to a second aspect of the present invention is the information processing method according to the first aspect of the present invention, wherein the training data is recorded together with information indicating whether or not the training data has been adopted as annotation data. An information processing method according to a third invention is the information processing method according to the second invention, wherein information indicating the first image group and the second image group is recorded together with recording the training data.
[0009] An information processing method according to a fourth invention is the information processing method according to the third invention, wherein the group of images serving as the training data candidates is selected for each case.
[0010] The information processing device of the fifth invention comprises a recording unit that records, when creating a first inference model, information obtained when an endoscope is inserted and during examination and indicates the history of the adopted training data, and a selection unit that compares a newly acquired group of endoscopic images with a first group of images at a time when an image of the affected area is included and a second group of images at a second other time, which are included in the historical video of the training data, and selects, based on the similarity, a group of images from the newly acquired group of endoscopic images to be candidates for training data for the inference model.
[0011] An information processing method according to a sixth aspect of the present invention is an information processing method in an information processing system capable of creating an inference model for an endoscope, the information processing system having an inference unit, a judgment unit, a classification unit, and a learning unit, and the information processing method includes determining whether a group of images obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine specific image features for images obtained continuously from the first endoscope, by the judgment unit, and generating an inference model based on the group of images obtained continuously from the first endoscope and obtained at the second timing. The inference unit is provided with a first inference model created by the learning unit based on images created by the second endoscope, and the judgment unit determines whether a group of images obtained in time series by the second endoscope is a group of images obtained at a first timing or a group of images obtained at a second timing.The classification unit selects and classifies a group of images from the newly acquired endoscope that will be candidates for training data for improving the first inference model based on the similarity between the group of images obtained at the first timing when the first inference model was created and a group of images from the newly acquired endoscope.
[0012] An information processing device according to a seventh aspect of the present invention is an information processing device in an information processing system capable of creating an inference model for an endoscope, wherein a learning unit is provided within the information processing system, and the information processing device includes a determination unit that determines whether a group of images obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine specific image features for images obtained continuously from the first endoscope, and a learning unit that determines whether a group of images obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing based on images obtained continuously from the first endoscope and created based on the image group obtained at the second timing. The information processing device has an inference unit that has a first inference model created by the learning unit, and the judgment unit determines whether a group of images obtained in time series by a second endoscope is a group of images obtained at a first timing or a group of images obtained at a second timing.The information processing device further has a classification unit that selects and classifies a group of images from the newly acquired image group from the endoscope as candidates for training data for improving the first inference model based on the similarity between the group of images obtained at the first timing when the first inference model was created and the group of images from the newly acquired endoscope.
[0013] An information processing method according to an eighth aspect of the present invention is an information processing method in an information processing system capable of creating an inference model for an endoscope, the information processing system having an inference unit, a determination unit, a classification unit, and a learning unit, and the information processing method includes determining whether a group of images obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine specific image features for images obtained continuously from the first endoscope, and creating an inference model based on the image group obtained continuously from the first endoscope and obtained at the second timing. The inference unit is provided with a first inference model created by the learning unit based on the acquired images, and the judgment unit determines whether a group of images acquired in time series by the second endoscope is a group of images acquired at a first timing or a group of images acquired at a second timing.The classification unit selects and classifies a group of images from the newly acquired endoscope that will be candidates for training data for an inference model different from the first inference model based on the similarity between the group of images acquired at the first timing when the first inference model was created and the group of images from the newly acquired endoscope.
[0014] An information processing device according to a ninth aspect of the present invention is an information processing device in an information processing system capable of creating an inference model for an endoscope, wherein a learning unit is provided within the information processing system, and the information processing device includes a determination unit that determines whether a group of images obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine specific image features for images obtained continuously from the first endoscope, and a learning unit that determines whether an inference model for an endoscope is created based on the image group obtained continuously from the first endoscope and the image group obtained at the second timing. The information processing device further has an inference unit having a first inference model created by a learning unit, and the judgment unit determines whether a group of images obtained in time series by a second endoscope is a group of images obtained at a first timing or a group of images obtained at a second timing, and the information processing device further has a classification unit that selects and classifies a group of images from the newly acquired image group from the endoscope as candidates for training data for an inference model different from the first inference model based on the similarity between the group of images obtained at the first timing when the first inference model was created and the group of images from the newly acquired endoscope.
[0015] The information processing method of the 10th invention acquires an endoscopic video during a specific examination, and uses frames of the acquired endoscopic video corresponding to teacher images as annotation information to obtain the video information as teacher data.After acquiring the endoscopic video during the specific examination, multiple endoscopic images including insertion or removal images are further acquired, and the endoscopic video including the insertion or removal images is input into a learning device, and learning is performed to create an inference model so that the frames of the teacher data are output.
[0016] The information processing method of the 11th invention is the same as the 10th invention, in which, during the specific examination, the endoscopic video is acquired using a first endoscope, and the endoscopic video of the multiple examination results is acquired using a second endoscope. The information processing method of the 12th invention is the same as the 11th invention, in which a first inference model is created based on a group of images obtained from the first endoscope at a second timing, and a group of images to be used as second training data candidates is selected from the group of images from the first endoscope or the second endoscope based on the similarity between the group of images obtained at the first timing when the first inference model was created and a newly acquired group of images from the first endoscope or the second endoscope. The information processing method of the 13th invention is the same as the 12th invention, wherein the first inference model is an inference model for displaying inference results for images obtained by the first endoscope, and the second inference model is an inference model for displaying inference results for images obtained by the second endoscope. The information processing method of the 14th invention is the 12th invention, in which, when an image is input into the first inference model, the reliability of the inference result is lower than a predetermined value, and it is determined that the input image is an image from the second endoscope.
[0017] The information processing device of the 15th invention comprises a first image acquisition unit that acquires an endoscopic video during a specific examination, a teacher data creation unit that acquires video information by using frames of the endoscopic video acquired by the first image acquisition unit that correspond to teacher images as annotation information, a second image acquisition unit that acquires a plurality of endoscopic images including insertion or removal images after acquiring the endoscopic video during the specific examination, and a request unit that inputs the endoscopic video including the insertion or removal images into a learning device and requests learning to create an inference model so that the frames of the teacher data are output. [Effects of the Invention]
[0018] According to the present invention, it is possible to provide an information processing device and an information processing method that can efficiently generate an inference model when data from sensors with different characteristics or data that can be considered an unknown situation is input. [Brief explanation of the drawings]
[0019] [Figure 1A] 1 is a block diagram showing the configuration of an information processing system and its peripheral systems according to an embodiment of the present invention; [Figure 1B] 1 is a block diagram showing the configuration of an information processing system according to an embodiment of the present invention and a portion of its peripheral system; [Figure 2] A diagram showing the relationship between image data and the generation of an inference model in an information processing system and its peripheral systems according to one embodiment of the present invention. [Figure 3A] A figure showing a flowchart showing the operation of creating a new inference model in an information processing system and its peripheral systems according to one embodiment of the present invention, and an example of an input image. [Figure 3B] A flowchart showing a modified example of the operation of creating a new inference model in an information processing system and its peripheral systems according to one embodiment of the present invention. [Figure 4] This figure shows a case where an existing inference model can correctly infer an input image in an information processing system according to one embodiment of the present invention and its peripheral systems, and a case where an existing inference model cannot correctly output an input image. [Figure 5A] 1 is a flowchart showing the operation of an endoscope 1 in an information processing system and its peripheral system according to an embodiment of the present invention. [Figure 5B] 1 is a flowchart showing the operation of an information processing system according to an embodiment of the present invention and its peripheral systems. [Figure 5C] 3 is a flowchart showing the operation of an endoscope 2 in an information processing system and its peripheral system according to one embodiment of the present invention. [Figure 6] 10 is a flowchart showing the operation of determining whether the information processing system according to one embodiment of the present invention is an improvement on the current model, a new model, or something else in the information processing system and its peripheral systems. [Figure 7] 1 is a flowchart showing the operation of a learning device in an information processing system and its peripheral systems according to an embodiment of the present invention. [Figure 8]This figure explains the generation of an inference model for inferring an image to be annotated, and inference using the generated inference model, in an information processing system and its peripheral systems according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0020] The information processing device and information processing method according to the present application efficiently collect new training data when creating a new or improved inference model. To this end, the features of past training data, so to speak, used in machine learning when creating a proven inference model, are utilized. The past training data used in creating a proven inference model was often created with great effort, as selecting images that contain the target object to be detected from among countless images already requires a great deal of effort.
[0021] Therefore, in one embodiment of the present invention, when creating an inference model similar to an inference model obtained by learning training data obtained by annotating some frames of a series of images, it is desirable to narrow down the number of images that are too large to annotate for the training data frames. Therefore, prior to annotation, annotation image frames are selected based on the creation history of the model inference model. In other words, when creating the model inference model, a step is provided in which image frames that were not selected as annotation targets from the series of images are referenced to select the newly obtained series of images as annotation candidate images for a new (or improved) inference model (e.g., see S5 and S7 in Figure 3A(a), S103 and S107 in Figure 3B, S29 and S39 in Figure 5B, etc.). Inference models with proven specifications will likely also have proven training data. Therefore, if creating a model with similar specifications, it is possible to reduce unnecessary effort by imitating the previous model and selecting candidate image frames for training data.
[0022] Inference models are used in a variety of situations. For example, with endoscopes, there are inference models with various specifications, such as inference models that provide navigation displays during the process to guide access to certain areas, inference models that alert doctors to potential lesions so they don't miss them, inference models that guide how to observe suspected lesions, and inference models that help identify them. The following examples will mainly focus on inference models for differentiation and observation, but it goes without saying that they can also be applied to inference with other specifications. All of these models select the necessary frames from videos of a series of examinations and procedures to create training data.
[0023] For example, even if the examination time is only 10 minutes, if the video is shot at 60 frames per second to emphasize smoothness, the total number of frames will be a massive 36,000. Furthermore, when creating an inference model for lesion detection for gastrointestinal endoscopy, the access images from the body cavity to the affected area inside the digestive tract are left unused, resulting in a huge number of frames. The same goes for images taken when the endoscope is removed after the affected area has been confirmed. Therefore, images of the affected area must be found from the remaining images after removing these portions. Incidentally, when creating inference models with other specifications, it is possible that the access images themselves are important.
[0024] Next, the remaining images are checked for the presence of affected areas, making sure not to miss multiple areas. By performing this checking operation, the images are classified into images without affected areas and images with affected areas, which are used to find and confirm the nature of affected areas. Normally, the images without affected areas make up the majority. Images with affected areas are selected, and a doctor or other expert looks at the frames that show affected areas and annotates them as affected areas. Even if it takes just one minute to check the affected areas, 3,600 frames will be the target. Annotating each frame (including writing the name of the disease, symptoms, and extent) also requires a huge amount of effort.
[0025] Because it is difficult to create an inference model that can be applied to many cases from a single case, it is necessary to devise an information processing process that narrows down the images to be annotated after removing access images and non-lesion confirmation images from the image group. Without effective narrowing down, doctors and specialists must perform a huge amount of work outside their area of expertise. In other words, without removing access images and non-lesion confirmation images, the endoscopic images remain overwhelming, preventing doctors from spending enough time on important tasks, such as accurately identifying lesions in the critical lesion confirmation images, or wasting time while trying to reach those tasks. This makes it impossible to quickly and efficiently create an inference model trained with high-quality training data.
[0026] The above describes an example of creating an inference model that identifies an affected area, distinguishes the type of lesion, and detects the extent of the affected area. However, this example is not limiting; other examples include creating inference models that correctly guide endoscopic operation and creating inference models that prevent oversight by finding easily overlooked affected areas. When creating such an inference model, images of scenes that should guide endoscopic operation and images that are easily overlooked become training data. In these cases, as with the reasons mentioned above, target images that serve as candidate training data corresponding to the key scenes must be selected from tens of thousands of frames, and asking experts to perform annotation is not efficient.
[0027] In other words, when creating an improved version of an existing inference model or creating a new inference model with similar specifications for a model other than the one originally intended, it is possible to create the inference model efficiently by referring to the image selection method used to create a proven inference model with similar specifications. Images acquired using an endoscope are collected through the same process for each examination, and even if information such as "this is an access image," "this is an image to confirm that there is no lesion," or "this is an image to confirm that there is a lesion" is not assigned to every frame of the endoscopic image in the process of creating the inference model, if there is information from the selection results in the process of creating a large amount of training data, this information can be used to easily narrow down which images can become training data.
[0028] In other words, when creating an inference model by learning from training data obtained by annotating some frames of a sequence of images, a process is required to find candidate training data from among many frames. However, if the condition is to create an inference model with pre-existing model specifications, it is sufficient to refer to the image frames that were not targeted for annotation from the original sequence of images used to create the training data for the model inference model, remove the candidate images that were not targeted for annotation from the newly obtained sequence of images, and select candidate annotation images from the remaining frames. Learning methods, systems, and programs that follow these steps can make use of the know-how gained when creating the model inference model.
[0029] For example, it is possible to create a special inference model that performs inference using images used as training data and the entire endoscopic image. Using the above-mentioned concept, annotation information is added to frames of endoscopic video information obtained during a specific examination that correspond to training images, and the video information obtained by this process is used as training data. Furthermore, similar annotations are added to endoscopic video images (which may include insertion or removal images) from multiple examination results to create training data. Once training data is created, an inference model can be created by learning, with endoscopic video including insertion or removal images as input and training data frames as output.
[0030] In the above-described method for creating an inference model, there is a possibility that a training data image to be acquired may be overlooked in the following two situations. In other words, it is assumed that a training data candidate may be overlooked in the following two situations. First, when using the intended endoscope (one for which an inference model has already been created), the acquired images, observation methods, and operation methods may be different for endoscopic images that are not typical patterns, such as rare cases that are completely different from those in the past or endoscopic images that involve unusual procedures, and therefore the training data may not be correctly inferred. Second, when using an endoscopic device that has a different operation method, image quality, etc. from the intended endoscopic device, the training data may not be correctly inferred. In such cases, it is desirable to relax the conditions for images to be used as training data to minimize oversight. Therefore, in an embodiment of the present invention, a method is adopted in which images with low priority are eliminated in order to efficiently handle cases with as few precedents as possible while taking into account unexpected cases. Furthermore, while it is preferable for experts such as doctors to annotate candidate training data when differentiating pathologies, an existing inference model can be used to assist with this. If the doctor gives the OK, the results can also be used as training data as is.
[0031] As an embodiment of the present invention, an example in which the present invention is applied to an information processing system and its peripheral systems will be described below. This system is originally designed to create and improve an inference model for an inspection device such as a first endoscope (e.g., an existing endoscope), but it also manages training data used when creating the inference model for the first endoscope. This is because it is preferable to manage training data and the endoscopic images that served as the source of the training data when improving an inference model, creating an inference model with different specifications, or clarifying the history of the inference model. However, this system is designed not only to improve existing models, but also to create inference models for a second endoscope or other inspection device different from the first endoscope or other inspection device. The inference model created here for the first endoscope or other inspection device determines specific image features, such as determining whether or not a lesion such as cancer is present, in images obtained from the first endoscope or other inspection device (see, for example, S25 in FIG. 5B).
[0032] To clarify the history of the inference model, we have described the management of training data and the endoscopic images that served as the basis for that training data. The original endoscopic images are, for example, images obtained during an endoscopic examination of a specific patient, recorded in a video format or, as appropriate, in a still image format taken by a physician during the examination. Because the images are obtained during the process of inserting an endoscope into a body cavity, performing the examination, and removing the endoscope from the body cavity, a video consisting of a sequence of several minutes or more of image frames or accompanying still images are recorded for each patient and case. This video includes a group of images in chronological order during insertion, examination, and removal. Images during the examination include images of the detection and examination of lesions, as well as images of the screening for the absence of lesions. In other words, this explains how image sequences along various time axes can be divided into multiple image groups. The multiple image groups may be simply referred to as first and second image groups.
[0033] In particular, if the image group during the examination is divided into two, the image group obtained in time series by the first endoscope may be the image group obtained at the first timing (unnecessary images during access), or the image group obtained at the second timing (necessary images when searching for an affected area such as a tumor). For example, when determining which of these is true (see, for example, S5 in FIG. 3A(a), S23 in FIG. 5B, and images Ps and Po in FIG. 3A(b)), these can be considered the first and second timings, but the image groups obtained at the time of insertion and removal can also be considered the third and fourth image groups. Furthermore, the insertion image group and the removal image group can also be considered to be screening and can be included in the first timing. Here, the description will be given assuming that the images obtained at the second timing include images of the affected area.
[0034] If an affected area is found in the image group of the first timing, it may be classified as the image group of the second timing, and if no affected area is found in the image group of the second timing, it may be reclassified as the image group of the first timing. The first timing and the second timing do not divide the time-series image group into two at a specific time point, but rather divide the time series into areas as multiple timings as appropriate, such as the first timing here, the second timing here, and then the first timing again.
[0035] The information system also works in conjunction with a proven learning unit (e.g., learning unit 18 in Fig. 1B) that has already created an inference model. This inference model is a first inference model for determining image features of the first endoscopic image, and has been trained using the training data that is the result of annotating the image group obtained at the second timing, as described above.
[0036] This information processing system has a proven track record of creating inference models for the first endoscope, and is therefore well-suited to creating inference models for other types of endoscopes, etc. Therefore, it may be possible to use a group of images suitable for creating the first inference model as needed to create a new inference model for a completely unknown second endoscope that is different from the first endoscope, or it may have a function for collecting and managing training data, and is equipped with a classification unit (e.g., classification unit 15) that classifies images so that training data can be efficiently created from a group of images acquired from the second endoscope.
[0037] As described above, in one embodiment of the present invention, images from a second endoscope different from the first endoscope can be classified. Therefore, even if the second endoscope is an unknown endoscope different from the first endoscope, the information processing system can efficiently collect images suitable for generating an inference model for the second endoscope. That is, efficiency is improved by applying the training data, technology, logic, know-how, etc., used when creating an inference model for the first endoscope. In other words, when creating training data for the first endoscope, frames that serve as training data are extracted from the large number of video frames collected for each examination, and an inference model has been created using that training data. Therefore, when a video of an examination is acquired from the second endoscope, frames that can be used as training data can be selected from the acquired video to create an inference model for the second endoscope (by learning the training data) using the same approach as when creating an inference model for the first endoscope.
[0038] That is, a training image is selected from a series of endoscopic images obtained during a single examination using the first endoscope, and the training image frames from the series of images are annotated and used as training data to train an inference model. Since a single examination alone will not provide enough training data, similar processing is performed on more examination results. The input information for learning is a group of endoscopic examination image frames including insertion and removal images, and the output information is the frames adopted as training data, and learning is performed to create an inference model. Once the inference model is created, a group of examination images from the second endoscope can be input into the inference model, and candidate training data can be output as a result of inference.
[0039] 1A and 1B are block diagrams illustrating the configuration of an information processing system and its peripheral systems according to an embodiment of the present invention. This system includes an information processing system 10, an imaging system 20, and a recording unit 30. FIG. 1A shows the entire system, while FIG. 1B shows the internal configuration of the control unit 11 and recording unit 30 within the information processing system 10. The imaging system 20 is included in an inspection device such as an endoscope and acquires image data such as endoscopic images. The information processing system 10 acquires the image data acquired by the imaging system 20, and the acquired image data is recorded in the recording unit 30. The acquired image data is annotated to create training data, and an inference model is generated using this training data. In a system utilizing an inference model 19a, it is preferable to be able to prove what training data was used to acquire the model. Transparency of the training data is also required, and the training data is recorded in the recording unit 30.
[0040] The information processing system 10 is provided in one or more servers or the like. The recording unit 30 is also provided in a server or the like. The information processing system 10 and the recording unit 30 may be provided in the same server, or may be provided separately. In this embodiment, the learning unit 18 is described as being provided in the information processing system 10, but it may also be provided outside the information processing system 10. The image system 20, the information processing system 10, and the recording unit 30 may each be connected via a communication network such as the Internet or an intranet, and may be capable of data communication.
[0041] As described above, the imaging system 20 is provided in an examination device such as an endoscope. The imaging system 20 has a control unit 21, a treatment unit 22, an image acquisition unit 23, and a display unit 24. In addition to the above-mentioned units, the imaging system 20 also has a communication unit (having a communication circuit) and can transmit and receive data to and from the information processing system 10 and the like through the communication unit.
[0042] The control unit 21 controls the entire imaging system 20. The control unit 21 is composed of one or more processors having a processing device such as a CPU, a memory that stores a program, etc., and can control each part in the imaging system 20 by executing the program.
[0043] The treatment unit 22 is inserted into the body or other part of the body to perform treatments to observe the interior. Industrial endoscopes and medical endoscopes are essentially imaging devices inserted into the interior of something to observe the interior. However, in reality, they contain not only imaging devices but also devices and mechanisms for insertion and observation. The treatment unit 22 adjusts the position and orientation of the imaging unit at the tip of the endoscope and, as necessary, requires functions such as irrigation and suction, and is equipped with dedicated tubes to meet these functional requirements. It also drives treatment tools (biopsy forceps, snares, high-frequency knives, syringe needles, etc.) used for biopsies and treatments. It also includes functions such as drug administration. Furthermore, the treatment unit 22 includes an illumination unit, which not only illuminates the interior but also switches light source characteristics (such as wavelength) to capture reflected images of the surface of the object or objects deeper within. Furthermore, if the treatment unit 22 has functions capable of staining or fluorescent observation, the treatment unit 22 may also have these corresponding functions.
[0044] The image acquisition unit 23 includes a photographing lens, an image sensor, an image acquisition control circuit, an image processing circuit, etc., and acquires image data. During an examination, the tip of the endoscope is inserted into the body, and the image acquisition unit 23 can acquire image data of the inside of the body, for example, the esophagus, stomach, or intestines. The image acquisition unit 23 may have functions for zooming, close-up photography, and acquiring 3D information. Furthermore, the image acquisition unit 23 may have multiple imaging units, and as described above, may acquire images with different characteristics of the same part in cooperation with switching of the irradiating light, etc., thereby enabling various image analyses.
[0045] The image acquisition unit 23 can continuously acquire images when the distal end of the endoscope is inserted into the body. In particular, endoscopic examinations, which observe the inside of the human body along the digestive tract and other areas, involve insertion through a body cavity, such as the oral cavity, nasal passage, or anus. Confirmation observations begin at the beginning of insertion, followed by actual examinations at various sites, and the endoscope is finally removed from the insertion point and the examination ends. Therefore, during endoscopic examinations, very similar images are obtained for each target examination. Driving along a path while checking the image results can be considered a technique similar to driving a car, but differs significantly in the above-mentioned respects. Therefore, the image acquisition unit 23 can continuously acquire images from the beginning of image acquisition when the distal end of the endoscope is inserted into the body until the examination is completed and the endoscope is withdrawn from the same location.
[0046] In Figure 1A, input images P1, P2, and P3 are sequentially captured image data (which may be expressed as a video, or individual frames of the video). Input image SP is a single frame from a video, or a still image intentionally captured by a doctor or medical professional for recording. Since many inference models are designed to determine what appears in which parts of a still image, the term "input image" is used to refer to the inference model. In most cases, video is used for inference, and each frame of the video is input to the inference model, so there is no need to refer to it as a still image, but we have used the term for clarity. Of course, inference may be performed from multiple frames, in which case the input image does not need to be a still image. However, using the term "still image" makes it easy to imagine applications such as recording ordinary photographic images with various annotations. Therefore, category information Pc is associated with this input image SP.
[0047] Category information Pc (a type of annotation) is information about the device that acquired the image, such as the endoscope manufacturer or model name. Category information may be assigned to each of images P1-P3. Alternatively, multiple images may be organized into a file and category information may be associated with the image file. Image quality and other characteristics vary depending on the sensors and image processing circuits within the image acquisition unit 23. By assigning category information Pc to image data, the image quality characteristics can be easily determined. The blank space below the category information Pc is used to record other annotations as metadata. Recording various additional information in this area and associating it with the image allows users to determine the type of image by simply looking at the metadata, without having to manually determine the image content. Images can also be searched for using other information, such as the shooting conditions and supplementary information. Of course, such metadata can also be used to improve the efficiency of learning and inference. If annotation information and other information are assigned during learning, advanced inference, such as multimodal methods, becomes possible.
[0048] The display unit 24 has a display or the like and can display the endoscopic image acquired by the image acquisition unit 23. When displaying this, the display content 25 inferred by the inference engine 19 in the information processing system 10, i.e., the inference result 25a and reliability 25b, are also displayed. The inference result 25a is displayed on the endoscopic image acquired by the image acquisition unit 23 so that the position of a lesion such as cancer can be seen.
[0049] The information processing system 10 includes a control unit 11 and an inference engine 19 equipped with an existing inference model 19a. The control unit 11 controls the entire information processing system 10. The control unit 11 is composed of one or more processors including a processing device such as a CPU, a memory storing a program, etc., and can control each unit in the information processing system 10 by executing the program. As shown in FIG. 1B, the control unit 11 includes an input determination unit 12, an image similarity determination unit 13, a metadata assignment unit 14, a classification unit 15, a first request unit 16, a second request unit 17, a learning unit 18, and a recording control unit 18A. To ensure transparency of the training data used to learn the existing inference model 19a, information such as the history and specifications is recorded in the recording unit 30.
[0050] The input determination unit 12 inputs image data acquired by the image acquisition unit 23. When the information processing system 10 and the image system 20 are connected via a communication network such as the Internet or an intranet, image data such as input images P1 to P3, SP, and category information are input via this communication network.
[0051] The input determination unit 12 also determines category information, metadata, etc. associated with the input image data. It may also determine whether an image is from a second endoscope based on an inference result obtained when an image acquired by the image acquisition unit 23 is input to the existing inference model 19a (see, for example, S23 in FIG. 5B). Because the existing inference model 19a is generated based on an image from a first endoscope, the reliability of the inference result for an image from a second endoscope is low. The input determination unit 12 functions as an input determination unit that determines whether an image is from a second endoscope based on the result obtained when an image is input to the first inference model (see, for example, S23 in FIG. 5B). The input determination unit 12 also functions as a video acquisition unit that acquires a new examination video when the learning unit is learning to create a second inference model different from the first inference model (see, for example, S21 in FIG. 5B).
[0052] Furthermore, if category information Pc is not associated with the input image, the input determination unit 12 determines the characteristics of the input image by taking into consideration various information such as the model of the imaging system 20, image quality (optical characteristics), the lighting conditions of the image, the angle of view of the image, and the shape of the image. The characteristics of the input image may also be determined based on the shape of a treatment tool that appears in the image, or, in the case of a video, the state of operation. This allows for determining what kind of device the image is from, what purpose the same device is being used for, and the skill of the doctor or medical professional operating it.
[0053] The image similarity determination unit 13 determines whether the image input by the input determination unit 12 is similar to any of the images recorded in the first-timing image section 31c and the second-timing image section 31d of the group of examination images 31 used during existing training data creation in the recording unit 30. The first-timing image section 31c and the second-timing image section 31d record which data was used as training data and which data was not used as training data when creating the existing inference model, as well as which data was used as annotation data and which data was not used. For example, if the existing inference model 19a is an inference model for differentiation, the image recorded in the first-timing image section 31c is an image acquired at the first timing, i.e., an image acquired when a doctor or other medical professional accesses a lesion such as a tumor. The first-timing image recorded in the first-timing image section 31c is an image that is not generally used as training data for generating an existing inference model (first inference model) for the first endoscope. The second timing image recorded in the second timing image section 31d is an image acquired at a second timing, i.e., an image acquired at a time when a doctor or the like is searching for and examining a lesion such as a tumor, or an image adopted as the lesion.
[0054] Due to the characteristics of endoscopes, safety, and the importance of visual confirmation during operation, images are recorded of the insertion process leading up to the specific site examination, as well as the removal process after the examination. Images are also recorded of the process of the endoscope being inserted deeper and then withdrawn. Because image changes are common regardless of the user, the object, or the equipment used, the above-mentioned similarity determination is effectively utilized to separate images along a time axis into images acquired at a first timing and / or a second timing, or even images acquired at a third and fourth timing. For example, an image similar to the image acquired at insertion can be obtained during removal, as the insertion passes through the same site, seemingly going back in time. This means that if the image at insertion is considered the first timing image, the image at removal can be classified as the same first timing image but different from the second timing image.
[0055] In addition, image classification is possible. Examples of similar images that are unique to images obtained with an endoscope include situations in which a doctor or other medical professional magnifies and observes the same lesion (affected area) such as a tumor, or performs special observations by changing the wavelength of the light source or the image processing characteristics. These images can also be classified, for example, as second-timing images and treated as similar images. In this case, images before and after magnification can be treated as one set, or images with and without special (light) observation can be treated as another set, and images in such sets can be treated as second-timing images and distinguished from first-timing images. As explained above, examination videos are evidence that record the constantly changing procedures of an examination, and can be easily organized as a series of still images classified along a time axis.
[0056] Therefore, the first timing corresponds to the time when a doctor or other medical professional is operating an endoscope or other device and moving it to a target location within the body. In this case, for example, if an inference model is assumed to detect the type of tumor in the affected area, since there are no candidate training data, this corresponds to the time when the doctor or other medical professional is accessing the affected area. Here, we will explain this case as an example (see image Ps in Figure 3A(b)). The images acquired at this first timing are not intended for observation; they are sufficient to efficiently access the affected area, and sufficient image quality is sufficient to show that the tip is being inserted into the body without causing internal damage. Sometimes, image acquisition cannot keep up with the endoscope insertion operation, resulting in blurred and low image quality. For this reason, the image quality is not appropriate for creating an inference model that performs observation and differentiation, and it is not recorded as an unnecessary image (see, for example, S3 in Figure 3A(a)).
[0057] In the case of an inference model obtained by learning training data obtained by annotating some frames of a sequence of images, a process of finding training data candidates from among many frames is required, which requires a great deal of effort. However, in this embodiment, if an inference model with specifications similar to a pre-existing model is to be created, the training data can be easily created by referring to the image frames that were not subject to annotation from the original sequence of images used to create the training data for the model inference model, removing non-target image candidates from the newly obtained sequence of images, and determining annotation candidate images from the remaining frames. In this way, images of low importance are determined by reflecting the results of timing selection from a sequence of images (a group of images obtained in chronological order), such as a video, when creating an inference model with the model specifications. The learning method, learning system, and program that follow the steps described above can utilize know-how, such as training data selection, when creating a model inference model.
[0058] However, when creating an inference model for safe insertion guidance, images that were removed as being of low importance may actually be important and should be recorded. In that case, the same concept of unnecessary images (images with low priority) can be applied. In other words, when creating an existing inference model with similar specifications, images that were not subject to annotation can be treated as unnecessary images.
[0059] On the other hand, the second timing occurs when a physician or other professional operates an endoscope or other device to observe or identify an object. The endoscope is positioned near the target site within the body, slowly moving, bending, and examining the tip of the endoscope (see image Po in Figure 3A(b)). During this second timing, the physician or other professional may slowly move and bend the tip of the endoscope, while also controlling the device by injecting water to remove dirt, applying suction, and performing other operations, such as changing the image quality, magnifying the image, tracing the affected area, changing the viewing angle, performing special observations, and performing staining and fluorescent observations. The images acquired during this second timing are near the target site, and tend to capture the same area continuously, with minimal blur and high image quality, and are in focus. Images acquired during this timing can be considered, at least, to be images acquired at a different timing than the first timing.
[0060] The images acquired at the first timing and the second timing are divided into groups according to the timing of image capture and recorded in the recording unit 30 as first or second images (see S5 and S7 in FIG. 3A(a) and S35, S37, and S29 in FIG. 5B). Of course, because they cannot be separated at a clear time point, there are cases where frames between groups cannot be correctly grouped. However, this is often not a problem when identifying candidates for training data from a large number of images. Even if the grouping is inappropriate, it is entirely possible to correct the grouping during the training data creation process. For example, if a specific frame in the second image becomes training data, similar frames can be extracted from the first image using a similar image search and reclassified as second images. The first image is the first-timing image recorded in the first-timing image section 31c, and the second image is a simplified representation of the second-timing image recorded in the second-timing image section 31d.
[0061] As described above, the image similarity determination unit 13 determines whether an input image is similar to the first-timing image recorded in the first-timing image section 31c or the second-timing image recorded in the second-timing image section 31d based on the image similarity. When creating an existing inference model, the first-timing image section 31c and the second-timing image section 31d record which data became training data and which data became annotation data. Therefore, using the images recorded here provides evidence of which data was adopted as training data and learned. These images can be used as reference for classification when new images arrive for future inference models. When creating an inference model for an existing endoscope, the first-timing image and the second-timing image are recorded in the first-timing image section 31c and the second-timing image section 31d. Subsequently, the information processing system 10 may input images from various endoscopes (see, for example, input images PI1 to PI3 in FIG. 2). Each time these images are input, the image similarity determination unit 13 determines whether or not they are similar to the first timing image and the second timing image, which record which data became training data or not, and which data became annotation data or not, when creating an inference model with similar specifications. The result of this determination is sent to the classification unit 15.
[0062] The metadata assigning unit 14 assigns metadata to the image data. As described below, category information 31a and specification information 31b are associated as metadata with the image data (first timing image, second timing image) of the inspection image group 31 at the time of creating existing training data. The category information 31a is, for example, endoscope manufacturer information, as well as the sensor, optical characteristics, and image processing circuit characteristics of the endoscope used in the image system 20. When creating an inference model, images for each endoscope manufacturer and model are collected, training data is created, and the inference model is created using this training data. This is because even if an inference model is suitable for a certain model, the reliability of inference may be low for other models. Therefore, the category information 31a is associated with the image. The specification information 31b is information on the specifications of the inference model when the first inference model (existing inference model) was created using the image file 32A. The specification information includes various information related to the inference model, such as what images were used to create the inference model and the structure of the inference model. The metadata adding unit 14 functions as an adding unit that adds metadata to each image in the image group.
[0063] In a system utilizing an existing inference model 19a, it is preferable to be able to prove what kind of training data was used to learn the model. Transparency of the training data is also required, and this information is recorded in the recording unit 30. The image files recorded in the recording unit 30 are also expected to include a time-series image group (e.g., video obtained during a specific endoscopic examination) that served as the source of the training data. The type of equipment used to obtain the image, the object captured, and the date and time of capture can be organized in metadata as category information 31a. A recording unit 30 or metadata indicating what kind of inference model was created from the image may also be included. While Figures 1A and 1B show specification information 31b for the inference model 19a, multiple inference models may be created from the same image. It may also be possible to record which parts (frames) of the image group served as training data.
[0064] For ease of understanding, this specification simplifies the description by assuming that information about the first timing image and the second timing image can be recorded in the first timing image section 31c and the second timing image section 31d. Alternatively, the images may be classified based on whether they have become training data candidates, or the images that have actually become training data may be designated as the second timing images. The images may be identified by the elapsed time or number of frames from the beginning of the video (a group of images obtained in time series), or a flag or the like may be attached to the image frame for management purposes. Images that have become training data, rather than training data candidates, may be classified and recorded. These may be designated as second timing images, or may be further classified as third timing images.
[0065] In this embodiment, the features of frames constituting a time-series image group, such as a video, are broadly divided into first and second timings to indicate that they can be recorded in at least two, i.e., multiple, categories. Although annotated images are not necessarily used as training data, images that have undergone the time-consuming process of annotation are considered to have a certain degree of importance. If they are made distinguishable, they can be used or referenced when creating other inference models. While FIG. 1B shows that only one inference model information is stored in the image file, multiple inference models may be created from the same image. In such cases, the first-timing image portion 31c and the second-timing image portion 31d are separately organized and recorded accordingly.
[0066] If category information Pc is assigned to the input image in the imaging system 20, the category information Pc may be used as the category information 31a. If category information Pc is not associated with the input image, the input determination unit 12 may assign the category information 31a by taking into consideration various information, such as the model of the imaging system 20, image quality (optical characteristics), the lighting conditions of the image, the angle of view of the image, and the shape of the image. Furthermore, the category information 31a may be determined and assigned based on the shape of a treatment tool that appears in the image, or, in the case of a video, the operation status. Furthermore, in FIG. 1B, specification information 31b is assigned only to the image file 32A of the examination image group 31 when creating existing training data. However, when an inference model is created using the examination image group 32, the specification information of the inference model may be recorded.
[0067] The classification unit 15 classifies the input image as a first timing image or a second timing image based on the determination result of the image similarity determination unit 13. Depending on the result of this classification, category information 32a indicating that the input image is a first timing image or a second timing image is attached to the input image. The first timing image may be usable when generating a second inference model for an unknown endoscope (an endoscope whose identity is unknown) that is not well known on the market.
[0068] The classification unit 15 functions as a classification unit (processor) that classifies images from newly acquired images from the first or second endoscope into images that are candidates for training data, using the images acquired at the first timing when the first inference model was created (see, for example, S35, S37, etc. in FIG. 5B). The classification unit classifies images that were not subject to annotation when an existing inference model with similar specifications was created using training data based on the images acquired in time series by the first endoscope as unnecessary images (see, for example, image Plo, etc. in FIG. 2). The classification unit 15 also functions as a classification unit that classifies examination images acquired from a second endoscope different from the first endoscope, using images that were not used as training data when the first inference model was created (see, for example, S35, S37, etc. in FIG. 5B).
[0069] Furthermore, the classification unit 15 classifies a group of images acquired from a second endoscope different from the first endoscope (an existing endoscope, an endoscope made by a specific manufacturer) using the features of the group of images acquired at the first timing when the first inference model was created. That is, the classification unit 15 uses the group of images acquired at the first timing when the first inference model was created, determines the features of this group of images (for example, differences in features between the group of images at the first timing and the group of images at the second timing), divides the images acquired from the second endoscope different from the first endoscope in terms of time, compares them with the same specifications, image quality, and other performance, finds differences in the features between the first and second images, and classifies annotation candidate images from all image frames of the group of inspection images when the existing training data was created.
[0070] The classification unit 15 functions as a classification unit that uses images, including a group of images that were used when the first inference model was created and were not used as training data, to determine their features and classify them as annotation candidate images from among images acquired from a second endoscope different from the first endoscope (see, for example, S35, S37, etc. in FIG. 5B).The second timing is determined according to the input-output relationship for which the inference model is required (the specifications of the inference model).
[0071] The recording control unit 18A controls the recording of images that have been judged by the input judgment unit 12, judged similar by the image similarity judgment unit 13, assigned metadata by the metadata assignment unit 14, and classified by the classification unit 15, into the recording unit 30. In controlling the recording of images into the recording unit 30, the group of images input when the existing teacher data was created is recorded in the test image group 31 used when creating the existing teacher data, and the group of images input during normal testing after the creation of the existing teacher data is recorded in the test image group 32. Since there can be multiple existing inference models using the existing teacher data, multiple image files 31A are recorded in the test image group 31 used when creating the existing teacher data, corresponding to each existing inference model. While FIG. 1A shows only one image file 31A, many image files are actually recorded.
[0072] In this way, new examination images are recorded in the examination image group 32 for the purpose of examination and its evidence, and the image file group 31A of the examination image group 31 at the time of creating the existing training data is also recorded (though they do not necessarily have to be in the same memory as long as they are linkable). Therefore, by comparing these images, it is easy to determine what was captured and what events occurred in the images (videos) of the examination image group 32 at what time. This concept will be shown later as an example of inference in Figure 8. In this way, newly acquired and recorded image groups can be easily classified by timing.
[0073] The first request unit 16 requests the learning unit 18 to create an inference model using the second timing image group acquired at the second timing and recorded in the second timing image section 31d of the inspection image group 31 at the time of creating existing teacher data. This inference model is used as the existing inference model 19a. Furthermore, after generating the existing inference model 19a, when an endoscopic image is input using a first endoscope (an existing endoscope, an endoscope made by a specific manufacturer), if the first request unit 16 finds a rare image that has never been seen before, the first request unit 16 requests the learning unit 18 to generate an improved inference model of the existing inference model 19a using this image (see, for example, S7 in Figure 3A(a) and S27 to S33 in Figure 5B, etc.). The first request unit 16 functions as a first request unit that requests annotation of a group of images for use as training data to generate a third inference model that improves the first inference model, based on the reliability of the results of inferring the same image using the first inference model (see Figure 4(a), S7 in Figure 3A(a), S107 and S109 in Figure 3B, and S27 to S33 in Figure 5B).
[0074] The first request unit 16 for requesting the creation of an existing inference model does not necessarily have to be provided within this system. However, this first request unit has a track record of creating existing inference models by selecting and rejecting teacher data, and is somehow linked to the information processing system 10. The results of the selection of teacher data are recorded as a group of inspection images 31 when creating existing teacher data, so this is shown in the figure to clearly show that the first request unit controlled the type of images used as teacher data at that time.
[0075] The second request unit 17 requests the learning unit 18 to create an inference model using images suitable for generating a second endoscopic inference model from among the images 32b recorded in the examination image group 32. As will be described later, the examination image group 32 contains image data acquired during endoscopic examinations using a first endoscope (an existing endoscope, an endoscope manufactured by a specific manufacturer) and a second endoscope after the existing inference model was created. Among these image data, annotations are added to the image data acquired by the second endoscope and used as training data. The inference model generated by the request from the second request unit 17 is an inference model used by a doctor or the like when performing an examination using the second endoscope, and this inference model may be set in the inference engine 19 in the image system 20 as the second inference model.
[0076] The second request unit 17 functions as a second request unit that classifies the classified images acquired from the second endoscope as images for generating an inference model for the second endoscope and requests annotation for the classified images to obtain an inference model that has the same specifications as the first inference model but corresponds to the images acquired by the second endoscope (see, for example, S43 in FIG. 5B). The first request unit 16 and / or the second request unit 17 functions as a selection unit that selects new training data candidates from a new examination video in accordance with the classified image group recorded in the recording unit (see, for example, S93 in FIG. 7). The second request unit 17 also functions as a second request unit that learns the results of annotating the examination images as training data and requests the generation of a second inference model (see, for example, S27 to S33 in FIG. 5B).
[0077] The learning unit 18 receives requests from the first request unit 16 and the second request unit 17 and generates an inference model. The learning unit 18 has an inference engine, inputs training data, and sets the weighting of neurons in the intermediate layer so that the inference result is the same as the annotated one.
[0078] Here, we will explain deep learning. "Deep learning" is a multi-layered version of the "machine learning" process using neural networks. A typical example is a "forward propagation neural network," which sends information from front to back and makes a judgment. The simplest forward propagation neural network has three layers: an input layer consisting of N1 neurons, a hidden layer consisting of N2 neurons given by parameters, and an output layer consisting of N3 neurons corresponding to the number of classes to be distinguished. The neurons in the input layer and hidden layer, and the hidden layer and output layer, are each connected by connection weights, and by adding bias values between the hidden layer and output layer, logic gates can be easily formed.
[0079] While a neural network can have three layers for simple classification, adding multiple intermediate layers allows it to learn how to combine multiple features during the machine learning process. In recent years, neural networks with nine to 152 layers have become practical in terms of training time, judgment accuracy, and energy consumption. It is also possible to use a "convolutional neural network," which performs a process called "convolution" to compress image features, operates with minimal processing, and is strong in pattern recognition. Furthermore, a "recurrent neural network" (fully connected recurrent neural network), which can handle more complex information and allows information analysis where the meaning changes depending on the order or sequence, can also be used, allowing information to flow bidirectionally.
[0080] To realize these technologies, conventional general-purpose arithmetic processing circuits such as CPUs and FPGAs (Field Programmable Gate Arrays) can be used. However, this is not a limitation. Because much of neural network processing involves matrix multiplication, processors specialized for matrix calculations, such as GPUs (Graphic Processing Units) and Tensor Processing Units (TPUs), can also be used. In recent years, such dedicated artificial intelligence (AI) hardware, called "Neural Network Processing Units (NPUs)," have been designed to be integrated and embedded with CPUs and other circuits, and are sometimes incorporated as part of the processing circuit.
[0081] Other machine learning methods include, for example, support vector machines and support vector regression. The learning here involves calculating the weights, filter coefficients, and offsets of a classifier, and other methods utilize logistic regression processing. When a machine is to make a judgment, a human must teach the machine how to make the judgment. While this embodiment employs a method for deriving image judgments through machine learning, it is also possible to use rule-based methods that apply rules acquired by humans through experience and heuristics.
[0082] The learning unit 18 can also be used to generate an existing inference model. When generating an existing inference model, the learning unit generates a first inference model for image feature determination of first endoscopic images by learning using as training data the results of annotating a group of images acquired at a second timing among images acquired in time series by the first endoscope to determine specific image features in images acquired continuously from the first endoscope. The group of images acquired at the first timing is assumed to be classified as images during access, while the group of images acquired at the second timing is assumed to be classified as images at the time of confirmation of a tumor or the like. For images in which a tumor or the like is confirmed, annotations may be made to indicate where the tumor is located within the image frame and the results of identifying the type of tumor. By training these annotated image frames, an existing inference model can be generated. The specifications of this existing inference model can be written as tumor differentiation. Again, learning may be performed to create inference models with other specifications, but in that case, the segmentation of the image group and the content of the annotations will be different.
[0083] Therefore, the learning unit that generates the existing inference model in this embodiment determines whether a group of images obtained in time series by the first endoscope is a group of images obtained at a first timing (images during access) or a group of images obtained at a second timing (images obtained when searching for a tumor, etc.) in order to determine specific image features for images obtained continuously from the first endoscope, and generates a first inference model for determining image features of the first endoscopic images that has been learned using the results of annotating the group of images obtained at the second timing as training data. In this embodiment, the learning unit 18 is provided in the information processing system 10, but the learning unit for generating the existing inference model does not need to be provided inside the information processing system 10 and may be provided externally as long as this learning unit and the information processing system 10 can cooperate.
[0084] The learning unit 18 uses as training data images selected by the classification unit from the group of images that are candidate training data, and generates a second inference model that has the same specifications as the first inference model but corresponds to images acquired by the second endoscope (see, for example, S89 in Figure 7). The second inference model generated here is assumed to have the same specifications as the existing inference model, and the specifications are set in the inference engine 19. If the existing inference model infers lesions such as cancer, the second inference model will have similar specifications for inferring lesions such as cancer from images acquired by the second endoscope. In such cases where the specifications are similar, the classification of images already used for training the first inference model can be used as a reference.
[0085] The learning unit 18 also functions as a learning unit that creates a first inference model for the endoscope by learning from a group of images included in the inspection video obtained from the endoscope as training data. The learning unit 18 functions as a learning unit that obtains a first inference model for determining image features of the first endoscopic image, which is learned from training data that includes annotation results for the first group of images obtained by the first endoscope in order to determine specific image features for the images obtained from the first endoscope. The first inference model is an inference model for displaying inference results for images obtained by the endoscope.
[0086] An existing inference model 19a is set in the inference engine 19, but in this example, it will be described assuming that it is used to distinguish between types of lesions. The imaging system 20 transmits images (P1 to P3, SP, etc.) acquired by the image acquisition unit 23 to the information processing system 10 (see, for example, S13 in FIG. 5A). When the inference engine 19 receives an image, it infers whether or not a lesion such as cancer is present, its location, etc., and transmits the inference result to the imaging system 20 (see, for example, S25 in FIG. 5B). When the imaging system 20 receives the inference result, it outputs an inference result 25a and reliability 25b to the imaging system 20. The imaging system 20 outputs the received inference result 25a and reliability 25b to the display unit 24. The inference engine 19 functions as an inference engine that sets a first inference model for determining image features of a first endoscopic image, which is learned using as training data the results of annotating a group of existing training data examination images obtained by the first endoscope in order to determine specific image features of images obtained from the first endoscope.
[0087] Next, the recording unit 30 shown in Figure 1B will be described. The recording unit 30 may be provided inside a server or the like that includes the information processing system 10, or may be provided inside a server or the like external to the information processing system 10. The recording unit 30 is an electrically rewritable non-volatile memory that can record a large amount of image data. The recording unit 30 can record a group of examination images 31 used when creating the existing training data, examination images 32 input from various endoscopes after creating the existing training data, and creation specification information 33 for the new inference model.
[0088] The examination image group 31 at the time of creating existing teacher data records image data when images are collected to generate an inference model suitable for an existing endoscope. As described above, a doctor or other medical professional performs an examination using an existing endoscope (first endoscope), and the image data acquired at this time and used to create the existing teacher data is recorded in the recording unit 30 as the examination image group at the time of creating the existing teacher data. To create this existing teacher data, second-timing images acquired at the second timing, i.e., the timing when the doctor or other medical professional finds the affected area, are used. However, as described above, during an examination, images at a so-called first timing can also be acquired from the time the endoscope is inserted into the body, the affected area is found, and the endoscope is removed from the body, i.e., other than the second timing. This first-timing image should be the same as the image recorded in the first-timing image section 31c of the examination image group 31 at the time of creating existing teacher data. Therefore, when a new endoscopic image is input, it is organized as a first-timing image with reference to the first-timing image.
[0089] When creating an inference model for an existing endoscope, this inference model is not necessarily a single model; multiple inference models may be created to suit various uses and models. Therefore, each time the image acquisition unit 23 acquires image data (including image data groups), the information processing system 10 generates an image file 32A. Multiple image files 31A are recorded in the examination image group 31 used when creating existing training data. Each image file 31A has category information 31a, specification information 31b, a first-timing image section 31c, and a second-timing image section 31d. The first-timing image recorded in the first-timing image section 31c corresponds to image data such as images lo, Pna, and second-timing images Pad1 to Pad3 (see FIG. 2) acquired by the imaging system 20.
[0090] The examination image group 32 records image data acquired during an endoscopic examination using a first endoscope (an existing endoscope, an endoscope manufactured by a specific manufacturer) or a second endoscope after images have been collected to create existing training data. As described above, the information processing system 10 may input images from an unknown endoscope in addition to existing endoscopes (see, for example, image PI3 in FIG. 2 ), and may also input images using an existing endoscope. The examination image group 32 corresponds to such images. A new inference model is generated using this examination image group 32. Furthermore, even with existing endoscopes, there may be images that are rarely seen. In such cases, the examination image group 32 is used to improve the existing inference model. For this reason, various information is associated with the image data acquired by the information processing system 10 and recorded in the recording unit 30 as the examination image group 32.
[0091] As will be described later, new inference model creation specification information 33 is information regarding specifications for creating an inference model when creating an inference model using image data from the inspection image group 32. When creating an inference model, specifications regarding the population of training data and specifications defining the configuration of the inference model are required. In addition, when improving an existing inference model, specification information 33a for additional training data is recorded.
[0092] It is desirable to properly manage specifications such as what kind of equipment is used, what kind of inputs are used, and what kind of output is produced using these records. This information manages the specifications of each inference model and the training data used at the time, so it is easy to refer to which inference model was used and what kind of inference was produced, for example, by recording it as evidence information, attaching it to an image, or when recording it. It is desirable to properly verify whether new or improved inference models have the same or superior performance compared to existing inference models. If a model no longer can distinguish between what it was able to distinguish with the existing model, it can no longer be considered an upgrade, so it is desirable to properly manage information such as these handovers and performance limitations in the recording area. If any problems arise, this recording area can be referenced. Of course, information about existing inference models can also be recorded here.
[0093] As described above, the recording unit 30 records a group of images acquired at a first timing and a group of images acquired at a second timing, which were acquired for creating a first inference model (see, for example, images Plo, Pad1 to Pad3, Pna, etc. in FIG. 2). That is, not only the group of images acquired at a second timing for creating a first inference model (existing inference model), but also the images acquired at the first timing are recorded. The recording unit 30 also functions as a recording unit that records at least a portion of the inspection video, along with the results of classifying the image group when teacher data is selected from image frames included in the inspection video by time information of the video frames (e.g., S35, S37, etc. in FIG. 5B). In this embodiment, when acquiring an inspection image for teacher data, the first timing and the second timing are determined, and the image data are recorded based on this determination result. The recording unit 30 shown in FIG. 1B may determine whether the inspection image group 32 is the first or second timing, and record this timing information together. Note that at least a portion of the inspection video described above includes annotation candidate images.
[0094] Next, with reference to FIG. 2, we will explain the collection of images and the generation of an inference model using these images in an information processing system according to one embodiment of the present invention and its peripheral systems. As described above, the information processing system 10 according to this embodiment can be connected to various endoscopic devices (image processing systems 20), and therefore a variety of image data is transmitted. The information processing system 10 can use this variety of image data to generate, in addition to existing inference models, improved inference models and new inference models. The improved inference model is an improvement of an existing inference model that has existed for some time. The new inference model is a completely new type of inference model that can accurately perform inference even when image data with image quality different from that of the previous model is input. This new inference model is generated using a group of inspection images 32 recorded in the recording unit 30.
[0095] In Figure 2, images Plo, Pad11 to Pad3, and Pna are images that are or were the subject of consideration when the information processing system 10 creates the existing inference model 19a. These images have a track record of use when creating the existing inference model (a specific inference model), and their history has been organized. When an inference model has been completed with a proven track record, it is organized to include very important information.
[0096] Although not shown in Figure 2, the information processing system 10 has an existing inference model 19a shown in Figure 1A, an input judgment unit 12 shown in Figure 1B, an image similarity judgment unit 13, a metadata assignment unit 14, a classification unit 15, a first request unit 16, a second request unit 17, a learning unit 18, and a recording control unit 18A.
[0097] The information processing system 10 first generates an existing inference model 19a based on input images. That is, when input images Plo, PI1 to PI3, and Pna are input, the input determination unit 12 determines that the images were acquired using an endoscope manufactured by a specific manufacturer based on information such as image category 1. The information processing system 10 then eliminates image Plo, which is of low importance among the input images, and rejects image Pna for some reason. It then requests that input images PI1 to PI3 be annotated to indicate the location of cancer or other abnormalities, thereby creating training data. The information processing system 10 uses training data based on a group of images acquired in time series by a first endoscope to create an existing inference model with similar specifications, and classifies images that were not subject to annotation as unnecessary images. Once the training data is created, the learning unit 18 uses this training data to generate an existing inference model 19a.
[0098] The group of images (images Plo, PI1-PI3, Pna, etc.) organized and used in this way is valuable know-how information, including the results of trial and error obtained up to now, such as which images were annotated and used as training data, and which images were annotated but not used as training data. In other words, whether an image can be a candidate for new training data can be determined using images that have not been used up to now. Even if collecting images similar to unused images would be completely useless, this system makes this determination possible. However, images other than those that are not supposed to be used may contain new information, so it is better to carefully examine images other than those that are not supposed to be used (i.e., images at the first timing, so to speak) if it is worth the effort. In this embodiment, this concept is utilized to efficiently collect images.
[0099] The generation of the existing inference model 19a will be described in detail below. Image Plo is an image of low importance, images Pad1 to Pad3 are images that have been used as existing training data, and image Pna is an image that was created as existing training data but has not been used. These images Plo, Pad1 to Pad3, and Pna are all images acquired using endoscopes from the specific manufacturer mentioned above. The reason for mentioning a specific manufacturer here is that endoscopes from different manufacturers generally have different sensors, optical systems, image processing, etc. Therefore, if image data acquired from various manufacturers are mixed, the reliability of the inference model decreases. Note that even if the manufacturers are different, as long as the sensors, optical systems, image processing, etc. are similar, it may be possible to perform inference using image data within a similar range using an existing inference model. Therefore, when creating a specific inference model, the information processing system 10 carefully selects and uses image data from endoscopes from a specific manufacturer to generate an inference model 19a that takes into account the characteristics of that model.
[0100] Existing inference models are the product of initial attempts, and are often created using training data selected through trial and error, manual labor, or visual inspection. Therefore, the logic and know-how of this selection process, as well as the resulting classified images, constitute intellectual property packed with information so rich it is difficult to describe in words. As mentioned above, unused image Plo is a low-importance image and has image data 51a and category information 51b. Because image Plo was acquired by the specific manufacturer mentioned above, its category information 61b is categorized as Category 1. For example, in the case of creating an inference model for lesion differentiation, image Plo may be an image acquired while a doctor is searching for a lesion using an endoscope (see, for example, image Ps in Figure 3A(b)). Images acquired during the search are often blurred due to the movement of the imaging unit, resulting in low image quality and a low likelihood of clearly capturing the lesion.
[0101] There are various specifications for inference models, but here we will continue the explanation assuming a specification that infers what a lesion is. In such cases, since few of the search images include a detailed view of the lesion, we will explain here using an example that is simplified and has a lower priority. Of course, if the inference model is intended to prevent oversights using search images, the search images should be given priority as training data, but here we will ignore such cases. Of course, as long as the image group is organized in accordance with the inference model specifications, when creating a new inference model that prioritizes search images, it is sufficient to give priority to the image group from the first timing as training data.
[0102] If an image is determined to be a search image based on the characteristics of the image, the features of the captured image, the features of image changes, etc., it may be classified into a group of images at a first timing that are not required when creating an inference model for discrimination. It can also be said that an image does not fit the results of an inference of which frames of a video became training data. As mentioned above, the idea behind this embodiment is to utilize the reason why such an image was rejected when creating an inference model when creating a subsequent inference model.
[0103] Images Pad1 to Pad3 are images acquired when a doctor or other medical professional uses an endoscope to closely examine a patient's body for possible lesions. In the case of an inference model for differential diagnosis, it is possible that training data will be created from the images acquired during the examination. Therefore, it is better not to initially classify these images as images from the first timing. These images are selected as images from the second timing, based on the characteristics of continuous observation of the same area, such as high-quality images with a stationary or minimal motion image capture unit, and images that are likely to depict a lesion. To improve visibility, the images may be sprayed with cleaning water, surface dirt may be removed by suction, stained, or the wavelength of light during observation or image representation may be changed using a light source or image processing. When initially classifying endoscopic images to create training data, images from the second timing are often selected according to an inference model creation manual. However, when creating training data for subsequent iterations, it is desirable to automate this manual image classification. Therefore, in this embodiment, the process of converting the initially acquired images into training data is referenced and used when creating subsequent inference models.
[0104] Image Pna is an image that was not adopted as existing training data. The category information 52e for this image Pna belongs to category 1 because it was acquired using an endoscope made by a specific manufacturer, the image quality is usable, and annotations regarding the location of lesions such as cancer are added. However, for some reason, this image was not adopted when generating the existing inference model 19a. One possible reason is that the reliability of generating an inference model using this image Pna was reduced, so it was excluded. Such images are unlikely to be treated as high priority even if similar images are obtained in the future. If necessary, such images may be classified as images of the first timing.
[0105] In this way, images Plo, Pad1 to Pad3, and Pna are images acquired using an endoscope made by a specific manufacturer, and of these images, images Pad1 to Pad3 are used by the information processing system 10 to perform deep learning and generate an existing inference model 19a. Note that in Figure 2, only one image Plo with low importance is listed, only three images Pad1 to Pad3 of existing training data that have been adopted are listed, and only one image Pna of existing training data that have not been adopted is listed. However, it goes without saying that in reality, many more images will be used.
[0106] As described above, in this embodiment, the images collected and organized for a proven inference model can be fully utilized when improving this inference model or when creating a new inference model for a different model. Whether a model has a proven track record or not is obvious because, in order to solve the problem of AI transparency in the future, the basis for the inference must be clearly stated in the inference results. Furthermore, because of transparency, the training data used when learning this proven inference model also needs to be made public as necessary, which is also obvious.
[0107] As described above, once the existing inference model 19a is generated using Pad1 to Pad3, the image system 20 can input the endoscopic images P1 to P3 and SP using this existing inference model 19a and determine the presence or absence of cancer, etc. If such an inference model does not perform as originally expected, it may be retrained.
[0108] After images Pad1 to Pad3 are selected through a painstaking process of trial and error and an existing inference model is generated, information processing system 10 then inputs image data 51b, 51c, and 51d from various endoscope devices. Category 1 is associated with each image data as category information 52b, 52c, and 52d, and annotation information 53b, 53c, and 53d is added to each image data. The location of a lesion, such as cancer, is added as annotation information 53b, 53c, and 53d.
[0109] The input image PI1 in Figure 2 has category information 62a of Category 1, and is an existing endoscope, so it is possible to use the existing inference model 19a to infer whether or not a lesion such as cancer is present. For example, each image frame from the beginning to the end of an endoscopic examination is input into the inference model in real time, and when a frame that is identified as having a lesion is detected, a lesion frame is displayed, a warning is issued, and correct identification is possible. After this inference is performed, the inference result IO1 and the reliability R1 at the time of inference are output. Furthermore, a similar efficacy may be displayed for a specific still image that a doctor inputs and requests identification.
[0110] The input image PI2 is an image acquired by an endoscope manufactured by a different manufacturer than the existing endoscopes, or by an endoscope manufactured by the same manufacturer but of a different model. This input image PI2 may be capable of being inferred using the existing inference model 19a for the input image PI1. In this case, for example, each image frame is input into the inference model in real time from the beginning to the end of the endoscopic examination. When a frame that indicates a lesion is detected, a lesion outline is displayed, a warning is issued, and correct identification is possible. A similar effect may also be displayed for a specific still image input by a physician for identification. The inference described above is performed on the input image, and the inference result IO2 and the current reliability R2 are output. If the reliability is low, this may indicate a lesion that the current inference model cannot confidently identify. Images belonging to Category 2 are collected, and a request is made to generate an improved inference model or a new inference model suitable for Category 2 (see, for example, S83 and S87 in FIG. 6). The images at this time are recorded as the examination image group 32.
[0111] Furthermore, input image PI3 is an image belonging to category 3, acquired by a completely unknown endoscope not known on the market, and has characteristics different from those of image data acquired up to that point. For this reason, for example, when each image frame of an endoscopic examination is input into the inference model in real time from start to finish, or when a doctor inputs a specific still image and requests differentiation for input image PI3, the existing inference model may not be able to output a reliable inference. Therefore, the second request unit 17 requests the generation of a new inference model. To make this request, new inference model creation specification information IF is created using images recorded in the examination image group 32, and this new inference model creation specification information IF includes the additional teacher data specification TD.
[0112] In this way, the image classification know-how gained when generating a proven inference model (the logic used to classify images Plo, Pad1-Pad3, and Pna, as well as the feature information of these images, etc.) can be used as a reference when creating an improved or new inference model. If such classification can be performed efficiently for the input images PI2 and PI3, it will be possible to immediately narrow down the images to be annotated by experts, and if annotation can be performed quickly, it will be possible to improve the inference model using training data containing the annotation results and quickly begin the learning process for a new inference model.
[0113] The steps described above are summarized below. The system was originally developed to incorporate an inference model. It involves a segmentation step, which segments the examination video obtained from a first endoscope according to timing, referred to as the first and second timings. It also involves an image classification step, which removes frames corresponding to the first timing from the training data candidates and adopts frames corresponding to the second timing as training data candidates. It also involves a recording step, which classifies and records the examination video for each segmentation timing. These steps were described as processing images Plo, Pad1-Pad3, and Pna. By utilizing the classification logic and feature information of these images, it becomes possible to efficiently create training data for a second inference model (which may be an improved inference model or a new model) that is different from but has similar specifications to the first inference model, by incorporating useful knowledge. Specifically, the system includes a step of acquiring new examination video to be used for inference by the second inference model, and a selection step of selecting training data candidates for the second inference model by classifying the new examination images using a classification that reflects the image classification results described above. This enables rapid development (learning) of inference models.
[0114] In addition, there is a step of training a first inference model that infers diseased area information contained in frames of the examination video using at least one of the training data candidates. An annotation step may be provided in which the training data at this time is the result of annotating frames corresponding to the second timing as the training data candidate. In other words, this can also be described as an apparatus and method that allows a learning unit that created a first inference model for endoscopy to continue further additional learning, refinement learning, etc. by learning a group of images contained in an examination video obtained from an endoscope as training data. In other words, if a recording unit is provided that records the examination video (or at least a portion thereof) along with the results of classification based on the time information of the video frames when training data is selected from image frames contained in the examination video, this recorded content becomes a valid past asset. In other words, when the learning unit additionally trains a second inference model that is different in version and assumed equipment from the first inference model but has similar specifications, a video acquisition unit that acquires new examination videos and a selection unit that efficiently selects new training data candidates from the new examination videos according to the group of classified images recorded in the recording unit may be provided.
[0115] In creating this new inference model, it would take a huge amount of time to start collecting new images belonging to category 3, so in this embodiment, we try to make as effective use as possible of images that have already been collected by a doctor or the like during an examination and recorded as examination images 32 in the recording unit 30. In other words, the examination images acquired by a doctor or the like during an examination are recorded, and these images are used (see, for example, S28, S29, S37, etc. in Figure 5B).
[0116] The new inference model creation specification information 33 and the additional teacher data specification 33a are specification information for generating a new inference model that infers images belonging to category 3. The second request unit 17 requests the learning unit 18 to generate an inference model along with this information. The characteristics of the inference model required for the endoscope and the type of affected area image for which inference is required can be determined based on the image data, changes in the image data, image features of objects captured in the captured image, or the inference results and reliability of existing inference models.
[0117] Next, the operation of generating a new inference model will be explained using the flowchart shown in Figure 3A. However, this new inference model is assumed to be an existing inference model with similar specifications, but with improved performance and usability for other devices. Therefore, the flow shown in Figure 3A shows the operation of generating a new inference model based on the specifications of an inference model with specific specifications that exists as the existing inference model 19a in Figure 1A. There may be multiple inference engines 19 in Figure 1A, or multiple existing inference models 19a. In this case, there may be a step of selecting which of the existing inference models to create a new inference model with specific specifications. Here, we will explain a situation in which inference is being performed with specific expectations for an inference model with existing specifications, so we will continue the explanation assuming that the specification selection has already been made.
[0118] The flow in Figure 3A is realized by the control unit 11 in the information processing system 10 controlling each unit in the information processing system 10 according to a program stored in memory. This flow can efficiently acquire training data for generating a new inference model from a large amount of images.
[0119] As described above, it is assumed that important training data candidates may be overlooked using simple image and situation assessment in the following two situations. In the first situation, endoscopic image data groups (frame groups) that are not typical patterns, such as rare cases that are completely different from those seen in the past or unusual procedures, may not be correctly identified and acquired as training data because the obtained image frames, observation methods, and operation methods are different. In the second situation, training data may not be correctly identified and acquired for an endoscopic device that has a different operation method, image quality, etc. from the intended endoscopic device (one for which an inference model has already been created). In other words, in these cases, the conditions for images to be used as training data must be relaxed as much as possible to avoid overlooking anything. Therefore, in this embodiment, a method is adopted in which images with low priority are eliminated in order to efficiently respond to such cases with as few precedents as possible while taking into account unexpected situations. Although we wrote that the second inference model has a different version and assumed equipment than the first inference model, but similar specifications, the first and second situations here are respectively a version upgrade and an expansion of assumed equipment, but assume similar specifications. Similar specifications are assumed to be those that can turn image frames obtainable under similar conditions into training data.
[0120] The first situation described above is considered when branching to Yes in step S1 in Figure 3A. Although the images are from a known device, if they are unprecedented, they are prioritized as training data candidates, as they may offer potential for improvement. The second situation described above is considered when branching to No in step S1 in Figure 3A. Because these images are not from a known endoscope, all images may be, so to speak, "unseen." However, even with these images, when used as an endoscope to examine a specific body part, images of the entire process—from insertion into a body cavity, access, identification of the affected area, and removal—should be captured. Therefore, the logic and inference used to create an inference model for a known endoscope can be effectively utilized to determine which images are unnecessary and which images provide training data candidates. In other words, images of low importance are first removed, and training data candidates are identified from the remaining images.
[0121] When the flow for creating a new inference model starts, first, it is determined whether or not it is compatible with an existing inference model (S1). Here, the input determination unit 12 inputs image data, etc. from the image acquisition unit 23 of the image system 20, and determines whether or not inference is possible using the existing inference model 19a. This determination is made based on the category information associated with the image data. Note that among the input images in Figure 2, images that belong to categories 2 and 3 may not have associated category information. In this case, it is sufficient to make the determination using other information, such as other attached data, information related to image quality, information related to the screen shape, etc. Alternatively, the determination may be made based on the reliability of the inference rather than on the category, etc.
[0122] If the result of the determination in step S1 is that the image is compatible with the existing inference model, only unseen images are collected as training data (S7). Since the result of the determination in step S1 is that the image is compatible with the existing inference model, it is usually highly likely that the image is previously seen (in other words, an image that is similar to previously acquired images or has the same image characteristics as previous images). This determination can be made by the similarity determination unit 13 comparing the features of the original endoscopic images (videos) used to create the proven inference model, if they are recorded (recorded as the group of examination images 31 used when creating the existing training data). A similarity determination is performed, and images above a certain threshold are determined to be similar images (images previously seen).
[0123] Furthermore, if the system records which parts of an image have previously been used as training data when generating an existing inference model, it can also determine whether the image is rare enough to be used as training data. Images that are determined to be dissimilar can be classified as "unseen images." For example, for endoscopic video information obtained during a specific examination, frames corresponding to training images in the endoscopic video are used as annotation information to generate training data. Similar training data is then applied to multiple endoscopic video images from multiple examinations to generate an existing inference model. Endoscopic video images, including insertion and removal images, are input into this existing inference model, and inference is performed using the frames used as training data as the output. Images that have never been seen before may be overlooked because they have never been used as training data. Therefore, images with intermediate reliability can be selected for this inference, and their similarity to previously used endoscopic images (or images used as training data) is assessed. Images with low similarity can be classified as rare data (rare images) and used as training data candidates. In other words, if the image determined in step S1 is a rare image that has never been seen before, it may be used to improve the reliability of the inference of the existing inference model. Therefore, in this step, we do not collect images that have been seen before, but only collect unique (rare) images that have never been seen before, so that these images can be used as training data.
[0124] Furthermore, when collecting rare, never-before-seen images, images that are not normal when judged as normal, but are not highly reliable when judged as lesions, may be collected as never-before-seen images. The collected images are recorded in the recording unit 30 as an examination image group 32. When recording, metadata indicating that the images are rare, never-before-seen images is added. A specific example of a method for determining whether an image is rare, never-before-seen, will be described later with reference to FIG. 4(a).
[0125] Next, we will explain what happens when it is determined in step S1 that an existing inference model cannot handle the situation. For example, for endoscopic video information obtained during a specific examination using a specific model of endoscope, frames from this endoscopic video corresponding to training images are used as annotation information to obtain training data. Similar training data is then generated for endoscopic video images from multiple examination results, generating an existing inference model. Endoscopic video including insertion and removal images may be input into this existing inference model, and inference may be performed in which the frames used as training data are output. Inference in this step is performed on images from endoscopes with different specifications and characteristics. As mentioned above, there are few frames that can be used as training data compared to the number of frames obtained during an endoscopic examination, so inference must be performed with a very high level of accuracy. Furthermore, the reliability is expected to be even lower for images with image quality that has never been used for learning, such as in this case. Rather, a preferred method would be to generate training data from a series of endoscopic images obtained during a specific examination using a specific model of endoscope by annotating frames that were not even candidates for training images, generate an inference model, input endoscopic images from multiple examination results into this inference model, and based on the inference results, select images that cannot be training data and carefully examine the rest as candidates for training data.
[0126] However, the flow shown in Figure 3A provides a simpler example, determining whether an image is of low importance (S3) and selecting the rest as training data candidates. Images of low importance are determined by the timing of selection and rejection from a sequence of images, such as videos, when creating an inference model with the model specifications. For example, if an inference model for observation or diagnosis is used as the model, access images are necessary regardless of the endoscope model, and are therefore acquired (or at least digitized). Using these images, a method can be considered for separating the images from other parts (non-access parts). Comparing images from completely different models may reduce the reliability of the comparison due to the disparity of the comparison targets. However, if the same model is used, it is easy to identify differences between access images and other images (confirmation images of no lesion or presence of lesion). These differences can be determined by applying the logic used to select specific training data candidate parts (frames) from known, proven endoscopic images, or by inference.
[0127] In other words, it is already known which frames were used as training data for a group of images acquired in time series to determine specific image features, such as lesion detection, from images continuously acquired from the first endoscope when a known inference model was created. In other words, it can be determined whether the image group was acquired at a first timing or a second timing. Here, it is known that the training data was acquired from the second timing. In other words, there is a learning unit that obtains a first inference model for determining image features of first endoscopic images, learned using the results of annotating the image group acquired at the second timing as training data. Therefore, it is only necessary to provide a classification unit for the learning unit that classifies an image group acquired from a second endoscope different from the first endoscope using the features of the image group acquired at the first timing when the first inference model was created.
[0128] In other words, if an information processing device that efficiently selects the above-mentioned candidate image groups for training data from all image (frame) data during an inspection can be linked to the above-mentioned learning unit, new training data can be provided to an already proven learning unit, enabling high-speed learning. In other words, the above-mentioned classification unit uses the image group obtained at a first timing when the first inference model was created, determines its features (for example, differences in the features of the first and second image groups), temporally divides images obtained from a second endoscope different from the first endoscope, compares them with the same specifications, image quality, and other performance, finds differences in the features of the first and second images, and classifies annotation candidate images from all image frames of a specific inspection image.
[0129] Here, we have explained the assumption that annotation candidate frames are selected automatically, and that the actual annotation is performed by a doctor or other expert to complete the training data, but the annotation itself can also be performed by the first inference model mentioned above. If a process is established in which the results are checked by a doctor, it becomes possible to create training data and an inference model quickly while ensuring quality.
[0130] Images taken by doctors and other medical professionals during screening (access) for endoscopic examination (see, for example, image Plo in Figure 2 and image Ps in Figure 3A(b)) are taken while the endoscope is inserted into the body and the imaging unit is being moved to search for lesions, etc. The images tend to blur and have low image quality, making them unsuitable as training data for use in improving existing inference models. In step S3, it is determined whether or not such images are suitable as training data. Note that this determination can be made not only on a single image basis, but also on a continuous image basis. In other words, even if the quality of a single image is low, if the continuous image is evaluated, it may be usable as training data, for example, for an inference model for displaying a guide during access.
[0131] As described above, we have determined that images are unlikely to be candidates for training data, i.e., images with low importance. Images other than these images are then collected as training data candidates (S5). Here, images determined to be of low importance in step S3 are eliminated, and image data with a high importance are collected as training data. These images correspond to the image Po during observation in Figure 3A(b) and should include images of lesions carefully observed (closely examined) by a doctor or other medical professional. Images that cannot be handled by existing inference models and that pose difficulties for conventional inference models to infer are valuable images for creating a new inference model. Therefore, the images collected here are recorded as part of the examination image group 32 (see Figure 1B). When recording these images, metadata is added indicating that they are unknown images for use in the new inference model.
[0132] If we can narrow down the image frames that can be used as training images from the tens of thousands of endoscopic frames for each examination, we can focus on which images should be annotated and used as training data, without having to consider the other frames. This will increase the efficiency of creating improved or new inference models.
[0133] Once images are collected in step S3 or step S5, a request for annotation 1 or annotation 2 is made (S9). Here, the first request unit 16 uses the images collected in step S7 to request annotation 1 for generating an improved inference model of the current state. This annotation 1 is primarily based on existing training data and is used to create additional training data for learning, so existing annotation tools can be used as is.
[0134] In step S9, the second request unit 17 requests annotation 2 for generating a new inference model using the images collected in step S5. This annotation 2 is intentionally separated from annotation 1 because it is unclear whether existing training data can be used (depending on the quality of the inference model), it is necessary to obtain output for this new, unknown endoscope device, and it is not clear whether an annotation tool with the same specifications as conventional ones will necessarily be sufficient. Once annotations 1 and 2 are completed, the learning unit 18 is requested to generate an improved inference model of the current situation or a new inference model. Once the annotation is requested, this flow ends.
[0135] In the flow of creating a new inference model, when inferring lesions such as cancer from an input image, it is determined whether or not inference is possible using an existing inference model, and if inference is possible, rare images that have never been seen before are collected as candidate training data (see S1 Yes → S7). The collected candidate training data is used as data to improve the existing inference model (see, for example, S81, S83, and S87 in Figure 6). Furthermore, for input images that cannot be inferred using an existing inference model, images with low importance are excluded and these images are used as candidate training data. The candidate training data collected here is used as data to generate a new inference model (see, for example, S81 No → S85 Yes → S91 in Figure 6).
[0136] The information processing system 10 receives a variety of images from various image processing systems 20 (including endoscopes), and some of these images cannot be inferred using existing inference models. In this embodiment, even images that cannot be inferred are recorded as an inspection image group 32 (see S5), so that these images can be efficiently used when generating a new inference model. Furthermore, even if an image can be inferred using an existing inference model, if it is different (dissimilar) from images previously acquired, this image is recorded as an inspection image group 32 so that it can be used when improving the existing inference model.
[0137] Next, a modified example of the operation of creating a new inference model will be described using the flowchart shown in Figure 3B. In the flow shown in Figure 3A, in step S1, it is determined whether an existing inference model can handle the image. If not, images for creating a new inference model are collected (see S3 and S5). On the other hand, if they can handle the image, rare images that have never been seen before are collected for improving the inference model (see S7). In the flow shown in Figure 3B, images for creating a new inference model are collected regardless of whether they can be handled by an existing inference model (see S101 and S103). After this collection, rare images that have never been seen before are collected for improving the inference model only if they can be handled by an existing inference model (S105 Yes) (see S109).
[0138] Compared to the flow of FIG. 3A, the flow of FIG. 3B performs the same processing as steps S3, S5, S1, S7, and S9 of FIG. 3A, respectively, as steps S101, S103, S105, S107, and S109 of FIG. 3B, and only the order of processing is different. Therefore, detailed explanations will be omitted and a brief explanation will be given.
[0139] 3B starts, first, an image is determined to determine whether it is an image of low importance (S101). Here, the classification unit 15 determines whether the image is of low importance based on the determination result of the similar image determination unit 13 performed on the input image. As described above, an image of low importance is determined by reflecting the results of selection, according to timing, from a sequence of images such as videos when an inference model with model specifications is created. For example, if an image is similar to an image that was not used when creating the inference model, it is determined to be of low importance.
[0140] Next, images other than those determined to have low importance are collected as training data (candidates) (S103). Here, images other than those determined to have low importance based on the determination result in step S101 are recorded as training data (candidates) in the recording unit 30. In other words, images selected from the group of images that will be training data candidates are recorded in the recording unit.
[0141] Once images other than those with low importance have been collected as training data (candidates), a determination is made as to whether an existing inference model can be used for the input image (S106), as in step S1. If the result of this determination is that the input image cannot be used, only unseen images are collected as training data (candidates), as in step S7. After the processing in steps S105 and S107 is completed, a request for annotation 1 or annotation 2 is made (S109), as in step S9. Once the request is complete, this flow ends.
[0142] As described above, in the flow in Figure 3B, images for creating a new inference model can be collected regardless of whether they are compatible with existing inference models, allowing for the collection of a wide variety of images.
[0143] Next, using Figure 4, we will explain an example of the processing that occurs after determining whether or not an image can be handled with an existing propulsion model in step S1 (see Figure 3A (S105 in Figure 3B)). Figure 4(a) is a diagram that explains a method for determining whether or not an image is unfamiliar in step S7 (S107 in Figure 3) when it is determined in step S1 (S105 in Figure 3B) that an image can be handled with an existing inference model.
[0144] In the example shown in FIG. 4(a), an inference engine is prepared in which two types of inference models, a normality detection AI (inference model) 19a and a lesion detection AI (inference model) 19b, are set, and these AIs (inference models) are used to perform inference on the same endoscopic image PI4. Here, the normality detection AI 19A determines whether the endoscopic image PI4 is normal, in other words, whether the endoscopic image PI4 is abnormal. Furthermore, the lesion detection AI 19B determines whether the endoscopic image PI4 contains a lesion. This lesion detection AI generally corresponds to an inference model that infers whether a lesion such as cancer is present. If this lesion detection AI is considered to be the first inference model, the normality detection AI corresponds to the third inference model.
[0145] The inference model set in the normality detection AI 19b is generated by learning using images without lesions as training data. As described above, this inference model can infer whether an input image is normal or not. Images for generating this inference model can be collected from images determined to be other than normal in step S27 of FIG. 5B, which will be described later.
[0146] If the endoscopic image PI4 is normal, i.e., if there is no abnormality, the normality detection AI 19A will determine that it is normal, and the lesion detection AI 19B will determine that there is no lesion. However, the example shown in Figure 4(a) is a case where the existing inference model is not suitable for inferring the endoscopic image PI4, and the inference result by the normality detection AI 19A is "abnormal," while the inference result by the lesion detection AI 19B is "no lesion." In other words, the judgment results by the two AIs are contradictory.
[0147] In this way, even if the endoscopic image PI4 can be inferred using an existing inference model, if the judgment results from the two AIs are contradictory, the endoscopic image PI4 can be said to be a rare image that has never been seen before. In this case, the endoscopic image PI4 can be recorded as an examination image 32 in the recording unit 30 as an unseen image, and annotation can be performed to create training data, which can then be used to improve the existing inference model.
[0148] When the similar image determination unit 13 determines that the image is a rare image containing no undesired features, the first request unit 16 requests annotation of the image, creates training data, and uses the training data to request the generation of an inference model. That is, the first request unit 16 functions as a first request unit that requests annotation of training data for generating a third inference model that improves the first inference model, based on the reliability of the inference result of the same image (e.g., image PI4) using a first inference model (e.g., normality detection AI 19a) (see, for example, S27 to S33 in FIG. 4(a) and FIG. 5B). The inference result of the first inference model may be determined to be reliable if, for example, it is inferred that the image was acquired using a first endoscope (e.g., an endoscope manufactured by a specific manufacturer) or if it is inferred that the inference model correctly identifies other lesions.
[0149] The example shown in FIG. 4(b) shows a case where it is determined in step S1 (S105 in FIG. 3B) that the existing inference model cannot handle the situation. In this case, the image change speed of the endoscopic image PI5 is detected to determine whether screening (access) is in progress, i.e., whether a doctor or other medical professional is moving the imaging unit in search of a lesion. In addition, it is determined whether the distance between the endoscopic image PI5 and the surface of the body is closer than a predetermined distance (object proximity determination). If the result of this determination is that the object distance is closer than the predetermined distance, it can be determined that a doctor or other medical professional is currently observing the body.
[0150] In this way, when it is determined that an existing inference model cannot be used, it is determined whether screening is in progress based on the image change speed, and whether observation is in progress based on the object distance. Images that are not important, such as images that are being screened, are filtered (i.e., unimportant images are excluded), and important images (e.g., images that are being observed) are recorded in the recording unit 30 as an inspection image group 32. These images are annotated and used as training data when creating a new inference model.
[0151] Next, specific operations of the information processing system and the endoscope as an image system according to this embodiment will be described with reference to the flowcharts shown in FIGS. 5A to 5C.
[0152] The flowchart shown in FIG. 5A explains the main operation of the endoscope 1. The information processing system 10 is connected to various endoscopes and acquires various information from the image acquisition unit. The endoscope 1 is an endoscope manufactured by a specific manufacturer, and the information processing system 10 has previously generated an existing inference model 19a based on images from this endoscope 1. The main operation of the endoscope 1 shown in FIG. 5A is realized by a control unit 21 in an image system 20 corresponding to the endoscope 1 controlling each unit in the image system 20.
[0153] 5A starts, an image is first acquired and displayed (S11). Here, the image acquisition unit 23 acquires an endoscopic image and displays this endoscopic image on the display within the display unit 24. While viewing this endoscopic image, a doctor or other medical professional operates the endoscope 1, moves the imaging unit of the endoscope 1 to a target area, and observes the target area.
[0154] Once the image is acquired and displayed, the image is then transmitted to the information processing system 10, and an inference result is acquired (S13). Here, the control unit 21 transmits the image acquired by the image acquisition unit 23 to the information processing system 10 via the communication unit. At this time, model information (information indicating the endoscope 1) is transmitted to the information processing system 10 together with the image information (see S23 in FIG. 5B). Having received the image, the information processing system 10 uses the existing inference model 19a of the inference engine 19 to infer whether or not a lesion such as a cancerous area is present, and returns this inference result to the endoscope 1 (see S25 in FIG. 5B). When the imaging system 20 receives this inference result, it also receives information about the reliability of the inference result.
[0155] Next, it is determined whether a highly reliable judgment result has been obtained (S15). Here, based on the result of the inference reliability received in step S13, the control unit 21 determines whether this reliability is higher than a predetermined value. If the result of this determination shows that the reliability is lower than the predetermined value, the process returns to step S11.
[0156] On the other hand, if the result of the determination in step S15 is high reliability, the determined position is then displayed in a frame (S17). In step S13, the inference engine 19 infers the presence or absence of a lesion such as cancer and, if present, its location, and the endoscope 1 receives this inference result. In this step, the display unit 24 displays the position of the lesion such as cancer inferred by the inference engine 19 in a frame. Of course, a display method other than a frame may be used. Also, in this step, an operation guide may be displayed to the operator of the endoscope. For example, a message such as "Possible bleeding" may be displayed. For this guide, an inference model for guide may be prepared, and the guide may be output using this inference model. After the determined position is displayed in a frame, the process returns to step S11.
[0157] Next, the main operation of the information processing system 10 will be described using the flowchart shown in FIG. 5B. As described above, the information processing system 10 connects to various endoscopes and acquires various information from the image acquisition unit. This information processing system determines whether the endoscopic images transmitted from the various endoscopes are images that can be used for an improved current inference model, images that can be used to create a new inference model, or other images. Then, depending on the determination result, the system requests annotation, creates training data, and requests the generation of an inference model. The main operation of the information processing system shown in FIG. 5B is realized by the control unit 11 of the information processing system 10 controlling each unit within the information processing system 10.
[0158] When the flow of the information processing system in Fig. 5B starts, first, the system enters an image acquisition standby state (S21). As described above, the information processing system 10 can be connected to various endoscopes, and endoscopic images are transmitted from each endoscope. In this step S21, the control unit 11 is in a standby state so that it can receive images from various endoscopes.
[0159] Next, when an image is acquired, it is determined whether it is a target image (S23). Here, the control unit 11 determines whether the image received by the information processing system 10 is an image to be used to determine the presence or absence of a lesion such as cancer. Here, the control unit 11 determines whether it is a target image based on the scene of the image (for example, whether it is an image of inside the body and whether it is an image taken while a doctor or the like is examining it). For example, when a doctor or the like is examining a target area (a region where a lesion such as an affected area or a tumor is likely to exist), the tip may stay in the same position or may search around the area. Whether or not an examination is being performed may be determined by analyzing the endoscopic image, or may be determined based on the operation state of the doctor or the like. If the result of this determination is that it is not a target image, the process returns to step S21.
[0160] On the other hand, if the result of the determination in step S23 is that the image is the target image, the inference result and reliability are transmitted (S25). Here, the inference engine 19 uses the existing inference model 19a to infer whether or not the target image contains a lesion such as cancer, and if so, its location. During the inference, the reliability of the inference theory is also calculated. When the inference by the inference engine 19 is completed, the inference result and its reliability are transmitted to the source of the image (see S13 in FIG. 5A).
[0161] Next, a determination is made as to whether the received target image is an improvement of the current state, a new model, or something else (S27). In this flow, the control unit 11 determines whether the image is an image that could not have been collected before, and uses the image that could not have been collected before as training data. Here, when generating an inference model, the control unit 11 first determines whether the received target image is suitable for generating an inference model that improves the current state, a new inference model, or something else. As described above, the information processing system 10 can connect to a variety of inspection devices such as endoscopes via the Internet, etc., and these endoscopes include endoscopes made by a specific manufacturer that were used when generating the existing inference model, endoscopes that are relatively similar to endoscopes made by this specific manufacturer, and completely unknown endoscopes that are not known on the market. In this step, even if the endoscope image is made by a manufacturer other than the specific manufacturer, a determination is made on the endoscopic image received in step S23 in order to generate an inference model that infers lesions such as cancer. The determination in step S27 will be described in detail below with reference to FIG. 6.
[0162] If the result of the determination in step S27 is that the image is neither for current state improvement nor for a new inference model, other processing is performed (S28). Here, images determined to be neither for current state improvement nor for a new inference model are recorded. As will be described later (see S81 and S89 in FIG. 6), since the image is from an existing model and the reliability is outside the 40-60% range, it is an image that clearly indicates whether or not it is a lesion such as cancer. Therefore, by recording this image and using it as training data, an AI can be created to infer that there is no abnormality. After the processing in step S28 is performed, the process returns to step S21.
[0163] On the other hand, if the result of the determination in step S27 is that the image is suitable for improving the current state, the image is then recorded as a candidate for training data (S29). Here, the control unit 11 records the image as an image suitable for generating an inference model for improving an existing state in the recording unit 30 as an examination image group 32. For example, rare images that doctors have never seen before can be useful in improving an existing inference model. When recording an image in the examination image group 32, it is advisable to record category information as metadata indicating that the image is for improving an existing state.
[0164] Next, an annotation is requested (S31). Here, the control unit 11 requests annotation in order to generate an inference model for improving the current state. For annotation, a doctor or other expert associates data indicating the location of a lesion such as cancer with image data. In addition to having a doctor or other expert manually add annotations, they may also be added automatically using AI or by an expert referring to the advice of AI.
[0165] Next, new teacher data is selected and learned to create an improved inference model (S33). Once annotations are added, new teacher data can be created, so the control unit 11 selects and discards new teacher data as appropriate. That is, from the new teacher data, teacher data suitable for generating an improved inference model is selected. After selecting the new teacher data, the second request unit 17 then requests the learning unit to generate an improved inference model. The learning is requested to the learning unit 18 within the information processing system 10, but if the information processing system 10 does not have a learning unit, an external learning unit is requested. Once the generation of the improved inference model has been requested, the process returns to step S21.
[0166] Returning to step S27, if the result of the determination is that the image is suitable for generating a new inference model, an examination image determination for existing training data is performed (S35). As described above, the examination image group 31 for existing training data is acquired at the first and second timings, and the image data is recorded when generating an inference model suitable for existing endoscopes. As described above, the second timing is the timing when a doctor or other professional closely examines the target area (see, for example, image Po in Figure 3A(b)), and the image is still and of high image quality with high clarity. In this step, the image similarity determination unit 13 determines whether the acquired image is similar to the examination image group 31 for existing training data.
[0167] Next, the control unit 11 records the images other than the group of test images for existing training data as a group of test images (S37). In this step, the control unit 11 records only the images that were determined not to be similar to the group of test images for existing training data as a group of test images in the recording unit 30 as test image group 32. It is highly likely that these images were acquired with a completely unknown endoscope that is not known on the market, and by performing inference using these images, it is possible to create an inference model suitable for this completely unknown endoscope.
[0168] Next, annotation is requested, and images that have not been annotated are added to the group of examination images (S39). Images determined to be part of the group of examination images in step S37 are images acquired using a completely unknown endoscope, etc., and therefore are images of an unknown type. Therefore, annotation is requested in order to generate a new inference model. This annotation can be done manually by a medical expert or automatically using AI, etc. Images that have not been annotated will not be used in creating the new inference model, but since there is a possibility that they will be used when creating an inference model on another occasion, these images are added to the recording unit 30 as the group of examination images 32. When recording, a tag is added to indicate that the images were unannotated.
[0169] Next, the control unit 11 records the annotated test image group as new training data (S41). When the control unit 11 finishes annotating the test image group recorded in step S37, the control unit 11 records the images as new training data in the recording unit 30 as test image group 32.
[0170] Next, new teacher data is selected, learning is performed, and a new inference model is created (S43). Once the new teacher data has been created, the control unit 11 then selects the new teacher data as appropriate. That is, from the new teacher data, teacher data suitable for generating the new inference model is selected. Once the new teacher data has been selected, the second request unit 17 then requests the learning unit to generate the new inference model. As in the case of generating an improved inference model, learning is requested from the learning unit 18 within the information processing system 10, but if the information processing system 10 does not have a learning unit, it is requested to an external learning unit. Once the generation of the new inference model has been requested, the process returns to step S21.
[0171] Next, the main operation of the endoscope 2 will be described using the flowchart shown in FIG. 5C. As described above, the information processing system 10 can be connected to various endoscopes and acquires various information from the image acquisition unit. Unlike the endoscope 1, the endoscope 2 is an endoscope manufactured by a manufacturer other than the specific manufacturer, and is the endoscope that generated the image PI2 or image PI3 in FIG. 2. The information processing system 10 has never generated an inference model based on an image from this endoscope 2. The main operation of the endoscope 2 shown in FIG. 5C is realized by the control unit 21 in the image system 20, which corresponds to the endoscope 2, controlling each unit in the image system 20.
[0172] When the flow of the endoscope 2 in Fig. 5C starts, first, an image is acquired and displayed (S51). Here, the image acquisition unit 23 in the endoscope 2 acquires an endoscopic image and displays this endoscopic image on the display in the display unit 24. While viewing this endoscopic image, the doctor or other medical professional operates the endoscope 2, moves the imaging unit of the endoscope 2 to the target area, and observes the target area.
[0173] Next, it is determined whether an inference assistance operation has been performed (S53). If the endoscope 2 is an unknown endoscope that is not known on the market, there is a possibility that the operator, such as a doctor, is unfamiliar with it. In such cases, the operator may request assistance through inference. To request inference assistance, for example, it is advisable to provide the endoscope 2 with an operation member for requesting an assistive operation. Furthermore, even if the operator, such as a doctor, does not perform a requesting operation, it may be automatically determined whether an inference assistance operation has been performed based on the analysis results of the endoscopic image and the operating state of the endoscope. Of course, it may also be determined by an inference model that infers whether an assistive operation is necessary. If the result of the determination in this step is that the operator has not requested an inference assistance operation, the system enters a standby state.
[0174] If the result of the determination in step S53 is that an inference assistance operation has been performed, then the image is transmitted to the information processing system 10, and the inference result is acquired (S55). Here, the control unit 21 of the endoscope 2 transmits the image acquired by the image acquisition unit 23 to the information processing system 10 via the communication unit. At this time, model information (information indicating the endoscope 2) may be transmitted to the information processing system 10 together with the image information. Having received the image, the information processing system 10 uses the existing inference model 19a of the inference engine 19 to infer whether or not a lesion such as a cancerous area is present, and returns this inference result to the endoscope 2 (see S25 in FIG. 5B). When receiving this inference result, the reliability of the inference result is also received.
[0175] Next, it is determined whether a highly reliable judgment result has been obtained (S57). Here, based on the result of the inference reliability received in step S55, the control unit 21 determines whether this reliability is higher than a predetermined value. If the result of this determination is low reliability, the process returns to step S53.
[0176] On the other hand, if the result of the determination in step S57 is high reliability, the determined position is then displayed in a frame (S59). In step S55, the inference engine 19 receives the inference result of whether or not there is a lesion such as cancer, and if there is a lesion, its position. In this step, the position of the lesion such as cancer inferred by the inference engine 19 is displayed in a frame. Of course, a display method other than a frame may be used. Once the determined position has been displayed in a frame, the process returns to step S51.
[0177] As described above, in the operation of the information processing system and the endoscopes 1 and 2, when the information processing system 10 receives images from the imaging system (endoscopes 1 and 2), it performs inference using the inference model 19a and returns the inference results to the imaging system (see S21 to S25). Therefore, even if the imaging system does not have an inference engine, it is possible to easily determine the presence or absence and location of a lesion such as cancer. Furthermore, if the acquired images are rare images that have never been seen before, even if they were acquired with an existing endoscope, these images are collected and an improved inference model is created (see S29 to S33).
[0178] Furthermore, in the operation of the information processing system, in the case of an image acquired with a completely unknown endoscope (for example, endoscope 2 shown in FIG. 5C) that is different from existing endoscopes, if the image is dissimilar to the first image, this image is collected as an inspection image and a new inference model is generated (see S35 to S43). In order to create an inference model suitable for images from an unknown endoscope, a large number of images must be collected, but this system makes it possible to efficiently use images collected daily from a variety of endoscopes.
[0179] Next, details of the operation of determining whether the current model is an improved or new model or other in step S27 (see FIG. 5B) will be described using the flowchart shown in FIG.
[0180] When the flow of FIG. 6 starts, first, it is determined whether model information, etc., is existing (S61). Model information is information such as the name of the manufacturer and the model name of the manufacturer, such as endoscope 1 in FIG. 5A or endoscope 2 in FIG. 5C. Based on this model information, it is determined whether the image received in steps S21 and S23 is an image acquired by an endoscope made by an existing manufacturer that was used to generate the existing inference model 19a. Since model information may be included in category information, etc., associated with the image, the control unit 11 determines the model information based on the category information. In this step, it may be determined whether or not the image relates to an existing model based on information such as the mode, accessories, and surgical procedure, as well as the model information. If tag data, etc., indicating model information, etc., is not attached to the image, it may be determined based on the characteristics or shape of the image.
[0181] If the result of the determination in step S61 is that the model information, etc., is from an existing manufacturer, it is next determined whether the reliability of the cancer determination is within a specific range (S63). In step S25 (see FIG. 5B), a lesion such as cancer is inferred, and the reliability of the inference at this time is determined. The reliability of the inference is not limited to 0% or 100%, but can also take intermediate values. When the reliability is 40% to 60%, when a lesion such as cancer is estimated using an existing inference model, it is difficult to say whether the inference result is correct or incorrect. When the reliability is 0%, the existing inference model is completely incompatible, while when the reliability is 100%, the existing inference model is completely compatible. When the reliability is within a specific range, it indicates that the endoscopic image to be inferred does not perfectly match the existing inference model, and there is room for improvement in the existing inference model.
[0182] Therefore, if the result of the determination in step S83 is that the reliability is not within the specific range (40 to 60%), the process goes to "otherwise" (S73) and proceeds to step S28 in Fig. 5B. On the other hand, if the result of the determination in step S83 is that the reliability is within the specific range, the process goes to S65 to improve the current state and proceeds to S29 in Fig. 5B.
[0183] Returning to step S61, if the model information does not indicate an existing endoscope, etc., the process next determines whether the model is completely unknown, taking into account differences in image quality characteristics and the surgical procedure (S67). In this step, it is determined whether the input image is an image acquired by an entirely unknown endoscope model, or whether it is an image that is not significantly different from an existing endoscope, or an endoscope model that is known on the market. Since the result of the determination in step S61 is No, it can be said that the input image is an image acquired by an endoscope that is not an existing model (the model used to generate the existing inference model). Therefore, the control unit 11 determines whether the model is completely unknown, taking into account differences in image quality characteristics and the surgical procedure. When taking into account differences in image quality characteristics and the surgical procedure, the model, image quality (optical characteristics), lighting, angle of view, operation, treatment tool, etc., and reliability are taken into account. In the case of surgical endoscopes, etc., different surgical procedures require different inference models, so it is possible to switch between a new inference model and an improved current model depending on the surgical procedure.
[0184] In step S67, if model information is associated with the endoscopic image received by the information processing system 10, this information is used to determine whether the endoscope model is not significantly different from the model used to generate the existing inference model, or whether it is a completely unknown model. Furthermore, the image quality (optical characteristics), lighting for the image, angle of view, etc., reflect the characteristics of the manufacturer, so the model can also be determined based on this information. Furthermore, since different endoscope manufacturers have unique operating methods and different treatment tools, etc., the model can be determined based on the acquired image (especially video). Furthermore, the model can also be determined based on the reliability of the inference.
[0185] If the result of the judgment in step S67 is that the endoscopic image was not acquired by a completely unknown model, it is designated as new model 1 (S69). Although the acquired endoscopic image is not of an existing model, it is not significantly different from the existing model, so it can be said that it is sufficient to generate a new inference model that merely improves the existing inference model. Therefore, in step S69, the generation of new model 1 is designated, and the current status is improved (S75), and the process proceeds to step S29 in Figure 5B.
[0186] On the other hand, if the result of the judgment in step S67 is that the endoscopic image was acquired using a completely unknown model, it is designated as new model 2 (S69). Since the acquired endoscopic image is from a model that is completely different from existing models, a new inference model is to be generated. Therefore, in step S71, the generation of new model 2 is designated, the new model is selected (S77), and the process proceeds to step S35 in Figure 5B.
[0187] 6, the model and other factors that captured the endoscopic image are determined based on the model information, reliability range, and image quality characteristics and surgical procedure, such as image quality, lighting, angle of view, operation processing, and treatment tool, and the inference model to be generated is then divided based on these determination results. In other words, images captured by a variety of endoscopes and collected in the information processing system 10 can be used efficiently according to the characteristics of these images.
[0188] Next, the operation of the learning device (information processing device) will be described using the flowchart shown in Fig. 7. This operation is realized by the learning unit 18 in the information processing system 10 controlling each unit in the information processing system 10 in accordance with a program stored in memory. For this flow, the learning unit 18 may have a control processor such as a CPU, or the control unit 11 may control each unit in the information processing system 10 to realize this flow. Furthermore, the learning unit 18 may of course be provided outside the information processing system 10 rather than within the information processing system 10.
[0189] When the learning flow of FIG. 7 starts, first, it is determined whether or not the image is a target scene (S81). A target scene is a scene in which a target site, such as a lesion such as cancer, is being observed, or a scene other than observation that is the subject of learning. Here, the information processing system 10 determines whether or not the image is a target image based on the scene of the image (for example, whether it is an image of inside the body, an image taken during a detailed examination by a doctor, etc.). As described above, this determination may be made based on the model information associated with the image, or may be made based on image quality (optical characteristics) or other information (see S67 of FIG. 6).
[0190] If the result of the determination in step S81 is that the scene is not of interest, the acquired endoscopic images are recorded as a first-timing image group (S83). In this case, the acquired endoscopic images are of a non-target scene, for example, images from the time the endoscope is inserted, to the time the target area is found, and then until the endoscope is removed, i.e., at the first timing. The information processing system 10 records the acquired images as first-timing images in the first-timing image section 31c of the recording unit 30.
[0191] On the other hand, if the result of the determination in step S81 is that the scene is the target, the image is recorded as a second timing image group (S85). In this case, since the image is of the target scene, it is often an image suitable for generating an inference model. If the image was captured with a model different from the one used to create the existing inference model, it will become an image for generating a new inference model. The information processing system 10 records the image as the second timing image in the second timing image section 31d of the recording unit 30.
[0192] Next, the annotation result is recorded (S87). Here, the learning unit 18 (control unit 11) requests that an annotation be added to the image indicating the position of a lesion such as cancer, and the annotated image, i.e., the training data, is recorded in the image file 31A. Note that the annotation may be added using AI or the like.
[0193] Next, learning is performed (S89). Here, machine learning is performed and weighting of the neural network is set so that when the annotated training data obtained in step S67 is input to the neural network of the inference engine, the location of a lesion such as cancer is output.
[0194] After the learning is performed in step S69, it is next determined whether the learning result has high reliability (S91). This determination can be made by inputting test data into the inference model and determining the range within which the error falls, or the amount of test data that falls within a specific error range, by comparing it with, for example, a predetermined reference value, to determine whether high reliability has been achieved.
[0195] If the result of the determination in step S71 is that high reliability cannot be ensured, images are selected (S93). Here, images that have reduced reliability are eliminated, and images that can improve reliability are added. After the image selection is complete, the process returns to step S81, and the above-mentioned operations are repeated.
[0196] On the other hand, if the determination in step S91 confirms that high reliability is ensured, the inference model, specifications, version, etc. are recorded (S95). Since a highly reliable inference model has been created, this inference model, the specifications used to generate this inference model, version information, etc. are recorded in the recording unit 30, etc. Once these are recorded, the inference model is determined and the operation of the learning device ends. The inference model created here can be used to infer lesions such as cancer when the information processing system 10 receives images of the same type as the images recorded as part of the examination image group.
[0197] 7, when the information processing system 10 performs learning to generate an inference model, it determines whether the input endoscopic image is a target scene (S81), and if it is a target scene, it collects this image as a second timing image (S85), and uses this image to create training data and an inference model (S87-S89). This allows the images collected by the information processing system 10 to be used efficiently.
[0198] If the scene is not the target, it is recorded as a first timing image. The first timing image is not directly used in the learning (S89) for creating an inference model, but contains information indicating what images can be used in the inference model in comparison with the second timing image. When creating a new inference model, this information can be used to select images for creating the new inference model from the test image group 32 that has been accumulated up to that point.
[0199] The concept of the present embodiment described above is that if an inference model with specific specifications exists and the history of the training data and the video from which it was created can be traced, then when improving or creating a new inference model with similar specifications, it is possible to easily find new training data candidate images by referencing the video from which the training data originated and the relationships between the image frames in the video that served as training data, and present these images in the annotation process. This concept can also be realized by an embodiment in which training data candidate image frames are inferred. Therefore, using FIG. 8, we will explain the generation of an inference model that selects one of a series of images (which may be a video or multiple still images) as an image for creating training data, and inference using this inference model.
[0200] As described above, in this embodiment, in order to efficiently collect new training data, we take advantage of the fact that when creating a proven inference model, images containing the object to be detected are selected from an infinite number of images. In other words, by annotating and learning from images selected for creating training data from a series of images, we can generate an inference model that extracts images that are candidates for annotation.
[0201] Figure 8 shows the learning method (see Figure 8(a)) for creating an inference model that selects annotation candidate images using inspection videos (still image frames) obtained up to now, and the input / output state of the inference model generated by this learning (Figure 8(b)). In other words, based on the relationship between proven inspection videos up to now and the training data within them, it is possible to create an inference model that can detect a limited number of frames from an inspection video consisting of a huge number of frames.
[0202] The series of test videos Pm1 to Pm3 are test images used as training data. If each of the test videos Pm1 to Pm3 is treated as a single image and one of the images is designated as the image to be annotated, it becomes possible to infer where the image is, similar to face detection technology in photographic images.
[0203] 8(a), in a series of examination videos Pm1, images Pmi1 are a group of images taken during insertion, images Pma1 are a group of images to be annotated, and images Pmr1 are a group of images taken during removal. Before machine learning, if there is an image (here, examination image Pma1) in the series of examination videos that has been annotated as containing a lesion such as an affected area, that image is annotated for creating an inference model that selects annotation candidate images.
[0204] However, since the series of inspection videos Pm1 to Pm3 contains a huge number of images, the amount of data becomes enormous. Therefore, if necessary, pixels in each frame may be thinned out, or the frames themselves or the entire image may be compressed, and then the images to be annotated may be selected and used as training data. Also, while FIG. 8(a) uses a video as an example, other timing information such as treatment information Ti1 and Ti2 may be included as part of the data in addition to the video data. This information, along with changes in the images, is also useful information for detecting features. Note that although FIG. 8(a) shows only three inspection videos, Pm1 to Pm3, it indicates that there are many inspection videos.
[0205] In this manner, in this embodiment, endoscopic video information obtained during a specific examination (in the figure, examination image Pm1 as training data) is used as training data by annotating frames of the endoscopic video corresponding to the training image (in the figure, examination image Pma1 to be annotated) as annotation information, and similar training data is created for endoscopic video images from multiple examination results (examination images Pma2 and Pma3 as training data in Figure 8(a)). Once the training data is created, endoscopic videos including insertion or removal images are input to the input layer In of the neural network NNW, and learning is performed to create an inference model so that the output layer Out becomes training data frames (Pma1 to Pma3).
[0206] Incidentally, in describing this embodiment, it has been written as if it were assumed that images of insertion and removal are included, but this is not a requirement. In other words, this follows the precedent that accuracy is improved by aligning images according to similar rules, such as detecting a facial image before detecting eyes, based on the concept of normalization. Therefore, a similar inference model can be created even for images that do not include insertion and removal, but include images that are candidates for training data, and include videos before and after them. In this case, if changes in procedures and imaging methods are converted into information and recorded in synchronization with the images, the amount of information required to infer images unique to the training image candidates increases, enabling better inference.
[0207] After generating an inference model using the method shown in Figure 8(a), the generated inference model is then set in the inference engine InEn shown in Figure 8(b). When a newly acquired inspection video Pmx is input to the input layer In of this inference engine InEn, it can infer which frames should be annotated and output annotation candidate images from the output layer Out. The inspection video may also include treatment time information Tix1, Tix2, etc. During inference, each image contained in the video is judged as a large image data set arranged in chronological order, which can be expressed as a panoramic composite image, and inference can be performed to find parts from this that are most suitable as training data. Furthermore, if treatment information can be effectively utilized, it can also be used during inference to infer annotation candidates.
[0208] In the method shown in Figure 8(a)(b), the inference model judges the image group at the first timing and the image group at the second timing, and infers and judges the image group at the second timing or the candidate training data frames therein. In particular, when similar affected areas are present in similar locations, the inspection video contains similar image information, enabling highly accurate inference.
[0209] Therefore, in this embodiment, the method for creating an inference model is to use endoscopic video information obtained during a specific examination (see, for example, examination images Pm1 to Pm3, etc.) as training data by annotating frames of this endoscopic video that correspond to training images, and then similarly converting endoscopic video images from multiple examination results into training data (see, for example, annotation images Pma1 to Pma3, etc.), inputting the endoscopic video including insertion or removal images into a learning device (see the input layer In of the neural network NNW), and learning is performed so that the frames of the training data become the output (see the output layer Out of the neural network), thereby creating an inference model.
[0210] As described above, in one embodiment of the present invention, an information processing device is capable of cooperating with a learning unit that determines whether a group of images acquired in time series by the first endoscope is an image group acquired at a first timing or an image group acquired at a second timing in order to determine specific image features of images acquired from the first endoscope, and obtains a first inference model for determining image features of the first endoscopic images, learned using the results of annotating the image group acquired at the second timing as training data. This information processing device has a classification unit that classifies a newly acquired group of images from the first or second endoscope using the image group acquired at the first timing when the first inference model (an existing inference model) was created. This makes it possible to efficiently collect images for generating a second inference model different from the first inference model.
[0211] Furthermore, the image group obtained in time series by the first endoscope is determined to be either an image group obtained at a first timing (images during access) or an image group obtained at a second timing (images obtained during detailed examination to find a lesion such as cancer), and the image group for generating the second inference model is selected from images obtained at the first timing, not just the second timing. Therefore, images obtained during access can also be used, and images for the second inference model can be collected efficiently.
[0212] In addition, in one embodiment of the present invention, the method includes a division step of dividing the examination video obtained from the first endoscope according to timing (e.g., first timing or second timing), an image classification step of excluding frames corresponding to the first timing from the training data candidates and adopting frames corresponding to the second timing as training data candidates (e.g., see S35, S37, etc. in FIG. 5B), a recording step of classifying and recording the examination video by division timing (e.g., see S37, etc. in FIG. 5B), and a step of learning a first inference model that uses at least one of the training data candidates to infer diseased area information contained in the frames of the examination video (e.g., see S89, etc. in FIG. 7). Therefore, it is possible to efficiently collect images for generating a second inference model different from the first inference model and generate the inference model.
[0213] In addition, one embodiment of the present invention includes an annotation step in which the results of annotating the frame corresponding to the second timing are used as candidate teacher data (see, for example, S85, S87, etc. in FIG. 7). In addition, one embodiment of the present invention includes a step in which, when creating a second inference model different from the first inference model, a new inspection video is acquired (see, for example, S81, etc. in FIG. 7), and a selection step in which candidate teacher data for the second inference model is selected by classifying the new inspection video using a classification that reflects the image classification results (see, for example, S85, S87, etc. in FIG. 7).
[0214] Although one embodiment of the present invention has been described with a focus on endoscopic images, the present invention can also be applied to information processing devices that generate inference models using images from various other inspection devices. In other words, the technology for selecting training image candidates from time-series image frames is expected to be utilized in a variety of fields. In this embodiment, the example of the insertion and removal of an endoscope from a body cavity, which is unique to endoscopes, was used to explain how video images obtained during the inspection process are classified and analyzed in chronological order by timing. However, such a so-called normalized procedure can be used in any field and is not limited to endoscopic images.
[0215] For example, in medical settings, applications to other imaging diagnostic devices such as ultrasound and radiology are conceivable, and applications to emergency assessment in operating rooms are also possible. There are inference models used to assess emergencies based on images acquired by surveillance cameras and the like, based on human behavior and movements, changes in the situation of crowds of people, and the like, and to provide supplementary information. Even in this example, the process of finding specific data from a large amount of data obtained over time is time-consuming. As such, training data is required to create an inference model, and this embodiment can be applied to any technology that selects candidate information to be used as training data from a large amount of information. Of course, this application is not limited to image information.
[0216] Furthermore, in the field of lesion detection, there are approaches from various angles, such as differentiation, prevention of oversight, educational use, etc. In addition to applications in these fields, training data is also required in applications such as insertion guides and treatment guides, and this embodiment can be applied when a technology is required to select candidate information to be used as training data from a large amount of information.
[0217] Of course, the concept of this embodiment can be applied to areas outside the medical field, such as the field of autonomous driving, where a vehicle enters and exits a highway interchange, where various situation determination inference models are required for video frames sandwiched between specific processes, such as roads, bridges, tunnels, and other structures, relationships with other vehicles, and relationships with signs, etc. Even when leaving evidence of the concept behind the inference model, by clarifying the method from data acquisition to selecting training data candidates, as in this embodiment, it is possible to prevent the system from becoming a black box.
[0218] Furthermore, in one embodiment of the present invention, the information processing system 10, the image system 20, and the recording unit 30 have been described as separate entities, but these two or all three may be configured as an integrated unit. Furthermore, in one embodiment of the present invention, the control units 11 and 21 have been described as devices including a CPU, memory, and the like. However, in addition to being configured as software using a CPU and a program, some or all of the units may be configured as hardware circuits, or may be configured as hardware such as gate circuits generated based on a programming language such as Verilog or VHDL (Verilog Hardware Description Language), or may be configured as hardware using software such as a DSP (Digital Signal Processor). Of course, these may be combined as appropriate.
[0219] Furthermore, the control units 11 and 21 are not limited to CPUs, and any element that functions as a controller may be used, and the processing of each of the above-described units may be performed by one or more processors configured as hardware. For example, each unit may be a processor configured as an electronic circuit, or each circuit unit in a processor configured as an integrated circuit such as an FPGA (Field Programmable Gate Array). Alternatively, a processor configured as one or more CPUs may execute the functions of each unit by reading and executing a computer program recorded on a recording medium.
[0220] Furthermore, in one embodiment of the present invention, the information processing system 10 has been described as having a control unit 11, an input determination unit 12, an image similarity determination unit 13, a metadata assignment unit 14, a classification unit 15, a first request unit 16, a second request unit 17, and a learning unit 18. However, these do not need to be provided within a single device, and the above-mentioned units may be distributed as long as they are connected by a communication network such as the Internet.
[0221] Furthermore, in recent years, artificial intelligence capable of making judgments based on various criteria has become increasingly popular, and it goes without saying that improvements such as collectively performing each branch of the flowchart shown here also fall within the scope of the present invention. If the user can input their opinion on the pros and cons of such control, the embodiment shown in this application can be customized to suit the user by learning their preferences.
[0222] Furthermore, among the technologies described in this specification, the controls mainly described in the flowcharts can often be set by a program, and may be stored on a recording medium or a recording unit. The recording method for this recording medium or recording unit may be recording at the time of product shipment, using a distributed recording medium, or downloading via the Internet.
[0223] Furthermore, in one embodiment of the present invention, the operation of this embodiment is explained using a flowchart, but the order of the processing procedure may be changed, any step may be omitted, steps may be added, and the specific processing content within each step may be changed.
[0224] Furthermore, even if the operational flow in the claims, specification, and drawings is explained using words expressing order such as "first" and "next" for convenience, this does not mean that it is necessary to perform the operation in this order unless otherwise specified.
[0225] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some of the components shown in the embodiments may be omitted. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]
[0226] 10 Information processing system, 12 Input judgment unit, 13 Image similarity judgment unit, 14 Metadata assignment unit, 15 Classification unit, 16 First request unit, 17 Second request unit, 18 Learning unit, 19 Inference engine, 19a Existing inference model, 20 Image system, 21 Control unit, 22 Processing unit, 23 Image acquisition unit, 24 Display unit, 25 Display content, 25a Inference result, 25b... Reliability, 30... Recording section, 31... Inspection image group for existing training data, 31A... Image file, 31a... Category information, 31b... Specification information, 31c... First timing image section, 31d... Second timing image section, 32... Inspection image group, 32a... Category information, 32b... Image, 33... New inference model creation specification information section, 33a... Additional training data specification,
Claims
1. When creating the first inference model, the data obtained during the insertion of the endoscope and the examination are recorded together with information indicating the history of the training data used, The newly acquired group of endoscopic images is compared with a first group of images at a timing including an image of the affected area and a second group of images at another second timing, both of which are included in the historical video of the training data, and based on the similarity, a group of images to be candidates for training data for an inference model is selected from the newly acquired group of endoscopic images. An information processing method comprising:
2. 2. The information processing method according to claim 1, wherein the teaching data is recorded together with information indicating whether or not the teaching data has been adopted as annotation data.
3. 2. The information processing method according to claim 1, wherein the information indicating the first image group and the second image group is recorded together with the teacher data.
4. 2. The information processing method according to claim 1, wherein the group of images serving as training data candidates is selected for each case.
5. a recording unit that records the training data acquired during the insertion of the endoscope and the examination when creating the first inference model, along with information indicating the history of the training data used; a selection unit that compares the newly acquired endoscopic image group with a first image group at a timing when an affected area image is included and a second image group at another second timing, both of which are included in the historical video of the training data, and selects an image group from the newly acquired endoscopic image group as candidate training data for an inference model based on the degree of similarity; An information processing device comprising:
6. An information processing method in an information processing system capable of creating an inference model for an endoscope, comprising: The information processing system includes an inference unit, a determination unit, a classification unit, and a learning unit, The information processing method includes: The determining unit determines whether an image group obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine a specific image feature for images obtained continuously from the first endoscope; providing the inference unit with a first inference model created by the learning unit based on images continuously obtained from the first endoscope and created based on a group of images obtained at the second timing; determining by the determining unit whether the image group obtained in time series by the second endoscope is an image group obtained at the first timing or an image group obtained at the second timing; The classification unit selects and classifies a group of images from the newly acquired image group from the endoscope as candidates for training data for improving the first inference model based on the similarity between the group of images obtained at the first timing when the first inference model was created and the group of images from the newly acquired endoscope. It is characterized by:
7. An information processing device in an information processing system capable of creating an inference model for an endoscope, a learning unit is provided within the information processing system, The information processing device includes: a determination unit that determines whether an image group obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine a specific image feature of images continuously obtained from the first endoscope; an inference unit having a first inference model created by the learning unit based on images obtained continuously from the first endoscope and created based on a group of images obtained at the second timing; and the determining unit determines whether an image group obtained in time series by the second endoscope is an image group obtained at a first timing or an image group obtained at a second timing; The information processing device further comprises: a classification unit that selects and classifies a group of images from the newly acquired image group from the endoscope as candidates for training data for improving the first inference model based on the similarity between the group of images obtained at the first timing when the first inference model was created and the group of images from the newly acquired endoscope; The present invention is characterized by having the following.
8. An information processing method in an information processing system capable of creating an inference model for an endoscope, comprising: The information processing system includes an inference unit, a determination unit, a classification unit, and a learning unit, The information processing method includes: The determining unit determines whether an image group obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine a specific image feature for images obtained continuously from the first endoscope; providing the inference unit with a first inference model created by the learning unit based on images continuously obtained from the first endoscope and created based on a group of images obtained at the second timing; determining by the determining unit whether the image group obtained in time series by the second endoscope is an image group obtained at the first timing or an image group obtained at the second timing; Based on the similarity between the group of images obtained at the first timing when the first inference model was created and the group of images obtained from the newly acquired endoscope, the classification unit selects and classifies a group of images from the newly acquired endoscope that will be candidates for training data for an inference model different from the first inference model. It is characterized by:
9. An information processing device in an information processing system capable of creating an inference model for an endoscope, a learning unit is provided within the information processing system, The information processing device includes: a determination unit that determines whether an image group obtained in time series by the first endoscope is an image group obtained at a first timing or an image group obtained at a second timing in order to determine a specific image feature of images continuously obtained from the first endoscope; an inference unit having a first inference model created by the learning unit based on images obtained continuously from the first endoscope and created based on a group of images obtained at the second timing; and the determining unit determines whether an image group obtained in time series by the second endoscope is an image group obtained at a first timing or an image group obtained at a second timing; The information processing device further comprises: a classification unit that selects and classifies a group of images from the newly acquired image group from the endoscope as candidates for training data for an inference model different from the first inference model based on the similarity between the group of images obtained at the first timing when the first inference model was created and the group of images from the newly acquired endoscope; The present invention is characterized by having the following.
10. Obtaining endoscopic video during specific examinations The video information obtained by using frames corresponding to the teaching images from the acquired endoscopic video as annotation information is used as teaching data; After acquiring the endoscopic video during the specific examination, a plurality of endoscopic images including insertion or removal images are further acquired; An endoscopic video including the insertion or removal image is input to a learning device, and learning is performed to create an inference model so that the frames of the training data are output. An information processing method comprising:
11. acquiring the endoscopic video using a first endoscope during the specific examination; The endoscopic videos of the plurality of examination results are acquired using a second endoscope.
11. The information processing method according to claim 10.
12. creating a first inference model based on a group of images obtained from the first endoscope at a second timing; selecting a group of images to be second training data candidates from the group of images from the first endoscope or the second endoscope based on the similarity between the group of images obtained at the first timing when the first inference model was created and the group of images newly obtained from the first endoscope or the second endoscope; 12. The information processing method according to claim 11.
13. The information processing method described in claim 12, characterized in that the first inference model is an inference model for displaying inference results for images obtained by the first endoscope, and the second inference model is an inference model for displaying inference results for images obtained by the second endoscope.
14. The information processing method described in claim 12, characterized in that if the reliability of the inference result when an image is input into the first inference model is lower than a predetermined value, it is determined that the input image is an image from the second endoscope.
15. a first image acquisition unit that acquires an endoscopic video during a specific examination; a teacher data creation unit, which uses video information obtained by using frames corresponding to teacher images from the endoscopic video acquired by the first image acquisition unit as annotation information; a second image acquisition unit that acquires a plurality of endoscopic images including insertion or removal images after acquiring the endoscopic video during the specific examination; a request unit that inputs an endoscopic video including the insertion or removal image into a learning device and requests the learning device to perform learning and create an inference model so that the frames of the training data are output; An information processing device comprising:
Citation Information
Patent Citations
Movement amount determining device, movement amount determining method, and movement amount determining program
JP2020067592A
Learning device, imaging device, ai information providing device, learning method and learning program
JP2020166744A
Image processing apparatus, method for controlling image processing apparatus, method for creating discriminator, identification method, identification device, device for creating discriminator, and discriminator
JP2021093142A
Signal analysis systems and methods for feature extraction and interpretation thereof
JP2019087221A