Display control device, display control method, display control system, and display control program

By utilizing both image and non-image data, the system enhances diagnostic accuracy in endoscopic imaging by creating multiple inference models that account for patient-specific factors, addressing the issue of inconsistent image quality and improving reliability in medical judgments.

WO2025173080A1PCT designated stage Publication Date: 2025-08-21OLYMPUS MEDICAL SYST CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/004839
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-13
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing endoscopic imaging systems face challenges in achieving accurate inference due to varying imaging conditions, leading to inconsistent image quality and reliance on training data that may not account for non-image information, resulting in unreliable diagnostic support.

Method used

The system incorporates both image and non-image information, such as text data, to create multiple inference models that enhance diagnostic accuracy by considering patient-specific factors like gender and race, using deep learning and neural networks to process endoscopic images and associated text data.

Benefits of technology

This approach enables more accurate and reliable diagnostic support by integrating diverse training data, improving the precision of medical judgments through comprehensive consideration of patient-specific factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024004839_21082025_PF_FP_ABST
    Figure JP2024004839_21082025_PF_FP_ABST
Patent Text Reader

Abstract

This image classification device for an endoscope comprises: a first inference model that has been trained on the basis of a first training data group including a plurality of pieces of first training data obtained by imparting, to each image, annotation information corresponding to the image and text information differing from the annotation information and pertaining to the image; and a control circuit that outputs, for display, a first inference result output by the first inference model when a prescribed image and prescribed text information are input into the first inference model.
Need to check novelty before this filing date? Find Prior Art

Description

Display control device, display control method, display control system, and display control program

[0001] The present invention relates to a display control device, a display control method, a display control system, and a display control program that support various types of judgments on images of organs and the like.

[0002] In recent years, medical systems that use AI (artificial intelligence) to make various judgments have been developed in the medical field. To enable highly accurate AI judgments, endoscopic images acquired during endoscopic examinations, for example, are used to build inference models used in AI judgments.

[0003] An endoscope is a device that is inserted into the body and enables observation of diseased areas that cannot be seen from the outside. An endoscope has an insertion section that is inserted into a body cavity, and an imaging device is provided, for example, at the tip of the insertion section. During an examination using an endoscope, a doctor sequentially displays images acquired by the imaging device at the tip of the insertion section inserted into the body, and adjusts the position of the tip of the insertion section while checking the displayed images to diagnose the patient's health or disease state.

[0004] Image information obtained from the process of inserting the endoscope into the body until it is removed is digitized and acquired. Learning is performed using a large amount of training data created using the acquired image information to build an inference model used for diagnosis, etc.

[0005] Japanese Patent Publication No. 2021-183017 discloses an information processing device that improves inference accuracy by selecting at least one inference model from multiple inference models based on at least one imaging condition of multiple imaging conditions.

[0006] Japanese Patent Application Laid-Open No. 2002-253539

[0007] However, imaging conditions vary greatly when imaging intraluminal areas using an endoscope, and the image quality of the obtained images often differs significantly from the image quality of the images used for learning to build an inference model. For this reason, even if the proposal in Patent Document 1 is used, inference may not be performed with sufficient accuracy. The present invention aims to provide a display control device, a display control method, a display control system, and a display control program that can build an inference model for determining the state inside the body based on training data that uses not only images but also information other than images, and display highly accurate inference results.

[0008] A display control device according to one aspect of the present invention comprises a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching to each image annotation information corresponding to each image and text information that is different from the annotation information and related to the image, and a control circuit that outputs for display a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model.

[0009] A display control method according to one aspect of the present invention includes a first teacher data creation step of creating first teacher data by attaching to an image annotation information corresponding to the image and text information that is different from the annotation information and related to the image; a first teacher data group generation step of repeating the first teacher data creation step to generate a first teacher data group including a plurality of the first teacher data; a first inference model generation step of generating an inference model trained based on the first teacher data group generated in the first teacher data group generation step; and a display step of displaying on a display device the first inference result output from the first inference model when a predetermined image and predetermined text information are input to the first inference model.

[0010] A display control system according to one aspect of the present invention comprises a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching to each image annotation information corresponding to each image and text information that is different from the annotation information and related to the image, an image generation device that provides a predetermined image to the first inference model, and a control circuit that outputs for display a first inference result output from the first inference model when the predetermined image and predetermined text information are input from the image generation device to the first inference model.

[0011] A display control device according to another aspect of the present invention has a control circuit that inputs a predetermined image and predetermined text information corresponding to the predetermined image into an inference model and outputs the inference results of the inference model for display.

[0012] A display control device according to another aspect of the present invention comprises: a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching, to each image, annotation information corresponding to each image and text information related to the image that is different from the annotation information; a second inference model trained based on a second group of training data including a plurality of second training data obtained by attaching the annotation information to each of the images; and a display control circuit that outputs for display in a distinguishable manner a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model, and a second inference result output from the second inference model when a specified image is input to the second inference model.

[0013] A display control device according to another aspect of the present invention comprises first and second inference models that receive at least an image as input, a display control circuit capable of displaying a first inference result of the first inference model and a second inference result of the second inference model superimposed on a display image, a detector that detects voice after the display control circuit displays the second inference result for the first image superimposed on the display image, and an interface circuit that converts the voice detected by the detector into text and provides the converted text and the first image to the first inference model, and after displaying the second inference result, the display control circuit displays the inference result of the first inference model based on the converted text and the first image superimposed on the display image.

[0014] A display control system according to another aspect of the present invention comprises: a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching, to each image, annotation information corresponding to each image and text information related to the image that is different from the annotation information; a second inference model trained based on a second group of training data including a plurality of second training data obtained by attaching, to each image, the annotation information; an image generation device that provides a predetermined image to the first and second inference models; and a display control circuit that outputs for display in a distinguishable manner a first inference result output from the first inference model when a predetermined image and predetermined text information are input to the first inference model, and a second inference result output from the second inference model when a predetermined image is input to the second inference model.

[0015] A display control system according to another aspect of the present invention comprises first and second inference models that receive at least an image as input, an image generation device that provides a predetermined image to the first and second inference models, a display control circuit capable of displaying a first inference result of the first inference model and a second inference result of the second inference model superimposed on a display image, a detector that detects voice after the display control circuit displays the second inference result when a first image is provided to the second inference model as the predetermined image superimposed on the display image, and an interface circuit that converts the voice detected by the detector into text and provides the converted text and the first image to the first inference model, and after displaying the second inference result, the display control circuit displays the inference result of the first inference model based on the converted text and the first image superimposed on the display image.

[0016] A display control method according to another aspect of the present invention creates a first group of teacher data including a plurality of first teacher data obtained by attaching, to each image, annotation information corresponding to each image and text information related to the image that is different from the annotation information, and creates a second group of teacher data including a plurality of second teacher data obtained by attaching the annotation information to each of the images, creates a first inference model trained based on the first group of teacher data and a second inference model trained based on the second group of teacher data, and outputs for display in a distinguishable manner a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model, and a second inference result output from the second inference model when a specified image is input to the second inference model.

[0017] A display control method according to another aspect of the present invention includes inputting at least an image into a first and a second inference model, displaying the second inference result for the first image superimposed on the display image, detecting speech, converting the detected speech into text, providing the converted text and the first image to the first inference model, and after displaying the second inference result, displaying the inference result of the first inference model based on the converted text and the first image superimposed on the display image.

[0018] A display control program according to another aspect of the present invention causes a computer to execute the following steps: create a first group of teacher data including a plurality of first teacher data obtained by attaching, to each image, annotation information corresponding to each image and text information related to the image that is different from the annotation information; create a second group of teacher data including a plurality of second teacher data obtained by attaching, to each image, the annotation information; create a first inference model trained based on the first group of teacher data; and create a second inference model trained based on the second group of teacher data. The program then outputs for display a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model, and a second inference result output from the second inference model when a specified image is input to the second inference model, in a distinguishable manner.

[0019] A display control program according to another aspect of the present invention causes a computer to execute the following steps: input at least an image into a first and a second inference model; display the second inference result for the first image superimposed on the display image; then detect speech; convert the detected speech into text; provide the converted text and the first image to the first inference model; and, after displaying the second inference result, display the inference result of the first inference model based on the converted text and the first image superimposed on the display image.

[0020] According to the present invention, it is possible to construct an inference model for determining the state of the body based on training data using not only images but also information other than images, and to display highly accurate inference results.

[0021] FIG. 1 is a block diagram showing an endoscope system including an inference device according to an embodiment of the present invention. FIG. 2 is an explanatory diagram showing an example of first teacher data. FIG. 3 is an explanatory diagram showing an example of the first teacher data. FIG. 4 is a block diagram showing a specific configuration for creating an inference model and presenting inference results. FIG. 5 is an explanatory diagram showing an example of recorded contents of a knowledge database. FIG. 6 is a flowchart showing the operation of an embodiment. FIG. 7 is a flowchart showing the operation of an embodiment. FIG. 8 is a flowchart showing the operation of an embodiment. FIG. 9 is an explanatory diagram showing an example of a display. FIG. 10 is an explanatory diagram showing an example of a display. FIG. 11 is an explanatory diagram showing an example of a display. FIG. 12 is an explanatory diagram showing an example of using multiple pieces of non-image information for annotation. FIG. 13 is an explanatory diagram showing an example of using multiple pieces of non-image information for annotation.

[0022] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0023] (Embodiment) FIG. 1 is a block diagram showing an endoscope system including a display control device according to an embodiment of the present invention.

[0024] In a conventional inference device employed in Patent Document 1 and the like, when video obtained by endoscopic examination is used to construct an inference model, training data is created by annotating images that include the object to be determined with information such as the object's position and attributes, such as a flag indicating that the object is present. Annotation is a process of attaching information called tags or metadata to each piece of data in any form, such as text, audio, image, or video. An inference model is created by learning using the large amount of training data created. When an image is input into such an inference model, information such as the presence or absence of the object and its location on the screen is inferred and output.

[0025] As such, conventional inference models can be built by collecting images and can be developed without the need for special medical equipment. This makes it easy to set up a development environment and establish procedures, such as contracts for image acquisition, and has led to their adoption by many research institutions. Furthermore, the accuracy of inference can be improved according to the images collected, and research and development into this is ongoing, making it possible to provide extremely high levels of diagnostic support.

[0026] However, in endoscopic examinations, it is not always possible to obtain images of the same quality (color, brightness, optical performance, differences from the shape of the lesion at the time of learning, etc.) as the images used for training. Imaging conditions vary, and it is often not possible to obtain images of the same quality. This can make it difficult to make reliable inferences.

[0027] Therefore, in this embodiment, when creating training data, annotation is performed using information other than images, such as text information. Here, the text information may be information indicating a character string (a group of information consisting of symbols representing characters), or information indicating the symbols assigned in advance for each type or attribute. Generally, doctors make medical diagnoses by comprehensively considering not only endoscopic images but also various factors other than endoscopic images. For example, the locations and types of colon tumors that are prone to develop in men and women may differ, and doctors may make diagnoses taking gender into account. Therefore, by creating training data using not only images but also various information other than images (hereinafter referred to as non-image information), such as text information related to the patient's gender, and then using an inference model obtained by learning from this training data, it is believed that more accurate inferences can be achieved.

[0028] On the other hand, non-image information such as text can be obtained from, for example, medical records that chronologically record a patient's condition, treatment details, and test results. However, such information may contain noise other than the necessary information and variations in age classifications. Furthermore, using medical records as non-image information may result in the patient's past information being adopted, ignoring the patient's changes over time. Furthermore, non-image information may contain subjective evaluations by the creator and variations in the granularity of the description, and in some cases, non-image information may not be obtainable. For this reason, when non-image information is used as training data, the quality of the training data may be unstable. Even when categorized information is used as non-image information, there are issues regarding how to select the population and the type of information to acquire in the first place, making it difficult to obtain non-image information based on a common standard. For this reason, the amount of training data may be uneven depending on the image, and an inference model that is effective for all images may not be obtained.

[0029] Therefore, in this embodiment, an inference model based on training data generated using images and an inference model based on training data generated using images and non-image information (e.g., text) are adopted. The inference results obtained by these inference models can be presented, enabling more optimal diagnostic support.

[0030] 1, the endoscopic system 1 includes an inference device 10, an endoscope 20, an operation input unit 25, a processor 30, and a display device 40. The endoscope 20 has an insertion portion (not shown) that is inserted into a body cavity, and an imaging device (not shown) is provided at the tip of the insertion portion. This imaging device includes an imaging element such as a CCD or CMOS sensor, and photoelectrically converts an optical image from a subject to obtain an imaging signal. This imaging signal is supplied to the processor 30.

[0031] The processor 30 performs predetermined signal processing on the input image signal, such as color adjustment processing, matrix conversion processing, noise removal processing, and various other signal processing. The processor 30 includes a display control circuit for supplying the endoscopic image obtained by the signal processing to the display device 40 to display the endoscopic image. The display device 40 has a display screen such as an LCD or EL, and displays the image supplied from the processor 30.

[0032] During endoscopic examination, the endoscopic images from the processor 30 are also supplied to the inference device 10 for identification, etc. That is, the endoscope 20 as an image generating device provides images to the first and second inference models described below that are provided in the inference device 10.

[0033] The inference device 10 as a display control device includes a control circuit 11, an input interface (IF) 12, an output interface (IF) 13, an external interface (IF) 14, an inference model unit 15, and a judgment result recording unit 16. The control circuit 11 comprehensively controls the entire inference device 10. The control circuit 11 may be configured by a processor using a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), or the like. The control circuit 11 may operate according to a program stored in a memory (not shown) to control each unit, or may realize some or all of its functions using hardware electronic circuits.

[0034] The input IF 12 receives information from the processor 30 and outputs it to the control circuit 11 and the determination result recording unit 16. The output IF 13 outputs information from the control circuit 11 to the processor 30. The external IF 14 exchanges information with the in-hospital network line 2, transmits information from the control circuit 11 to the in-hospital network line 2, and supplies information from the in-hospital network line 2 to the control circuit 11.

[0035] The inference model unit 15 includes multiple networks. Each network is trained using a large amount of training data, for example, through deep learning, to determine the network design so that an output corresponding to each input is obtained. An inference model is constructed for each network. The inference model unit 15 trains by receiving training data from, for example, an external device outside the endoscope system 1.

[0036] Deep learning is a multilayered version of the machine learning process using neural networks. A typical example is a forward propagation neural network, which sends information from front to back and makes a judgment. In its simplest form, it requires three layers: an input layer consisting of m1 neurons, a hidden layer consisting of m2 neurons determined by parameters, and an output layer consisting of m3 neurons corresponding to the number of classes to be discriminated. The neurons in the input and hidden layers, and those in the hidden and output layers, are connected by connection weights, and a bias value is added between the hidden and output layers, making it easy to form logic gates. While three layers are sufficient for simple discrimination, increasing the number of hidden layers makes it possible to learn how to combine multiple features during the machine learning process. In recent years, neural networks with 9 to 152 layers have become practical due to their training time, judgment accuracy, and energy consumption.

[0037] The network N1 used for machine learning may be any of a variety of well-known networks. For example, R-CNN (Regions with CNN features) or FCN (Fully Convolutional Networks) using CNN (Convolution Neural Network) may be used. This involves a process called "convolution" that compresses image features, operates with minimal processing, and is strong in pattern recognition. Furthermore, a "recurrent neural network" (fully connected recurrent neural network) that can handle more complex information and allows information analysis whose meaning changes depending on the order or sequence of information may be used, allowing information to flow bidirectionally.

[0038] To realize these technologies, conventional general-purpose arithmetic processing circuits such as CPUs and FPGAs can be used, but because much of the processing in neural networks involves matrix multiplication, GPUs and Tensor Processing Units (TPUs), which are specialized for matrix calculations, may also be used.In recent years, such dedicated artificial intelligence (AI) hardware, called "neural network processing units (NPUs)," have been designed to be integrated and embeddable with CPUs and other circuits, and may even become part of the processing circuit.

[0039] Furthermore, inference models may be obtained by employing various well-known machine learning techniques, not limited to deep learning. For example, techniques such as support vector machines and support vector regression are available. Here, learning involves calculating the weights, filter coefficients, and offsets of a classifier; other techniques include using logistic regression processing. When a machine is to make a judgment, a human must teach the machine how to make the judgment. In this embodiment, a method is used in which image judgments are derived using machine learning. However, rule-based techniques that apply rules acquired by humans through experience or heuristics to make specific judgments may also be used.

[0040] The endoscope system 1 is connected to an in-hospital network line 2 via an external IF 14. The in-hospital network line 2 connects various devices (not shown) within the hospital. For example, the in-hospital network line 2 is connected to a medical record database (DB) device 5. The medical record DB device 5 contains various information related to patients who are the subject of examinations using the endoscope 20. The in-hospital network line 2 is connected to a learning database (DB) device 4 via an external network line 3. The learning DB device 4 contains a database that stores a huge amount of training data used to build an inference model. The learning DB device 4 may be provided by the manufacturer that builds the endoscope system 1.

[0041] The images used for the training data are, for example, observation images of patients taken at a medical facility. For example, medical images such as endoscopic images acquired by an endoscope 20 or the like may be used as the images used for the training data. The imaging signals acquired by the endoscope are supplied to a processor and processed to obtain the images used for the training data. Note that the images used for the training data may be a series of continuously acquired images or may be still images.

[0042] When a target site such as a lesion is included in each frame of a still image or a series of continuously acquired images, annotation information indicating the area of ​​the image portion in which the target site (lesion site) is depicted, such as a frame image indicating the area of ​​the target site, is added to each image in accordance with the observation findings. That is, the annotation information is information indicating the area of ​​the lesion depicted in the image. Furthermore, annotation information such as the name of the target site or the type of lesion may be added to each image. That is, training data is obtained by adding annotation information obtained by image diagnosis to the image. Hereinafter, training data annotated based on images in this manner is referred to as second training data. That is, the second training data is obtained by attaching annotation information corresponding to each image to each image. Learning is performed to construct an inference model using a second training data group including multiple second training data.

[0043] Here, annotation methods may include information indicating whether a target site (lesion) is captured in each image, and in addition to this information, information such as the site name of the target site and the type of lesion may be added to each image. In other words, annotation methods include various methods, such as a method of adding the specific position of the target site in an image (Classification), a method of adding information about the target site captured in an image to the image itself, and a method of adding information collectively to a series of continuously acquired images (or to a section in the case of a video). Examples of annotations (Classification) that indicate whether an object is captured in an image include a method of showing the area in an image where the object is captured using a rectangle (Object detection) and a method of showing the area in an image where the object is captured using a free-form curve (Semantic segmentation).

[0044] In this embodiment, the second training data is not only created by annotating based on images, but also by annotating using images, annotation information corresponding to the images, and non-image information related to the images. Hereinafter, training data annotated using not only images but also non-image information is referred to as first training data. The non-image information is annotation information related to the images but different from annotation information corresponding to the observation findings (image diagnosis) for the images, such as information on age, gender, race, etc. The non-image information may be, for example, text information. For example, the text information may be patient information related to an observation image of a patient. Patient information may be information based on at least one of the patient's gender, age, and race. That is, the first training data is obtained by adding, to each image, annotation information corresponding to the image and text information related to the image that is different from the annotation information. Learning to build an inference model is performed using a first training data group including multiple first training data.

[0045] 2 and 3 are explanatory diagrams for explaining examples of the first training data.

[0046] FIG. 2 shows an example of training data for cases of colon tumors. As mentioned above, colon tumors tend to be of different types and occur in different locations in men and women. Therefore, the learning DB device 4 creates second training data by annotating endoscopic images of the colon with frame images indicating the tumor area, and creates first training data by further annotating the second training data with annotation information indicating male or female. FIG. 2 shows such first training data. Images P1a-P4a showing male cases are annotated with frame images P1aL-P4aL indicating the tumor area and with information indicating male. An inference model constructed using this training data is a male model suitable for diagnosing male colon tumors. Images P1b-P4b showing female cases are annotated with frame images P1bL-P4bL indicating the tumor area and with information indicating female. The inference model constructed using this training data will be a female model suitable for diagnosing colorectal tumors in women.

[0047] Here, to create a male model, pre-learning may be performed using training data of female cases, and then final training may be performed using training data of male cases. Also, to create a female model, pre-learning may be performed using training data of male cases, and then final training may be performed using training data of female cases.

[0048] 3 shows an example of training data for different races. Because different cases tend to occur for each race, an inference model for each race is also created. Specifically, the learning DB device 4 creates second training data by annotating the endoscopic image with a frame image or the like indicating the lesion, and also creates first training data by further annotating the second training data by adding annotation information indicating race, for example, whether the person is Japanese or foreign.

[0049] For example, the annotation information may be text such as "Japanese male," or a code may be assigned to information such as Japanese male or Japanese female, and the code may be used as the annotation information.

[0050] FIG. 3 shows such first training data, in which images P1c to P4c showing cases of Japanese patients are annotated with box images P1cL to P4cL showing the lesions, as well as information indicating that the patient is Japanese. An inference model constructed using this training data is a model for Japanese patients suitable for diagnosing Japanese patients. Furthermore, images P1d to P4d showing cases of foreign patients (e.g., Americans) are annotated with box images P1dL to P4dL showing the lesions, as well as information indicating that the patient is foreign (e.g., Americans). An inference model constructed using this training data is a model for foreign patients (e.g., Americans) suitable for diagnosing foreign patients (e.g., Americans).

[0051] Here, to create a model for Japanese people, pre-training may be performed using training data of foreign cases, and then final training may be performed using training data of Japanese cases. Also, to create a model for Americans, pre-training may be performed using training data of Japanese cases, and then final training may be performed using training data of foreign cases.

[0052] Furthermore, for example, in addition to models for Japanese people and Americans, it is also possible to create models for Japanese women, Japanese men, American women, American men, etc. In this case, similar to the above, in order to create a Japanese women model, training data for cases other than Japanese women may be pre-trained, and then final training may be performed using training data for Japanese women cases.

[0053] In this way, the first training data is created by adding text information (non-image information), such as male / female type, to the second training data, which is an image and a case tag such as a frame coordinate indicating the area of ​​the target area in the image, added to the image as annotation information.

[0054] The learning DB device 4 creates and stores a large amount of second training data, each of which includes multiple images and annotation information added to each of these images. Meanwhile, non-image information can be acquired from the medical record DB device 5. The learning DB device 4 reads out various pieces of information about patients recorded in the medical record DB device 5 as non-image information, and annotates some of the second training data based on the read non-image information. For example, if the second training data is an image of a male case of colon tumor, the learning DB device 4 adds text information indicating that the patient is male to the training data to obtain first training data.

[0055] The control circuit 11 of the inference device 10 acquires first teacher data and second teacher data from the learning DB device 4 via the external network line 3, the in-hospital network line 2, and the external IF 14. The control circuit 11 supplies the acquired first and second teacher data to each network of the inference model unit 15, respectively, for learning. This allows multiple inference models based on the second teacher data and multiple inference models based on the first teacher data to be constructed.

[0056] During an endoscopic examination, the processor 30 outputs endoscopic images acquired during the examination to the inference device 10. The control circuit 11 of the inference device 10 receives the endoscopic images from the processor 30 via the input IF 12 and provides the received endoscopic images to the inference model unit 15 to perform inference. The control circuit 11 outputs the inference results of the inference model unit 15 to the processor 30 via the output IF 13. The processor 30, which serves as a display control circuit, adds the inference results of the inference model unit 15 to the endoscopic images from the endoscope 20 and provides them to the display device 40. In this way, the display device 40 displays the endoscopic images and the inference results for these endoscopic images.

[0057] Fig. 4 is a block diagram showing a specific configuration for creating an inference model and presenting the inference results. In Fig. 4, the same components as in Fig. 1 are assigned the same reference numerals. Fig. 4 mainly shows the configuration of the control circuit 11 and the inference model unit 15 of the endoscope system 1.

[0058] As shown in FIG. 4 , the learning DB device 4 stores a second teacher data group including multiple second teacher data sets based on images with annotation information added. The learning DB device 4 includes a text listing unit 4a, which reads non-image information related to the images of the second teacher data with annotation information from the medical record DB device 5 and generates first teacher data by adding the non-image information to the second teacher data. For example, if the image of the second teacher data is an image of patient A, the text listing unit 4a reads information about patient A from the information stored in the medical record DB device 5. If patient A is, for example, male, the text listing unit 4a adds text information indicating maleness to the second teacher data to generate first teacher data. The first and second teacher data groups are created for each case.

[0059] The learning DB device 4 may include a text surplus / deficiency determination unit 4b. The text surplus / deficiency determination unit 4b determines whether or not the text information necessary for creating the first training data is insufficient. If the text information is insufficient, the text surplus / deficiency determination unit 4b notifies the control circuit 11 of the endoscope system 1 that the necessary text information is insufficient.

[0060] The text excess / deficiency determination unit 4b also determines whether there is too much text information for creating the first training data. If there is excessive text information, for example, if five years' worth of data is required but ten years' worth of data exists, it may be difficult to uniformly compile the excess text information during learning, or potentially irrelevant information may be used in learning. Furthermore, for example, if learning is intended to use data from the past three years, but information covering the entire period from birth to the present exists, learning may not be in line with the latest trends. Therefore, if there is excessive text information, the text excess / deficiency determination unit 4b notifies the control circuit 11 of the endoscope system 1 that the text information is excessive. Note that all text information may be used during annotation, assuming that selection will be made during learning.

[0061] The learning DB device 4 may also have a knowledge database (DB) 4c. The knowledge DB 4c stores knowledge information, which is information on medical findings. The learning DB device 4 may be configured to create first training data using the knowledge information recorded in the knowledge DB 4c.

[0062] The knowledge DB 4c is created in advance. For example, the knowledge DB 4c may be dynamically created by using information stored in electronic medical records and classifying similar cases using multivariate analysis or AI.

[0063] FIG. 5 is an explanatory diagram for explaining an example of the recorded contents of the knowledge database.

[0064] The knowledge DB 4c stores various types of knowledge information. In the example of FIG. 5, disease types are classified for each examination item, and case characteristics are recorded by age and gender. For example, case characteristics include information indicating the probability of occurrence in which part of the body. For example, if the knowledge DB 4c contains knowledge information indicating that, for example, men of a certain age have an extremely high incidence of disease A in a lower gastrointestinal endoscopy, the learning DB device 4 may create first training data by adding text information indicating that, for images obtained by a lower gastrointestinal endoscopy performed on men of a certain age, the incidence of disease A is extremely high as annotation information.

[0065] The inference model unit 15 of the endoscope system 1 has multiple inference models. First and second teacher data groups are provided to each inference model of the inference model unit 15 from the learning DB device 4. Note that, although the example of FIG. 4 illustrates a first inference model configured by the network N1 and a second inference model configured by the network N2, multiple inference models are provided in the inference model unit 15 for each case. That is, the inference model unit 15 includes multiple first inference models configured by the multiple networks N1 and multiple second inference models configured by the multiple networks N2, and each first teacher data group for each case is supplied to each network N1, and each second teacher data group for each case is supplied to each network N2.

[0066] A first inference model is constructed by providing first training data from a first training data group to network N1 and having network N2 learn from it, and a second inference model is constructed by providing second training data from a second training data group to network N2 and having network N2 learn from it. In this way, before the actual endoscopic examination, a second inference model is constructed based on the second training data obtained by image diagnosis of the image and the second training data, and a first inference model is constructed based on the first training data to which non-image information has been added.

[0067] The control circuit 11 includes a first information presentation unit 11a, a second information presentation unit 11b, a DB improvement unit 11c, and an input information unit 11d. During an endoscopic examination, the input information unit 11d provides endoscopic images received via the input IF 12 to networks N1 and N2 of the inference model unit 15. The input information unit 11d also provides non-image information, such as text information, about the patient undergoing the endoscopic examination to the network N1. The second inference model of the inference model unit 15 performs inference processing on the input endoscopic images and outputs the inference results (hereinafter referred to as the second inference results) to the control circuit 11. The first inference model of the inference model unit 15 also performs inference on the input endoscopic images and non-image information and outputs the inference results (hereinafter referred to as the first inference results) to the control circuit 11.

[0068] In this embodiment, the control circuit 11 has a function of presenting both the first inference result and the second inference result. That is, the first information presenter 11a outputs the first inference result based on the first inference model for display. The second information presenter 11b outputs the second inference result based on the second inference model for display. The first and second inference results from the first information presenter 11a and the second information presenter 11b are provided to the processor 30, which then displays the first and second inference results on the display screen of the display device 40. When two display devices are present, the processor 30 may provide the first and second inference results to separate display devices for display. In this manner, the endoscope system 1 constitutes a display control system that controls the display of the first and second inference results.

[0069] The DB improvement unit 11c of the control circuit 11 can also improve the database of the medical record DB device 5 based on the doctor's operation. The doctor checks the first and second inference results on the display device 40. The doctor can input observations (such as the presence or absence of a tumor) including their opinions on the first and second inference results as judgment results using the operation input unit 25. The operation input unit 25 can be configured with various input devices, such as a mouse, keyboard, tablet, GUI (Graphical User Interface), and microphone as a detector. For example, the doctor can determine whether the first or second inference result is correct. Alternatively, the doctor may determine that neither the first nor second inference result is correct. Furthermore, the doctor can input information that he or she considers to be correct regarding the area of ​​the lesion, symptoms, etc., as judgment results, regardless of the first and second inference results. The DB improvement unit 11c provides the endoscopic images obtained during the examination and the doctor's judgment results to the judgment result recording unit 16 (see FIG. 1) for recording.

[0070] The image and the judgment result recorded in the judgment result recording unit 16 are supplied to the medical record DB device 5 via the external IF 14 and recorded in the report database (DB) 5a of the medical record DB device 5. The DB improvement unit 11c instructs the text list unit 4a to re-learn as necessary, for example, when a doctor judges that the first and second inference results are incorrect. For example, the doctor can annotate the displayed image by adding a frame image indicating the area of ​​the lesion, or input text information corresponding to the patient.

[0071] When a re-learning instruction is given, the text listing unit 4a changes the annotated images used in the first training data or changes the text information by referencing the information in the report DB 5a. For example, it is possible to create new first training data based on annotations by a doctor. By re-learning the networks N1 and N2 using the changed first training data and second training data in this way, the inference accuracy of the first and second inference models is improved.

[0072] The endoscope system 1 may also include a text query unit 17, which is an interface circuit. When the doctor's voice is acquired by the microphone of the operation input unit 25, the text query unit 17 can capture this voice and convert the doctor's voice into text using, for example, known voice recognition technology, and provide the text to the learning DB device 4. The learning DB device 4 can create new first training data by adding the text information from the text query unit 17 to the annotated image. For example, it is possible to create the first training data based on the words uttered by the doctor during the endoscopic examination and then perform re-learning using the first training data created during the endoscopic examination, thereby improving the accuracy of the first inference model.

[0073] Although the above description has been given of an example in which the first and second inference models are constructed by learning using annotated first and second teacher data, it is also possible to adopt large language models (LLMs) and perform learning without annotation to obtain the first and second inference models. That is, in this embodiment, a display control device is obtained that has a control circuit that inputs a predetermined image and predetermined text information corresponding to this predetermined image into a first inference model and outputs the inference result of the first inference model for display.

[0074] Next, the operation of the embodiment configured as above will be described with reference to Figures 6 to 11. Figures 6 to 8 are flowcharts for explaining the operation of the embodiment. Figures 9 to 11 are explanatory diagrams showing display examples.

[0075] FIG. 6 shows the process flow up to implementing the inference model in the inference model unit 15. Data collection is performed in S1. As described above, for example, endoscopic images are collected, and annotations based on image diagnosis are applied in the learning DB device 4 to obtain second training data (S2). The learning DB device 4 also adds text information, etc. to the second training data using information from the medical record DB device 5 to obtain first training data. The first training data and the second training data are provided to the network, where machine learning is performed (S3). In this way, a first inference model based on the first training data and a second inference model based on the second training data are constructed by the network. In this case, the images and annotation information used to train the first and second inference models are common, thereby reducing development effort and shortening the development period. Of course, the images and annotation information used to train the first and second inference models may be different from each other.

[0076] It is also possible to construct the first inference model by learning using only pre-limited images, for example, images of Japanese men in their twenties.

[0077] In S4, model evaluation of the constructed inference model is performed. Machine learning is repeated until the model evaluation results are good, and a highly reliable inference model is constructed. In S5, the constructed first and second inference models are implemented in the inference model unit 15. In the example of Figure 4, learning is performed in networks N1 and N2 of the inference model unit 15, and the inference models are implemented as soon as learning is completed.

[0078] FIG. 7 shows the processing flow for an endoscopic examination. In S11, various information is input. For example, patient information about the patient undergoing the endoscopic examination, examination and procedure information, date and time information, doctor information, and equipment information are input. Next, the control circuit 11 determines whether screening is being performed (S12). Screening refers to the act of imaging a relatively wide area to search for a lesion, for example, a state of observation performed while the endoscope is being withdrawn. Furthermore, a mode in which a specific area is imaged in detail and examined, for example by imaging or zooming in on a specific area, is called an examination mode.

[0079] The inference model unit 15 of the inference device 10 may have an inference model for determining whether the mode is screening or examination. In this case, the inference device 10 can also determine whether the mode is screening or examination. For example, assume that the insertion section of the endoscope 20 is advancing within a lumen. In this case, the image portion of the inner part of the lumen (the inner part of the lumen in the longitudinal direction of the lumen), where illumination light from the endoscope tip does not reach, has a low-brightness lumen cross-sectional shape (often approximately circular). As the insertion section advances within the lumen, this image portion is located approximately at the center of the endoscopic image, and continuous images are obtained in which the lumen wall pattern moves toward the periphery of the image. Furthermore, if the endoscope tip is bent toward the target wall surface from a state in which the low-brightness lumen cross-sectional shape portion representing the inner part of the lumen is located at the center of the image, the low-brightness portion that was in the center moves to the periphery of the image, and this state can be determined by image analysis. Furthermore, when observing a specific patterned area for a long period of time by moving closer and further away or changing the viewing direction, similar patterns such as blood vessels and irregularities of lesions are captured in the continuous images, and the observation state can be determined based on changes in these patterns.

[0080] In the next step S13, it is determined which part of the body the input image depicts. The inference model unit 15 of the inference device 10 may be equipped with an inference model for determining the part of the body. In this case, the inference device 10 is also capable of determining the part of the body. For example, the inference device 10 determines each part of the human body from endoscopic images sequentially acquired by the endoscope during the insertion and removal process of the endoscope, i.e., from the start of insertion of the endoscope to the completion of removal.

[0081] In S14, the input information unit 11d provides the image acquired during the endoscopic examination to the second inference model of the network N2 corresponding to the region indicated by the image. The second inference result of the second inference model is supplied to the control circuit 11, and the second information presentation unit 11b provides the second inference result to the processor 30 for display on the display device 40 or the like (S15). Next, it is determined whether or not the inference model unit 15 has a first inference model (related text AI) obtained by learning first training data created using text related to the patient undergoing the endoscopic examination (S16). If not, the process proceeds to S19. If present, in S17, the image acquired during the endoscopic examination is provided to the first inference model of the network N1 corresponding to the region indicated by the image. The first inference result of the first inference model is supplied to the control circuit 11, and the first information presentation unit 11a provides the first inference result to the processor 30 for display on the display device 40 or the like (S18).

[0082] In this way, if related text AI exists, both the first inference result and the second inference result are displayed on the display device 40 or the like.

[0083] 9 and 10 are explanatory diagrams showing examples of display in this case, and Fig. 11 is an explanatory diagram showing an example of the display contents of the display areas R1L1 and R1L2 in Fig. 10.

[0084] 9 , a display region R1 is provided approximately in the center of the display screen M1 of the display device 40. For example, an endoscopic image during an examination is displayed in the display region R1. In the example of FIG. 9 , frame images RL1 and RL2 indicating the region of an object such as a lesion are displayed on the endoscopic image displayed in the display region R1. The frame image RL1 indicates a first inference result, and the frame image RL2 indicates a second inference result. These frame images RL1 and RL2 may be color-coded, for example, to clearly indicate whether they represent the first or second inference result.

[0085] Here, the frame images RL1 and RL2 are illustrated as ellipses, but they may also be rectangles or free-form curves. Furthermore, the first inference model and the second inference model may output the reliability of the inference results, and the display method of the frame images RL1 and RL2 may be changed depending on the reliability of the first inference result and the reliability of the second inference result. For example, possible methods include displaying the frame images when the reliability is above a certain level, or changing the color depending on the level of reliability.

[0086] Surrounding the display area R1 are a display area R2 for displaying patient information, a display area R3 for displaying examination information, and a display area R4 for displaying inference model identification. The display area R2 displays bibliographic information about the patient, such as the patient's name. The display area R3 displays various information about the examination, such as the date and time, the doctor's name, past information (such as medical records), and input findings. The display area R4 displays the results of inference using the first and second inference models in text. For example, the display area R4 displays a message such as, "According to the first inference model, there is a 70% probability that a tumor exists in the red-framed area (RI1)." The display area R4 may also display information indicating which inference model was used to obtain the inference, such as, "The first inference result was inferred using an inference model corresponding to a Japanese male in his twenties during a colonoscopy."

[0087] 10, a display area R1L1, which is a first information display area displaying a first inference result, and a display area R1L2, which is a second information display area displaying a second inference result, are arranged side by side at approximately the center on the display screen M1 of the display device 40. Fig. 11 shows an example of display in these display areas R1L1 and R1L2, in which the same endoscopic image acquired by an examination is displayed in each of the display areas R1L1 and R1L2. An image portion PL of a lesion area is present in this endoscopic image. A frame image PL1 indicating the lesion area is displayed in the display area R1L1 as the first inference result. Furthermore, a frame image PL2 indicating the lesion area is displayed in the display area R1L2 as the second inference result.

[0088] In Figure 10, display areas R2, R3, R4L1, and R4L2 are provided around display areas R1L1 and R1L2. The display areas R2 and R3 in Figure 10 display the same content as the display areas R2 and R3 in Figure 9. Furthermore, display areas R4L1 and R4L2 in Figure 10 display the content of display area R4 in Figure 9, divided into a first inference model and a second inference model. For example, display area R4L1 may display an explanation such as "First inference result based on a first inference model based on an image and text," and display area R4L2 may display an explanation such as "Second inference result based on a second inference model based on an image."

[0089] A doctor makes a diagnosis by referring to the lesion in the endoscopic image displayed on the display screen M1 of the display device 40 and the first and second inference results indicating the lesion area. As described above, a second inference model based on an image may produce a highly accurate second inference result. However, if the quality of the images used for training differs from that of the images acquired during the examination, the inference accuracy may be degraded. Therefore, by adopting a first inference model based on an image and text, a second inference result that compensates for the shortcomings of the second inference result may be obtained. However, the accuracy of the first inference result may be reduced due to inconsistencies in the training data caused by noise, fluctuations, the creator's subjectivity, etc. In this embodiment, the doctor can simultaneously view the first and second inference results, allowing for a more optimal diagnosis based on the doctor's judgment.

[0090] After checking the first and second inference results, the surgeon, such as a doctor, may perform input operations to create a report. In S19 of FIG. 7, a report is created based on the doctor's input operations. Furthermore, the inference model is improved as necessary. The doctor inputs which of the first inference result and the second inference result he or she has determined to be correct. Alternatively, the doctor may input that he or she has determined that neither the first nor the second inference result is correct. Furthermore, the doctor may input information that he or she considers to be correct regarding the area of ​​the lesion, symptoms, etc., as the judgment result. The DB improvement unit 11c provides the endoscopic image obtained by the examination and the doctor's judgment result to the judgment result recording unit 16 for recording.

[0091] The DB improvement unit 11c provides the images and judgment results recorded in the judgment result recording unit 16 to the medical record DB device 5 to update the report DB 5a. Specifically, the report DB 5a records the observation images acquired during the surgeon's observation, the observation findings attached by the surgeon to the observation images, which indicate the lesion area depicted on the observation images, and the patient information linked to the observation images. Depending on the content recorded in the report DB 5a, re-learning of the networks N1 and N2 is performed. The images and observation findings recorded in the report DB 5a are shared with the learning DB device 4 via the external network line 3 without including any information identifying individual patients. In this case, in addition to the images and observation findings, data such as the patient's gender, age, and nationality (race) are shared as related information.

[0092] FIG. 8 shows an example of a specific process for improving the inference model in S19 (S31) of FIG.

[0093] In S41, based on the doctor's input operation, the DB improvement unit 11c acquires observation images and finding information linked to patient information. The DB improvement unit 11c then links the observation images and finding information with the patient's gender, age, and nationality (race) from the patient information (S42). The information acquired by the DB improvement unit 11c is recorded in the judgment result recording unit 16 and then provided to and recorded in the report DB 5a. When a certain amount of new images and associated information is accumulated in the report DB 5a, the learning DB device 4, under the control of the control circuit 11, modifies the inference model of the inference device. That is, the learning DB device 4 performs additional annotation (re-annotation) on the first inference model using annotation information derived from the images and observation findings, as well as text information such as gender, age, and nationality (race), as new learning data. In addition, the learning DB device 4 performs additional annotation (re-annotation) on the second inference model using the observation image and annotation information derived from the observation findings, for example, area information indicating the area of ​​the lesion depicted in the observation image, as new learning data.

[0094] The first and second training data obtained by these re-annotations are respectively provided to the networks N1 and N2 of the inference model unit 15 for re-learning, thereby obtaining first and second inference models that reflect the doctor's input operations.

[0095] 7, the control circuit 11 determines whether the test has been completed. If the test has been completed, the control circuit 11 ends the process. If the test has not been completed, the control circuit 11 returns the process to S12.

[0096] If screening is not performed in S12 (NO in S12), that is, when the system transitions to examination mode, imaging suitable for examination, such as magnified observation, is performed (S21). The subsequent processes of S22 to S27 are the same as the processes of S13 to S18, respectively. In the examination mode, the difference from screening is that not only the site but also the affected area is determined, and first and second inference models for examination that are different from the first and second inference models for screening are used.

[0097] In the example of FIG. 7 , the presence or absence of additional text is determined in S28. For example, a doctor may specify additional text by operating the operation input unit 25. This additional text is supplied to the inference model unit 15 (S29). For example, assume that the first inference result of the first inference model used when a woman in her 30s is specified as the patient does not produce a satisfactory inference result. In this case, the doctor inputs additional text information, such as "thin build," to improve the inference accuracy. In this case, the inference model unit 15 supplies the endoscopic image acquired during the examination to a third inference model based on the text "woman in her 30s, thin build" (S29). In this case, the control circuit 11 supplies the third inference result obtained by the inference of the third inference model to the processor 30, which then displays the result on the display screen of the display device 40 (S30).

[0098] Although the control circuit 11 has been described as displaying the first and second inference results simultaneously, it may also be controlled to first display the second inference result corresponding to the input of image A to the second inference model on the display device 40. The doctor or other medical professional may refer to the second inference result and make a voice call for an inference that is considered more effective. The text query unit 17 then provides the text based on the doctor or other medical professional's voice input and image A to the first inference model. This allows the control circuit 11 to display the second inference result from the first inference model superimposed on the display image. In this case, both the second inference result displayed first and the first inference result displayed later may be displayed, or only the first inference result may be displayed.

[0099] Next, after the process of S31 is performed, the process proceeds to S20. Note that the process of S31 is the same as the process of S19.

[0100] (Example of Non-Image Information) FIGS. 12 and 13 are explanatory diagrams showing an example in which a plurality of pieces of non-image information are used for annotation.

[0101] The example of Figure 12 shows that a second inference model (image AI) is created using second training data created by adding annotation information to an image. Guide a is displayed based on the inference results of the created second inference model. Furthermore, a first inference model (image + interview AI) is created using annotations that add interview information such as patient gender and age differences as text information A to the second training data. Guide b is displayed based on the inference results of the created first inference model. Furthermore, another first inference model (image + (interview + past examination) AI) is created using annotations that add past examination information as text information B to the first training data. Guide c is displayed based on the inference results of the created another first inference model.

[0102] Fig. 13 shows an example of the display of guides a, b, and c. In Fig. 13, an endoscopic image acquired by an examination is displayed in a display area R on a display screen M1. An image portion PL of a lesion is present in this endoscopic image. Frame images PLa, PLb, and PLc indicating the region of the lesion are displayed as guides a, b, and c, respectively. A doctor can make an accurate diagnosis by referring to the endoscopic image and the frame images PLa, PLb, and PLc, which are guides a, b, and c.

[0103] As described above, in this embodiment, when creating training data, annotation is performed using non-image information, such as text, in addition to images with annotation information added. By learning the training data obtained in this manner, a highly accurate inference model can be obtained that comprehensively considers various factors other than endoscopic images. Furthermore, both the first inference result of the first inference model based on images and non-image information and the second inference result of the second inference model based on images are presented, enabling doctors to make more accurate diagnoses. Furthermore, re-learning using the doctor's diagnosis results is possible, allowing the construction of an inference model that enables even more accurate diagnoses.

[0104] The present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some of the components shown in the embodiments may be omitted. Furthermore, components from different embodiments may be appropriately combined.

[0105] Furthermore, among the technologies described herein, many of the controls and functions, mainly those described in the flowcharts, can be set by a program, and the above-described controls and functions can be realized by a computer reading and executing the program. The program can be recorded or stored, in whole or in part, as a computer program product on a portable medium such as a flexible disk, CD-ROM, or nonvolatile memory, or on a storage medium such as a hard disk or volatile memory, and can be distributed or provided at the time of product shipment, via a portable medium, or via a communication line. A user can easily realize the display control device, display control method, display control system, and display control program of this embodiment by downloading the program via a communication network and installing it on a computer, or by installing it on a computer from a recording medium.

Claims

1. A display control device having: a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching to each image annotation information corresponding to each image and text information that is different from the annotation information and related to the image; and a control circuit that outputs for display a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model.

2. A display control device as described in claim 1, further comprising a second inference model trained based on a second group of training data including a plurality of second training data obtained by attaching the corresponding annotation information to each of the images, wherein the control circuit outputs for display the second inference result output from the second inference model when a specified image is input to the second inference model and the first inference result.

3. The display control device described in claim 2, wherein the control circuit outputs for display the second inference result output when the specified image is input to the second inference model, and then, when the specified text information is input to the first inference model, outputs for display the first inference result output from the first inference model in a manner distinguishable from the second inference result.

4. A display control device as described in claim 2, wherein the image is an observation image of a patient taken at a medical facility, the annotation information is information indicating an area of ​​a lesion depicted in the image, and the text information is patient information related to the image, which is the observation image of the patient.

5. The display control device according to claim 2, wherein the patient information is information based on at least one of the patient's gender, age, and race.

6. The display control device described in claim 2, wherein the control circuit controls the first inference model to be re-annotated based on an observation image acquired during observation by the surgeon, area information indicating the area of ​​the lesion depicted in the observation image which is an observation finding attached by the surgeon to the observation image, and patient information linked to the observation image, and controls the second inference model to be re-annotated based on the observation image and the area information indicating the area of ​​the lesion depicted in the observation image.

7. A display control method comprising: a first teacher data creation step of creating first teacher data by attaching to an image annotation information corresponding to the image and text information different from the annotation information and related to the image; a first teacher data group generation step of repeating the first teacher data creation step to generate a first teacher data group including a plurality of the first teacher data; a first inference model generation step of generating an inference model trained based on the first teacher data group generated in the first teacher data group generation step; and a display step of displaying on a display device a first inference result output from the first inference model when a predetermined image and predetermined text information are input to the first inference model.

8. A display control system comprising: a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching to each image annotation information corresponding to each image and text information different from the annotation information and related to the image; an image generation device that provides a predetermined image to the first inference model; and a control circuit that outputs for display a first inference result output from the first inference model when the predetermined image and predetermined text information are input from the image generation device to the first inference model.

9. The display control system of claim 8, further comprising a second inference model trained based on a second group of training data including a plurality of second training data obtained by attaching the corresponding annotation information to each of the images, wherein the control circuit outputs for display the second inference result output from the second inference model when a specified image is input to the second inference model, and the first inference result.

10. The display control system of claim 8, wherein the control circuit controls the first inference model to be re-annotated based on an observation image acquired by the surgeon during observation, area information indicating the area of ​​the lesion depicted in the observation image, which is an observation finding attached by the surgeon to the observation image, and patient information linked to the observation image, and controls the second inference model to be re-annotated based on the observation image and the area information indicating the area of ​​the lesion depicted in the observation image.

11. A display control device having a control circuit that inputs a predetermined image and predetermined text information corresponding to the predetermined image into an inference model and outputs the inference results of the inference model for display.

12. A display control device having: a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching, to each image, annotation information corresponding to each image and text information different from the annotation information and related to the image; a second inference model trained based on a second group of training data including a plurality of second training data obtained by attaching, to each image, the annotation information; and a display control circuit that outputs, for distinguishable display, a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model, and a second inference result output from the second inference model when a specified image is input to the second inference model.

13. The display control device described in claim 12, wherein the display control circuit outputs for display the second inference result output when the specified image is input to the second inference model, and then, when the specified text information is input to the first inference model, outputs for display the first inference result output from the first inference model in a manner distinguishable from the second inference result.

14. A display control device comprising: first and second inference models that receive at least an image as input; a display control circuit capable of displaying a first inference result of the first inference model and a second inference result of the second inference model superimposed on a display image; a detector that detects audio after the display control circuit displays the second inference result for the first image superimposed on the display image; and an interface circuit that converts the audio detected by the detector into text and provides the converted text and the first image to the first inference model, wherein the display control circuit displays the inference result of the first inference model based on the converted text and the first image superimposed on the display image after displaying the second inference result.

15. A display control system comprising: a first inference model trained based on a first group of training data including a plurality of first training data obtained by attaching, to each image, annotation information corresponding to each image and text information different from the annotation information and related to the image; a second inference model trained based on a second group of training data including a plurality of second training data obtained by attaching, to each of the images, the annotation information; an image generation device that provides predetermined images to the first and second inference models; and a display control circuit that outputs, for display purposes, a first inference result output from the first inference model when a predetermined image and predetermined text information are input to the first inference model, and a second inference result output from the second inference model when a predetermined image is input to the second inference model, in a distinguishable manner.

16. A display control system comprising: first and second inference models that receive at least an image as input; an image generation device that provides predetermined images to the first and second inference models; a display control circuit that can display a first inference result of the first inference model and a second inference result of the second inference model superimposed on a display image; a detector that detects speech after the display control circuit displays the second inference result when the first image is provided to the second inference model as the predetermined image superimposed on the display image; and an interface circuit that converts the detected speech into text and provides the converted text and the first image to the first inference model; wherein the display control circuit displays the inference result of the first inference model based on the converted text and the first image superimposed on the display image after displaying the second inference result.

17. A display control method comprising: creating a first group of training data including a plurality of first training data obtained by attaching, to each image, annotation information corresponding to each image and text information different from the annotation information and related to the image; and creating a second group of training data including a plurality of second training data obtained by attaching, to each of the images, the annotation information; creating a first inference model trained based on the first group of training data and a second inference model trained based on the second group of training data; and outputting for display in a distinguishable manner a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model, and a second inference result output from the second inference model when a specified image is input to the second inference model.

18. A display control method comprising: inputting at least an image into a first and a second inference model; displaying the second inference result for the first image superimposed on the display image; detecting speech; converting the detected speech into text; providing the converted text and the first image to the first inference model; and, after displaying the second inference result, displaying the inference result of the first inference model based on the converted text and the first image superimposed on the display image.

19. A display control program for causing a computer to execute the following steps: create a first group of teacher data including a plurality of first teacher data obtained by attaching, to each image, annotation information corresponding to each image and text information related to the image that is different from the annotation information; create a second group of teacher data including a plurality of second teacher data obtained by attaching, to each image, the annotation information; create a first inference model trained based on the first group of teacher data; and create a second inference model trained based on the second group of teacher data; and output for display in a distinguishable manner a first inference result output from the first inference model when a specified image and specified text information are input to the first inference model, and a second inference result output from the second inference model when a specified image is input to the second inference model.

20. A display control program for causing a computer to execute the following steps: input at least an image into a first and a second inference model; after displaying the second inference result for the first image superimposed on the display image, detect audio; convert the detected audio into text, provide the converted text and the first image to the first inference model; and after displaying the second inference result, display the inference result of the first inference model based on the converted text and the first image superimposed on the display image.

Citation Information

Patent Citations

  • Ultrasonic diagnostic device and analysis device

    JP2021007512A

  • Image generation device and image generation method

    JP2022076940A

  • Diagnosis assistance device, ultrasound endoscope, diagnosis assistance method, and program

    WO2024004524A1