Estimation learning device and estimation learning method

CN115428011BActive Publication Date: 2026-09-29OLYMPUS CORPORATION(JP)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180003949.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-12
Publication Date
2026-09-29
Estimated Expiration
2041-03-12

AI Technical Summary

Technical Problem

在示教数据的生成时需要人力,导致耗费较大的成本

Benefits of technology

[0029]根据本发明,能够提供如下的估计用学习装置和估计用学习方法:不限于预先假定的类别的数据,在未知类别中,即便在数据的特性相对于之前蓄积的数据而改变的情况下,也能够进行适当的估计。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115428011B_ABST
    Figure CN115428011B_ABST
Patent Text Reader

Abstract

Provided are an estimation-use learning device and an estimation-use learning method that can make appropriate estimation in an unknown category without being limited to data of a category assumed in advance, even when the characteristics of data have changed with respect to data accumulated up to now. Image data from a first image acquisition device (S1a, S5a) is processed in accordance with the difference in image input characteristics when relearning an estimation model for a second image acquisition device that differs from the first image acquisition device, and is set as teaching data (S3a, S7a). The estimation model is learned using the teaching data obtained by annotating the image data, and thus an estimation model is obtained (S9).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an estimation learning apparatus and an estimation learning method that collects data from users and uses that data to generate an estimation model. Background Technology

[0002] In machine learning such as deep learning, teaching data is generated and then used for deep learning. Generating teaching data requires human effort, resulting in significant costs. Therefore, methods for collecting high-quality teaching data at low cost have been proposed. For example, in Patent Document 1, retrieval conditions are generated to collect data related to a specific domain from reference data related to that specific domain using a first feature vector. Then, data is collected using these retrieval conditions, and a second feature vector is calculated for the collected data. If the similarity between the first and second feature vectors is within a specified range, the data collected using the retrieval conditions is extracted as teaching data.

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2018-124617 Summary of the Invention

[0006] The problem the invention aims to solve

[0007] According to the data collection method described in Patent Document 1, teaching data can be collected at low cost. However, the data collection method in Patent Document 1 presupposes the collection of data in a specific field. On the other hand, the estimation model generated using the teaching data is not limited to data in a pre-assumed specific field (specific category), and its application scope is broadened to unknown categories (unknown fields), which sometimes require estimation.

[0008] Therefore, in the unknown category, if an estimation model is generated using data with characteristics different from previous data, it is also possible to estimate the data in that unknown category. However, generating an estimation model that corresponds to the unknown category requires collecting data that matches its characteristics, which is time-consuming and costly.

[0009] The present invention was made in view of the above circumstances, and its object is to provide an estimation learning device and estimation learning method that are not limited to data of a pre-assumed category, and can make appropriate estimations even when the characteristics of the data change relative to previously accumulated data in an unknown category.

[0010] means for solving problems

[0011] To achieve the above objectives, the estimation learning device of the first invention includes: an input unit that inputs image data from a first image acquisition device; and a learning unit that learns using teaching data obtained by annotating the image data to obtain an estimation model. The estimation learning device also includes an image processing unit that, when relearning the estimation model for a second image acquisition device with image input characteristics different from the first image acquisition device, processes the image data obtained from the first image acquisition device according to the differences in the image input characteristics and sets it as the teaching data.

[0012] The estimation learning device of the second invention is based on the first invention described above, wherein the image processing unit processes the first object image data contained in the image data obtained from the first image acquisition device in a manner suitable for the second object image data contained in the image data obtained from the second image acquisition device.

[0013] The learning device for estimation in the third invention is based on the first invention described above, wherein the image input characteristics are caused by at least one of the following differences: the specifications and performance of the camera sensor, the optical characteristics for imaging, the specifications and performance of the image processing, and the type of illumination light.

[0014] The fourth invention, based on the first invention described above, includes an image processing unit that modifies the annotation of the same image in such a way that the image data obtained from the first image acquisition device in the teaching data becomes teaching data corresponding to the differences in the image input characteristics.

[0015] The fifth invention's estimation learning device, based on the first invention described above, obtains existing teaching data from the first image acquisition device, and the image processing unit performs image processing on the existing teaching data according to the characteristics of the image data from the second image acquisition device.

[0016] The estimation learning device of the sixth invention is based on the first invention described above. The image data obtained from the first image acquisition device is existing teaching data. The image processing unit selects and discards the existing teaching data based on the characteristics of the image data from the second image acquisition device.

[0017] The estimation learning device of the seventh invention is based on the fifth invention described above, wherein the image processing unit processes the image data obtained from the first image acquisition device in the teaching data in a manner that adapts it to the image data from the second image acquisition device.

[0018] The estimation learning device of the eighth invention is based on the first invention described above, wherein the image data from the second image acquisition device belongs to an unknown category.

[0019] The estimation learning device of the 9th invention, based on the 8th invention described above, automatically determines whether it belongs to the unknown category through artificial intelligence, or the user of the 2nd image acquisition device manually sets whether it belongs to the unknown category.

[0020] The estimation learning device of the 10th invention, based on the 8th invention described above, determines whether an image belongs to the unknown category based on the model information of the second image acquisition device and / or an image estimated as a reference image from image data from the second image acquisition device.

[0021] The estimation learning device of the 11th invention is based on the first invention described above. The image data obtained from the first image acquisition device is existing teaching data. When the purpose of the estimation model is different, the image processing unit performs image processing on the existing teaching data or selects whether to use the existing teaching data according to the purpose.

[0022] The estimation learning device of the 12th invention is based on any one of the inventions 1 to 11 described above, wherein the image data from the first image acquisition device and the image data from the second image acquisition device are endoscopic image data.

[0023] In the estimation learning method of the 13th invention, image data from the first image acquisition device is input.

[0024] When learning the estimation model for a second image acquisition device with characteristics different from the first image acquisition device, the image data obtained from the first image acquisition device in the teaching data is processed and set as teaching data, and the estimation model is obtained by learning the teaching data obtained by annotating the image data.

[0025] The estimation learning device of the 14th invention includes: an input unit that inputs image data from a first image acquisition device; and a learning unit that learns using teaching data obtained by annotating the image data to obtain an estimation model, wherein the estimation learning device includes an image processing unit that, when customizing the estimation model for a second image acquisition device used under conditions different from the first image acquisition device, processes the image data obtained from the first image acquisition device by including selection or annotation corresponding to the differences in image acquisition characteristics and sets it as the teaching data.

[0026] In the estimation learning method of the 15th invention, image data from the first image acquisition device is input.

[0027] When customizing the estimation model for a second image acquisition device used under different conditions than the first image acquisition device, the image data obtained from the first image acquisition device is processed according to the differences in image acquisition characteristics, including selection or annotation, and set as the teaching data. The teaching data obtained by annotating the image data is used for learning, thereby obtaining the estimation model.

[0028] The effects of the invention

[0029] According to the present invention, an estimation learning device and an estimation learning method can be provided that are not limited to data of a pre-assumed category, and can make appropriate estimations even when the characteristics of the data change relative to previously accumulated data in an unknown category. Attached Figure Description

[0030] Figure 1 This is a block diagram illustrating the main electrical structure of a learning estimation device according to an embodiment of the present invention.

[0031] Figure 2 This figure shows an example of guided display using an estimation model in a learning estimation device according to an embodiment of the present invention.

[0032] Figure 3 This is a flowchart illustrating the operation of generating an estimation model in a learning estimation apparatus according to one embodiment of the present invention.

[0033] Figure 4 This is a flowchart illustrating the operation of a camera device in cooperation with an estimation learning device according to an embodiment of the present invention.

[0034] Figure 5 This is a flowchart illustrating the operation of generating a corrected estimation model in a learning estimation apparatus according to one embodiment of the present invention.

[0035] Figure 6 This figure illustrates a scenario where image data different from that previously existed is input into a learning estimation device according to one embodiment of the present invention.

[0036] Figure 7 This is a flowchart illustrating the action of determining whether AI correction is needed in a learning estimation device according to one embodiment of the present invention. Detailed Implementation

[0037] An estimation learning device according to one embodiment of the present invention collects image data and generates teaching data by annotating the image data. An estimation model is generated using a parent set composed of this teaching data. If the parent set of the teaching data used as the basis for generating the estimation data uses, for example, high-quality image data, then inputting low-quality image data into the estimation model may result in unreliable estimation. Furthermore, it is assumed that the estimation model for guided display is generated based on image data acquired by a highly skilled person (expert) using an image acquisition device. In this case, when a less skilled person (less proficient) uses the image acquisition device, even if they wish to obtain operational guided display based on the estimation model, they may not be able to perform appropriate guided display.

[0038] Thus, if the characteristics of the data used as the basis for generating the estimation model differ from the characteristics of the data input during actual estimation, a highly reliable estimation may not be possible. In such cases, it would be feasible to collect data with the same characteristics as the data input during actual estimation, similar to how the previously proven estimation model was generated, but this would be time-consuming and costly. Therefore, in this embodiment, when generating an estimation model that estimates data of unknown categories that differ from the previously proven estimation model due to differences in the equipment used or the user's skill level, the data is processed in a way that matches the characteristics of the previously accumulated teaching data with those of the unknown categories to generate the estimation model.

[0039] Here, regarding the classification of content, let's take the easiest-to-understand example. If it's an estimation model used to detect specific objects within an image by acquiring and estimating image data, even if similar images are obtained from image acquisition devices of different specifications, the image data acquired from these devices is treated as an unknown category due to differences in image quality, etc. Furthermore, the content reflected in the acquired images sometimes differs, and the objects to be identified also sometimes differ. Differences in the operator or robot performing the image acquisition, and the person or equipment handling the objects reflected in the image, lead to differences in how the image changes. Therefore, it can be said that the data becomes an unknown category different from the assumption.

[0040] The data processing described above includes image processing, annotation correction, and processes that affect data selection and the specifications of the estimation model. The specification is referred to as the estimation model specifications because, considering factors such as proficiency, the expected estimation results sometimes differ between experts and those other than experts. However, even in such cases, by utilizing the techniques in this embodiment, valuable teaching data used in generating previously proven estimation models can be easily retained.

[0041] Here, when the data is image data, the data processing includes image processing of the existing teaching data. Image processing includes various methods such as increasing or decreasing the number of pixels, changing brightness (luminance values), changing wavelength (color signals), and changing the field of view. Furthermore, data processing includes selecting teaching data from the existing teaching data that is included in the parent set. That is, unsuitable image data can be excluded, and new data can be selected and added from the image data. For example, teaching data may also include test data used to determine the completion status of the estimation model. Even if the test data is valid when generating the existing estimation model, it may sometimes be used in addition to the test data when generating an estimation model corresponding to an unknown category. This includes a selection process. Furthermore, if there are situations where certain operations or tool handling methods can be detected based on image information, teaching data obtained through skilled operators can be removed, and teaching data obtained through less skilled operators can be added.

[0042] Furthermore, when using surveillance cameras to collect images, the same camera images are sometimes used to identify criminals (where facial features are important), and sometimes for investigating congestion (where facial features are less important; in fact, from a personal information perspective, it is sometimes best not to know facial features). Thus, it can be seen that, depending on the intended use, even with the same images, the required quality, specifications, or processing of the teaching data can sometimes change depending on the estimation required. Similarly, even with medical images, the processing can differ in preventing the omission of lesions and in ensuring rigorous diagnosis. That is, even if the original estimation model can appropriately estimate the input data, when the purpose or object of the estimation model differs, the input data becomes data of an unknown category, requiring different processing of the teaching data or different learning methods.

[0043] Furthermore, the target audience of the estimation model can also be considered as an unknown category. For example, in medical devices used to diagnose cancer, the location or type of cancer varies depending on factors such as region, race, gender, and age. Moreover, regarding regional differences, it is expected that variations in hospital systems or medical equipment, physician skills, medical affiliations, or trends in patient data will lead to different assumed categories, sometimes falling under an unknown category. Therefore, the teaching data used to generate the estimation model can be modified appropriately based on the intended use of the model.

[0044] Hereinafter, with the aid of accompanying drawings, examples of the application of the present invention in an image estimation learning system are described as an embodiment of the present invention. Figure 1The image estimation learning system shown consists of an image estimation learning device 1 and a camera device 6.

[0045] The image estimation learning device 1 can be a stand-alone computer or similar device, or it can be configured within a server. When the image estimation learning device 1 is a stand-alone computer, it can be connected to the camera device 6 via a wired or wireless connection. Alternatively, when the image estimation learning device 1 is configured within a server, it can be connected to the camera device 6 via an information communication network such as the Internet.

[0046] Furthermore, the imaging device 6 can be a device installed on medical equipment such as an endoscope to photograph objects such as diseased areas, or it can be a device installed on scientific equipment such as a microscope to photograph objects such as cells, or it can be a device such as a digital camera whose primary purpose is image capture. Regardless of the type of device, in this embodiment, the imaging device 6 can be a device whose primary function is imaging, or it can be a device that also performs imaging functions to perform other primary functions. The following mainly describes the case where the imaging device 6 is an endoscope and the image acquisition device outputs endoscopic image data.

[0047] The camera device 6 includes an image estimation device 2, an image acquisition device 3, a guide unit 5, and a control unit 7. Additionally, regarding... Figure 1 The example shown is illustrated by demonstrating that the various devices of the camera device 6 are integrally constructed. However, it is also possible to construct them separately in different devices and connect them via information communication networks such as the Internet or dedicated communication networks. For example, the image estimation device 2 can also be separately constructed from the camera device 6 and connected via the Internet or the like. Furthermore, in Figure 1 The device includes an operation unit (input interface), a communication unit (communication circuit), a recording unit (for example, recording of image data acquired in the image acquisition device 3), an information acquisition device, and various components, circuits, and devices for enabling the camera device 6 to function, which are not shown in the figure.

[0048] The image acquisition unit 3 includes various imaging circuits such as an optical lens, an imaging element, an imaging control circuit, and an imaging signal processing circuit, acquiring and outputting image data of the object. Additionally, it may include components for exposure control during imaging (e.g., shutter, aperture), an exposure control circuit, and a lens drive device, focus detection circuit, and focus adjustment circuit for focusing the optical lens. Furthermore, the optical lens may be a zoom lens.

[0049] Image acquisition device 3a and image acquisition device 3b are disposed within image acquisition device 3. Figure 1The description includes both image acquisition device 3a and image acquisition device 3b, as they are largely functionally identical. Therefore, it is assumed that they can share subsequent components for ease of explanation. As described above, either device is mounted on the camera device 6. The image acquisition device 3 is used as either image acquisition device 3a (e.g., a reusable endoscope) or image acquisition device 3b (e.g., a disposable endoscope). That is, the image acquisition devices 3 are used separately depending on the region, facility, or case (object), and it is also assumed that the users are different. However, it is assumed that the same guidance functions can be effectively utilized, therefore, there is a possibility that the estimation model and other systems can be shared. Furthermore, it is assumed that the estimation model itself is customized, including the user interface or required guidance.

[0050] Furthermore, in this embodiment, it is explained that either image acquisition device 3a or image acquisition device 3b is disposed within the image acquisition device 3, but this does not preclude the image acquisition device 3 from having both image acquisition device 3a and image acquisition device 3b. This is because, depending on the situation, there may be cases where multiple devices are used. For example, image data from image acquisition device 3a belongs to a known category, while image data from image acquisition device 3b belongs to an unknown category. Moreover, the characteristics of the image data output from image acquisition device 3a and image acquisition device 3b are different. These characteristics include image quality, light source, field of view, etc. For example, if image acquisition device 3b has fewer pixels or lower resolution optical lenses compared to image acquisition device 3a, the image quality will be different. Furthermore, even if the user, object, or usage environment is different, the category of the data becomes unknown. In conjunction with this category, the estimation model itself is customized, including a user interface or required guidance.

[0051] Information used to distinguish the aforementioned differences can be obtained, for example, by comparing information on the model of each image processing device (in addition, information on peripheral systems such as light sources or processing instruments, as described later, can also be included, and additional information can be used) with databases. This information can also be obtained by sending data recorded in the device's built-in memory or in systems within the device's operating environment to the image estimation learning device 1, or by using information manually input by the user. Alternatively, model information can be omitted, and symbols or values ​​representing the detection and processing performance of the image processing device can be used. Furthermore, information (data) about the operating environment and objects such as patients can also be obtained and determined through communication with each device, or by obtaining and using manually input information through communication. Teaching data can be selected and processed based on the differences in these acquired supplementary data. Moreover, the desired estimation model changes depending on the tools or devices used, or the skill or performance of the person or robot handling them, and constraints. Therefore, this information can also be obtained from information recorded in memory, manual input, or sensor information. In addition to obtaining memory information, if the image data itself or the state (scene) of the moving image is analyzed, it is also possible to determine the unknown category that cannot be handled in the assumed estimation model.

[0052] Furthermore, the imaging device 6 has a light source, and when photographing an object illuminated by this light source, the resulting image varies depending on the wavelength or light distribution characteristics of the light source. Alternatively, either image acquisition device 3a or 3b may be capable of observation using narrow band imaging (NBI), in which case the characteristics of image acquisition device 3a and image acquisition device 3b differ.

[0053] Furthermore, the field of view varies depending on the focal length of the optical system of the image acquisition device 3. With a telephoto lens, an image can be obtained that is magnified despite having a narrow angle. Conversely, with a short focal length lens, an image can be obtained that is reduced in size despite having a wide angle. When the optical system is a zoom lens, the image varies significantly depending on the set focal length.

[0054] Furthermore, the image acquisition device 3a may also have a distance (distribution) detection function 3D (3Daa). If it has this 3D capability, the image acquisition device 3a differs from the image acquisition device 3b in this respect. 3D and similar devices (3aa) capture images of objects in three dimensions to obtain three-dimensional image data, but in addition to three-dimensional images, they can also acquire depth information by capturing reflected light or ultrasound waves. The three-dimensional image data can be used to detect the object's position in space, such as its distance from the camera device 3. For example, if the camera device 6 is an endoscope, when a doctor inserts the endoscope into the body and performs operations, if the camera unit is 3D, it can determine the positional relationship between the body part and the surgical instrument, as well as the three-dimensional shape of the body part, enabling three-dimensional display. Furthermore, even without strictly acquiring depth information, depth information can be calculated based on the size relationship between the background and nearby objects.

[0055] Image data and other data acquired in the image acquisition device 3 (image acquisition device 3a or image acquisition device 3b), and data that becomes a candidate group of teaching data, are output to the recording unit 4 in the image estimation learning device 1 and recorded as teaching data group A 4a. In this case, a memory may also be provided in the camera device 6 to store the image data acquired in the image acquisition device 3.

[0056] Additionally, an information acquisition device can also be configured within the image acquisition device 3. This information acquisition device is not limited to acquiring image data; it can also acquire information related to the object, such as patient-related information obtained from an electronic medical record or information related to the equipment used for diagnosis or treatment. For example, when a doctor performs a procedure using an endoscope, the information acquisition device acquires information such as the patient's name, gender, and the location within the body where the endoscope is inserted. Furthermore, in addition to acquiring information from the electronic medical record, the information acquisition device can also acquire audio data from the diagnosis or treatment, as well as medically relevant data such as body temperature, blood pressure, and heart rate data. This data can also be output to the image estimation learning device 1.

[0057] When assessing risks, the aforementioned factors can be incorporated to improve reliability. In this embodiment, an example using images is primarily described, but estimations can also be performed based on the numerical data described above. Furthermore, the need to customize teaching data collected under specific conditions for different environments is similar to the situation with image-based estimations. Therefore, it can be seen that the method in this embodiment is not limited to images and is an effective solution for general data. Just as the status quo based on dynamic images is important, the time-varying nature of this data can also be considered using the same approach as the method of this embodiment. In the following embodiments, estimation models considering such time-varying nature will be described as examples. Estimations using static images or individual data are simpler and are therefore not illustrated, but should be generally understood from the above description.

[0058] Image estimation device 2 takes image data acquired by image acquisition device 3 as input, performs estimation using estimation model generated by image estimation learning device 1, and outputs guidance display to guidance unit 5 based on the estimation result. Image estimation device 2 has image input unit 2IN, estimation modification unit 2SL, estimation unit 2AI, and estimation result output unit 2OUT. In addition, the term "estimation model" sometimes includes what kind of guidance (display or sound) is output to the user.

[0059] The image input unit 2IN receives image data output from the image acquisition device 3. This data is a time-series data consisting of multiple frames, which is continuously input to the image input unit 2IN. Furthermore, non-image information such as sound or data obtained from other sensors can be referenced as needed. Moreover, it is not limited to image input; it can also be a data input unit. Furthermore, the image input to the input unit can be each frame of a continuously acquired image, or multiple frames can be processed together. Such learning can be performed assuming an estimation engine that estimates based on multiple frames.

[0060] The estimation unit 2AI has an estimation engine, and the estimation model generated by the image estimation learning device 1 is set in the estimation engine. The estimation engine, like the learning unit 1c described later, has a neural network in which the estimation model is set. The estimation unit 2AI inputs the image data input from the image input unit 2IN to the input layer of the estimation engine, and performs estimation in the intermediate layer of the estimation engine. The estimation result is output to the guidance unit 5 by the estimation result output unit 2OUT.

[0061] The estimation modification unit 2SL modifies the estimation model used in the estimation unit 2AI. As mentioned above, the characteristics of the image acquisition devices 3A and 3B are different. When an estimation model generated based on data from the image acquisition unit 3A is set in the estimation unit 2AI, for example, in other environments, even if data from the image acquisition device 3B is input into the estimation unit 2AI expecting the function of the aforementioned estimation model, proper estimation may not be possible and guided display may not be possible. When data with different characteristics is input into the image input unit 2IN in this way, the control unit 7 entrusts the generation of a corrected estimation model suitable for the data output by the image acquisition device 3B. The estimation modification unit 2SL modifies the estimation model in the estimation unit 2AI to this corrected estimation model.

[0062] In other words, the corrected estimation model described above is learned through corrected teaching data. That is, even if the image data and other data are the same, when generating teaching data based on that data, by changing the processing method (processing and correction of the teaching data) or the selection method, it is possible to generate estimation models (corrected estimation models) of different specifications and performance.

[0063] The guidance unit 5 includes a display screen, etc., which displays an image of the object acquired by the image acquisition device 3. Furthermore, it performs guidance display based on the estimation results output by the estimation result output unit 2OUT.

[0064] The control unit 7 is a processor that includes a CPU (Central Processing Unit) 7a, a memory 7b, and peripheral circuitry. The control unit 7 controls the various devices or parts within the camera device 6 according to the program stored in the memory 7a.

[0065] The image estimation learning device 1 uses image data acquired by the image acquisition device 3 to perform machine learning (including deep learning) and generate an estimation model. The image estimation learning device 1 includes an image input unit 1b, a learning unit 1c, an image processing unit 1d, a learning result utilization unit 1e, a teaching data selection unit 1f, and a recording unit 4.

[0066] The recording unit 4 is a non-volatile memory capable of electrical rewriting, used to record image data or various information data output from the image acquisition device 3 within the imaging device 6. The various data recorded by the recording unit 4 are output to the image input unit 1b. The recording unit 4 can store teaching data group A 4a and teaching data group B 4b. Furthermore, test data such as verifying the strength of the estimation model can also be recorded in the recording unit 4. Even without pre-recording the test data itself, a portion of the teaching data recorded by the recording unit 4 can be extracted and used as test data.

[0067] Teaching data group A 4a is a teaching data group based on time series data acquired by image acquisition device 3a. As described later, teaching data group B 4b is teaching data generated by processing the already recorded teaching data group A 4a when generating an estimation model for an unknown category. As the recording unit 4, different teaching data groups are recorded depending on the characteristics of the image acquisition device. The recording unit 4 records both the candidate group of teaching data sent from image acquisition device 3 and the teaching data group with annotations described later. Furthermore, teaching data that is not used may also be used when generating teaching data for an unknown category, and may also be recorded in the recording unit 4, not limited to the teaching data adopted by the teaching data selection unit 1f.

[0068] Control unit 1a processes the time-series data acquired in image acquisition device 3 (described later) Figure 4 In step S35, annotations are applied to the image estimation (sent by the learning device 1), thereby generating a set of teaching data. For example, as described later... Figure 2 The image shows bleeding during the insertion of an endoscope into the body. Figure 2 In (a), the bleeding expands; on the other hand, in Figure 2 In (b), the bleeding decreases. In this case, teaching data can be generated by annotating how the bleeding changes in time series data ID1 and ID2. This annotation can be done automatically or manually as needed. Furthermore, during customization, it can also be done automatically, taking into account or reflecting the results of manual annotation. Even when annotation is done automatically, it can be checked manually, and based on the results, steps can be added to prompt repeating this process.

[0069] When assigning annotations, this is sometimes done by changing the "processing and correction of teaching data" or the "selection and omission of teaching data." This is because, for example, even with... Figure 2 Similar to the bleeding shown, the degree of remediation after bleeding varies depending on the availability of tools, personnel, or skills to immediately address the bleeding (making such information available). That is, even if the teaching data obtained from a system with comprehensive capabilities, including skills or tools, is judged and annotated as "no bleeding," it is best to assign annotations with more stringent judgments when generating estimated guidance for systems with weaknesses in skills or tools.

[0070] Information such as skill level classification can also be included. It can be used as information for customization based on manually inputted skill information, pre-registered records, and historical data. It can also be based on trends in acquired images. For example, in photos taken by professional photographers and those taken by beginners, besides differences in equipment, there are differences in composition, exposure, and focus; skill assessment can be based on these differences. When the image is moving, this trend becomes even stronger, as the way equipment is used is reflected in the image. Methods that simultaneously capture sound and use it as a reference also exist. Furthermore, it's possible to determine the equipment used based on image distortion or blur.

[0071] That is, the estimation learning device in this embodiment has: an input unit that inputs image data from a first image acquisition device; and a learning unit that learns based on teaching data obtained by annotating the image data to obtain an estimation model. The estimation learning device includes an image processing unit that, when performing customized learning (or re-customized learning) on ​​the estimation model for a second image acquisition device with image input characteristics different from the first image acquisition device, processes the image data obtained from the first image acquisition device in the teaching data into teaching data corresponding to the differences in the image acquisition characteristics (including changes in annotations) and sets it as teaching data.

[0072] For example, the outcomes of post-bleeding care differ between scalpels with and without hemostasis capabilities; therefore, it's best to incorporate this difference into the estimation model. Differences in instrument specifications or performance can be determined based on pre-input information or the characteristics of the instrument's image in the captured images. For instance, it's ideal to distinguish between images of procedures using instruments without hemostasis and those using instruments with hemostasis capabilities. By learning from these differences, one image can be used as the other, generating teaching data through processing and correction. That is, by processing and correcting images with hemostasis capabilities, a guiding estimation model for procedures without hemostasis can be generated. Since surgeries vary greatly depending on individual constitution and affected area, collecting ideal teaching data is not always easy; therefore, this method facilitates the creation of highly reliable estimation models.

[0073] The image input unit 1b inputs teaching data group A 4a, which is acquired by the image acquisition device 3a and recorded in the recording unit 4. The teaching data group 4a input to the image input unit 1b is annotated. The input teaching data group A 4a is output to the learning unit 1c and the image processing unit 1d. During learning, data other than image data can be used, not limited to image data. Furthermore, when the learning device relearns and generates a corrected estimation model, the teaching data group B 4b obtained by processing the teaching data group A 4a is input to the image input unit 1b. The image input unit 1b functions as an input unit (input interface) for inputting image data from the first image acquisition device (see, for example, reference). Figure 3 S1, S5, Figure 5 (S1a, S5a).

[0074] The image processing unit 1d processes the teaching data input from the image input unit 1b using an image processing circuit or program. As described above, the characteristics of the image acquisition devices 3a and 3b are different. Therefore, even if the learning unit 1c generates an estimation model based on the image data acquired by the image acquisition device 3a, it is impossible to make an appropriate estimation, even if an estimation model is generated using the image data acquired by the image acquisition device 3b. Therefore, the image processing unit 1d performs image processing on the image data input to the image input unit 1b, converting it in the same way as the image data acquired by the image acquisition device 3A. The image data processed by the image processing unit 1d is output to the learning unit 1c. Furthermore, it is then used... Figure 6 Describe the detailed processing of the image, and its subsequent use. Figure 5 Describe the generation of the corrected estimation model.

[0075] The image processing unit 1d functions as an image processing unit (image processing processor) that, when relearning the estimation model for a second image acquisition device with characteristics different from the first image acquisition device, processes the image data obtained from the first image acquisition device in the teaching data and sets it as the teaching data (e.g., reference data). Figure 5 S1a~S7a, Figure 6 (b) The above-mentioned image input characteristics are caused by at least one of the following differences: camera sensor specifications, performance, camera optical characteristics, image processing specifications, performance, and type of illumination light.

[0076] The image processing unit processes the first object image data contained in the image data obtained from the first image acquisition device in the teaching data in a manner suitable for the second object image data contained in the image data obtained from the second image acquisition device (e.g., referring to...). Figure 5 S1a~S7a, Figure 6 (b)). The image processing unit includes changing the annotation of the same image in a manner that makes the image data obtained from the first image acquisition device in the teaching data corresponding to the differences in image input characteristics.

[0077] Furthermore, image input devices are used in conjunction with certain operations, and sometimes the content of the image or the obtained image changes depending on changes in the environment, the object, or even the tools used simultaneously. In this case, it is assumed that the image input characteristics have changed, and "processing" is performed corresponding to the difference in image input characteristics. This "processing" corresponding to the difference in image input characteristics includes not only the type of image processing or correction method, but also corrections such as the content of annotations as part of the teaching data or the estimation start trigger timing related to the specification method of the learning results using the processed teaching data. This is because it is also assumed that "processing" is performed that corresponds to changes in the usage environment and conditions of the image acquisition device of the second specification that actually utilizes the estimation model (including not only the device's performance, specifications, environment, and peripheral systems, but also the object being processed here, accessories and other peripheral equipment, handling tools, operators, etc.).

[0078] Here, the image data obtained from the first image acquisition device is existing teaching data. That is, the teaching data adopted by the teaching data selection unit 1f is stored in the recording unit 4. The image processing unit performs image processing on the existing teaching data based on the characteristics of the image data from the second image acquisition device (i.e., the characteristics are different from those of the first image acquisition device). Figure 5 (S1a, S5a). Furthermore, the image processing unit selects from existing teaching data based on the characteristics of the image data from the second image acquisition device (see S1a, S5a). Figure 5 (S13). The image processing unit processes the image data obtained from the first image acquisition device in the teaching data in a manner that adapts it to the image data from the second image acquisition device (e.g., referring to...). Figure 5 (S1a, S5a, S13).

[0079] Furthermore, depending on the intended use of the estimation model, the image processing unit can perform image processing on existing teaching data or select from existing teaching data based on that intended use. Alternatively, when customizing the estimation model for a second image acquisition device used under conditions different from the first image acquisition device, the image processing unit can process the image data obtained from the first image acquisition device, including selections or annotations corresponding to the differences in image acquisition characteristics, and set it as teaching data. For example, if an estimation model is generated using teaching data based on image data acquired by a skilled user operating the camera device 6, even if an unskilled user uses the estimation model for operation guidance, it may sometimes be difficult to provide proper guidance. In such cases, inappropriate image data can be excluded by the image processing unit 1d or the teaching data selection unit 1f, and new image data can be selected and added, or the image can be appropriately corrected. Since they can also cooperate, communication can also occur between the image processing unit 1d and the teaching data selection unit 1f. Furthermore, the image processing unit can also process and select image data or teaching data considering the intended use of the estimation model. For example, as an estimation model for diagnosing cancer, factors such as race, gender, and age can be considered when selecting and processing image data.

[0080] Furthermore, when generating an estimation model through learning, the learning unit 1c determines the reliability of the learning result. This determination can be made by preparing test data and judging whether the output of the estimation model when the test data is input is a previously known correct solution (e.g., referring to...). Figure 3 and Figure 5 (S11). Test data used for judgment can be selected from the recording unit 4. Alternatively, when acquired from outside the learning device, appropriate test data can be selected by the teaching data selection unit 1f and processed by the image processing unit 1d as needed. When selecting test data that can be appropriately processed, the image processing unit 1d and the teaching data selection unit 1f should preferably cooperate. Furthermore, since the test data verifies the actual performance of the actual equipment used, it is best to select images actually acquired on the actual equipment. When the teaching data selection unit 1f selects a suitable image and the image processing unit 1d processes it, the appropriate image selected by the teaching data selection unit 1f can also be processed.

[0081] Like the estimation unit 2AI, the learning unit 1c also possesses an estimation engine to generate an estimation model. The learning unit 1c uses image data input from the image input unit 1b or image data processed by the image processing unit 1d to generate an estimation model through machine learning such as deep learning. Deep learning will be described later. The learning unit 1c functions as a learning unit (learning engine) that learns from teaching data obtained by annotating image data, thereby obtaining an estimation model (e.g., referring to...). Figure 3 and Figure 5 (S9).

[0082] The teaching data selection unit 1f determines the reliability of the estimation model generated in the learning unit 1c, and decides whether to use it as teaching data based on the determination result. That is, if the reliability is low, it is not used as teaching data when generating the estimation model, and only teaching data with high reliability is used. The learning unit 1c finally generates the estimation model based on the teaching data adopted by the teaching data selection unit 1f. In addition, the teaching data adopted by the teaching data selection unit 1f is pre-recorded as teaching data group A 4a in the recording unit 4. Depending on the situation, the teaching data selection unit 1f may also have a memory in which the adopted teaching data is pre-recorded.

[0083] The estimated model generated in the learning unit 1c is output to the learning result utilization unit 1e. The learning result utilization unit 1e sends the generated estimated model to the estimation engine of the image estimation unit 2AI, etc.

[0084] Here, we will explain deep learning. "Deep learning" refers to the multi-layered structure of the "machine learning" process that uses neural networks. A representative example is the "forward propagation neural network," which processes information by feeding it from front to back for judgment. The simplest forward propagation neural network consists of three layers: an input layer with N1 neurons, an intermediate layer with N2 neurons (specified by parameters), and an output layer with N3 neurons corresponding to the number of classes to be classified. The neurons in the input layer and the intermediate layer, and the intermediate layer and the output layer, are connected by connection weights. Bias values ​​are applied to the intermediate and output layers, thus easily forming logic gates.

[0085] Neural networks can have as few as three layers for simple discrimination, but by using multiple intermediate layers, they can learn combinations of multiple features during machine learning. In recent years, from the perspective of learning time, judgment accuracy, and energy consumption, 9 to 152 layers have become practical. Furthermore, "convolutional neural networks," which perform a process called "convolution" to compress image features and perform action pattern recognition with minimal processing, can also be used. Additionally, "recurrent neural networks" (fully coupled recurrent neural networks) can be used to process more complex information and allow bidirectional information flow corresponding to information analysis where meaning changes according to order or sequence.

[0086] To implement these technologies, conventional computing circuits such as CPUs or FPGAs (Field Programmable Gate Arrays) can be used. However, they are not limited to these. Since neural network processing mostly involves matrix multiplication, processors specifically designed for matrix computation, known as GPUs (Graphics Processing Units) or Tensor Processing Units (TPUs), can also be utilized. In recent years, such artificial intelligence (AI) dedicated hardware, the "Neural Network Processing Unit (NPU)," has sometimes been designed to be integrated and assembled with other circuits such as CPUs as part of the processing circuitry.

[0087] Furthermore, machine learning methods include, for example, support vector machines and support vector regression. Here, learning involves calculating the weights, filtering coefficients, and biases of the recognizer; alternatively, methods utilizing logistic regression are also available. When enabling a machine to make certain decisions, humans need to teach the machine how to make those decisions. In this embodiment, a method for deriving image decisions through machine learning is used; however, if the method involves deriving annotation results from teaching data, a rule-based method adapting to rules obtained by humans through experience and heuristics can also be used.

[0088] The control unit 1a is a processor that includes a CPU (Central Processing Unit) 1aa, a memory 1ab, and peripheral circuitry. The control unit 1a controls the various units within the image estimation learning device 1 according to the program stored in the memory 1ab. For example, the control unit 1a assigns annotations (see reference) to image data output from the image acquisition device 3. Figure 3 S3, S7, Figure 5 (S3a, S7a).

[0089] Next, use Figure 2(a) and (b) illustrate the use of an endoscope for procedures as examples of image acquisition and guided display based on those images. The endoscope has... Figure 1 The camera device 6 shown here therefore includes an image acquisition device 3, an image estimation device 2, and a guide unit 5.

[0090] Figure 2 (a) illustrates an example where, during endoscopic procedures, bleed (BL) develops internally and expands to become enlarged hemorrhagic bleed (BLL). The endoscopic image acquisition device 3 continuously collects image data at predetermined time intervals during the physician's procedures, and the control unit 1a records this image data as a teaching data candidate group in the memory of the imaging device 6. Figure 2 In example (a), bleeding occurs at time T=0, and at time T=T1a, it can be identified as bleeding enlargement. In this case, image data ID1 starting from the moment 5 seconds prior to time T=0 is recorded as the image at the time of bleeding enlargement. If the image estimation learning device 1 assigns an annotation to the collected image data ID1 indicating that the bleeding has enlarged after time T=T1a, it becomes teaching data at the time of bleeding enlargement. In this embodiment, the annotation is performed in the image estimation learning device 1 (see...). Figure 3 S3, Figure 6 (S3a) But annotation can also be performed in the camera device 6, and the annotated teaching data can be sent to the image estimation learning device 1.

[0091] Figure 2 (b) illustrates an example where, during endoscopic procedures, bleeding occurred internally but subsequently decreased in size. The endoscopic image acquisition device 3 and... Figure 2 Similarly, in example (a), image data is collected at predetermined time intervals during the processing, and the control unit 1a records this image data as a candidate group of teaching data in the memory of the camera device 6. Figure 2 In example (b), bleeding occurs at time T=0, and at time T=T1b, it can be identified as bleeding shrinkage. In this case, image data ID2 from the moment 5 seconds prior to T=0 is also collected as the image at the time of bleeding shrinkage. If the image estimation learning device 1 assigns an annotation to the collected image data ID2 indicating that the bleeding has shrunk after time T=T1b, it becomes teaching data for bleeding shrinkage. In this embodiment, the annotation is performed in the image estimation learning device 1 (see...). Figure 3 S7 Figure 6While S7a), annotation can also be performed in the camera device 6, and the annotated teaching data can be sent to the image estimation learning device 1. By analyzing images acquired in a continuous time sequence (moving images) in this way, various useful information can be obtained.

[0092] exist Figure 2 In (a) and (b), time T=0 is the time at which the user notices the bleeding, but the behavior or phenomenon that causes the bleeding mostly occurs at a time before time T=0. Therefore, in this embodiment, when an event occurs (e.g., the bleeding expands, the bleeding shrinks, etc.), trigger information is generated, and time is traced back from that specific time point to collect data and organize causal relationships. By analyzing images acquired in a continuous time sequence (moving images) in this way, various effective information can be obtained.

[0093] By collecting a large amount Figure 2 Examples like (a) and (b), with annotations, can generate a large amount of teaching data, which can be processed as big data. Learning Department 1c uses this large amount of teaching data to generate an estimation model. This estimation model can estimate the impact of bleeding at time T=0, after a specified time (in... Figure 2 In examples (a) and (b), is the bleeding in T1a or T1b enlarged or reduced?

[0094] If such an estimation model is generated and the estimation unit 2AI of the imaging device 6 is set to this estimation model, then the future can be predicted based on the image acquired by the image acquisition unit 3. That is, if the imaging device 6 detects bleeding at a time T=0 that is not yet at time T=1, then... Figure 2 As shown in (a) and (b), teaching data (or teaching data candidates) obtained based on image data from the timing point up to a point retraceable to a predetermined time (T = -5 sec) is input into the estimation model, thereby enabling prediction of whether the bleed will expand or shrink. If the prediction (estimation) indicates that the bleed is expanding, a notice display Ga is shown in the guidance section 5 of the imaging device 6. Conversely, if the prediction (estimation) indicates that the bleed is shrinking, a guidance Go is displayed indicating that the bleed is not a problem.

[0095] Next, use Figure 3 The flowchart shown illustrates the process in Figure 2 The generation of the estimation model used in (a) and (b). The image estimation process is implemented by the CPU 1aa of the control unit 1a in the learning device 1 according to the program stored in the memory 1ab.

[0096] when Figure 3The process of generating the estimation model shown begins by first collecting images of the bleeding process (S1). As described above, the imaging device 6 collects images from the continuous images acquired by the image acquisition device 3. Figure 2 Image (a) shows the increase in the area of ​​the hemorrhage during the period from time T = -5 to T = T1a. Specifically, in the above... Figure 2 In (a), the control unit 7 performs image analysis of the image data, and if it determines that the bleeding is expanding, it generates trigger information (see reference). Figure 4 S27), traced and recorded images of the expanded hemorrhage (refer to S27). Figure 4 (S29). The image of the trace record is temporarily recorded in the memory of the camera device 6. In this step S1, the control unit 1a of the image estimation learning device 1 collects process images of the bleeding expansion from the camera device 6, etc., and temporarily stores them in the recording unit 4.

[0097] After collecting process images of bleeding expansion in step S1, the image data is annotated with "bleed expansion" (S3). Here, the control unit 1a applies the annotation "bleed expansion" to each collected image data, and records the annotated image data as teaching data A group 4a in the recording unit 4.

[0098] Next, images of the bleeding process as it shrinks are collected (S5). As described above, the imaging device 6 collects images from the continuous images acquired by the image acquisition device 3. Figure 2 Image (b) shows the reduction in the area of ​​the hemorrhage during the period from time T = -5 to T = T1b. Specifically, in the above... Figure 2 In (b), the control unit 7 analyzes the image data and, if it determines that the bleeding has shrunk, generates a trigger (see reference). Figure 4 S27), traced back to the reduced image of the bleeding (refer to S27). Figure 4 (S29). The image of the trace record is temporarily recorded in the memory of the camera device 6. In this step S5, the control unit 1a of the image estimation learning device 1 collects process images of the bleeding reduction from the camera device 6, etc., and temporarily stores them in the recording unit 4.

[0099] After collecting process images of bleeding reduction in step S5, the image data is annotated with "bleed reduction" (S7). Here, the control unit 1a applies the annotation "bleed reduction" to each collected image data, and records the annotated image data as teaching data A group 4a in the recording unit 4.

[0100] exist Figure 3In the illustrated process, an image of the reduced bleeding is collected after the bleeding has expanded. However, in practice, steps S1 to S7 are appropriately selected and executed based on whether bleeding has occurred in the image collected by the image acquisition device 3, and whether the extent of bleeding has expanded or decreased if it has occurred.

[0101] Next, an estimation model is generated (S9). Here, the teaching data with annotations generated by the camera device 6 applied in steps S3 and S7 is recorded as teaching data group A 4a, and this teaching data is input to the image input unit 1b. The learning unit 1c in the image estimation learning device 1 uses the teaching data to generate an estimation model. This estimation model can make a prediction such as "bleeding will expand after 0 seconds" when an image is input.

[0102] After generating the estimation model, the reliability is determined (S11). Here, the learning unit 1c inputs the image data used to confirm the reliability of the answer, which is known in advance, into the estimation model, and determines the reliability based on whether the output in this case is the same as the answer. If the reliability of the generated estimation model is low, the proportion of answers that match is low.

[0103] In estimating the prediction of such treatments, it is desirable to reflect the skills of the doctor performing the treatment, the differences in treatment equipment, etc. However, in most cases, image data intended to be collected when generating the estimation model, as literally set as teaching data, is readily available for collection, showing the treatment process of a skilled doctor using high-quality equipment. However, for treatments performed by inexperienced individuals using unconventional equipment, it is important to provide guidance; therefore, it is desirable to be able to handle unconventional cases. Furthermore, the introduction of entirely new treatment equipment may also be an unconventional situation, and with such equipment, there are often many inexperienced users initially. Moreover, the degree of inexperience varies greatly, and in many cases, this may also be an unconventional situation. In short, it is desirable to provide highly reliable guidance for inexperienced individuals using unseen equipment, and the estimation learning system in this embodiment can handle such situations.

[0104] In this way, a highly reliable estimation model is generated through learning. Therefore, an estimation learning device is primarily provided, which has a learning unit that learns using teaching data obtained by annotating image data from a first image acquisition device (e.g., image acquisition device 3a), and obtains an estimation model through this learning. Furthermore, when generating an estimation model for a second image acquisition device (e.g., image acquisition device 3b) with characteristics different from the first image acquisition device, the teaching data collected for the first image acquisition device can also be effectively utilized.

[0105] That is, when effectively using teaching data for learning in order to acquire a second image with image input characteristics different from the first image acquisition device, the image processing unit performs image processing as follows and generates an estimation model for the second image acquisition device with different image input characteristics. In this image processing, the image data obtained from the first image acquisition device in the teaching data is processed according to the differences in image acquisition characteristics and set as teaching data. The differences in image input characteristics can be caused by differences in the specifications and performance of the camera sensor, the optical characteristics of the camera, the specifications and performance of the image processing, and the type of illumination light.

[0106] Furthermore, considering the differences in image acquisition devices, it is also possible that differences exist in other devices as well. For example, the objects being photographed in such an environment may naturally present different appearances. Therefore, the image processing unit described above can also process the image data of the first object contained in the image data obtained from the first image acquisition device and the image data of the second object contained in the image data obtained from the second image acquisition device in the teaching data in a manner that is compatible.

[0107] For example, consider assigning all the features of images of similar objects detected by the second image acquisition device to the objects reflected in the image obtained by the first image acquisition device to generate new teaching data. For instance, if an image of a zebra is unavailable during a presentation, a horse image can be used as a substitute by marking stripes on it. Although this only changes the color and pattern, differences in other shape features can be corrected for utilization. Furthermore, for example, when using teaching data for a processing device with a quadrilateral front end for a processing device with a circular front end, a structure where the front end of the processing device is more rounded can be selected as teaching data, and the image can be modified by correcting for differences in the features of that front end shape for learning purposes.

[0108] However, while image processing broadens the application scope of objects, it may not always meet the desired specifications (e.g., guidance functions matched to the user's skill level). Therefore, in such cases, it's not enough to simply process the image; the selection and omission of teaching data, the content of annotations, and the processing (adjustment or modification) of methods can also be performed. Furthermore, the method of displaying the estimation results can be modified. Alternatively, it's possible to customize the process by issuing a warning to skilled users at a specific reliability point in image acquisition, but in other cases, even with lower reliability, observing safety and issuing a warning at a higher reliability point.

[0109] That is, the estimation learning device in this embodiment includes: an input unit that inputs image data from a first image acquisition device; and a learning unit that learns using teaching data obtained by annotating the image data to obtain an estimation model. Furthermore, it includes an image processing unit that, when relearning the estimation model for a second image acquisition device with image input characteristics different from the first image acquisition device, processes the image data obtained from the first image acquisition device by changing the reliability judgment level corresponding to the difference in image input characteristics generated based on the user's skill, and sets it as the teaching data. The user's skill is known in aspects such as hand tremors, slow movement, and response speed to changes in specific scenarios. Based on this difference in skill, differences in image input characteristics (or differences in temporal image data changes) are generated. Here, the difference in the way image data changes over time is conceptually represented as a difference in image input characteristics.

[0110] Furthermore, as test data used to determine the reliability of the following estimation model, since it is an estimation model for the second image acquisition device, data from the second image acquisition device can also be used, wherein the estimation model is an estimation model generated for the second image acquisition device using teaching data generated through correction or processing.

[0111] Furthermore, when learning from unfamiliar props, combining multiple props with similar shapes can increase the probability of success, or props with shapes similar to the unfamiliar parts can be used for learning. Alternatively, images of treatment devices captured in previous teaching data can be replaced, or parts of their shapes can be altered for learning. To ensure safety during learning, consider using the bleeding timer from previous treatment device teaching data earlier in time, or strictly controlling the extent of bleeding.

[0112] When teaching skills to beginners, the first step is to explain strategies for changing the reliability level. However, in addition to this method, other approaches are considered, such as emphasizing the jitter of the prop's movement based on previous teaching data or advancing the timing. It is also possible to advance the timing of the bleeding effect from previous skill teaching data and advance the timing of the guidance activation.

[0113] If the determination result in step S11 is that the reliability is lower than the specified value, the teaching data (S13) is selected for rejection. In cases of low reliability, sometimes selecting teaching data can improve reliability. Therefore, in this step, the teaching data selection unit 1f removes image data where there is no causal relationship. For example, teaching data where there is no causal relationship between the cause and effect of bleeding expansion / contraction is removed. Regarding this process, an estimation model for estimating causality can be prepared in advance, automatically excluding teaching data with low causal relationships. Furthermore, the overall conditions of the teaching data can be changed. After selecting the teaching data, the process returns to step S9, and the estimation model is generated again.

[0114] On the other hand, if the determination result in step S11 is, for example, that reliability is verified by prioritizing data obtained in the assumed system and the reliability is OK, then the estimation model is sent (S15). Here, since the generated estimation model meets the reliability benchmark, the teaching data selection unit 1f determines the teaching data candidate to be used at this time as the teaching data. Furthermore, the learning result utilization unit 1e sends the generated estimation model to the camera device 6. After receiving the estimation model, the camera device 6 sets the estimation model for the estimation unit 2AI. After sending the estimation model, the estimation model generation process ends. In addition, if the sent estimation model is sent along with information such as specifications, then during the estimation by the camera device, control can be achieved that also reflects whether estimation is performed using a single image, or determination is performed using multiple images, or the degree of time difference (frame rate, etc.). Other information can also be processed.

[0115] In this process, the learning device inputs image data from the image acquisition device 3 (S1, S5), annotates the image data to generate teaching data (S3, S7), and obtains an estimation model by learning from the generated teaching data (S9). Specifically, in the images continuously acquired from the image acquisition device 3 in a time series, image data from a specific time point to a subsequent time point is annotated (S3, S7, S13) to become teaching data (S11, S13). Thus, within the continuously output image data, time series image data is acquired by tracing back from a specific time point where certain events occurred (e.g., bleeding expansion, bleeding contraction), and annotated to become teaching data candidates. An estimation model is generated by learning from these teaching data candidates, and if the generated estimation model has high reliability, the teaching data candidate is set as teaching data.

[0116] In other words, this process generates an estimation model using data traced back to a specific time when certain events occurred. Specifically, it generates an estimation model that can predict the future based on events that cause outcomes at a specific time—that is, based on causal relationships. Using this estimation model, even minor behaviors or phenomena that users might not notice can be predicted without omission; for example, it can alert or warn users in the event of an accident. Furthermore, even if users have concerns that they notice, this message can be conveyed if the concerns do not escalate to a serious level.

[0117] The image estimation learning device 1 in this process can collect teaching data sets 4A from a large number of camera devices 6. Therefore, it is possible to generate teaching data using a large amount of data, and thus generate a highly reliable estimation model. In addition, in this embodiment, when an event occurs, data is collected that is narrowed down to the scope related to that event, thereby enabling the efficient generation of the estimation model.

[0118] In this process, the image estimation learning device 1 collects candidate image data sets that can serve as teaching data from the camera device 6, and annotates these image data sets with bleed expansion, etc. (see S3, S7). However, the camera device 6 can also perform these annotations to generate teaching data sets, and the learning unit 1c can use these teaching data sets to generate an estimation model. In this case, the annotation process can be omitted in the image estimation learning device 1. In this case, the process is implemented collaboratively by the control unit 1a in the image estimation learning device 1 and the control unit 7 in the camera device 6.

[0119] Next, use Figure 4 The flowchart shown illustrates the operation of the camera device 6. This operation is performed by controlling the various devices and components within the camera device 6 via the control unit 7. An example of the camera device 6 being installed within an endoscope is also described. Furthermore, routine operations such as power on / off switching are omitted in this flowchart.

[0120] when Figure 4 At the start of the process shown, firstly, image capture and display are performed (S21). Here, when the image acquisition device 3 acquires image data at predetermined time intervals (determined by the frame rate), the image data is displayed in the guide section 5. For example, if the camera device 6 is installed inside the endoscope, the image inside the body acquired by the camera element installed at the front end of the endoscope is displayed in the guide section 5. This display is updated every predetermined time interval determined by the frame rate. The guidance method can also be divided into beginner and advanced user types, etc., according to the technology described in this specification, and can be changed according to the user. It is also assumed that the method may change depending on the object or the usage environment.

[0121] Next, it is determined whether AI correction is needed (S23). The estimation model mounted on the estimation unit 2AI sometimes becomes unsuitable because the characteristics of the image data have changed due to the equipment used (including the camera device 6A) being changed to the image acquisition device 6B or the version being upgraded. Furthermore, it may become unsuitable for other reasons as well. In such cases, it is preferable to correct the estimation model set on the estimation unit 2AI. Therefore, in this step, the control unit 7 determines whether the estimation model needs correction.

[0122] In cases where the estimation model becomes unsuitable due to reasons such as changes in the equipment used, it is preferable to generate the estimation model using image data from that device. However, if the data from that device is limited, it is impossible to collect a sufficient amount of data to generate an estimation model. Therefore, in this embodiment, a corrected estimation model is generated by processing the image data collected so far. This is then used... Figure 7 To describe in detail the actions required to determine whether AI correction is needed.

[0123] When the determination result in step S23 indicates that the AI ​​needs correction, the generation of a corrected estimation model is then commissioned and obtained (S25). Here, the camera device 6 commissions the image estimation learning device 1 to generate a corrected estimation model, and obtains the estimation model after it is generated. When commissioning the corrected estimation model, information such as the parts that need correction can also be sent. That is, as described above, in this embodiment, the teaching data that has already been used is processed by applying it to a new device, and the corrected estimation model is generated using the processed teaching data. Then it is used... Figure 5 This section describes the detailed process of generating the corrected estimation model.

[0124] If the corrected estimation model is obtained, or if the determination in step S23 is that no AI correction is needed, the next step is to determine whether it is trigger information (S27). For example, in the case of using Figure 2 In the event described in (a) and (b), such as bleeding occurring and expanding during treatment, trigger information is generated. In this example, the control unit 7 can output trigger information after analyzing the image data acquired by the image acquisition device 3 and determining that the bleeding has expanded. Furthermore, this image analysis can be performed using AI with an estimation model, or the trigger information can be output by the doctor manually operating a specific button, etc.

[0125] If the determination result in step S27 indicates that trigger information has been generated, a predetermined time retrospective recording is performed (S29). Here, the image data acquired by the image acquisition device 3 is recorded in the image data storage memory within the imaging device 6 after a predetermined time. Typically, all image data acquired by the image acquisition device 3 is pre-recorded in the memory, and image data within a predetermined time period retrospectively determined based on the generation of trigger information is assigned predetermined metadata and temporarily recorded in the teaching data candidate group. If no trigger information exists, the control unit 7 may appropriately eliminate the image data candidate group. Figure 2 In the examples shown in (a) and (b), the specific timing is the point in time when the bleeding expands, and the tracing time is from a specified time (e.g., T = -1 sec) to T = -5 sec. Furthermore, if image data from T = 0 to T = T1a is added to the image data set, the process of bleeding expansion can also be included for learning. The starting point of the tracing record can be the time when the trigger information was generated, or it can be an earlier time. The tracing time should be appropriately determined in a way that includes the range of causes from which a causal relationship can be found. The timing of the cause varies depending on reliability; the longer the tracing time, the lower the reliability. However, for beginners, a less reliable timing can be used for safety. These also apply to image processing.

[0126] After the traceability record is performed in step S29, or if the determination result in step S27 indicates that no trigger information exists, image estimation is then performed (S31). Here, the image data acquired by the image acquisition device 3 is input to the image input unit 2IN of the image estimation device 2, and the estimation unit 2AI performs estimation. When the estimation result output unit 2OUT outputs the estimation result, the guidance unit 5 provides guidance based on the output result. For example, such as... Figure 2 As shown in (a) and (b), estimation can be performed at time T = -5 seconds, and a display can be made indicating that bleeding will begin 5 seconds later (T = 0). Furthermore, if bleeding occurs at time T = 0, based on the estimation result of whether the bleeding is enlarging or shrinking, either Ga or Go is displayed. Additionally, when multiple image estimation devices, such as image estimation device 2a, are provided in addition to image estimation device 2, multiple estimations can be performed. For example, in addition to anticipating bleeding, other estimations can be performed.

[0127] Furthermore, when estimating images, the estimation can be supplemented not only by image data but also by the doctor's voice during diagnosis or treatment. Additionally, the reliability of the equipment used for diagnosis or treatment can be estimated, and if the reliability is lower than a specified value, higher-reliability equipment can be recommended. Moreover, sometimes the treatment instruments used may become noise (obstructing the view in the image); therefore, image estimation can also be used to process the image of the treatment instruments.

[0128] After image estimation, the next step is to determine whether to output teaching data candidates (S33). Here, the control unit 7 determines whether trace recording was performed in step S29. If trace recording was performed, the image data at that time is stored as teaching data candidates in the memory of the camera device 6. If the determination result is that trace recording was not performed, the process returns to step S21.

[0129] If the determination result in step S33 is "yes", then teaching data candidates are output (S35). Here, the control unit 7 outputs the teaching data candidate set stored in the memory of the imaging device 6 to the image estimation learning device 1. In addition, when the image estimation learning device 1 receives the teaching data candidate set, it records it in advance in the recording unit 4. After outputting the teaching data candidates in step S35, the process returns to step S21.

[0130] Furthermore, in this embodiment, the camera device 6 performs a determination of whether the bleeding expands or shrinks (see reference). Figure 4 (S27). However, this determination can also be performed in the control unit 1a of the image estimation learning device 1. That is, the enlargement / reduction of bleeding can be determined based on changes in the shape or size of the blood color occupying the screen, and can be detected either logically or through estimation. Furthermore, the determination of enlargement / reduction can be intentionally changed according to the customization of the teaching data. Alternatively, for the safety of beginners, the teaching data can be digitized as an image annotated as enlargement even if it does not enlarge. This situation also manifests as image processing.

[0131] Furthermore, regarding the trigger information in step S27, an example of bleeding occurring inside the body during endoscopy is given. However, this implementation can also be applied to situations other than bleeding. For example, in cases where body temperature or weight can be measured by wearable sensors, trigger information can be generated when body temperature rises sharply, and previous body temperature data, weight data, or other data (including image data) can be retrospectively recorded. If this data is sent as teaching data to an estimation learning device, an estimation model can be generated.

[0132] Furthermore, in step S35, a set of candidate teaching data generated based on the trace records is sent to the estimation learning device. This set of candidate teaching data can trace not only image data sets recorded in the same device (camera device 6), but also detection data from other devices to investigate causal relationships.

[0133] Furthermore, in step S23, it is determined in the imaging device whether AI correction is needed. However, the determination of whether AI correction is needed can also be made in the image estimation learning device 1. If the image processing unit 1d (or the control unit 1a) detects that the teaching data set input into the image input unit 1b of the image estimation learning device 1 has different characteristics (including uses) from the previously accumulated teaching data set, it is determined that AI correction is needed.

[0134] Furthermore, even on first-time users, by generating teaching data by annotating normal images as normal, and using this teaching data for learning, it is possible to estimate abnormalities. For example, this could be a determination unit that identifies abnormalities such as lesions, colors, and shapes based on images of the stomach, or it could be configured as an AI that needs to determine the nature of the abnormality when it is identified as abnormal.

[0135] Alternatively, one could use the reliability of the "normal" judgment based on the currently available AI (or the reliability of the "abnormal" judgment) to determine whether AI correction is needed. If the reliability of this judgment is below a certain level, it is determined to be the first time it has been seen, and AI correction is implemented.

[0136] Furthermore, as a learning example for generating the estimation model, we illustrate causal relationship-guided estimation in the highly advanced medical field. However, this implementation is not limited to the medical field and can also be applied to guided estimation. In practice, most commonly used estimation models are used to identify what is observed in images, such as various person detection and action detection by surveillance cameras or obstacle detection by vehicle cameras—these are image detection types. The technique described in this implementation for improving estimation performance by eliminating variations in the input image data patterns is also effective in detection and identification types.

[0137] That is, the image estimation learning device has: an input unit that receives image data from a first image acquisition device; and a learning unit that learns using teaching data annotated from the image data to obtain an estimation model. In such an image estimation learning device, its application scope is broadened, and for useful estimation models, it can effectively overcome various constraints and be used in various fields. However, due to constraints, it is sometimes difficult to immediately collect useful teaching data. Therefore, the estimation learning device in this embodiment includes an image processing unit that, when customizing the estimation model for a second image acquisition device used under different conditions than the first image acquisition device (not creating a completely different estimation model, but expecting the same specifications that already have proven performance), processes the image data obtained from the first image acquisition device in the teaching data, including selection or annotation corresponding to the differences in image acquisition characteristics, and sets it as teaching data. With this design, even if useful teaching data cannot be immediately collected, a useful estimation model can be generated.

[0138] Furthermore, it is not necessary to limit the use of teaching data obtained solely from the first image acquisition unit or to the direct use of teaching data. For example, even if the data is not obtained from the first image acquisition unit, information about anomalies of the object published in papers or reported elsewhere can be used. For instance, if information such as a tumor exists as an anomaly of the object, the size of the tumor image can be distorted or enlarged / reduced to correct the teaching data. Color correction, distortion of similar image regions, etc., can also be performed as needed to re-database the teaching data. If such processing is performed based on information obtained from the usage environment of the second image acquisition unit, taking into account potential conditions, reliability is further improved.

[0139] Next, in the explanation Figure 5 Before the process of generating the corrected estimation model shown, use Figure 6 The processing of image data is explained. Figure 6 (a) and Figure 2 Similarly, (a) shows the situation after bleeding occurred and expanded during endoscopic treatment.

[0140] Figure 6 (b) and Figure 6 Case (a) similarly illustrates the situation after the bleeding has expanded. In this example, the imaging device 6 uses an image acquisition device (e.g., image acquisition device 3b) with characteristics different from those of the image acquisition device 3a. Because the number of pixels of the imaging element of the image acquisition device 3b is smaller, the image data ID3 that can be acquired is significantly different from the image data ID1. Therefore, in the estimation model generated by accumulating image data with characteristics equivalent to image data ID1, even if the input... Figure 6 Image data like (b) can only produce low-reliability estimates. Furthermore, even when using a parent set that mixes image data ID3 with previously accumulated image data to generate an estimation model, only low-reliability estimation models can be generated.

[0141] The characteristics described here are based on the specifications and performance of the image reading device, the objects being processed, peripheral devices such as accessories, and associated cooperating devices, and may vary depending on their usage environment. In other words, image input characteristics arise from differences in the specifications and performance of the camera sensor, the optical characteristics of the camera, the specifications and performance of the image processing, and the type of illumination light. Of course, such factors can sometimes vary depending on the user's mode settings; in such cases, these factors can also be considered.

[0142] Therefore, in this embodiment, the image processing unit 1d processes (corrects) the previously used image data based on the differences in image input characteristics, adjusting it to match... Figure 6 (b) The same image data rank (refer to) Figure 5 (S1a, S5a). Then, the corrected image data is annotated (refer to S1a, S5a). Figure 5 S3a and S7a), generate the estimation model (refer to S3a and S7a). Figure 5 (S9). Furthermore, when processing (correcting) the image data of previously used teaching data based on differences in image input characteristics, if there is no need to change the annotation of the teaching data, only the image data is processed (corrected).

[0143] Next, use Figure 5 The flowchart shown illustrates the actions involved in generating the corrected estimation model. In step S25 (refer to...) Figure 4 This process is executed when a request is made from the camera device 6 to generate an estimation model after the image estimation learning device 1 has been corrected. This process is implemented by controlling the various parts within the image estimation learning device 1 through the control unit 1a of the image estimation learning device 1. This process is an example of the image estimation learning device 1 generating a corrected estimation model based on an image with enlarged or reduced bleed.

[0144] when Figure 5 The process of generating the corrected estimation model, as shown, begins by first collecting process images (S1a) during hemorrhage expansion. If using... Figure 2 As described in (a), in cases where bleeding expands during treatment, the image processing unit 1d collects image data from the recording unit or camera device 6 within the image estimation learning device 1, and corrects this image data. Figure 2In the example shown in (a), image data is collected between T = -5 seconds and T = -1 second. Then, the image processing unit 1d processes the collected image data to form a... Figure 6 The image data ID3 shown in (b) is corrected in the same manner as the image data of the same level as the image data output by the image acquisition device 3b.

[0145] Furthermore, as mentioned above, since it is desirable to obtain a more reliable estimation model by customizing it according to the user's device usage, environment, objects, etc., a process such as obtaining desired specifications (customization request) can also be performed in step S1a. In conjunction with the customization request, image selection, image correction, annotation modification, etc., are performed to reconstruct (process, edit, manipulate) the teaching data appropriately.

[0146] For example, in most cases, image data acquired from a first-specification (including factors listed below) image acquisition device is rarely consistent with image data acquired from a second-specification image acquisition device, not only in terms of device performance, specifications, environment, and peripheral systems, but also in terms of the objects being processed, peripheral equipment such as accessories, handling instruments, and operators. Therefore, it is difficult to directly utilize an estimation model in a second-specification image acquisition device, which is an estimation model learned and obtained by annotating image data from a first-specification image acquisition device. Thus, the estimation learning device of this embodiment includes an image processing unit that, when customizing and learning an estimation model for a second (specification) image acquisition device with image input characteristics different from a first (specification) image acquisition device, processes the image data acquired from the first image acquisition device in the teaching data according to the differences in image acquisition characteristics (including not only device performance, specifications, environment, and peripheral systems, but also objects being processed, peripheral equipment such as accessories, handling instruments, and operators) and sets it as teaching data. By optimizing the teaching data through the image processing unit, an estimation model that can also be used in the image acquisition device of the second specification can be generated.

[0147] As an example of correction when using treatment devices, a method is proposed as follows: When it can be determined that the shape of the treatment device has changed, a set of teaching data containing the treatment device with the closest shape is used to perform geometric transformation of the image or a non-linear transformation such as stretching / scaling. Furthermore, in conjunction with this transformation, the annotation information representing the treatment device portion is also transformed. In the case of a sharp-shaped tip during geometric transformation, the annotation is weighted in the direction prone to bleed, or annotated in a way that results in "bleed" based on the weighted determination. Alternatively, the influence of other AIs (shape change effect prediction AIs) learned from teaching data with different shape differences on the shape change can be determined, and methods reflecting these results can be used.

[0148] In most cases, the image processing unit described above aims to make the most efficient use of the first object image data (many of which have proven performance) contained in the image data obtained from the first image acquisition device in the teaching data. Therefore, it processes the image data according to the image acquisition characteristics, including not only the performance, specifications, environment, and peripheral systems of the device, but also the object being processed, accessories and other peripheral equipment, handling tools, operators, etc., so as to make it suitable for the second object image data contained in the image data obtained from the second image acquisition device.

[0149] In addition, the image processing unit 1d (or in cooperation with the teaching data selection unit 1f) takes into account safety factors, such as the use of image sensors with poor detection performance or treatment devices with poor operability, the user's proficiency, and the patient or affected area, and issues warnings as soon as possible, and prioritizes the use or collection of teaching data with similar factors.

[0150] After collecting and correcting the image in step S1a, the next step is to annotate "bleed expansion" and the timing (S3a). Here, the control unit 1a annotates the image data with the intention of "bleed expansion" and the timing for acquiring the image data, for use as teaching data. Specifically, the control unit 1a reselects image data, changes the weighting, or performs customized measures on previous image data or processed image data, and annotates the image data with the intention of "bleed expansion" and the timing for acquiring the image data, for use as teaching data. This customized measure can also be described as "processing". Furthermore, even if an image is set to "bleed reduction", if the bleed area does not shrink within a certain time, it is re-annotated as "bleed expansion". Such a change as converting images obtained by skilled operators into teaching data for beginners can also be called "processing". In addition, the teaching data after correcting (processing) the image data and annotating it can also be recorded in advance as teaching data group B 4b in the recording unit 4.

[0151] Next, images of the bleeding process as it shrinks are collected (S5a). For example, using... Figure 2 As explained in (b), in the case of reduced bleeding during treatment, the image processing unit 1d collects an image at this time from the recording unit or the camera device 6 within the image estimation learning device 1. Figure 2 In the example shown in (b), images are collected between T = -5 seconds and T = -1 second. Then, the image processing unit 1d corrects the collected image data to make it consistent with... Figure 6 Image data of the same level as image data ID3 (e.g., image data output by image acquisition device 3b) shown in (b).

[0152] After collecting and correcting the image in step S5a, the next step is to annotate "bleed reduction" and set the timing (S7a). Here, the control unit 1a annotates the image data with the intention of "bleed reduction" and sets the timing for acquiring the image data, making it a candidate for teaching data. In addition, the teaching data after correcting (processing) the image data and annotating it can also be recorded in advance as teaching data group B 4b in the recording unit 4.

[0153] After annotating the image data and generating teaching data in steps S3a and 7a, and... Figure 3 Similarly, an estimation model is generated (S9). Here, the learning unit 1c uses the teaching data annotated in steps S3a and S7a to generate an estimation model. This estimation model is generated based on the input... Figure 6 In the case of image data ID3S shown in (b), it is possible to achieve a prediction such as "the bleeding will expand after 0 seconds".

[0154] After generating the estimated model, determine whether the reliability is OK (S11). Here, with Figure 3 Similarly, the learning unit 1c inputs image data used to verify the reliability of the known answer into the estimation model, and determines the reliability based on whether the output in this case is the same as the answer. When the reliability of the generated estimation model is low, the proportion of consistent answers is low.

[0155] In this step, test data is input, and it is determined whether the expected estimation result is output. This test data is preferably matched to the specifications, environment, and conditions of the second-specification image acquisition device that actually utilizes the estimation model (including not only the device's performance, specifications, environment, and peripheral systems, but also the objects being processed, accessories, peripheral equipment, handling tools, and operators). Here, it is desirable to prioritize the use of data obtained under the specifications, environment, and conditions of the second-specification image acquisition device. However, such data is often not readily available. Therefore, based on the differences in image acquisition characteristics, including not only the device's performance, specifications, environment, and peripheral systems, but also the objects being processed, accessories, peripheral equipment, handling tools, and operators, the data obtained from the first image acquisition device is processed and utilized to make it suitable for image data of the second object contained in the image data obtained from the second image acquisition device. Of course, the determination can also be adopted by manually inputting the determination the user wants to make.

[0156] If the determination result in step S11 is that the reliability is lower than the specified value, then... Figure 3 Similarly, the teaching data is selected for selection (S13). In cases of low reliability, sometimes selection of teaching data can improve reliability. Here, in this step, image data that does not have a causal relationship is removed. After selecting the teaching data, return to step S9 to generate the estimation model again.

[0157] On the other hand, if the determination result in step S11 is that the reliability is OK, then... Figure 3 Similarly, the estimated model is sent (S15). Here, the generated estimated model meets the reliability benchmark, therefore, the teaching data selection unit 1f determines the teaching data candidates to be used during estimation as teaching data. Furthermore, the learning result utilization unit 1e sends the generated estimated model to the camera device 6. When the camera device 6 receives the estimated model, it sets the estimated model for the estimation unit 2AI. After sending the estimated model, the estimated model generation process ends.

[0158] Thus, in Figure 5 In the process of generating the corrected estimation model shown, image data used as teaching data is collected (S1a, S5a), and the image processing unit 1d corrects (processes) the collected image data (S3a, S7a). Since the characteristics of the image data and other data that become the estimation object have changed, the correction of the estimation model is requested. Therefore, in this process, in order to generate an estimation model corresponding to the characteristics of the new data, correction is performed so that the accumulated data is suitable for the characteristics of the new data. Therefore, it is possible to generate an estimation model quickly and at low cost without having to recollect data with new characteristics.

[0159] Next use Figure 7 The flowchart shown illustrates step S23 ( Figure 4 The process of determining whether AI correction is needed (refer to) will be explained. This process is executed by the CPU 7a within the camera device 6 controlling the various components within the camera device 6 according to the program stored in the memory 7b.

[0160] When it begins Figure 7 When determining whether AI correction is needed, the system first checks whether the image acquisition device has model information (S41). Here, for the image acquisition device 3, it checks whether detailed model information is available. Model information includes, for example, the number of pixels, frame rate, resolution, focal length, and distance to the object. Furthermore, if the image acquisition device 3 is integrated into the camera device 6, model information is easily obtained. However, even if it is a separate device, model information can be obtained via information communication networks such as the Internet, or by referring to databases as needed.

[0161] Here, for simplicity, the differences in specifications and performance of image acquisition devices are illustrated. However, as mentioned above, since the purpose is to customize the device according to the user's usage, environment, object, etc., in order to obtain a more reliable estimation model, it is also possible to determine the desired specifications (customization request). For example, even if the model is the same, differences in the equipment used, the user's strength, or the object can allow for the same processing as the model information based on manual input results or information recorded in the recording unit.

[0162] If the determination result in step S41 includes the model information of the image acquisition device, then the correction method is obtained based on the image quality information DB based on the model information, and the correction method is determined (S43). Here, the control unit 7 determines the correction method performed in steps S1a and S5a. For example, if the number of pixels of the camera element is small, the number of pixels in the acquired image data can be multiplied or divided according to the pixel ratio (gap removal, addition, etc.). Such processing is also considered processing, but in addition, the processing method of the image as teaching data is sometimes also considered processing.

[0163] Here, we will continue to explain in detail the method for determining poor image quality. If the determination result in step S41 is that there is no model information, we will determine whether a reference scene image exists (S45). The reference scene image is an image obtained when photographing an object to determine whether the characteristics of image data, etc., are different. That is, when determining whether to correct the AI, it is best to determine whether the image data used to generate the current estimation model is the same as the image data input at this time. For this purpose, it is easy to know if images obtained by photographing the same object are compared. However, it is usually difficult to photograph exactly the same object, so photographing similar objects is sufficient. For example, with an endoscope, even if the equipment or the patient is different, the image obtained when inserted from the mouth into the esophagus is largely the same; therefore, this image can be used as the reference scene. As an example other than an endoscope, a blue sky sometimes serves as a benchmark for camera performance; in addition to white images and gray images, there are also benchmark images for performance determination. Even without preparing a special image, if an image with known characters or patterns or a standardized image is photographed, changes in peripheral light intensity, aberration information, etc., can be obtained based on the differences from the original shape, etc.

[0164] If the determination result in step S45 is not a reference scene image, an image of the reference scene is estimated (S47). Since no image exists that can serve as a reference scene, a substitute image must be found from the images acquired by the image acquisition device 3. Even if the substitute image is not a reference scene to a certain degree, it is desirable to be similar to the degree to which the characteristics of the image data can be determined by comparing two images. For example, in endoscopic examinations, instruments are sometimes used, and the shapes of these instruments are mostly similar. In this case, the image in the acquired image that shows the shape of the instrument is estimated as the reference scene. Furthermore, not only the shape of the instrument, but also the way the instrument appears in the frame (its position, etc.) can be used as a judgment criterion when estimating the image of the reference scene. Not limited to endoscopes, even microscopes or cameras, and the equipment used often have similar shapes or colors; therefore, the inclusion of these instruments can also be determined and compared.

[0165] After estimating the reference scene image in step S47, or if the determination result in step S45 is that a reference scene image exists, the next step is to determine whether a difference from the reference image is permissible (S49). As described above, when comparing the two images, if the characteristics of the image data are not different, there is no need to modify the estimation model. Here, it is determined whether the characteristics of the image data acquired from the image acquisition device 3 are different to the extent that the estimation model must be modified. Furthermore, even for images of the same location, it is determined whether the difference is significant.

[0166] If the determination result in step S49 is that the difference from the reference image is within an acceptable range, then the branch proceeds to "No" (S55). If the difference between the image acquired by the image acquisition device 3 and the reference scene is not significant at this moment, then the estimation model does not need to be corrected, therefore the branch proceeds to "No" and continues to the next step. Figure 4 Step S27.

[0167] On the other hand, if the determination result in step S49 is that the difference from the reference image is not within the acceptable range, then a correction method is determined based on the characteristics of the image (S51). Since the correction method varies depending on the degree of difference between the image acquired by the image acquisition device 3 at this moment and the reference scene, the control unit 1a can determine the correction method based on the degree of difference, etc. For example, if the number of pixels is different, a method can be used to increase or decrease the number of pixels in the accumulated image so that it becomes the same number of pixels as the image acquired by the image acquisition device 3 at this moment. In addition to differences in the performance of the optical system, the camera sensor, and image processing, differences in frame rate, field of view, and illumination light can also be addressed using the same method.

[0168] After determining the correction method in step S43 or S51, the branch proceeds to "Yes" (S53). Since the image acquired by the image acquisition device 3 differs significantly from the reference scene at this moment, the estimation model needs to be corrected, so the branch proceeds to "Yes" and continues... Figure 4 Step S25.

[0169] Thus, in Figure 7 In the process shown for determining whether AI correction is needed, if the model information of the image acquisition device exists, the correction method for the image data to be corrected is determined based on the model (see S41, S43). On the other hand, if the model information of the image acquisition device does not exist, a reference scene image or an image estimated as the reference scene is used to determine whether correction of the estimation model is needed (see S49). If it is determined that correction is needed, the correction method is determined based on the characteristics of the image (S51). In this process, the need for correction and the method to be used are determined based on the model information of the image acquisition device, the reference scene image, etc. However, there are various factors for determining whether AI correction is needed; therefore, this information can be added, and the determination itself can be performed by AI.

[0170] As described above, this section explains the customization of the camera unit's performance, functions, and specifications, particularly in cases where the equipment's operating environment varies, and the processing of teaching data to match this customization. However, it is possible to provide systems, apparatus, and methods that not only process image quality and features to match customization requirements (image correction), but also reconstruct appropriate teaching data (processing, editing, or manipulation) while simultaneously making image selections or correcting annotations.

[0171] To reiterate, in most cases, image data from image acquisition devices of the first specification (including factors listed below) is rarely consistent with that of image acquisition devices of the second specification, not only in terms of device performance, specifications, environment, and peripheral systems, but also in terms of the objects being processed, peripheral equipment such as accessories, handling tools, and operators. In such cases, it often becomes difficult to directly utilize the estimation model, which is learned using teaching data annotated from image data from image acquisition devices of the first specification. This embodiment solves this problem.

[0172] To address the aforementioned issues, this embodiment includes an image processing unit. When customizing and learning an estimation model for a second (specification) image acquisition device with image input characteristics different from the first (specification) image acquisition device, this image processing unit processes the image data obtained from the first image acquisition device according to the differences in image acquisition characteristics (including not only the device's performance, specifications, environment, and peripheral systems, but also the objects being processed, accessories, peripheral equipment, handling tools, and operators), and sets this as teaching data. By optimizing the teaching data through this image processing unit, an estimation model corresponding to the second image acquisition device can be generated. Figure 7 This illustrates the acquisition of information that forms the basis of the image processing, and provides an example of the processing.

[0173] In other words, the image processing unit generally aims to utilize the first object image data (many of which have proven performance) contained in the image data obtained from the first image acquisition device as efficiently as possible. Therefore, it processes the image data based on image acquisition characteristics, including not only the device's performance, specifications, environment, and peripheral systems, but also the object being processed, accessories, peripheral equipment, handling tools, and operators, to make it suitable for the second object image data contained in the image data obtained from the second image acquisition device. By performing such a design, for example, it is possible to optimally detect a specific object from the image in a way that matches the performance of the device. This becomes an effective technique in in-image object detection or segmentation, which is an important category of image estimation models.

[0174] Furthermore, in order to estimate accidents such as "bleeding expansion" or "bleeding contraction," when making predictive annotations on timing, it is preferable to perform processing that is not just about differences in image quality. Regarding image sensors with poor detection performance, the above-mentioned correction (processing) methods can be used to address the issue. However, depending on the situation of using treatment instruments with poor operability, low user proficiency, and the patient or affected area, it is also advisable to consider safety aspects and issue warnings as soon as possible, and to prioritize the use or collection of teaching data based on similar factors.

[0175] In the aforementioned cases, the control unit 1a performs customized measures on reselected image data, image data with altered weights, or previous image data or processed image data, and annotates it with the intention of "bleed expansion" and the timing of acquiring the image data, for use as teaching data. This customized measure can also be described as "processing." Furthermore, even if an image is set to "bleed reduction," if the bleed area does not shrink within a certain time, it is re-annotated as "bleed expansion." Such changes, such as digitizing teaching data from images processed by experienced users to generate estimation models for beginners, can also be called "processing."

[0176] In addition, Figure 7 The determination of whether AI correction is needed is performed in the camera device 6. However, this determination is not limited to the camera device 6; it can also be performed in the image estimation learning device 1. In this case, when acquiring image data from the camera device 6, information such as device model information can also be acquired and utilized. Furthermore, a database of reference scene images can be prepared in advance, and the presence of reference scene images can be determined by comparing it with the image data from the camera device 6. Of course, in accordance with human input and other customized requirements, this can be reflected not only in image quality and feature processing (image correction), but also in the selection of images or the correction of annotations, and appropriate reconstruction (processing, editing, manipulation) of teaching data can also be performed. If it is written as AI correction, it includes not only the performance, specifications, environment, and peripheral systems of the device, but also the objects being processed, accessories and other peripheral equipment, handling tools, operators, etc. In the case of the image estimation engine, since the input data is based on images, the characteristics of image acquisition can be considered in a broad sense.

[0177] Alternatively, AI correction may be required when the data from image acquisition device 3 belongs to an unknown category. In this case, artificial intelligence can automatically determine whether the data belongs to an unknown category, or the user of the second image acquisition device (e.g., image acquisition device 3b) can manually set whether the data belongs to an unknown category. Furthermore, the determination of whether the data belongs to an unknown category can also be based on the model information of the second image acquisition device (e.g., image acquisition device 3b) and / or on an image estimated as a reference image from the image data from the second image acquisition device.

[0178] As explained above, in one embodiment of the present invention, image data from the first image acquisition device (e.g., referring to...) is input. Figure 5 In S1a and S5a), when relearning the estimation model for a second image acquisition device with characteristics different from the first image acquisition device, the image data obtained from the first image acquisition device in the teaching data is processed and set as teaching data (e.g., referring to...). Figure 5 S3a and S7a) are used to obtain the estimated model by learning from teaching data obtained by annotating image data (e.g., referring to S3a and S7a). Figure 5 (S9). Therefore, it is not limited to data of a pre-assumed category, but can also make appropriate estimates for unknown categories when the characteristics of the data have been changed for previously accumulated data. That is, when dealing with unassumed data, by processing some previously accumulated data, it is also possible to generate an estimation model that can estimate unassumed data.

[0179] Furthermore, in one embodiment of the present invention, image data from the first image acquisition device (e.g., referring to...) is input. Figure 5 In S1a and S5a), when the estimation model is customized for use with a second image acquisition device under conditions different from the first image acquisition device, the image data obtained from the first image acquisition device is processed according to the differences in image acquisition characteristics, including selection or annotation, and set as teaching data (e.g., reference). Figure 5 S3a and S7a) obtain the estimation model by learning from the teaching data obtained by annotating the image data. Therefore, it is not limited to data of a pre-assumed category, and can make appropriate estimations even when the characteristics of the data are changed for previously accumulated data, even in unknown categories. That is, even when processing data that is not assumed, by selecting or processing the previously accumulated data from the first image acquisition device, an estimation model that can also be estimated for data that is not assumed can be generated.

[0180] Here, it is written as "data outside the assumption," which refers to data from "devices outside the assumption" where sufficient teaching data cannot be collected, or data from "environment outside the assumption." That is, it is "data outside the assumption" that results from image acquisition characteristics. These characteristics include not only the device's performance, specifications, environment, and peripheral systems, but also the objects being processed, peripheral equipment such as accessories, handling instruments, and operators. Therefore, by selecting or processing teaching data or images in accordance with these outside factors, known data can be utilized to the maximum extent, expanding the expected area of ​​AI, reducing constraints from equipment or users, and creating a safe and secure world.

[0181] As described above, in one embodiment of the present invention, valuable teaching data can be processed and utilized according to the situation. Therefore, a system capable of responding immediately to required situations can be constructed, enabling the use of advanced AI in various global circumstances to achieve a safe and secure society for all. Furthermore, it can assist in trouble-free output during consumer use and entertainment, effectively supporting high-quality content or creations. Thus, various types of data or high-quality data supported and made available with the assistance of AI become effective teaching data, supporting the realization of such a world.

[0182] In another embodiment of the present invention, the imaging device 6 sends only the image data acquired by the image acquisition device 6 to the image estimation learning device 1. However, it is also possible to generate teaching data by annotating the imaging device 6 and then send the teaching data to the image estimation learning device 1. In this case, if AI correction is required, the image estimation learning device 1 can process the teaching data to correct the estimation model. Furthermore, the imaging device 6 determines whether AI correction is required (see [reference]). Figure 4 (S23), but not limited to the camera device 6, the image estimation learning device 1 may also determine whether AI correction is needed. For example, the image estimation learning device 1 may analyze image data sent from various camera devices 6, compare it with known data, and perform AI correction if it determines that the characteristics (the differences are due to factors, including not only the performance, specifications, environment, and peripheral systems of the device, but also the objects being processed, accessories, peripheral equipment, handling tools, operators, etc.) are different.

[0183] Furthermore, in one embodiment of the present invention, a model is generated by learning using teaching data generated from image data. However, the teaching data is not limited to image data; it can also be generated based on other data, such as important time-series data like body temperature or blood pressure.

[0184] Furthermore, while one embodiment of the present invention primarily describes logic-based decision-making, it is not limited to this; decision-making can also be performed using machine learning estimation. Any of these methods can be used in this embodiment. Additionally, a hybrid decision-making process can be performed, partially utilizing the advantages of each method.

[0185] Furthermore, in one embodiment of the present invention, the control unit 7 or control unit 1a is described as a device composed of a CPU and memory, etc. However, in addition to being configured in software form by a CPU and a program, each part or all of the units may be composed of hardware circuits, or it may be a hardware structure such as gate circuits generated based on a programming language described by a hardware description language (Verilog), or it may utilize a hardware structure using software such as a DSP (Digital Signal Processor). Of course, they can also be appropriately combined.

[0186] Furthermore, the control unit is not limited to a CPU; it can be any element that performs the functions of a controller, or it can be the processing of the aforementioned units performed by one or more processors configured as hardware. For example, each unit can be a processor configured as an electronic circuit, or it can be a circuit unit within a processor composed of integrated circuits such as FPGAs (Field Programmable Gate Arrays). Alternatively, a processor composed of one or more CPUs can also perform the functions of each unit by reading and executing a computer program recorded on a recording medium.

[0187] Furthermore, in one embodiment of the present invention, an image estimation learning device 1 is described as having a control unit 1a, an image input unit 1b, a learning unit 1c, an image processing unit 1d, a learning result utilization unit 1e, a teaching data selection unit 1f, and a recording unit 4. However, it is not necessary to integrate them into a single device; for example, if connected via a communication network such as the Internet, the aforementioned components can be separated. Similarly, an imaging device 6 is described as having an image estimation unit 2, an image acquisition device 3, and a guiding unit 5. However, it is not necessary to integrate them into a single device; for example, if connected via a communication network such as the Internet, the aforementioned components can be separated.

[0188] Furthermore, in recent years, improvements such as the use of artificial intelligence capable of simultaneously determining various judgment criteria and simultaneously executing the branches of the flowchart shown here also fall within the scope of this invention. If the user can evaluate the quality of such control input, the system can learn the user's preferences and customize the implementation shown in this application in a direction suitable for that user.

[0189] Furthermore, most of the controls described in this specification, which are primarily illustrated with flowcharts, can be set via programming, and sometimes are recorded on recording media or in a recording unit. The method of recording to this recording media or recording unit can be recorded at the time of product shipment, using distributed recording media, or downloaded via the Internet.

[0190] Furthermore, in one embodiment of the present invention, a flowchart is used to illustrate the actions in this embodiment, but the order of the processing steps can be changed. In addition, any step can be omitted, steps can be added, and the specific processing content within each step can be changed.

[0191] Furthermore, even if the sequence of actions in the claims, description, and drawings is described using words such as "firstly" or "next" for convenience, it does not mean that the actions must be performed in that order unless otherwise specified.

[0192] This invention is not directly limited to the embodiments described above. During implementation, structural elements can be modified and made more specific without departing from its spirit. Furthermore, various inventions can be formed through appropriate combinations of the multiple structural elements disclosed in the above embodiments. For example, several structural elements, including all those shown in the embodiments, can be deleted. Additionally, structural elements from different embodiments can be appropriately combined.

[0193] Label Explanation

[0194] 1…Image estimation learning device, 1a…Control unit, 1aa…CPU, 1ab…Memory, 1b…Image input unit, 1c…Learning unit, 1d…Image processing unit, 1e…Learning result utilization unit, 1f…Teaching data selection unit, 2…Image estimation device, 2IN…Image input unit, 2SN…Estimation change unit, 2AI…Estimation unit, 2OUT…Estimation result output unit, 3…Image acquisition device, 3a…Image acquisition device, 3aa…3D, etc., 3b…Image acquisition device, 4…Recording unit, 4a…Teaching data group A, 4b…Teaching data group B, 5…Guidance unit

Claims

1. A learning device for estimation, comprising: The input unit receives image data from the first image acquisition device; and The learning unit uses teaching data obtained by annotating the image data to learn and thus obtain an estimation model. Its features are, The estimation learning device includes an image processing unit that, when relearning the estimation model for a second image acquisition device with image input characteristics different from the first image acquisition device, processes the image data obtained from the first image acquisition device according to the differences in image input characteristics and sets it as the teaching data. The learning unit relearns using the processed teaching data to obtain an estimation model that can estimate data from the second image acquisition device as input.

2. The estimation learning device according to claim 1, characterized in that, The image processing unit processes the first object image data contained in the image data obtained from the first image acquisition device in a manner suitable for the second object image data contained in the image data obtained from the second image acquisition device.

3. The estimation learning device according to claim 1, characterized in that, The difference in image input characteristics is caused by at least one of the following: camera sensor specifications, camera sensor performance, camera optical characteristics, image processing specifications, processing performance, detection performance, and type of light.

4. The estimation learning device according to claim 1, characterized in that, The image processing unit includes modifying the annotation of the same image in such a way that the image data obtained from the first image acquisition device in the teaching data becomes teaching data corresponding to the differences in the image input characteristics.

5. The estimation learning device according to claim 1, characterized in that, The image data obtained from the first image acquisition device is existing teaching data. The image processing unit performs image processing on the existing teaching data based on the characteristics of the image data from the second image acquisition device.

6. The estimation learning device according to claim 1, characterized in that, The image data obtained from the first image acquisition device is existing teaching data. The image processing unit selects and discards existing teaching data based on the characteristics of the image data from the second image acquisition device.

7. The estimation learning device according to claim 1, characterized in that, The image processing unit processes the image data obtained from the first image acquisition device in the teaching data in a manner that adapts it to the image data from the second image acquisition device.

8. The estimation learning device according to claim 1, characterized in that, The image data from the second image acquisition device belongs to an unknown category.

9. The estimation learning device according to claim 8, characterized in that, The unknown category is determined automatically by artificial intelligence, or it is manually set by the user of the second image acquisition device.

10. The estimation learning device according to claim 8, characterized in that, The determination of belonging to the unknown category is based on the model information of the second image acquisition device and / or the image estimated as the reference image from the image data from the second image acquisition device.

11. The estimation learning device according to claim 1, characterized in that, The image data obtained from the first image acquisition device is existing teaching data. Depending on the intended use of the estimation model, the image processing unit may perform image processing on the existing teaching data or select from the existing teaching data based on that intended use.

12. The estimation learning device according to any one of claims 1 to 11, characterized in that, The image data from the first image acquisition device and the image data from the second image acquisition device are endoscopic image data.

13. An estimation learning method, characterized in that, Input image data from the first image acquisition device. When learning the estimation model for a second image acquisition device with image input characteristics different from the first image acquisition device, the image data obtained from the first image acquisition device is processed and set as teaching data. The teaching data is used for learning, thereby obtaining an estimation model that can estimate data from the second image acquisition device as input.

14. A learning device for estimation, comprising: The input unit receives image data from the first image acquisition device; and The learning unit uses teaching data obtained by annotating the image data to learn and thus obtain an estimation model. Its features are, The estimation learning device includes an image processing unit that, when customizing the estimation model for a second image acquisition device used under conditions different from the first image acquisition device, processes image data obtained from the first image acquisition device by performing selection or annotation based on differences in image acquisition characteristics and sets it as the teaching data. The learning unit uses the processed teaching data to learn, thereby obtaining an estimation model that can estimate using image data from the second image acquisition device as input.

15. An estimation learning method, characterized in that, Input image data from the first image acquisition device. When customizing the estimation model for a second image acquisition device used under conditions different from the first image acquisition device, the image data obtained from the first image acquisition device is processed according to the differences in image acquisition characteristics, including selection or annotation, and set as teaching data. The teaching data is used for learning, thereby obtaining an estimation model that can estimate data from the second image acquisition device as input.

16. A learning device for estimation, comprising: The input unit receives image data from the first image acquisition device; and The learning unit uses teaching data obtained by annotating the image data to learn and thus obtain an estimation model. Its features are, The estimation learning device includes an image processing unit that, when relearning the estimation model for a second image acquisition device with image input characteristics different from those of the first image acquisition device, processes the image data obtained from the first image acquisition device according to the differences in the image input characteristics and sets it as the teaching data, wherein the second image acquisition device outputs image data belonging to an unknown category.

17. The estimation learning device according to claim 16, characterized in that, The processing performed by the image processing unit includes the selection and / or image processing of the teaching data.

18. The estimation learning device according to claim 16, characterized in that, The determination that the image data belongs to the unknown category is based on the image data obtained as a result of image acquisition. The result of image acquisition is based on the differences in device performance, specifications, environment, peripheral systems, as well as the differences in the object being processed, peripheral equipment, handling tools, and operators.

19. The estimation learning device according to claim 16, characterized in that, The determination that the image data belongs to the unknown category is based on the model information of the second image acquisition device and / or the image estimated as a reference image from the image data from the second image acquisition device.

20. The estimation learning device according to claim 16, characterized in that, Regarding image data that belongs to the unknown category, differences in the skill of the physician performing the treatment and / or the treatment instruments used by the physician are reflected in the estimation of the prediction for the treatment.

21. An estimation learning method, characterized in that, Input image data from the first image acquisition device. When relearning the estimation model for a second image acquisition device with image input characteristics different from the first image acquisition device, the image data obtained from the first image acquisition device is processed according to the differences in image input characteristics and set as teaching data, wherein the output of the second image acquisition device is image data belonging to an unknown category. The teaching data is used for learning to obtain an estimated model.

Citation Information

Patent Citations

  • Teacher data collection apparatus, teacher data collection method and program

    JP2018124617A

  • Learning system, learning device, learning method, learning program, teacher data creation device, teacher data creation method, teacher data creation program, terminal device, and threshold value changing device

    CN108351986A

  • Machine learning device, teacher data generation device, inference model, and teacher data generation method

    JP2020035094A