Program, information processing apparatus, and information processing method

The program addresses the challenge of identifying learning model discrepancies by using a language generation model to explain causes and solutions, improving accuracy and efficiency in inspections and reviews.

JP2025100091APending Publication Date: 2025-07-03H U GROUP HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023217197
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing learning models struggle with identifying the cause of discrepancies between their output and human judgments, requiring significant labor and expertise, especially for low-frequency causes, and existing systems only provide warnings without explaining the reasons for inaccuracies.

Method used

A program that acquires prompts about the differences between model and user results, using a language generation model to output the cause and potential solutions, integrating with a classification and detection model to facilitate understanding and improve model accuracy.

Benefits of technology

Enables identification of the cause and solution for discrepancies, enhancing model accuracy through user-AI collaboration, even with inexperienced personnel, and improving work efficiency in inspections and report reviews.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100091000001_ABST
    Figure 2025100091000001_ABST
Patent Text Reader

Abstract

To provide a program, an information processing apparatus, and an information processing method configured to output a solution and the cause of a difference between a result output by a learning model and a result intended by a user.SOLUTION: A program causes a computer to execute the processes of: acquiring, when there is a difference between a first result obtained from an image recognition model trained to execute recognition processing on an input image and a second result intended by a user for the input image, a prompt including the first result and a question to ask the cause of the difference; and inputting, when the prompt is input, the acquired prompt to a language generative model configured to output the cause of difference between results, to output the cause.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a program, an information processing apparatus, and an information processing method that output the cause of incompleteness in a learning model or its user.

Background Art

[0002] In recent years, learning models have been used as main components of artificial intelligence (AI). To improve the accuracy of a learning model, it is necessary to create a large number of teacher data, but it is not easy to create. Therefore, when a certain level of accuracy is obtained, operation is started, and in the operation, the results output by the learning model are appropriately confirmed by a person. Then, when there is a difference between the results output by the learning model and the results judged by a person, the cause of the difference is identified, and the parameters of the learning model are adjusted or the learning model is relearned to improve the accuracy.

[0003] However, it is not easy for inexperienced personnel to identify the cause of the difference between the results output by the learning model and the results judged by a person, and it requires a great deal of labor. In addition, it is assumed that many types of causes with low occurrence frequencies will accumulate, and it may not be possible to obtain sufficient cost-effectiveness, such as requiring personnel for information maintenance and management.

[0004] Regarding the difference in results, Patent Document 1 discloses that in a visual recognition support device that supports visual recognition by a driver of a vehicle, when it is determined that the driver's visual recognition is inappropriate, warning information for causing a warning to be generated by a warning output device is output.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] However, the visual recognition support device disclosed in Patent Document 1 only generates a warning and does not output the cause of inappropriate visual inspection. The present invention has been made in view of such a situation. Its object is to provide a program, an information processing apparatus, and an information processing method that output the cause and solution when the result output by the learning model differs from the result intended by the user.

Means for Solving the Problems

[0007] When the first result by an image recognition model learned to execute recognition processing on an input image differs from the second result by a user for the input image, a program according to an aspect of the present application acquires a prompt including a question asking about the first result and the cause of the difference, and causes a computer to execute a process of outputting the cause by inputting the acquired prompt to a language generation model that outputs the cause of the difference when the prompt is input.

Effects of the Invention

[0008] According to one aspect of the present invention, when the result output by the learning model differs from the result intended by the user, it becomes possible to output the cause and the solution.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Mode for Carrying Out the Invention

[0010] (Embodiment 1) The following embodiments will be described with reference to the drawings. FIG. 1 is an explanatory diagram showing a configuration example of an information processing system. The information processing system 100 includes an information processing apparatus 1, a camera 2, and an interactive service 3. The information processing apparatus 1 and the interactive service 3 are communicably connected by a network such as a public telephone network or a mobile phone communication network.

[0011] The information processing apparatus 1 is configured using a personal computer, a server computer, or the like. The information processing apparatus 1 includes a processing unit (control unit) 11, an input / output unit 12, a storage unit 13, a communication unit 14, an operation unit 15, and a display unit 16. The processing unit 11 is configured using an arithmetic processing device such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), or a GPU (Graphics Processing Unit), a ROM (Read Only Memory), and a RAM (Random Access Memory).

[0012] In the processing unit 11, by the processing unit 11 reading and executing the program 1P stored in the storage unit 13, an acquisition unit 11a, a determination unit 11b, a display output unit 11c, a result acquisition unit 11d, a detection unit 11e, a model update unit 11f, a registration update unit 11h, etc. (see FIG. 7) are realized as software functional units.

[0013] The input / output unit 12 performs input / output of data with the camera 2. The input / output unit 12 is connected to the camera 2 via, for example, a signal line, and performs input / output of data by serial communication, parallel communication, or the like via the signal line. The input / output unit 12 transmits data such as control commands given from the processing unit 11 to the camera 2, and gives the image data input from the camera 2 to the processing unit 11.

[0014] The storage unit 13 is configured using a non-volatile memory element such as a flash memory or an EEPROM (Electrically Erasable Programmable Read Only Memory). The storage unit 13 stores various programs executed by the processing unit 11 and various data necessary for the processing of the processing unit 11. The storage unit 13 stores a cause DB 131, an inquiry DB 132, and a training DB 133. The storage unit 13 also stores a classification model M1 (image recognition model) and a detection model M2. In the present embodiment, the storage unit 13 stores a program 1P executed by the processing unit 11. The storage unit 13 may also store data of images captured by the camera 2.

[0015] In the present embodiment, the program 1P is written into the storage unit 13, for example, at the manufacturing stage of the information processing apparatus 1. Further, for example, the program 1P may be acquired by the information processing apparatus 1 through communication of what is distributed by a remote server device or the like. Further, for example, the program 1P may be provided in a form recorded on a recording medium such as a memory card or an optical disk, and the information processing apparatus 1 may read the program 1P from the recording medium and store it in the storage unit 13. Further, for example, the program 1P recorded on the recording medium may be read by a writing device and written into the storage unit 13 of the information processing apparatus 1. The program 1P may be provided in a form of distribution via a network or in a form recorded on a recording medium.

[0016] The communication unit 14 communicates with various devices via a network N such as a mobile phone communication network, a wireless LAN (Local Area Network), or the Internet. In the present embodiment, the communication unit 14 communicates with the dialogue service 3. The communication unit 14 transmits the data given from the processing unit 11 to other devices and gives the data received from other devices to the processing unit 11.

[0017] The operation unit 15 is a hardware keyboard, a mouse, etc. The display unit 16 includes a liquid crystal display panel, an organic EL (electro Luminescence) display panel, or the like. The operation unit 15 and the display unit 16 may be integrated to form a touch panel display. Note that the information processing apparatus 1 may perform display on an external display device.

[0018] Note that the information processing apparatus 1 may be configured as a multi-computer composed of a plurality of computers, a virtual machine virtually constructed by software, or a quantum computer. Further, the functions of the information processing apparatus 1 may be realized by a cloud service.

[0019] The camera 2 has a lens, an image sensor, etc., and acquires image data of a subject image through the lens. The camera 2 performs shooting according to an instruction from the information processing apparatus 1 and outputs the acquired image data to the information processing apparatus 1. The camera 2 may be configured to be externally attached to the information processing apparatus 1 or may be built into the information processing apparatus 1.

[0020] The dialogue service 3 includes a dialogue server 31 and an LLM 32. Under the control of the dialogue server 31, the dialogue service 3 provides an interactive AI (Artificial Intelligence) service using the LLM 32 (Large Language Models), a language generation model. The LLM 32 is a natural language processing model trained using a large amount of text data. The dialogue service 3 is configured using Transformer, BERT, GPT-3, ChatGPT, BARD, etc. Note that the information processing device 1 may be equipped with the LLM 32 and provide the dialogue service 3.

[0021] The images handled in this embodiment are those obtained by photographing the object to be confirmed with the camera 2 in an inspection where the inspection results are visually confirmed among clinical inspections performed at the medical site. In the following description, a case will be described where the results visually confirmed and determined by a person are compared with the results determined using the classification model M1 in an inspection using a test kit based on immunoelectrophoresis.

[0022] Next, the database used by the information processing apparatus 1 will be described. FIG. 2 is an explanatory diagram showing an example of the cause DB. The cause DB 131 stores, in association with each other, information related to inspections, the causes thereof, solutions, etc. in cases where there are differences (disagreement cases) between visual inspections and the determinations of the classification model M1 in past inspections. The cause DB 131 includes an inspection ID column, an inspection information column, a disagreement content column, a category column, a cause column, a detailed cause column, a solution column, and a solution rate column. The inspection ID column stores an inspection ID that can uniquely identify an inspection. The inspection information column stores information related to the inspection, such as inspection target information including the determination value of the classification model M1 and its confidence level (Probability) and the output value of the detection model M2 (Probability of abnormality). The disagreement content column stores the content of the disagreement in the determination. The category column stores the category of the cause of the disagreement. For example, the categories are five: "user", "image", "AI", "client, etc.", and "others". "User" indicates that there was a factor with the visual inspector. For example, the visual inspector had insufficient skills or was careless. "Image" indicates that there was a factor with the captured image. For example, the lighting conditions were inappropriate, a shadow was reflected, the shooting angle was inappropriate, the image quality was inappropriate, or there was noise in the image. "AI" indicates that it was due to the learning data of the classification model M1 and the classification model M1 clearly misrecognized. "Client, etc." indicates that an object outside the scope was photographed. "Others" includes bugs, power outages, etc. The cause column stores the cause of the disagreement. The detailed cause column stores the detailed content of the cause of the disagreement. The solution column stores the solution for each cause. The solution rate column stores the solution rate for each solution.

[0023] FIG. 3 is an explanatory diagram showing an example of an inquiry DB. The inquiry DB 132 stores information necessary when inquiring about the cause of a discrepancy or the like. The inquiry DB 132 includes a content column, a destination name column, a destination column, a method column, a response time column, and a resolution rate column. The content column stores the content of the discrepancy. The destination name column stores the name of the inquiry destination. The destination column stores the destination. When the inquiry method is e-mail, the destination is an e-mail address. When the inquiry method is chat, the destination is @+(account name). The method column stores the inquiry method. The response time column stores the average value of the response time from when an inquiry is sent to the destination until a response thereto is received. The resolution rate column stores the probability that the cause of the discrepancy has been resolved by the response from the destination.

[0024] FIG. 4 is an explanatory diagram showing an example of a training DB. The training DB 133 stores training data for the classification model M1. As a result of pursuing the cause of the discrepancy, if it is determined that the error of the classification model M1 is the cause, training data is created and the classification model M1 is relearned. The training DB 133 includes an image column and a determination result column. The image column stores the input image or the file name of the input image. The determination result column stores the correct value (label) of the determination result.

[0025] FIG. 5 is an explanatory diagram showing a configuration example of a classification model. The classification model M1 is configured using various object detection algorithms and object classification algorithms (both neural networks) such as, for example, CNN (Convolution Neural Network), R-CNN (Region-based CNN), Fast R-CNN, Faster R-CNN, Mask R-CNN, SSD (Single Shot Multibook Detector), YOLO (You Only Look Once). The classification model M1 shown in FIG. 5 has an input layer, a feature extraction unit M11, a fully connected layer M12, and an output layer M13.

[0026] In the classification model M1 used for the test using immunoelectrophoresis, in the prediction process, a captured image including the test results of a specimen from a healthy person and the test results of a specimen from a patient is input through the input layer. The input image may be a captured image of the test kit, or an image obtained by performing a reduction process on the captured image. The input layer has the number of input nodes corresponding to the number of pixels of the input image. Each pixel in the captured image of the test kit is input to each input node of the input layer, and the captured image input through the input layer is input to the feature extraction unit M11. The feature extraction unit M11 includes a convolutional layer and a pooling layer. The captured image input to the feature extraction unit M11 has its image features extracted by filter processing or the like in the convolutional layer to generate a feature map, and is compressed in the pooling layer to reduce the amount of information. The convolutional layer and the pooling layer are provided repeatedly in multiple layers, and the feature maps generated by the multiple convolutional layers and pooling layers are output to the fully connected layer M12. The fully connected layer M12 is provided in multiple layers. Based on the input feature map, the output values of the nodes in each layer are calculated using various functions, thresholds, etc., and the calculated output values are sequentially input to the nodes in the subsequent layer. The fully connected layer M12 finally outputs the output value to the output layer M13 in the subsequent stage by sequentially inputting the output values of the nodes in each layer to the nodes in the subsequent layer.

[0027] The output layer M13 is composed of multi-class classifiers such as the softmax function, random forest, and Gradient Boosting. In the classification model M1, the output layer M13 has 12 softmax functions corresponding to each of the 12 protein components to be inspected, the first softmax function M131, the second softmax function M132, and so on. For example, the first softmax function M131 is the softmax function for IgG, and the second softmax function M132 is the softmax function for IgA. In the prediction process, the output values from the fully connected layer M12 are input into the respective softmax functions M131, M132, and so on. The softmax functions M131, M132, and so on take the output values from the input fully connected layer M12 as inputs (arguments), calculate the output values using a predetermined function, and output the calculated output values from each output node. One softmax function M131, M132, and so on has 5 output nodes. Since the output layer M13 is provided with 12 softmax functions M131, M132, and so on, it has a total of 60 output nodes. For example, the 5 output nodes connected to the first softmax function M131 for IgG are nodes that output the test results of the test target (patient sample) with respect to the test criteria (reference values based on the test results of samples from healthy individuals) in the input image for IgG. Each node corresponds to the detection level of IgG (level 0 to level 4). Specifically, for example, node 2 outputs the probability of being judged normal at level 3 (the test target is about the same as the test criteria). Node 1 outputs the probability of being judged at level 2 (the test target is slightly decreased compared to the test criteria). Node 0 outputs the probability of being judged at level 1 (the test target is decreased compared to the test criteria). Node 3 outputs the probability of being judged at level 4 (the test target is slightly increased compared to the test criteria). Node 4 outputs the probability of being judged at level 5 (the test target is increased compared to the test criteria). Each of the 5 output nodes outputs an output value between 0 and 1.0, and the sum of the output values (probabilities) output from the 5 output nodes is 1.0.That is, in the output example shown in FIG. 5, as the IgG level, an output value of 0.85 is output from node 2 corresponding to "normal", and output values from 0 to 0.09 are output from the other output nodes 0, 1, 3, and 4. Therefore, it indicates that the detected amount of IgG to be inspected here should be judged as "normal" with respect to the inspection standard.

[0028] Similarly, the five output nodes connected to the second softmax function M132 for IgA are nodes that output the inspection results of the inspection target with respect to the inspection standard in the input image for IgA. Specifically, for example, node 2 outputs the probability that the inspection target should be judged as normal. Node 1 outputs the probability that the inspection target should be judged as slightly decreased compared to the inspection standard. Node 0 outputs the probability that the inspection target should be judged as decreased compared to the inspection standard. Node 3 outputs the probability that the inspection target should be judged as slightly increased compared to the inspection standard. Node 4 outputs the probability that the inspection target should be judged as increased compared to the inspection standard. Therefore, in the output example shown in FIG. 5, as the IgA level, an output value of 0.85 is output from node 3 corresponding to "slightly increased", and output values from 0 to 0.09 are output from the other output nodes 0, 1, 2, and 4. Therefore, it indicates that the detected amount of IgA to be inspected here should be judged as "slightly increased" with respect to the inspection standard.

[0029] The classification model M1 learns using the training data stored in the training DB133. The training data includes a photographed image and a determination result. The photographed image is an image including the inspection results (inspection standards) of specimens from healthy individuals and the inspection results (inspection targets) of specimens from patients. The determination result is the correct label. Here, the correct label is information indicating the inspection results (any one of the five patterns of decrease, slightly decreased, normal, slightly increased, and increased) of the inspection target with respect to the inspection standard in the photographed image for each of the 12 types of inspection items.

[0030] FIG. 6 is an explanatory diagram showing a configuration example of a detection model. When estimating the cause of the difference between visual determination and the determination of the classification model M1, the detection model M2 is used to determine whether there is a problem with the captured image. The detection model M2 can adopt any of three types of machine learning models: 1. a machine learning model that performs supervised learning, 2. a machine learning model that performs self-supervised learning, and 3. a general-purpose AI model. FIG. 6A shows a configuration example of the detection model M2 when adopting a learning model that performs supervised learning or self-supervised learning. FIG. 6B shows a configuration example of the detection model M2 when adopting a general-purpose AI model.

[0031] The case of adopting a machine learning model that performs supervised learning as the detection model M2 will be described. The detection model M2 is a learning model that outputs an interpretation (such as an inspection score) of the image when the image is input. The detection model M2 is configured using various object detection algorithms and object classification algorithms (both neural networks) such as CNN, R-CNN, Fast R-CNN, Faster R-CNN, Mask R-CNN, SSD, YOLO (You Only Look Once), etc.

[0032] The detection model M2 learns using training data (teacher data) including an image and an interpretation of the image. The image is a normal captured image including inspection results from specimens of healthy individuals and inspection results from specimens of patients, or a captured image with a problem caused by operations during inspection or imaging. The interpretation of the image is an inspection score, which is at 12 levels here.

[0033] The case of adopting a machine learning model that performs self-supervised learning as the detection model M2 will be described. The detection model M2 is a learning model that outputs an interpretation (normal or abnormal) of an input image. The detection model M2 is composed of a CNN or the like. The detection model M2 is trained using training data including an image and an interpretation of the image. The image is a normal captured image including examination results of specimens from healthy subjects and examination results of specimens from patients, or an artificially created image including artificially created abnormalities obtained by adding artificial noise (overwriting lines and dots) or image conversion (extreme changes in brightness and contrast, etc.) to a normal captured image. The interpretation of the image is a value indicating whether the input image is normal or abnormal.

[0034] The detection model M2 is trained using training data consisting of an image and a label indicating whether the image is normal or abnormal. When a normal image is input to the detection model M2, the detection model M2 is trained to output a value indicating normality. Also, when an abnormal image artificially created from a normal image is input to the detection model M2, the detection model M2 is trained to output a value indicating abnormality.

[0035] The case of adopting a general-purpose AI as the detection model M2 will be described. For example, as shown in FIG. 6B, the detection model M2 is composed of an Image Captioning model M21 and a text interpretation model M22. When the Image Captioning model M21 receives an image as input, it outputs text that describes the image. The Image Captioning model M21 is configured using, for example, the SmartLens Image captioning API, the Microsoft Describe Image API, ClipCap, BLIP, or the like. The description text of the image output by the Image Captioning model M21 is input to the text interpretation model M22. The text interpretation model M22 interprets the description text. If the input description text contains words or expressions indicating problems (lighting, shadows, blurring, etc.) related to the image itself, the text interpretation model M22 outputs "problematic", and if not, it outputs "no problem".

[0036] Next, an image inspection using an information processing system will be described. Here, it is an inspection using the immunoelectrophoresis method as described above. The visual judge (user) prepares a specimen collected from a patient and a specimen collected from a healthy person, and drops them into the respective dropping areas of two test kits. Then, the visual judge applies voltage to the two test kits simultaneously. After applying voltage for a predetermined time, anti-human whole serum is added, and the protein components in the specimen and the anti-human whole serum are allowed to freely diffuse and an immunoprecipitation reaction is caused. After removing the unreacted protein components, staining is performed using a staining solution. After drying, imaging is performed. The visual judge takes pictures of the electrophoresis results of the specimen of the healthy person and the electrophoresis results of the specimen of the patient with the camera 2 and makes a determination. In addition, when the determination work cannot be performed immediately, the photographed images of the electrophoresis results of the specimen of the healthy person and the photographed images of the electrophoresis results of the specimen of the patient are stored in the storage unit 13 or the like, and the photographed images are read out and displayed on the display unit 16 at the time of the determination work, and the determination may be made.

[0037] The operation of the information processing device 1 during the determination work will be described. FIG. 7 is an explanatory diagram showing the relevance between the functional parts of the information processing device. The acquisition unit 11a acquires an image from the camera 2 via the input / output unit 12. In addition to acquiring an image from the camera 2, the acquisition unit 11a may acquire an image stored in the storage unit 13 or the like. The acquisition unit 11a inputs the acquired image to the classification model M1. The determination unit 11b outputs a determination result based on the output of the classification model M1 to the detection unit 11e. The acquisition unit 11a also outputs the acquired image to the display output unit 11c. The display output unit 11c displays the image on the display unit 16 to prompt the visual judge to make a visual determination. The visual judge inputs the determination result via the operation unit 15. The result acquisition unit 11d acquires the determination result and outputs it to the detection unit 11e. The detection unit 11e detects the difference between the determination result (the first result) by the classification model M1 and the determination result (the second result) by the visual observation of the visual judge. When the detection unit 11e does not detect a difference, the detection unit 11e outputs the image and the determination result to the model update unit 11f. The model update unit 11f stores the image and the determination result as training data in the training DB133 and performs re-learning of the classification model M1.

[0038] When the detection unit 11e detects a difference, it outputs the image, the output of the classification model M1, and the difference in the determination results between the classification model M1 and the visual inspector to the estimation unit 11g. The estimation unit 11g estimates the cause. The cause is estimated according to the following procedure.

[0039] The estimation unit 11g determines whether the visual inspector was careless based on the output value of the classification model M1. When the output value is a confidence level that takes a value between 0 and 1, if it is equal to or greater than a threshold value (for example, 0.5), it is interpreted that the determination result (the first result) by the classification model M1 is more reliable than the determination result (the second result) by visual inspection of the visual inspector, and as the cause of the difference in these determination results, "Cause candidate 1: Possibility of visual error" (user's judgment error) is determined. The estimation unit 11g inputs the image used for the determination to the detection model M2 and determines whether there is a problem with the image. When the output of the detection model M2 is the confidence level (between 0 and 1) that the image is abnormal, if it is equal to or greater than a threshold value (for example, 0.5), it is interpreted that the captured image itself contains a problem, and as the cause of the difference in the determination results, "Cause candidate 2: Problem with the captured image" is determined.

[0040] If neither cause candidate 1 nor cause candidate 2 is determined, the estimation unit 11g reads out cause candidates (cases) from the cause DB131 as cause candidates other than the above. The candidates to be read out may be 10 typical causes (cases) extracted from all past mismatch cases.

[0041] The LLM32 can be used for extracting typical causes. For example, the following prompt is sent to the dialogue service 3.

[0042] You are an expert in cause analysis. When the prediction result of the AI and the determination of the visual inspector are different, the following cause candidates have been raised. Please classify them into about 4 categories and summarize them. *** "Problem of the visual inspector" "Problem with the captured image" "Problem of interpretation" "Problem of the system" "Lack of experience of the inspector" "Noise in the captured image" "Bug" "Power outage" "Device failure"

[0043] The dialogue service 3 returns the following answers to the above prompts.

[0044] The above candidate causes indicate factors that may be considered when explaining the differences between different prediction results and the judgments of visual judges. These factors are categorized and summarized for each factor. 1. Factors related to the judge: · Problems with visual judgment · Lack of experience of the judge These factors are problems related to the judge himself / herself, and... 2. Factors related to the image: · Problems with the captured image · Noise in the captured image These factors are... … 3. Factors related to interpretation: …

[0045] Based on the answers from the dialogue service 3, it is possible to determine typical causes.

[0046] The estimation unit 11g estimates the cause using the dialogue service 3. Prior to estimating the cause, the estimation unit 11g performs Few-examplers (few exemplars) for the purpose of improving the estimation accuracy of the dialogue service 3. Specifically, the correspondence relationship between typical causes and their inspection information is input as the condition ("role" of "system") of the dialogue service 3 to cause the LLM32 to perform Few-shot learning (learning with a small amount of data). The following is an example of the input.

[0047] system=“AI-1 IgA Score=-2 (Probability=0.4), AI-2 Probability of abnormality =0.3, Info / Sample / Comment=Insufficient samples, cause category=II. Problems with the captured image, cause=Insufficient samples”

[0048] AI-1 corresponds to the classification model M1. Probability is the confidence level output by the classification model M1. AI-2 corresponds to the detection model M2. Probability is the confidence level that the image output by the detection model M2 is abnormal.

[0049] For the number of cause candidates extracted from the cause DB131, the estimation unit 11g inputs the condition ("role" of "system") to the dialogue service 3. When using 10 typical causes, the estimation unit 11g inputs the condition ("role" of "system") 10 times to the dialogue service 3.

[0050] Next, after inputting the inspection information as a question ("role" of "user") of the dialogue service 3, finally, input the question about the cause to estimate the cause. The cause category is selected by checking if there is a corresponding one in the list of past causes. The following shows an input example.

[0051] user = "AI-1 IgM Score=-2 (Probability=0.2), AI-2 Probability of abnormality =0.1, Info / Sample / Comment=Lot change of inspection reagent, cause category and cause=?"

[0052] The estimation unit 11g obtains the answer from the dialogue service 3. The following shows an answer example.

[0053] assistant = "Cause category=II. Problem with imaging image, Cause=The staining is too light to see clearly"

[0054] The estimation unit 11g displays the estimated cause output by the dialogue service 3 on the display unit 16. After the estimated cause is displayed on the display unit 16, if the visual inspector inputs a question via the prompt, the dialogue service 3 can be interactively responded to, and the cause and its details can be narrowed down in a dialogue format. The visual inspector refers to the estimated cause and attempts to resolve the cause. For example, three actions can be taken as the attempt (solution method). 1. Re-visual inspection, 2. Re-photography, 3. Inquiry.

[0055] 1. In re-visual inspection, the visual inspector visually inspects the image again and makes a determination. The visual inspector re-enters the determination result via the operation unit 15. The result acquisition unit 11d acquires the determination result and outputs it to the detection unit 11e. The subsequent operations are the same as described above, but the operation of the detection unit 11e is different.

[0056] In the attempt to resolve the cause, when the detection unit 11e determines that there is no difference between the determination result by the classification model M1 and the determination result by the visual inspection of the visual inspector, it outputs that fact to the registration update unit 11h. The registration update unit 11h outputs the image and the determination result to the model update unit 11f. Also, the registration update unit 11h stores the determined cause in the cause DB 131 and updates the resolution rate for each cause.

[0057] 2. In re-photography, settings of the camera 2, change of lighting, etc. are performed, and the image of the inspection kit is photographed again. The acquisition unit 11a acquires the image from the camera 2 via the input / output unit 12. The subsequent operations are the same as in the case of 1. Re-visual inspection.

[0058] 3. In the inquiry, following the displayed advice, the visual inspector instructs the dialogue service 3 to make an inquiry via mail or chat. The dialogue service 3 executes mail sending / receiving and chat sending / receiving, asynchronously answers the operator with the inquiry result, and prompts for reconfirmation. The visual inspector refers to the answer and attempts to resolve the cause.

[0059] When the cause of the inconsistency is displayed as "possibility of visual error", if the visual inspector does not understand the reason, for example, enter a question such as "Since I don't understand the reason for the visual error, may I ask you to inquire?" into the dialogue service 3. In response, since User ID2 made a similar determination, obtain an answer (inquiry proposal) from the dialogue service 3 such as "Shall I send an email inquiry to User ID2?" In response, the visual inspector enters a request into the dialogue service 3 such as "Please create the text for sending an email inquiry to User ID2. I want a reply by mid-morning." In response, the dialogue service 3 outputs the text of the email. The dialogue service 3 and the email application may be linked, and the text and destination output by the dialogue service 3 may be passed to the email application, and the creation and sending of the email may be performed without going through the operation of the visual inspector.

[0060] The operation regarding the trial is the same as described above, but the operation of the registration update unit 11h is added. When the cause is resolved after the inquiry, the registration update unit 11h updates the response time and resolution rate stored in the inquiry DB 132.

[0061] FIG. 8 is a flowchart showing an example of a business process procedure. The business process is a process performed by the information processing apparatus 1 when a visual inspector visually determines an inspection result. The processing unit 11 of the information processing apparatus 1 acquires an image (step S1). The processing unit 11 displays the acquired image on the display unit 16 (step S2). The processing unit 11 inputs the image into the classification model M1 and obtains a determination result (step S3). The visual inspector refers to the image displayed on the display unit 16 and performs a visual determination. The visual inspector inputs the determination result via the operation unit 15. The processing unit 11 receives the determination result of the visual inspector (step S4). The processing unit 11 determines whether the determination result of the classification model M1 matches the determination result of the visual inspector (step S5). When the processing unit 11 determines that the determination result of the classification model M1 matches the determination result of the visual inspector (YES in step S5), the processing unit 11 stores the image and the determination result in the training DB133 (step S6). The processing unit 11 performs re-learning of the classification model M1 (step S7). When the processing unit 11 determines that the determination result of the classification model M1 does not match the determination result of the visual inspector (NO in step S5), the processing unit 11 performs cause estimation (step S8).

[0062] FIG. 9 is a flowchart showing an example of a procedure for cause estimation processing. The processing unit 11 determines whether the cause is due to the visual inspector (step S31). If the output value of the classification model M1 is a confidence level taking a value between 0 and 1 and this value is equal to or greater than a threshold value (for example, 0.5), the processing unit 11 determines that the cause is due to the visual inspector. Otherwise, the processing unit 11 determines that the cause is not due to the visual inspector. When the processing unit 11 determines that the cause is due to the visual inspector (YES in step S31), it sets "possibility of visual error" as the cause (step S32). When the processing unit 11 determines that the cause is not due to the visual inspector (NO in step S31), it determines whether the cause is due to the image (step S33). The processing unit 11 inputs the image used for the determination to the detection model M2. If the output of the detection model M2 is a confidence level taking a value between 0 and 1 that the image is abnormal and this value is equal to or greater than a threshold value (for example, 0.5), the processing unit 11 determines that it is due to the image. Otherwise, the processing unit 11 determines that the cause is not due to the image. When the processing unit 11 determines that the cause is due to the image (YES in step S33), it sets "problem with the captured image" as the cause (step S34). When the processing unit 11 determines that the cause is not due to the image (NO in step S33), the processing unit 11 reads out cause candidates (step S35). For example, the processing unit 11 reads out 10 typical causes extracted from all past mismatch cases from the cause DB 131. The processing unit 11 causes the LLM 32 to perform few-shot learning based on the cause candidates (step S36). The processing unit 11 estimates the cause (step S37). It inputs a cause question to the dialogue service 3 and obtains the cause estimated by the dialogue service 3. The processing unit 11 returns the process.

[0063] The processing unit 11 displays the estimated cause on the display unit 16 (step S9). The visual inspector refers to the estimated cause and attempts to resolve the cause. When the visual inspector performs re-visual inspection, the determination result is input again via the operation unit 15. Upon receiving the input, the processing unit 11 determines whether the attempt is a re-visual inspection determination (step S10). If the processing unit 11 determines that it is a re-visual inspection determination (YES in step S10), the process returns to step S4. If the processing unit 11 determines that it is not a re-visual inspection determination (NO in step S10), it determines whether to perform re-photography (step S11). If the processing unit 11 determines that it is re-photography (YES in step S11), the process returns to step S1. If the processing unit 11 determines that it is not re-photography (NO in step S11), an inquiry is made (step S12). The processing unit 11 instructs the interactive service 3 to make an inquiry by email or chat according to the instruction input of the visual inspector. The processing unit 11 receives the answer and displays it on the display unit 16 (step S13). The visual inspector refers to the answer and attempts to resolve the cause. The processing unit 11 returns the process to step S10.

[0064] FIG. 10 is an explanatory diagram showing an example of a determination result screen. The determination result screen d01 is a screen that displays the determination result of the visual inspector and the determination result of the classification model M1. When the two determination results do not match, it is a screen that displays the estimated cause and prompts for attempts to solve the cause. The determination result screen d01 includes a selection item display d011, a determination result display d012, an estimated cause display d013, an output area d014, an input area d015, and an imaging button d016. The selection item display d011 indicates the selected item among a plurality of inspection items. In the example of FIG. 10, "IgA (immunoglobulin A)" is selected. The determination result display d012 displays the determination result of the visual inspector and the determination result of the classification model M1. Also, the difference is displayed. In the example of FIG. 10, since the determination value of the visual inspector (inspector) is lower than the determination value of the classification model M1, a filled triangle is displayed as the difference. When the determination value of the visual inspector is higher than that of the classification model M1, an unfilled triangle is displayed. The estimated cause display d013 displays the cause (estimated cause) for which the determination result of the visual inspector and the determination result of the classification model M1 do not match. The output area d014 displays the output of the prompt. The "prompt" here is not the prompt in the conventional information technology field, "a short sequence of characters or symbols indicating that the system is in a state where it can accept input in a command-line interface where the user types commands in text to the computer", but the prompt in prompt engineering. The input area d015 is an area for inputting an instruction to the dialogue service 3. When the imaging button d016 is selected, imaging by the camera 2 is performed.

[0065] The present embodiment has the following effects. When there is a difference between the result output by the learning model and the result intended by the user, it becomes possible to output the cause and the solution. As a result, even if the accuracy of the determination by an inexperienced person or AI (classification model M1) is low, it becomes possible for the inexperienced person and AI to cooperate to achieve a higher accuracy than their respective accuracies. By registering the result of the visual determination, it is possible to create learning data for improving the accuracy of the classification model M1.

[0066] When estimating the cause using the dialogue service 3, the reason for not passing the image to the dialogue service 3 is that in the dialogue service 3 using the LLM32, it does not accept images as a specification. Also, even if the dialogue service 3 can accept images, no effect can be expected. Therefore, instead of passing the image itself to the dialogue service 3, the feature amount of the image that can be expressed in text is passed. Note that this is not the case when the dialogue service 3 becomes multimodal and effects such as an improvement in the accuracy of cause estimation are expected even when an image is passed.

[0067] (Embodiment 2) This embodiment relates to the inspection of surgical instruments. In medical institutions, in order to clean and sterilize surgical instruments contaminated with blood or the like after surgery, in the operation of collecting surgical instruments, the type and number of surgical instruments are visually confirmed (inspected). Thereby, it is indirectly confirmed that no surgical instrument remains in the body, and it is confirmed that the instruments have been received. The process of the person in charge of collecting surgical instruments performing the inspection is being streamlined using image judgment AI.

[0068] In this embodiment, in order to ensure the reliability of the inspection results, the judgment result of the person in charge and the judgment result of the image judgment AI are compared and verified. If there is a difference between the two results, reconfirmation and verification work are performed.

[0069] The reasons for the difference between the judgment result of the person in charge and the judgment result of the image judgment AI are considered to be somewhat limited. Examples of the reasons include the influence of contaminants on imaging, surgical instruments overlapping, a part of the instrument being hidden, imaging errors, and misidentifications. However, since it takes time for inexperienced personnel to identify the cause, the inspection takes time. Therefore, if the cause can be easily understood, the confirmation work becomes easier and the work efficiency improves dramatically. Note that since surgical instruments are collected after inspection, this inspection time has a great impact on work efficiency when collecting a large number within a limited time from multiple facilities.

[0070] It is assumed that the cause of the difference is often revealed by reconfirmation. Therefore, the cause of the difference is stored in the cause DB131 as natural language of meta information (examiner name, set consisting of individual types or multiple combinations of surgical instruments, quantity, etc.). If the cause of the difference cannot be found in the re-inspected products or the like, an inquiry is made.

[0071] The business process in this embodiment is the same as that in Embodiment 1 (Fig. 8). Hereinafter, it will be described with reference to Fig. 8. When the collection of the instruments is completed, the person in charge takes a picture of the collected instruments with the camera 2. The processing unit 11 of the information processing apparatus 1 acquires the image (step S1). The processing unit 11 displays the acquired image on the display unit 16 (step S2). The processing unit 11 inputs the image to the image determination AI (corresponding to the classification model M1) and obtains a determination result including the type and quantity of the instruments (step S3). The person in charge visually checks the collected instruments and inputs the type and quantity of the instruments. The processing unit 11 acquires the determination result of the person in charge (step S4). The processing unit 11 determines whether the results match (step S5). If the processing unit 11 determines that the results match (YES in step S5), the image and the results are stored in the training DB133 (step S6), and re-learning is performed (step S7). If the processing unit 11 determines that the results do not match (NO in step S5), it reads the cause DB131, and the dialogue service 3 asks about the presumed cause (corresponding to steps S8 and S9). The processing unit 11 uses the past data stored in the cause DB131 to train the dialogue service 3 by few-shot learning. The data includes the output of the image determination AI such as the type and quantity of the instruments, and the cause of the mismatch. Thereafter, the processing unit 11 inputs the output of the current image determination AI and a question asking about the cause of the mismatch to the dialogue service 3, and obtains the cause estimated from the dialogue service 3. If the cause is known, the person in charge re-verifies it or tells the verification result to the dialogue service 3 to record the progress (corresponding to steps S10 to S13).

[0072] In this embodiment, even an inexperienced operator can perform inspection in the same amount of time as an expert, improving work efficiency. As a result, it becomes possible to efficiently collect instruments from multiple facilities within a limited time.

[0073] (Embodiment 3) This embodiment relates to the proofreading and review of clinical test result reports. In a clinical test result report, multiple test results may be described in a single report, and some reports are described in natural language. In such reports, due to carelessness of a pathologist or the like, there may be contradictions in the description content or content that does not match existing knowledge. Therefore, an inspector checks the content, requests confirmation from the pathologist, and makes corrections as necessary. Such work requires skill and reading ability, so the confirmation work is not easy. In this embodiment, the work is targeted at detecting contradictions in the content described in the report or contradictions with past content.

[0074] In this embodiment, when the input data is a paper report, the area to be character-recognized is a rectangle (quadrilateral). A finder pattern mark for positioning is printed at three vertices among the four corners of the rectangular area included in each page of the report. The finder pattern mark is a characteristic mark of a quadrilateral without duplication within the object, such as at three vertices among the four corners of a QR code (registered trademark). When recognizing an image obtained by photographing the report with the camera 2, it becomes possible to detect the finder pattern mark by a feature amount extraction technique such as SIFT (Scale Invariant Feature Transform) or AKAZE (Accelerated KAZE), and detect the rectangular area where character recognition should be performed from the detection result.

[0075] FIG. 11 is a flowchart showing another example of the procedure of business processing. The person in charge of proofreading and reviewing the report designates the data format of the report via the operation unit 15 of the information processing apparatus 1. The processing unit 11 receives the data format (step S51). The data format is assumed to be paper, image data, or text data (including word processing data in which text data is embedded, etc.). The processing unit 11 determines whether the designated data format is electronic data (step S52). When the processing unit 11 determines that it is not electronic data (NO in step S52), it acquires an image of the report (step S53). The image is acquired by photographing the paper report with the camera 2 or by an optical scanner. When the processing unit 11 determines that it is electronic data (YES in step S52), it determines whether it is text data (step S54). When the processing unit 11 determines that it is not text data, that is, it is image data (NO in step S54), it proceeds to step S55. When the processing unit 11 determines that it is text data (YES in step S54), it advances the process to step S56. In the case where there are a plurality of description areas in the report, the text data is read separately for each. The processing unit 11 performs character recognition using OCR (Optical Character Recognition) technology and acquires text data (step S55). When the processing unit 11 inputs the text data, it inputs the acquired text data to a learning model trained to output a determination result including parts to be corrected, etc., and acquires the determination result output by the learning model (step S56). The processing unit 11 displays the determination result on the display unit 16 (step S57). The determination result is an indication of a contradiction in the content or a part that contradicts the past content. The determination may be made by the dialogue service 3. The person in charge (visual judge) inputs the determination result. The processing unit 11 acquires the determination result (step S58). The processing unit 11 determines whether the two determination results match (step S59). When the processing unit 11 determines that the determination results match (YES in step S59), it stores the image and the determination result in the training DB 133 (step S60). The processing unit 11 performs re-learning of the learning model (step S61) and ends the process.

[0076] When the processing unit 11 determines that the determination results do not match (NO in step S59), it estimates the cause (step S62). The processing unit 11 searches the cause DB 131 via the dialogue service 3 and estimates the cause. The processing unit 11 uses the past data stored in the cause DB 131 to train the dialogue service 3 by few-shot learning. The data includes the determination result of the learning model including the part to be corrected, etc., and the cause of the inconsistency. Then, the processing unit 11 inputs the determination result of the current learning model and a question asking about the cause of the inconsistency to the dialogue service 3, and obtains the cause estimated by the dialogue service 3. The processing unit 11 displays the estimated cause on the display unit 16 (step S63). The person in charge identifies the cause of the difference. Based on the identified cause, rework is performed. The processing unit 11 determines whether it is a re-visual inspection determination (step S64). When the processing unit 11 determines that it is a re-visual inspection determination (YES in step S64), the process returns to step S58. When the processing unit 11 determines that it is not a re-visual inspection determination (NO in step S64), it determines whether to re-acquire the image (step S65). When the processing unit 11 determines to re-acquire the image (YES in step S65), the process returns to step S53. When the processing unit 11 determines that the image is not to be re-acquired (NO in step S65), an inquiry is made (step S66). The processing unit 11 makes an inquiry by email, chat, voice message, etc. via the dialogue service 3 based on the instruction from the person in charge. The processing unit 11 receives the reply and displays it on the display unit 16 (step S67). The person in charge refers to the reply and attempts to resolve the cause. The processing unit 11 returns the process to step S64. After the final determination is confirmed through re-trial, the processing unit 11 registers the presence or absence of contradiction (completion or non-completion of confirmation) in the database, and registers the cause of the contradiction in the cause DB 131.

[0077] This embodiment has the following effects. It becomes possible to efficiently identify and correct the contradictory parts. Also, when the person in charge visually checks all the reports, if the image determination AI determines the presence or absence of contradiction in advance and the person in charge preferentially checks the reports determined to have contradictions, it is considered that the work efficiency is improved.

[0078] In addition, in the second and third embodiments, the detection model M2 may also be used to determine whether the cause of the discrepancy is due to the image.

[0079] The technical features (constituent elements) described in each embodiment can be combined with each other, and by combining them, new technical features can be formed. The embodiments disclosed this time should be considered as illustrative in all respects and not restrictive. The scope of the present invention is shown not by the above meaning, but by the scope of claims, and it is intended that all modifications within the meaning and scope equivalent to the scope of claims are included. In addition, although the scope of claims uses a format (multi-claim format) that describes claims that cite two or more other claims, it is not limited to this. It may be described using a format that describes a multi-claim (multi-multi-claim) that cites at least one multi-claim.

Explanation of Reference Numerals

[0080] 100: Information processing system 1: Information processing apparatus 11: Processing unit 11a: Acquisition unit 11b: Determination unit 11c: Display output unit 11d: Result acquisition unit 11e: Detection unit 11f: Model update unit 11g: Estimation unit 11h: Registration update unit 12: Input / output unit 13: Storage unit 131: Cause DB 132: Inquiry DB 133: Training DB M1: Classification model M11: Feature extraction unit M12: Fully connected layer M13: Output layer M2: Detection model M21: Image Captioning Model M22: Text Interpretation Model 14: Communication Unit 15: Operation Unit 16: Display Unit 1P: Program 2: Camera 3: Dialogue Service 31: Dialogue Server 32: LLM N: Network

Claims

1. When the first result by an image recognition model learned to perform recognition processing on an input image and the second result by a user for the input image are different, obtain a prompt including a question asking about the first result and the cause of the difference, Input the obtained prompt into a language generation model that outputs the cause of the difference when the prompt is input, thereby outputting the cause A program for causing a computer to execute the process.

2. The first result includes an output value and a confidence level of the output value The program according to claim 1.

3. Obtain a plurality of past cases including the cause of the difference between the first result and the second result and the past cases including the first result, Based on the cases, create a prompt including the first result and the cause, Input the created prompt into the language generation model to train the language generation model, After training, input a prompt including a question asking about the first result and the cause of the difference into the language generation model The program according to claim 1 or claim 2.

4. When the confidence level included in the first result is equal to or higher than a predetermined threshold, the cause is determined to be a user's judgment error The program according to claim 2.

5. When an image is input, input the input image into a detection model that outputs the confidence level that the image includes a problem that hinders the recognition processing, and obtain the confidence level, When the obtained confidence level is equal to or higher than a predetermined threshold, the cause is determined to be an error of the image recognition model The program according to claim 1 or claim 2.

6. Obtain a plurality of past cases including the cause of the difference between the first result and the second result, the first result, and the confidence level output by the detection model, Based on the cases, create a prompt including the first result, the confidence level, and the cause, Input the created prompt into the language generation model to train the language generation model, After training, input a prompt including the first result, the confidence level output by the detection model, and a question asking about the cause of the difference into the language generation model The program according to claim 5.

7. When an input indicating that there is no solution for resolving the difference between the first result and the second result is received, input a prompt asking for an inquiry into the language generation model, Output an inquiry proposal including the information of the inquiry destination obtained from the language generation model The program according to claim 1 or claim 2.

8. Obtain the first result and the second result, By comparing the first result and the second result, determine whether they are different The program according to claim 1 or claim 2.

9. After outputting the cause, when the first result and the second result match with the retrial, associate the first result and the cause when they were different, and store them as the case The program according to claim 3.

10. When the error in the first result is eliminated by the retrial after outputting the cause, store the input image and the second result as training data for the image recognition model The program according to claim 1 or claim 2.

11. An information processing apparatus including a control unit, When the first result by an image recognition model learned to execute recognition processing on an input image and the second result by a user for the input image are different, the control unit obtains a prompt including a question asking for the first result and the cause of the difference, When the control unit inputs the prompt, it inputs the obtained prompt to a language generation model that outputs the cause of the result difference, The control unit outputs the cause obtained from the prompt Information processing apparatus.

12. A computer, When the first result by an image recognition model learned to execute recognition processing on an input image and the second result by a user for the input image are different, obtain a prompt including a question asking for the first result and the cause of the difference, When the prompt is input, by inputting the obtained prompt to a language generation model that outputs the cause of the result difference, output the cause Information processing method.

Citation Information

Patent Citations

  • Visual recognition support device, method and program

    JP2018151765A