Information processing device, information processing method, and information processing program
The information processing apparatus and method stabilize abnormality detection accuracy by using a generation model with multimodal prompts and user interactions, automating detection and improving precision.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- OUEN CO LTD
- Filing Date
- 2026-01-14
- Publication Date
- 2026-07-23
AI Technical Summary
Existing abnormality detection methods in products suffer from variations in inspection accuracy due to operator skill levels and experience, necessitating a solution to stabilize and improve detection precision.
An information processing apparatus and method that utilizes a generation model to analyze target object images, incorporating multimodal prompts and user interactions to enhance detection accuracy by identifying and displaying abnormalities, and allowing for model parameter adjustments based on feedback.
The solution automates abnormality detection, improves accuracy, reduces costs, and enhances user understanding of detection precision, facilitating quick response and record management.
Smart Images

Figure JP2026000871_23072026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Information Processing Method, and Information Processing Program
[0007] ,
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.
[0002] Japanese Patent Application Laid-Open No. 2022-071675 discloses a technique for improving the accuracy of estimating an equipment state or a dangerous state that violates a safety manual based on a site image captured by a camera.
[0003] By the way, it is known to perform abnormality detection of an object such as an article in order to guarantee the quality of a product and maintain and improve customer satisfaction.
[0004] Here, in abnormality detection by a human, there are problems such as variations in inspection accuracy depending on the skill level and experience of an operator. Therefore, it is required to suppress variations in inspection accuracy in abnormality detection of an object.
[0005] Therefore, an object of the present disclosure is to provide an information processing apparatus, an information processing method, and an information processing program that can allow a user to grasp the inspection accuracy of a generation model before performing abnormality detection when performing abnormality detection using a generation model capable of outputting the presence or absence of an abnormality of a target object from a target object image showing the target object.
[0006] The information processing apparatus according to the first aspect includes a processor, and the processor receives an input of a target object image showing a target object to be subjected to abnormality detection, and inputs a multi-modal prompt including the input specific target object image and a predetermined instruction sentence to a generation model capable of outputting the presence or absence of an abnormality of the specific target object shown in the specific target object image, and causes a display unit to display the presence or absence of an abnormality of the specific target object shown in the specific target object image, which is an output result of the generation model.
[0007] In the first embodiment of the information processing device, the processor receives an input of an object image showing an object to be detected for anomaly. The processor then inputs a multimodal prompt, which includes the received specific object image and a predetermined instruction, into the generation model, and causes the display unit to show whether or not there is an anomaly in the specific object shown in the specific object image, which is the output result of the generation model. Thus, according to the information processing device, by inputting the multimodal prompt into the generation model before performing anomaly detection using the generation model, the user can understand the inspection accuracy of the generation model before performing anomaly detection.
[0008] The second embodiment of the information processing apparatus is the same as the first embodiment, wherein the generation model identifies the type of the specific object from the characteristics of the input specific object image, infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific object.
[0009] In the second embodiment of the information processing device, the generation model identifies a specific object type from the features of a specific object image input, and infers whether or not there is an abnormality corresponding to that type according to instructions indicated by a predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific object. As a result, the information processing device can automate everything from object type identification to abnormality detection using the generation model, thereby improving the accuracy of abnormality detection, increasing work efficiency, and reducing costs.
[0010] The third embodiment of the information processing apparatus is the same as the second embodiment, wherein the generation model, when it infers that there is an abnormality in the specific object, generates at least one of image data and text data indicating the nature of the abnormality, and the processor causes the display unit to display at least one of the image data and text data indicating the nature of the abnormality in the specific object, which is the output result of the generation model.
[0011] In the third embodiment of the information processing device, when the generation model infers that there is an abnormality in a specific object, it generates at least one of image data and text data indicating the nature of the abnormality. The processor then causes the display unit to display at least one of the image data and text data indicating the nature of the abnormality in the specific object, which are the output results of the generation model. As a result, the information processing device displays the nature of the abnormality in the specific object on the display unit as at least one of the image and text, allowing the user to intuitively and concretely grasp the problem area, facilitating quick response and record management.
[0012] The fourth embodiment of the information processing apparatus is one of the first to third embodiments, wherein the processor receives input of supplementary text that supplements the characteristics of the specific object and a reference image showing at least one of a state in which the predetermined object is normal and a state in which it is abnormal, and inputs a multimodal prompt to the generation model that includes the image of the specific object, the predetermined instruction sentence, and at least one of the supplementary text and the reference image that was received as input.
[0013] In the fourth embodiment of the information processing device, the processor receives input of supplementary text that supplements the characteristics of a specific object and a reference image showing at least one of the states in which the given object is free from abnormalities and when it is abnormal. The processor then inputs a multimodal prompt to the generative model, which includes the image of the specific object, a predetermined instruction, and at least one of the received supplementary text and reference image. As a result, the information processing device deepens the understanding of the generative model by adding specific information regarding the characteristics of the object and whether or not it has abnormalities, enabling more accurate anomaly detection and reducing false positives and missed detections.
[0014] The fifth embodiment of the information processing apparatus is the same as the fourth embodiment, wherein the input of the auxiliary text is performed in an interactive format between the user and the generation model.
[0015] In the fifth embodiment of the information processing device, the input of auxiliary text is performed through an interactive format between the user and the generative model. This allows the information processing device to flexibly acquire necessary information by supplementing the characteristics of a specific object through an interactive format with the user, enabling anomaly detection that aligns with the user's intentions.
[0016] The sixth embodiment of the information processing device is one of the first to fifth embodiments, wherein the processor receives input of an abnormality definition indicating the criteria for determining whether or not an abnormality exists in the specific object.
[0017] In the sixth embodiment of the information processing device, the processor accepts input of an abnormality definition that indicates the criteria for determining whether or not a particular object is abnormal. As a result, by accepting input of an abnormality definition, the information processing device can set flexible judgment criteria according to the user's application, thereby improving the detection accuracy of the generation model and increasing its adaptability to specific applications.
[0018] The seventh embodiment of the information processing apparatus is any one of the first to sixth embodiments, wherein the processor receives at least one input of supplementary text that supplements the characteristics of the specific object, an image showing that an object of the same type as the specific object is free from abnormalities, and an image showing that an object of the same type is abnormal, and adjusts at least one of the parameters of the generation model relating to a multimodal prompt to be input to the generation model and the output content based on the supplementary text, the image showing no abnormalities, and the image showing abnormalities that have been received as input.
[0019] In the seventh embodiment of the information processing apparatus, the processor receives at least one input: supplementary text that supplements the characteristics of a specific object, an image showing that an object of the same type as the specific object is free from abnormalities, and an image showing that an abnormality exists in the same type of object. Based on the supplementary text, the image showing no abnormalities, and the image showing abnormalities that the processor has received, the processor adjusts at least one of the multimodal prompts to be input to the generation model and the parameters of the generation model related to the output content. As a result, the information processing apparatus can optimize the abnormality detection of the generation model for a specific object and specific conditions by adjusting the parameters of the generation model based on information about the characteristics and abnormal state of the specific object. Furthermore, by adjusting the multimodal prompts based on information about the characteristics and abnormal state of the specific object, the generation model can more accurately understand the characteristics and abnormal state of the specific object, thereby improving the accuracy of the output results.
[0020] The information processing device of the eighth embodiment is an information processing device of any one of the first to seventh embodiments, wherein the processor receives user feedback on the determination of whether or not there is an abnormality in the specific object by the generative model, and performs learning of the generative model based on the feedback.
[0021] In the eighth embodiment of the information processing device, the processor receives user feedback on the generative model's determination of whether or not a particular object has an anomaly. Based on this feedback, the processor performs training on the generative model. As a result, the information processing device can continuously correct the validity of the anomaly detection results based on multimodal prompts with user involvement, and can further enhance the reliability and explainability of the anomaly detection results while gradually improving the inspection accuracy of the generative model in line with actual operation.
[0022] The ninth aspect of the information processing method involves a computer receiving an input of an object image showing an object subject to anomaly detection, inputting a multimodal prompt including the specific object image received as input and a predetermined instruction sentence into a generation model capable of outputting whether or not there is an anomaly in the object shown in the object image, and displaying the output result of the generation model, which is whether or not there is an anomaly in the specific object shown in the specific object image, on a display unit.
[0023] In the ninth aspect of the information processing method, the computer performs a process to receive input of an object image showing an object to be detected for anomalies. The computer also inputs a multimodal prompt, which includes the received specific object image and a predetermined instruction, into a generation model, and performs a process to display on the display unit whether or not there is an anomaly in the specific object shown in the specific object image, which is the output result of the generation model. Thus, according to the information processing method, by inputting the multimodal prompt into the generation model before performing anomaly detection using the generation model, the user can understand the inspection accuracy of the generation model before anomaly detection is performed.
[0024] The information processing program of the tenth embodiment receives an input of an object image showing an object to be detected for anomaly detection, inputs a multimodal prompt including the specific object image received as input and a predetermined instruction sentence into a generation model capable of outputting whether or not there is an anomaly in the object shown in the object image, and causes a computer to execute a process that displays on a display unit whether or not there is an anomaly in the specific object shown in the specific object image, which is the output result of the generation model.
[0025] In the tenth embodiment of the information processing program, the computer is instructed to perform a process to receive input of an object image showing the object to be detected for anomalies. The computer then inputs a multimodal prompt, which includes the received specific object image and a predetermined instruction, into a generation model, and performs a process to display on the display unit whether or not there is an anomaly in the specific object shown in the specific object image, which is the output result of the generation model. As a result, according to the information processing program, by inputting the multimodal prompt into the generation model before performing anomaly detection using the generation model, the user can understand the inspection accuracy of the generation model before anomaly detection is performed.
[0026] As described above, the information processing device, information processing method, and information processing program related to this disclosure allow the user to understand the inspection accuracy of the generation model before performing anomaly detection when performing anomaly detection using a generation model capable of outputting whether or not an object is abnormal from an image of the object in which the object is shown.
[0027] This is a block diagram showing the hardware configuration of an information processing device. This is a block diagram showing the storage configuration of an information processing device. This is a block diagram schematically showing the functional configuration of the CPU of an information processing device. This is a flowchart showing the flow of a specific process executed by an information processing device. This is the first example of a display shown on the display unit of an information processing device. This is a specific example 1 of the reference image. This is a specific example 2 of the reference image. This is the second example of a display shown on the display unit of an information processing device. This is the third example of a display shown on the display unit of an information processing device. This is the fourth example of a display shown on the display unit of an information processing device. This is the fifth example of a display shown on the display unit of an information processing device. This is the sixth example of a display shown on the display unit of an information processing device.
[0028] The information processing device 20 according to this embodiment will be described below. Figure 1 is a block diagram showing the hardware configuration of the information processing device 20. As an example, the information processing device 20 may be a general-purpose computer device such as a server computer or a PC (Personal Computer), or a mobile terminal such as a smartphone or tablet terminal.
[0029] As shown in Figure 1, the information processing device 20 includes a CPU (Central Processing Unit) 21, a ROM (Read Only Memory) 22, a RAM (Random Access Memory) 23, storage 24, an input unit 25, a display unit 26, and a communication unit 27. Each component is connected to the others via a bus 28 so as to be able to communicate with each other. The information processing device 20 is an example of the "information processing device" of this disclosure.
[0030] The CPU 21 is a central processing unit that executes various programs and controls various parts. Specifically, the CPU 21 reads a program from the ROM 22 or storage 24 and executes the program using the RAM 23 as a working area. The CPU 21 controls each of the above components and performs various calculations according to the program stored in the ROM 22 or storage 24. The CPU 21 is an example of the "processor" of this disclosure.
[0031] ROM 22 stores various programs and data. RAM 23 temporarily stores programs or data as a working area.
[0032] The storage 24 consists of storage devices such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory, and stores various programs and various data.
[0033] The input unit 25 includes, for example, a pointing device such as a mouse, various buttons, a keyboard, a microphone, and a camera, and is used for various types of input.
[0034] The display unit 26 is, for example, a liquid crystal display and displays various information. The display unit 26 may also function as an input unit 25 by employing a touch panel method. The display unit 26 is an example of the "display unit" in this disclosure.
[0035] The communication unit 27 is an interface for communicating with other devices. For such communication, a wired communication standard such as Ethernet (registered trademark) or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) may be used.
[0036] Next, the configuration of the storage 24 of the information processing device 20 will be described. Figure 2 is a block diagram showing the configuration of the storage 24 of the information processing device 20.
[0037] As shown in Figure 2, the storage 24 stores the information processing program 24A and the generation model 24B.
[0038] The information processing program 24A is a program that causes the CPU 21 to execute various processes described later. When executing the information processing program 24A, the information processing device 20 uses the hardware resources shown in Figure 1 to execute the processes based on the information processing program 24A. The information processing program 24A is an example of an "information processing program" in this disclosure.
[0039] The generative model 24B is a so-called generative AI (Artificial Intelligence). The generative model 24B is a model that can output whether or not there are abnormalities in an item shown in an item image that is the subject of visual inspection. An example of a component of the generative model 24B is CLIP (Internet search <URL: https: / / trail.tu-tokyo.ac.jp / ja / blog / 22-12-02-clip / >). The generative model 24B is a multimodal generative AI that appropriately uses components that can associate images and text, such as CLIP, and image generators such as diffusion models and GANs, and text decoders such as transformers. However, these are just examples and are not limiting. The generative model 24B is input with a multimodal prompt that includes text data representing text and image data representing images. The generative model 24B performs inference according to instructions indicated by multimodal prompts and outputs the inference results in data formats such as image data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0040] Next, the functional configuration of the CPU 21 will be described. FIG. 3 is a block diagram showing an example of the functional configuration of the CPU 21.
[0041] As shown in FIG. 3, the CPU 21 includes an input unit 21A, a processing unit 21B, and an output unit 21C.
[0042] The input unit 21A acquires the user input received by the information processing apparatus 20. Specifically, the input unit 21A acquires at least one of the input data of text, voice, image, and video received by the information processing apparatus 20.
[0043] The processing unit 21B performs a specific process using the generation model 24B. Specifically, the processing unit 21B inputs a predetermined multimodal prompt to the generation model 24B and obtains an output result.
[0044] The output unit 21C outputs the output result of the generation model 24B, that is, the result of the specific process, to the display unit 26. As a result, the result of the specific process is displayed on the display unit 26.
[0045] FIG. 4 is a flowchart showing the flow of the specific process executed by the information processing apparatus 20. The specific process is performed by the CPU 21 reading the information processing program 24A from the storage 24, expanding it in the RAM 23, and executing it. The specific process is performed by the CPU 21 functioning as the input unit 21A, the processing unit 21B, and the output unit 21C described above. As an example, the specific process is performed when the user executes a predetermined application.
[0046] In step S10 shown in FIG. 4, the CPU 21 acquires various information received the user input. Then, the CPU 21 proceeds to step S11. Here, the various information at least includes an image (hereinafter, "specific article image") showing a specific article to be inspected for appearance. Hereinafter, the article will be described by taking "screw" as an example (see FIG. 5 etc.), but the type of the article is not limited to this. For example, as the article, various products such as "fabric products", "cosmetics (e.g., foundation)", and "substrates" can be used in addition to "screws".
[0047] In step S11, based on the various information acquired in step S10, the CPU 21 generates a multimodal prompt for outputting whether there is an abnormality in the specific article shown in the specific article image. Then, the CPU 21 proceeds to step S12.
[0048] As an example, the CPU 21 generates the following multimodal prompt including at least the specific article image acquired in step S10 and a predetermined instruction sentence. Note that the instruction sentence shown in the following multimodal prompt is merely an example and is not limited thereto.
[0049] 「Multimodal Prompt」 Instruction sentence: Please output whether there is a defect in the specific article shown in the following specific article image. The output format should be at least one of image data and text data. Specific article image: Screw AAA.jpg
[0050] In step S12, the CPU 21 inputs the multimodal prompt generated in step S11 into the generation model 24B and obtains the output result by the generation model 24B. Then, the CPU 21 proceeds to step S13.
[0051] In step S13, the CPU 21 outputs the output result of the generation model 24B in step S12 to the display unit 26. As a result, the result of the specific process is displayed on the display unit 26. Then, the CPU 21 ends the specific process.
[0052] Next, a display example according to this embodiment will be described. FIG. 5 is a first display example displayed on the display unit 26 of the information processing apparatus 20. Specifically, FIG. 5 shows a verification screen for verifying the inspection accuracy by the generation model 24B.
[0053] The verification screen shown in FIG. 5 shows GUI (Graphical User Interface) buttons 40 to 46 and a verification area 50.
[0054] The GUI button 40 is a button for selecting a generation model 24B for verifying the inspection accuracy of the visual inspection. In this embodiment, it is assumed that multiple generation models 24B with different parameters related to the output content (e.g., generation model 24B-1, generation model 24B-2, etc.) are provided. The user can select the desired generation model 24B by operating the GUI button 40. If the GUI button 40 is not operated, a pre-set generation model 24B will be selected. Note that at least one model may be selected as the target of the pre-set generation model 24B, and accuracy verification may be performed on multiple models, or the model with the optimal accuracy and verification result may be automatically selected.
[0055] The GUI button 41 is a button for selecting at least one image of a specific item to be inspected for abnormalities. The image of the specific item may also have data attached indicating the location and characteristics of the abnormality. After operating the GUI button 41, the user selects the desired image of the specific item by selecting an image from the storage 24 or by taking a picture with the camera of the information processing device 20. The verification area 50 shown in Figure 5 displays the screw image 60 selected as the specific item image. If multiple images of specific items are selected, they may be displayed in a list, or they may be viewable sequentially using an image scrolling function. The verification area 50 may also display at least one of the automatically selected images to be judged, for purposes such as re-verifying the image to be judged during actual visual inspection.
[0056] The GUI button 42 is a button for setting the detection conditions for anomalies by the generation model 24B. These detection conditions include a threshold, the nature of the anomaly, and its location. The user can set these detection conditions by operating the GUI button 42. As a result, the CPU 21 accepts the detection conditions set by the user as an anomaly definition that indicates the criteria for determining whether or not a particular item is an anomaly. This anomaly definition may be freely entered by the user, or it may be defined by the user modifying a suggestion from the software. Furthermore, if data indicating the correct answer for an anomaly is added to the image of the specific item, the conditions that result in the highest judgment accuracy may be automatically searched for, for example.
[0057] GUI button 43 is a button for entering supplementary text that complements the characteristics of a specific item. The supplementary text is text that enhances the specificity of the item, such as a specific product name and defect information. The user can enter the supplementary text by operating the GUI button 43. The input of the supplementary text may be in the form of the user directly typing into a text box, or it may be in the form of the user modifying suggestions from the software.
[0058] The GUI button 44 is a button for inputting a reference image of a good product that shows no abnormalities in a specified item. The specified item may be the same type of item as the specific item, or it may be a different type of item. After operating the GUI button 44, the user selects the desired reference image of a good product by selecting an image from the storage 24 or by taking a picture with the camera of the information processing device 20.
[0059] GUI button 45 is a button for inputting a reference defective product image that shows a condition where a predetermined item has an abnormality. After operating the GUI button 45, the user selects a desired reference defective product image by selecting an image from the storage 24 or by taking a picture with the camera of the information processing device 20. Furthermore, the user also inputs text indicating the location and content of the abnormality in the selected reference defective product image.
[0060] GUI button 46 is a button for performing verification of the inspection accuracy of the visual inspection by the generation model 24B. When GUI button 46 is operated, the CPU 21 generates a multimodal prompt based on the settings based on the operation of GUI buttons 40 to 45. The CPU 21 then inputs the multimodal prompt to the generation model 24B and outputs whether or not there is an abnormality in the specific item (e.g., screw) shown in the specific item image (e.g., screw image 60) displayed in the verification area 50. If multiple specific item images are selected, the contents of multiple images can be viewed by displaying a list of the contents of whether or not there is an abnormality or by using an image scrolling function. In addition, if data indicating the correct answer for abnormality is attached to the specific item image, statistical information such as the accuracy of each image and the overall judgment accuracy may also be displayed. In addition, when detection conditions are automatically searched for and suggested, inputting the abnormality detection conditions in GUI button 42, inputting the auxiliary text in GUI button 43, and inputting the reference good product image and reference defective product image in GUI buttons 44 and 45 are optional. Even without these inputs, the generation model 24B can output whether or not there is an abnormality in the specific item based on the operation of GUI button 46.
[0061] Figure 6 is a specific example of a reference image showing a given item in both its normal and abnormal states. As an example, in Figure 6, the given item is a "screw" of the same type as a specific item.
[0062] Figure 6A shows screw image 62, which indicates a screw without any abnormalities. Figure 6B shows screw image 64, which indicates a screw with an abnormality. For example, if the user operates the GUI button 45 and selects screw image 64 as the reference image of a defective product, they can input text indicating the location and nature of the abnormality, such as "Bent tip of screw."
[0063] Figure 7 shows a second display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 7 shows the state after the GUI button 46 has been operated on the verification screen shown in Figure 5.
[0064] In the verification screen shown in Figure 7, text data 65 indicating whether or not there is an abnormality in the screw shown in the screw image 60 is displayed below the screw image 60 within the verification area 50. Specifically, the text data 65 says "No abnormalities," indicating that there is no abnormality in the screw shown in the screw image 60.
[0065] Figure 8 shows a third display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 8 shows a state in which a screw image 70, which represents a screw different from the screw image 60, is displayed as a specific item image in the verification area 50.
[0066] Figure 9 shows a fourth display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 9 shows the state after the GUI button 46 has been operated on the verification screen shown in Figure 8.
[0067] In the verification screen shown in Figure 9, text data 75 indicating whether or not there is an abnormality in the screw shown in the screw image 70 is displayed below the screw image 70 within the verification area 50. Specifically, the text data 75 reads, "There is an abnormality. The tip of the screw is bent," indicating that there is an abnormality in the screw shown in the screw image 70 and the nature of that abnormality.
[0068] Here, the display method of the screw image 70 shown in Figure 9 differs from the display method of the screw image 70 shown in Figure 8. Specifically, in the screw image 70 shown in Figure 9, a heat map is superimposed on the screw image to indicate abnormal areas. Note that in Figure 9, for illustrative purposes, the heat map is represented by a pattern of diagonal hatching and dots. In the screw image 70, the area 70A from the head of the screw to near the tip is represented by a first color (diagonal hatching in the figure), and the tip area 70B is represented by a second color different from the first color (dots in the figure). The second color is a predetermined color (e.g., red, yellow) that indicates the area with an abnormality.
[0069] As described above, the verification screen shown in Figure 9 indicates, through the display mode of the screw image 70 and the text data 75, that there is an abnormality in the screw shown in the screw image 70 and the details of the abnormality.
[0070] Figure 10 shows a fifth display example shown on the display unit 26 of the information processing device 20. As an example, Figure 10 shows an adjustment screen for adjusting at least one of the multimodal prompts input to the generation model 24B and the parameters of the generation model 24B related to the output content. The adjustment screen is displayed when a predetermined operation is performed on the verification screen.
[0071] The adjustment screen shown in Figure 10 displays GUI buttons 80 to 83. GUI button 80 is a button for selecting the generation model 24B to be adjusted for the multimodal prompt and parameters. The user can select the desired generation model 24B by operating GUI button 80.
[0072] GUI button 81 is a button for selecting the training images to be used for adjusting the multimodal prompt and parameters. The training images include images of good products showing no abnormalities in a specific item (e.g., a screw) and an item of the same type (e.g., a screw), and images of defective products showing abnormalities in the same item. After operating GUI button 81, the user selects at least one desired good product image and one desired defective product image by selecting a desired image from storage 24 or by taking a picture with the camera of the information processing device 20. The user may also input text, text data, or structured data such as Json or YAML indicating the location and content of the abnormality in the selected defective product image.
[0073] The GUI button 82 is a button for entering supplementary text that provides details about the characteristics of a specific item. The user can enter supplementary text by operating the GUI button 82.
[0074] GUI button 83 is a button for adjusting at least one of the multimodal prompts and parameters. When GUI button 83 is operated, the CPU 21 adjusts at least one of the multimodal prompts and parameters based on the settings based on the operation of GUI buttons 80 to 82. The adjustment of the multimodal prompts or parameters is performed as appropriate using known techniques.
[0075] Figure 11 shows a sixth display example shown on the display unit 26 of the information processing device 20. As an example, Figure 11 shows the operation screen during actual operation of the visual inspection.
[0076] The operation screen shown in Figure 11 shows GUI buttons 90-93 and a verification area 50. In actual operation, an image of an item captured by a camera (not shown) is displayed in the verification area 50, and the information processing device 20 automatically determines whether or not there is an abnormality each time an image is taken. In the verification area 50 shown in Figure 11, a screw image 100, which represents a screw captured by the camera, is displayed.
[0077] The GUI button 90 is for selecting the generation model 24B to be used in the actual operation of the visual inspection. The user can select the desired generation model 24B by operating the GUI button 90. However, the selection is not limited to the user; for example, the model with the highest accuracy in the verification results may be automatically selected.
[0078] The GUI button 91 is a button for setting the anomaly detection conditions according to the generation model 24B. The user can set these detection conditions by operating the GUI button 91.
[0079] GUI button 92 is a button for selecting whether to enable or disable online learning. Users can switch online learning on or off by operating GUI button 92. In this embodiment, online learning refers to a function that adjusts the parameters of the generative model 24B during visual inspection to improve anomaly detection performance. Specifically, this includes methods such as performing learning by assuming the first XX items during visual inspection are good products, and receiving user feedback on the judgment of good / defective products on the spot and performing learning in real time or retrospectively. With this configuration, in addition to the user explicitly selecting images to train the model at times other than when visual inspection is actually in operation, the generative model 24B can continuously improve its accuracy by automatically incorporating feedback.
[0080] GUI button 93 is a button for starting the actual operation of the visual inspection using the generation model 24B. When GUI button 93 is operated, the CPU 21 generates a multimodal prompt based on the settings based on the operation of GUI buttons 90 to 92. The CPU 21 then inputs the multimodal prompt into the generation model 24B and outputs whether or not there is an abnormality in the item shown in the item image (e.g., screw image 100) of the item displayed in the verification area 50.
[0081] In the operation screen shown in Figure 11, text data 105 indicating whether or not there is an abnormality in the screw shown in the screw image 100 is displayed below the screw image 100 within the verification area 50. Specifically, the text data 105 is "There is an abnormality. The tip of the screw is bent," indicating that there is an abnormality in the screw shown in the screw image 100 and the nature of the abnormality. The nature of the abnormality may be written in conjunction with the location of the abnormality, for example, but the display location is not particularly limited. Furthermore, there are no particular restrictions on the nature of the abnormality itself, such as using quantitative or qualitative expressions to describe the degree of the abnormality.
[0082] Furthermore, the display method for screw images 100 that have been determined to have an abnormality is the same as that for screw image 70 shown in Figure 9, where a heat map is superimposed on the screw image to indicate the abnormal area. In the screw image 100, the region 100A from the head of the screw to near the tip is represented in the first color, and the region 100B at the tip is represented in the second color.
[0083] Note that the GUI buttons shown in Figure 11 represent only a portion of the operation screen, and other GUI buttons actually exist. For example, above GUI button 90, there is an image input source selection button. This image input source selection button is used to select the source from which to receive images, such as images saved in a designated folder or images taken directly from the camera.
[0084] As explained above, in the information processing device 20, the CPU 21 receives input of an image of an item that is the subject of the visual inspection. The CPU 21 also inputs a multimodal prompt to the generation model 24B, which includes the received image of the specific item and a predetermined instruction. The CPU 21 then displays on the display unit 26 whether or not there is an abnormality in the specific item shown in the image of the specific item, which is the output result of the generation model 24B. Thus, with the information processing device 20, by inputting the multimodal prompt to the generation model 24B before performing the visual inspection using the generation model 24B, the user can understand the inspection accuracy of the generation model 24B before the visual inspection is performed. Furthermore, with the information processing device 20, by inputting the multimodal prompt to the generation model 24B during the actual operation of the visual inspection, the user can understand the inspection accuracy of the generation model 24B while the visual inspection is being performed.
[0085] Furthermore, the generation model 24B identifies the type of a specific item from the characteristics of the input specific item image, and infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific item. As a result, the information processing device 20 can automate everything from item type identification to abnormality detection using the generation model 24B, thereby improving the accuracy of visual inspection, increasing work efficiency, and reducing costs.
[0086] Furthermore, if the generation model 24B infers that there is an abnormality in a particular item, it generates at least one of image data and text data indicating the nature of the abnormality. Then, in the information processing device 20, the CPU 21 displays at least one of the image data and text data indicating the nature of the abnormality of the particular item, which are the output results of the generation model 24B, on the display unit 26. As a result, with the information processing device 20, the nature of the abnormality of the particular item is displayed on the display unit 26 in the form of at least one of the image and text, allowing the user to intuitively and concretely grasp the problem area, making it easier to take quick action and manage records.
[0087] Furthermore, in the information processing device 20, the CPU 21 receives input of supplementary text that supplements the characteristics of a specific item, and a reference image showing at least one of the states in which the specified item is free from abnormalities and in which it is abnormal. The CPU 21 then inputs a multimodal prompt to the generation model 24B that includes the image of the specific item, a predetermined instruction sentence, and at least one of the received supplementary text and reference image. As a result, the information processing device 20 deepens the understanding of the generation model 24B by adding specific information regarding the characteristics of the item and whether or not there are abnormalities, enabling more accurate abnormality detection and reducing false detections and missed detections.
[0088] Furthermore, in the information processing device 20, the CPU 21 accepts input of an abnormality definition that indicates the criteria for determining whether or not a particular item is abnormal. As a result, the information processing device 20 can accept input of an abnormality definition, allowing for flexible judgment criteria to be set according to the user's application, improving the detection accuracy of the generation model 24B and increasing its adaptability to specific applications.
[0089] Furthermore, in the information processing device 20, the CPU 21 receives at least one input: supplementary text that supplements the characteristics of a specific item, an image of a good product showing that an item of the same type as the specific item is free of defects, and an image of a defective product showing that an item of the same type has defects. Based on the supplementary text, the image of a good product, and the image of a defective product that the CPU 21 has received as input, the CPU 21 adjusts at least one of the multimodal prompts to be input to the generation model 24B and the parameters of the generation model 24B related to the output content. As a result, the information processing device 20 can optimize the abnormality detection of the generation model 24B for a specific item and specific conditions by adjusting the parameters of the generation model 24B based on information about the characteristics and abnormal state of a specific item. In addition, the information processing device 20 can improve the general detection accuracy for visual inspection of items by optimizing for multiple items. Furthermore, the information processing device 20 can improve the accuracy of the output results by adjusting the multimodal prompts based on information about the characteristics and abnormal state of a specific item, allowing the generation model 24B to understand the characteristics and abnormal state of the specific item more accurately.
[0090] (Other) Although embodiments of the present disclosure have been described in detail above with reference to the attached drawings, the technical scope of the present disclosure is not limited to these examples. It is clear that a person with ordinary skill in the art of the present disclosure may conceive of various modifications or alterations within the scope of the technical idea set forth in the claims, and it is understood that these modifications or alterations also fall within the technical scope of the present disclosure.
[0091] Furthermore, the effects described in the above embodiments are descriptive or illustrative, and are not limited to those described in the above embodiments. In other words, the technology relating to this disclosure may produce other effects that would be obvious to a person of ordinary skill in the art of this disclosure from the descriptions in the above embodiments, in addition to or in lieu of the effects described in the above embodiments.
[0092] The processing described in the above embodiment can also be implemented using dedicated hardware circuits. In this case, it may be executed on one piece of hardware or on multiple pieces of hardware.
[0093] In the above embodiment, the input of auxiliary text may be performed in an interactive format between the user and the generation model 24B. This allows the information processing device 20 to flexibly acquire the necessary information for the generation model 24B by supplementing the characteristics of a specific item in an interactive format with the user, enabling anomaly detection that is in line with the user's intentions.
[0094] In the above embodiment, the verification screen shown in Figure 5, etc., allows the input of supplementary text and reference images based on the user's selection, but the timing of these inputs is not limited. For example, instead of a configuration that allows input at any time the user desires, a configuration may be adopted in which the generation model 24B requests the user to input at least one of the supplementary text and reference images when predetermined conditions are met.
[0095] For example, the "case where the predetermined conditions are met" described above can be defined as the case where the difference in the confidence level (%) of whether or not an anomaly is present, inferred by the generative model 24B without input of auxiliary text and reference images, is less than or equal to a predetermined value.
[0096] Furthermore, the system may be configured such that, when certain conditions are met, the generation model 24B requests the user to input either supplementary text or a reference image, and when specific conditions are met, the generation model 24B requests the user to input the other supplementary text or reference image.
[0097] For example, the "case where specific conditions are met" described above can be defined as the case where the difference in the confidence level of whether or not an anomaly is inferred by the generation model 24B after inputting either the auxiliary text or the reference image is less than or equal to a predetermined value.
[0098] With the above configuration, the information processing device 20 can deepen its understanding of the generation model 24B by adding specific information regarding the characteristics of the item and whether or not there are abnormalities when the reliability of the output of the generation model 24B is low, thereby enabling more accurate abnormality detection and reducing false detections and missed detections.
[0099] Furthermore, based on the content of the supplementary text entered by the user, the generative model 24B may suggest to the user what kind of supplementary text to input to improve the reliability of the output. For example, when the generative model 24B accepts supplementary text input in a dialogue format, if the content of the supplementary text entered by the user is "product name", it may respond with a suggestion such as "Please also enter a specific example of the defect". In addition, when the generative model 24B accepts supplementary text input in a dialogue format, it may also make suggestions such as "The product name is foundation, and the defect is uneven color".
[0100] In the above embodiment, the method of representing the heatmap superimposed on the screw image is not particularly limited. For example, multiple colors corresponding to the degree of abnormality may be represented in the heatmap. Alternatively, instead of using a heatmap to indicate information about the abnormality, marks indicating the location and area of the abnormality may be used. Furthermore, the degree of abnormality may be indicated by adding numerical values or text.
[0101] The configuration described in the above embodiment is not limited to the visual inspection of goods, but can be extended to general "anomaly detection" such as hazard prediction (unsafe behavior by workers, illness), foreign object contamination, and suspicious person detection.
[0102] In the above embodiment, an example was described in which an image of an item to be inspected is received as input, but the system is not limited to this. For example, instead, an image of the item may be received as input.
[0103] In the above embodiment, the generation model 24B can output not only a binary output of "presence or absence of an abnormality" but also "degree of abnormality." Here, "degree of abnormality" is one element that constitutes "content of the abnormality." "Content of the abnormality" includes, for example, "location of the abnormality," "degree of abnormality," and "nature of the abnormality (type of abnormality)."
[0104] In the above embodiment, in addition to the case in which the generation model 24B directly generates image data and text data indicating whether or not there is an abnormality in the item, the embodiment also includes a case in which the generation model 24B outputs the degree of abnormality as numerical data in map format, and internally executes a process to generate image data and text data indicating whether or not there is an abnormality based on that numerical data and the input image.
[0105] In the above embodiment, the CPU 21 may save the output results of the generation model 24B, accept user modifications to the saved output results, and train the generation model 24B using the modified data. This creates a loop for improving the accuracy of the generation model 24B, enabling continuous improvement of anomaly detection performance.
[0106] In the above embodiment, the "characteristics of the article" entered as auxiliary text includes not only information indicating what the article itself is (e.g., foundation), but also information about specific abnormalities to be detected (e.g., type: uneven color, location: on powder, degree: very strong, etc.).
[0107] In the above embodiment, the dialogue format between the user and the generation model 24B when inputting auxiliary text is not particularly limited and may be, for example, a QA format (e.g., What is the product? → Foundation → What are the specific defect characteristics you want to detect? → Color unevenness) or a suggestion dialogue format (e.g., Is the product foundation? → Yes). Furthermore, the input format for auxiliary text may not depend on the dialogue format, and the generation model 24B may simply propose the auxiliary text entirely and present it in a format that the user can modify.
[0108] In each of the above embodiments, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, Programmable Logical Device, etc.).
[0109] Furthermore, the processor operations in each of the above embodiments may not be performed by a single processor, but may also be performed by multiple processors located in physically separate locations working together. Alternatively, some or all of the operations performed by specific multiple processors in each of the above embodiments may be integrated and performed by a single processor. In addition, the order of the processor operations is not limited to the order described in each of the above embodiments, and may be changed as appropriate.
[0110] Furthermore, although the above embodiment describes an embodiment in which the information processing program 24A is pre-stored (installed) in the storage 24, the invention is not limited thereto. The information processing program 24A may be provided in the form of being recorded on a recording medium such as a CD-ROM (Compact Disk Read Only Memory), DVD-ROM (Digital Versatile Disk Read Only Memory), and USB (Universal Serial Bus) memory. Alternatively, the information processing program 24A may be provided in the form of being downloaded from an external device via a network N. The technology of this disclosure can also be applied to programs and program products.
[0111] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0112] The disclosure of Japanese Patent Application No. 2025-007126 is incorporated herein by reference in its entirety.
Claims
1. Information processing device comprising a processor, the processor receiving an input of an object image showing an object to be detected for anomaly, inputting a multimodal prompt including the specific object image received as input and a predetermined instruction sentence to a generation model capable of outputting whether or not there is an anomaly in the object shown in the object image, and causing a display unit to display whether or not there is an anomaly in the specific object shown in the specific object image, which is the output result of the generation model.
2. The information processing apparatus according to claim 1, wherein the generation model identifies the type of the specific object from the characteristics of the input specific object image, infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific object.
3. The information processing apparatus according to claim 2, wherein the generation model, when it infers that there is an abnormality in the specific object, generates at least one of image data and text data indicating the nature of the abnormality, and the processor causes the display unit to display at least one of the image data and text data indicating the nature of the abnormality in the specific object, which is the output result of the generation model.
4. The information processing apparatus according to claim 1, wherein the processor receives input of supplementary text that supplements the characteristics of the specific object and a reference image showing at least one of a state in which the predetermined object is normal and a state in which it is abnormal, and inputs a multimodal prompt to the generation model, which includes the image of the specific object, the predetermined instruction sentence, and at least one of the supplementary text and the reference image that was received as input.
5. The information processing apparatus according to claim 4, wherein the input of the auxiliary text is performed in an interactive format between the user and the generation model.
6. The information processing apparatus according to claim 1, wherein the processor receives input of an abnormality definition indicating criteria for determining whether or not a particular object is abnormal.
7. The information processing apparatus according to claim 1, wherein the processor receives at least one input of supplementary text that supplements the characteristics of the specific object, an image showing that an object of the same type as the specific object is free from abnormalities, and an image showing that an object of the same type has abnormalities, and adjusts at least one of the parameters of the generation model relating to a multimodal prompt to be input to the generation model and the output content based on the supplementary text, the image showing no abnormalities, and the image showing abnormalities that have been received as input.
8. The information processing apparatus according to claim 1, wherein the processor receives user feedback on the determination of whether or not there is an abnormality in the specific object by the generative model, and performs learning of the generative model based on the feedback.
9. An information processing method in which a computer performs the following steps: receiving input of an object image showing an object to be detected for anomaly detection; inputting a multimodal prompt including the specific object image received as input and a predetermined instruction sentence into a generation model capable of outputting whether or not there is an anomaly in the object shown in the object image; and displaying the output result of the generation model, which is whether or not there is an anomaly in the specific object shown in the specific object image, on a display unit.
10. An information processing program for causing a computer to execute a process that receives an input of an object image showing an object to be detected for anomaly detection, inputs a multimodal prompt including the specific object image received as input and a predetermined instruction sentence into a generation model capable of outputting whether or not there is an anomaly in the object shown in the object image, and displays the output result of the generation model, which is whether or not there is an anomaly in the specific object shown in the specific object image, on a display unit.