Information processing device, information processing method, and information processing program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- OUEN CO LTD
- Filing Date
- 2026-02-17
- Publication Date
- 2026-07-30
Smart Images

Figure 2026123816000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Patent Document 1 discloses a technique for improving the accuracy of estimating an equipment state or a dangerous state that violates a safety manual based on a field image captured by a camera.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] By the way, it is known to perform abnormality detection of an object such as an article in order to guarantee the quality of a product and maintain and improve customer satisfaction. Here, in abnormality detection by a human, there are problems such as variations in inspection accuracy depending on the skill level and experience of an operator. Therefore, it is required to suppress variations in inspection accuracy in abnormality detection of an object.
[0005] Therefore, an object of the present disclosure is to provide an information processing apparatus, an information processing method, and an information processing program that can allow a user to grasp the inspection accuracy by a generation model before performing abnormality detection when performing abnormality detection using a generation model capable of outputting the presence or absence of an abnormality of an object from an object image showing the object.
Means for Solving the Problems
[0006] The first embodiment of the information processing device includes a processor, which receives an input of an object image showing an object subject to anomaly detection, inputs a multimodal prompt including the specific object image received as input and a predetermined instruction sentence to a generation model capable of outputting whether or not there is an anomaly in the object shown in the object image, displays the output result of the generation model, which is whether or not there is an anomaly in the specific object shown in the specific object image, on a display unit, and if predetermined conditions are met, prompts the user from the generation model to input supplementary text that supplements the characteristics of the specific object received as input, receives the input of the supplementary text from the user, and inputs a multimodal prompt including the specific object image, the predetermined instruction sentence, and the supplementary text received as input to the generation model.
[0007] In the first embodiment of the information processing device, the processor receives an input of an object image showing an object to be detected for anomaly. The processor then inputs a multimodal prompt, which includes the received specific object image and a predetermined instruction, into the generation model, and causes the display unit to show whether or not there is an anomaly in the specific object shown in the specific object image, which is the output result of the generation model. Thus, according to the information processing device, by inputting the multimodal prompt into the generation model before performing anomaly detection using the generation model, the user can understand the inspection accuracy of the generation model before performing anomaly detection.
[0008] The second embodiment of the information processing apparatus is the same as the first embodiment, wherein the generation model identifies the type of the specific object from the features of the input specific object image, infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific object.
[0009] In the second embodiment of the information processing device, the generation model identifies a specific object type from the features of a specific object image input, and infers whether or not there is an abnormality corresponding to that type according to instructions indicated by a predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific object. As a result, the information processing device can automate everything from object type identification to abnormality detection using the generation model, thereby improving the accuracy of abnormality detection, increasing work efficiency, and reducing costs.
[0010] The third embodiment of the information processing apparatus is the same as the second embodiment, wherein the generation model, when it infers that there is an abnormality in the specific object, generates at least one of image data and text data indicating the nature of the abnormality, and the processor causes the display unit to display at least one of the image data and text data indicating the nature of the abnormality in the specific object, which is the output result of the generation model.
[0011] In the third embodiment of the information processing device, when the generation model infers that there is an abnormality in a specific object, it generates at least one of image data and text data indicating the nature of the abnormality. The processor then displays at least one of the image data and text data indicating the nature of the abnormality in the specific object, which are the output results of the generation model, on the display unit. As a result, the information processing device displays the nature of the abnormality in the specific object on the display unit as at least one of the image and text, allowing the user to intuitively and concretely grasp the problem area, facilitating quick response and record management.
[0012] The fourth embodiment of the information processing apparatus is any one of the first to third embodiments, wherein the processor receives input of the auxiliary text and a reference image showing at least one of a state in which a predetermined object is normal and a state in which it is abnormal, and inputs a multimodal prompt to the generation model, which includes the specific object image, the predetermined instruction sentence, and at least one of the auxiliary text and the reference image that was received as input.
[0013] In the fourth embodiment of the information processing device, the processor receives input of supplementary text that supplements the characteristics of a specific object and a reference image showing at least one of the states in which the given object is free from abnormalities and when it is abnormal. The processor then inputs a multimodal prompt to the generative model, which includes the image of the specific object, a predetermined instruction, and at least one of the received supplementary text and reference image. As a result, according to the information processing device, by adding specific information regarding the characteristics of the object and whether or not it has abnormalities, the generative model's understanding is deepened, enabling more accurate anomaly detection and reducing false positives and missed detections.
[0014] The fifth embodiment of the information processing apparatus is the fourth embodiment of the information processing apparatus, wherein the input of the auxiliary text is performed in an interactive format between the user and the generation model.
[0015] In the fifth embodiment of the information processing device, the input of auxiliary text is performed through an interactive format between the user and the generative model. This allows the information processing device to flexibly acquire necessary information by supplementing the characteristics of a specific object through an interactive format with the user, enabling anomaly detection that aligns with the user's intentions.
[0016] The sixth embodiment of the information processing device is one of the first to fifth embodiments, wherein the processor receives input of an abnormality definition indicating the criteria for determining whether or not an abnormality exists in the specific object.
[0017] In the sixth embodiment of the information processing device, the processor accepts input of an anomaly definition that indicates the criteria for determining whether or not a particular object is abnormal. As a result, by accepting input of an anomaly definition, the information processing device can set flexible judgment criteria according to the user's application, thereby improving the detection accuracy of the generation model and increasing its adaptability to specific applications.
[0018] The seventh embodiment of the information processing apparatus is any one of the first to sixth embodiments, wherein the processor receives at least one input of the auxiliary text, an image showing that there is no abnormality in an object of the same type as the specific object, and an image showing that there is an abnormality in the object of the same type, and adjusts at least one of the parameters of the generation model relating to a multimodal prompt to be input to the generation model and the output content based on the input of at least one of the auxiliary text, the image showing no abnormality, and the image showing abnormality.
[0019] In the seventh embodiment of the information processing apparatus, the processor receives at least one input: supplementary text that supplements the characteristics of a specific object, an image showing that an object of the same type as the specific object is free from abnormalities, and an image showing that an abnormality exists in the same type of object. Based on the supplementary text, the image showing no abnormalities, and the image showing abnormalities received as input, the processor adjusts at least one of the multimodal prompts to be input to the generation model and the parameters of the generation model related to the output content. As a result, the information processing apparatus can optimize the abnormality detection of the generation model for a specific object and specific conditions by adjusting the parameters of the generation model based on information about the characteristics and abnormal state of the specific object. Furthermore, the information processing apparatus can improve the accuracy of the output results by enabling the generation model to understand the characteristics and abnormal state of the specific object more accurately by adjusting the multimodal prompts based on information about the characteristics and abnormal state of the specific object.
[0020] The information processing method of the eighth embodiment is a computer that receives an input of an object image showing an object subject to anomaly detection, inputs a multimodal prompt including the specific object image received as input and a predetermined instruction sentence to a generation model capable of outputting whether or not there is an anomaly in the object shown in the object image, displays the output result of the generation model, which is whether or not there is an anomaly in the specific object shown in the specific object image, and if predetermined conditions are met, prompts the user from the generation model to input supplementary text that supplements the characteristics of the specific object received as input, receives the input of the supplementary text from the user, and inputs a multimodal prompt including the specific object image, the predetermined instruction sentence, and the supplementary text received as input to the generation model.
[0021] In the eighth aspect of the information processing method, the computer performs a process to receive input of an object image showing the object to be detected for anomaly. The computer also inputs a multimodal prompt, which includes the received specific object image and a predetermined instruction, into a generation model, and performs a process to display on the display unit whether or not there is an anomaly in the specific object shown in the specific object image, which is the output result of the generation model. Thus, according to the information processing method, by inputting the multimodal prompt into the generation model before performing anomaly detection using the generation model, the user can understand the inspection accuracy of the generation model before anomaly detection is performed.
[0022] The information processing program according to the ninth aspect receives an input of an object image showing an object to be subjected to abnormality detection, inputs a multimodal prompt including the input specific object image and a predetermined instruction sentence into a generation model capable of outputting the presence or absence of an abnormality of the object shown in the object image, causes a display unit to display the presence or absence of an abnormality of a specific object shown in the specific object image that is the output result of the generation model, and when a predetermined condition is satisfied, requests the user from the generation model to input an auxiliary text that supplements the features of the specific object for which the input has been received, receives the input of the auxiliary text from the user, and inputs a multimodal prompt including the specific object image, the predetermined instruction sentence, and the input auxiliary text into the generation model, and causes a computer to execute the process.
[0023] In the information processing program according to the ninth aspect, the computer is caused to execute a process of receiving an input of an object image showing an object to be subjected to abnormality detection. Further, the computer inputs a multimodal prompt including the input specific object image and a predetermined instruction sentence into a generation model, and executes a process of causing a display unit to display the presence or absence of an abnormality of a specific object shown in the specific object image that is the output result of the generation model. Thereby, according to the information processing program, by inputting the multimodal prompt into the generation model before performing abnormality detection using the generation model, it is possible to allow the user to grasp the inspection accuracy by the generation model before performing the abnormality detection.
Advantages of the Invention
[0024] As described above, in the information processing apparatus, information processing method, and information processing program according to the present disclosure, when performing abnormality detection using a generation model capable of outputting the presence or absence of an abnormality of an object from an object image showing the object, it is possible to allow the user to grasp the inspection accuracy by the generation model before performing the abnormality detection.
Brief Description of the Drawings
Embodiments for Carrying Out the Invention
[0026] Hereinafter, the information processing apparatus 20 according to the present embodiment will be described. FIG. 1 is a block diagram showing the hardware configuration of the information processing apparatus 20. As an example, a general-purpose computer device such as a server computer or a PC (Personal Computer), or a mobile terminal such as a smartphone or a tablet terminal is applied to the information processing apparatus 20.
[0027] As shown in Figure 1, the information processing device 20 includes a CPU (Central Processing Unit) 21, ROM (Read Only Memory) 22, RAM (Random Access Memory) 23, storage 24, input unit 25, display unit 26, and communication unit 27. Each component is connected to the others via a bus 28 so as to be able to communicate with each other. The information processing device 20 is an example of the "information processing device" of this disclosure.
[0028] The CPU 21 is a central processing unit that executes various programs and controls various components. Specifically, the CPU 21 reads a program from the ROM 22 or storage 24 and executes the program using the RAM 23 as a working area. The CPU 21 controls each of the above components and performs various calculations according to the program stored in the ROM 22 or storage 24. The CPU 21 is an example of a "processor" in this disclosure.
[0029] ROM22 stores various programs and data. RAM23 temporarily stores programs or data as a working area.
[0030] Storage 24 includes HDD (Hard Disk Drive) and SSD (Solid). It consists of a storage device such as a State Drive or flash memory, and stores various programs and data.
[0031] The input unit 25 includes, for example, a pointing device such as a mouse, various buttons, a keyboard, a microphone, and a camera, and is used for various types of input.
[0032] The display unit 26 is, for example, a liquid crystal display and displays various information. The display unit 26 may also function as an input unit 25 by employing a touch panel method. The display unit 26 is an example of the "display unit" in this disclosure.
[0033] The communication unit 27 is an interface for communicating with other devices. For such communication, a wired communication standard such as Ethernet® or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi® may be used.
[0034] Next, the configuration of the storage 24 of the information processing device 20 will be described. Figure 2 is a block diagram showing the configuration of the storage 24 of the information processing device 20.
[0035] As shown in Figure 2, the storage 24 stores the information processing program 24A and the generation model 24B.
[0036] The information processing program 24A is a program that causes the CPU 21 to execute various processes described later. When executing the information processing program 24A, the information processing device 20 uses the hardware resources shown in Figure 1 to execute the processes based on the information processing program 24A. The information processing program 24A is an example of an "information processing program" in this disclosure.
[0037] The generative model 24B is a so-called generative AI (Artificial Intelligence). The generative model 24B is a model that can output whether or not there are abnormalities in an item shown in an image of an item that is subject to visual inspection. An example of a component of the generative model 24B is CLIP (Internet Search Engine Provider).<URL: https: / / trail.t.u-tokyo.ac.jp / ja / blog / 22-12-02-clip / > ) are examples. Generative models 24B include components that can associate images and text, such as CLIP, and image generation and transformers such as diffusion models and GANs. Using text decoders such as mer as appropriate, a multimodal generation AI This is how it is constructed. However, these are merely examples and not limiting. The generative model 24B is input with multimodal prompts, including text data representing text and image data representing images. The generative model 24B performs inference according to the instructions shown by the multimodal prompts and outputs the inference results in data formats such as image data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0038] Next, we will describe the functional configuration of CPU21. Figure 3 is a block diagram showing an example of the functional configuration of CPU21.
[0039] As shown in Figure 3, the CPU 21 includes an input unit 21A, a processing unit 21B, and an output unit 21C.
[0040] The input unit 21A acquires user input received by the information processing device 20. Specifically, the input unit 21A acquires at least one input data from text, audio, images, and video received by the information processing device 20.
[0041] The processing unit 21B performs specific processing using the generation model 24B. Specifically, the processing unit 21B inputs a predetermined multimodal prompt to the generation model 24B and obtains the output result.
[0042] The output unit 21C outputs the output result of the generation model 24B, that is, the result of the specific processing, to the display unit 26. As a result, the result of the specific processing is displayed on the display unit 26.
[0043] Figure 4 is a flowchart showing the flow of a specific process executed by the information processing device 20. The specific process is performed when the CPU 21 reads the information processing program 24A from the storage 24, loads it into the RAM 23, and executes it. The specific process is performed by the CPU 21 functioning as the input unit 21A, processing unit 21B, and output unit 21C. As an example... The specific process is performed when the user executes a designated application.
[0044] In step S10 shown in Figure 4, the CPU 21 acquires various information received from the user input. Then, the CPU 21 proceeds to step S11. Here, the various information includes at least an image showing a specific item that is the subject of the visual inspection (hereinafter referred to as the "specific item image"). In the following explanation, "screws" will be used as an example of the item (see Figure 5, etc.), but the type of item is not limited to this. For example, in addition to "screws," various products such as "fabric products," "cosmetics (e.g., foundation)," and "circuit boards" can be used as the item.
[0045] In step S11, the CPU 21 generates a multimodal prompt to output whether or not there is an abnormality in the specific item shown in the specific item image, based on the various information acquired in step S10. Then, the CPU 21 proceeds to step S12.
[0046] As an example, the CPU 21 generates a multimodal prompt like the one below, which includes at least the image of a specific item acquired in step S10 and a predetermined instruction. Note that the instruction shown in the following multimodal prompt is merely an example and is not limited to it.
[0047] "Multimodal prompt" Instructions: Output whether or not the specific item shown in the following image is defective. The output format should be at least one of image data and text data. Specific item image: screwAAA.jpg
[0048] In step S12, the CPU 21 inputs the multimodal prompt generated in step S11 to the generation model 24B and obtains the output result from the generation model 24B. Then, the CPU 21 proceeds to step S13.
[0049] In step S13, the CPU 21 outputs the output result of the generation model 24B in step S12 to the display unit 26. As a result, the display unit 26 displays the result of the specific process. Then, the CPU 21 terminates the specific process.
[0050] Next, an example of the display according to this embodiment will be described. Figure 5 is a first example of the display shown on the display unit 26 of the information processing device 20. Specifically, Figure 5 shows a verification screen for verifying the inspection accuracy by the generation model 24B.
[0051] The verification screen shown in Figure 5 displays GUI (Graphical User Interface) buttons 40-46 and the verification area 50.
[0052] GUI button 40 is a button for selecting a generation model 24B to verify the inspection accuracy of the visual inspection. In this embodiment, multiple generation models 24B with different parameters related to the output content (e.g., generation model 24B-1, generation model 24B-2, etc.) are provided. The user can select the desired generation model 24B by operating the GUI button 40. If the GUI button 40 is not operated, a pre-configured generation model 24B will be selected. Note that at least one model may be selected as the target of the pre-configured generation model 24B, and accuracy verification may be performed on multiple models, or the model with the optimal accuracy and verification result may be automatically selected.
[0053] GUI button 41 is a button for selecting at least one image of a specific item to be inspected for abnormalities. The image of the specific item may also have data attached indicating the location and characteristics of the abnormality. After operating GUI button 41, the user accesses storage 2 A desired image of a specific item is selected by selecting an image from within 4 or by taking a picture with the camera of the information processing device 20. The screw image 60 selected as the specific item image is displayed in the verification area 50 shown in Figure 5. If multiple images of specific items are selected, they may be displayed in a list, or they may be viewable sequentially using an image scrolling function. In addition, at least one of the automatically selected images to be judged may be displayed in the verification area 50 for purposes such as re-verifying the image to be judged during actual operation of the visual inspection.
[0054] GUI button 42 is a button for setting anomaly detection conditions by the generation model 24B. These detection conditions include a threshold, the nature of the anomaly, and its location. The user can set these detection conditions by operating GUI button 42. As a result, the CPU 21 accepts the detection conditions set by the user as an anomaly definition that indicates the criteria for determining whether or not a particular item is an anomaly. This anomaly definition may be freely entered by the user, or it may be defined by the user modifying a suggestion from the software. Furthermore, if data indicating the correct answer for an anomaly is attached to the image of the specific item, the conditions that result in the highest judgment accuracy may be automatically searched for, for example.
[0055] GUI button 43 is a button for entering supplementary text that complements the characteristics of a specific item. The supplementary text is text that enhances the specificity of the item, such as a specific product name and defect information. The user can enter the supplementary text by operating GUI button 43. The input of the supplementary text may be in the form of the user directly typing into a text box, or it may be in the form of the user modifying suggestions from the software.
[0056] GUI button 44 is a button for inputting a reference image of a good product that shows no abnormalities in a specified item. The specified item may be the same type of item as the specific item, or it may be a different type of item. After operating the GUI button 44, the user selects the desired reference image of a good product by selecting the desired image from the storage 24 or by taking a picture with the camera of the information processing device 20.
[0057] GUI button 45 is a button for inputting a reference defective product image that shows a condition where a specified item has an abnormality. After operating GUI button 45, the user selects the desired reference defective product image by selecting an image from storage 24 or by taking a picture with the camera of the information processing device 20. Furthermore, the user also inputs text indicating the location and content of the abnormality in the selected reference defective product image.
[0058] GUI button 46 is used to verify the inspection accuracy of the visual inspection performed by the generation model 24B. When GUI button 46 is operated, the CPU 21 generates a multimodal prompt based on the settings configured based on the operation of GUI buttons 40 to 45. The CPU 21 then inputs this multimodal prompt into the generation model 24B to output whether or not there is an abnormality in the specific item (e.g., screw image 60) shown in the verification area 50. If multiple specific item images are selected, the contents of multiple images can be viewed using a list display of the presence or absence of abnormalities or an image scrolling function. If data indicating the correct answer for abnormalities is attached to the specific item image, statistical information such as the accuracy of individual images and the overall judgment accuracy may also be displayed. Note that, among GUI buttons 40 to 45, when detection conditions are automatically searched and suggested, inputting abnormality detection conditions in GUI button 42, inputting auxiliary text in GUI button 43, and inputting reference good product images and reference defective product images in GUI buttons 44 and 45 are optional. Even without these inputs, the generation model 24B can output whether or not there is an abnormality in the specific item based on the operation of the GUI button 46.
[0059] Figure 6 is a specific example of a reference image showing a given item in both its normal and abnormal states. As an example, in Figure 6, the given item is a "screw" of the same type as a specific item.
[0060] Figure 6(A) shows screw image 62, which indicates a screw without any abnormalities. Figure 6(B) shows screw image 64, which indicates a screw with an abnormality. For example, if the user operates the GUI button 45 and selects screw image 64 as the reference image of a defective product, they can input text indicating the location and nature of the abnormality, such as "Bent tip of screw."
[0061] Figure 7 shows a second display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 7 shows the state after the GUI button 46 has been operated on the verification screen shown in Figure 5.
[0062] In the verification screen shown in Figure 7, text data 65 indicating whether or not there is an abnormality in the screw shown in the screw image 60 is displayed below the screw image 60 within the verification area 50. Specifically, the text data 65 says "No abnormalities found," indicating that there is no abnormality in the screw shown in the screw image 60.
[0063] Figure 8 shows a third display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 8 shows a state in which a screw image 70, which represents a screw different from the screw image 60, is displayed as a specific item image in the verification area 50.
[0064] Figure 9 shows a fourth display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 9 shows the state after the GUI button 46 has been operated on the verification screen shown in Figure 8.
[0065] In the verification screen shown in Figure 9, text data 75 indicating whether or not there is an abnormality in the screw shown in the screw image 70 is displayed below the screw image 70 within the verification area 50. Specifically, the text data 75 reads, "There is an abnormality. The tip of the screw is bent," indicating that there is an abnormality in the screw shown in the screw image 70 and the nature of that abnormality.
[0066] Here, the display method of the screw image 70 shown in Figure 9 differs from the display method of the screw image 70 shown in Figure 8. Specifically, in the screw image 70 shown in Figure 9, a heat map is superimposed on the screw image to indicate abnormal areas. Note that in Figure 9, for illustrative purposes, the heat map is represented by a pattern of diagonal hatching and dots. In the screw image 70, the area 70A from the head of the screw to near the tip is represented by a first color (diagonal hatching in the figure), and the tip area 70B is represented by a second color different from the first color (dots in the figure). The second color is a predetermined color (e.g., red, yellow) that indicates the area with an abnormality.
[0067] As described above, the verification screen shown in Figure 9 indicates, through the display of the screw image 70 and the text data 75, that there is an abnormality in the screw shown in the screw image 70 and the details of that abnormality.
[0068] Figure 10 shows a fifth display example shown on the display unit 26 of the information processing device 20. As an example, Figure 10 shows an adjustment screen for adjusting at least one of the multimodal prompts input to the generation model 24B and the parameters of the generation model 24B related to the output content. The adjustment screen is displayed when a predetermined operation is performed on the verification screen.
[0069] The adjustment screen shown in Figure 10 displays GUI buttons 80 to 83. GUI button 80 is a button for selecting the generation model 24B to which the multimodal prompt and parameters will be adjusted. The user can select the desired generation model 24B by operating GUI button 80.
[0070] GUI button 81 is a button for selecting the training images to be used for adjusting the multimodal prompts and parameters described above. The training images include images of good products showing no abnormalities in a specific item (e.g., a screw) and an item of the same type (e.g., a screw), and images of defective products showing abnormalities in the same item. After operating GUI button 81, the user selects at least one desired good product image and one desired defective product image by selecting a desired image from storage 24 or by taking a picture with the camera of the information processing device 20. The user may also input text, text data, or structured data such as JSON or YAML that indicates the location and content of the abnormality in the selected defective product image.
[0071] GUI button 82 is a button for entering supplementary text that adds details about the characteristics of a specific item. Users can enter supplementary text by operating GUI button 82.
[0072] GUI button 83 is a button for adjusting at least one of the multimodal prompts and parameters. When GUI button 83 is operated, the CPU 21 adjusts at least one of the multimodal prompts and parameters based on the settings based on the operation of GUI buttons 80 to 82. The adjustment of the multimodal prompts or parameters is performed as appropriate using known techniques.
[0073] Figure 11 shows a sixth display example shown on the display unit 26 of the information processing device 20. As an example, Figure 11 shows the operation screen during actual operation of the visual inspection.
[0074] The operation screen shown in Figure 11 shows GUI buttons 90-93 and a verification area 50. In actual operation, an image of an item captured by a camera (not shown) is displayed in the verification area 50, and the information processing device 20 automatically determines whether or not there is an abnormality each time an image is taken. In the verification area 50 shown in Figure 11, a screw image 100, which represents a screw captured by the camera, is displayed.
[0075] GUI button 90 is for selecting the generation model 24B to be used in the actual operation of visual inspection. By operating this GUI button 90, the user can select the desired generation model 24B. However, the selection is not limited to the user; for example, the model with the highest accuracy in the verification results may be automatically selected.
[0076] GUI button 91 is a button for setting the anomaly detection conditions by the generation model 24B. The user can set these detection conditions by operating GUI button 91.
[0077] GUI button 92 is a button for selecting whether to enable or disable online learning. Users can switch online learning on or off by operating GUI button 92. In this embodiment, online learning refers to a function that adjusts the parameters of the generative model 24B during visual inspection to improve anomaly detection performance. Specifically, this includes methods such as performing learning by assuming the first XX items during visual inspection are good products, or receiving user feedback on good / defective product judgments on the spot and performing learning in real time or retrospectively. With this configuration, in addition to the user explicitly selecting images to train the model at times other than during actual visual inspection, the generative model 24B can continuously improve its accuracy by automatically incorporating feedback.
[0078] GUI button 93 is a button to start the actual operation of visual inspection using the generation model 24B. When the GUI button 93 is operated, the CPU 21 generates a multimodal prompt based on the settings configured based on the operation of GUI buttons 90 to 92. The CPU 21 then inputs the multimodal prompt into the generation model 24B, which outputs whether or not there is an abnormality in the item shown in the item image (e.g., screw image 100) of the item displayed in the verification area 50.
[0079] In the operation screen shown in Figure 11, text data 105 indicating whether or not there is an abnormality in the screw shown in the screw image 100 is displayed below the screw image 100 within the verification area 50. Specifically, the text data 105 is "There is an abnormality. The tip of the screw is bent," indicating that there is an abnormality in the screw shown in the screw image 100 and the nature of the abnormality. The nature of the abnormality may be written in conjunction with the location of the abnormality, for example, but the display location is not particularly limited. Furthermore, there are no particular restrictions on the nature of the abnormality itself, such as using quantitative or qualitative expressions to describe the degree of the abnormality.
[0080] Furthermore, the display method for screw images 100 that have been determined to have an abnormality is the same as for screw image 70 shown in Figure 9, where a heat map is superimposed on the screw image to indicate the abnormal area. In the screw image 100, the region 100A from the head of the screw to near the tip is represented in the first color, and the region 100B at the tip is represented in the second color.
[0081] Note that the GUI buttons shown in Figure 11 represent only a portion of the operation screen, and other GUI buttons actually exist. For example, above GUI button 90, there is an image input source selection button. This image input source selection button allows you to select the source from which to receive images, such as images saved in a designated folder or images taken directly from the camera.
[0082] As explained above, in the information processing device 20, the CPU 21 receives input of an image of an item that is the subject of the visual inspection. The CPU 21 also inputs a multimodal prompt to the generation model 24B that includes the received image of the specific item and a predetermined instruction. The CPU 21 then displays on the display unit 26 whether or not there is an abnormality in the specific item shown in the image of the specific item, which is the output result of the generation model 24B. Thus, according to the information processing device 20, by inputting the multimodal prompt to the generation model 24B before performing the visual inspection using the generation model 24B, the user can understand the inspection accuracy of the generation model 24B before the visual inspection is performed. Furthermore, according to the information processing device 20, by inputting the multimodal prompt to the generation model 24B during the actual operation of the visual inspection, the user can understand the inspection accuracy of the generation model 24B while the visual inspection is being performed.
[0083] Furthermore, the generation model 24B identifies the type of a specific item from the characteristics of the input specific item image, and infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific item. As a result, the information processing device 20 can automate everything from item type identification to abnormality detection using the generation model 24B, thereby improving the accuracy of visual inspection, increasing work efficiency, and reducing costs.
[0084] Furthermore, if the generation model 24B infers that there is an abnormality in a particular item, it generates at least one of image data and text data indicating the nature of the abnormality. Then, in the information processing device 20, the CPU 21 displays at least one of the image data and text data indicating the nature of the abnormality of the particular item, which are the output results of the generation model 24B, on the display unit 26. As a result, according to the information processing device 20, the nature of the abnormality of the particular item is displayed on the display unit 26 in the form of at least one of the image and text, allowing the user to intuitively and concretely grasp the problem area, making it easier to take prompt action and manage records.
[0085] Furthermore, in the information processing device 20, the CPU 21 receives input of supplementary text that supplements the characteristics of a specific item, and a reference image showing at least one of the states in which the specified item is free from abnormalities and in which it is abnormal. The CPU 21 then inputs a multimodal prompt to the generation model 24B that includes the image of the specific item, a predetermined instruction sentence, and at least one of the received supplementary text and reference image. As a result, the information processing device 20 deepens the understanding of the generation model 24B by adding specific information regarding the characteristics of the item and whether or not it is abnormal, enabling more accurate anomaly detection and reducing false detections and missed detections.
[0086] Furthermore, in the information processing device 20, the CPU 21 accepts input of an abnormality definition that indicates the criteria for determining whether or not a particular item is abnormal. As a result, the information processing device 20 can accept input of an abnormality definition, allowing for flexible judgment criteria to be set according to the user's application, improving the detection accuracy of the generation model 24B and increasing its adaptability to specific applications.
[0087] Furthermore, in the information processing device 20, the CPU 21 receives at least one input: supplementary text that supplements the characteristics of a specific item, an image of a good product showing that an item of the same type as the specific item is free of defects, and an image of a defective product showing that an item of the same type has defects. Based on the supplementary text, the image of a good product, and the image of a defective product that the CPU 21 has received as input, it adjusts at least one of the multimodal prompts to be input to the generation model 24B and the parameters of the generation model 24B related to the output content. Thus, according to the information processing device 20, by adjusting the parameters of the generation model 24B based on information about the characteristics and abnormal state of a specific item, the abnormality detection of the generation model 24B can be optimized for a specific item and specific conditions. In addition, according to the information processing device 20, by optimizing for multiple items, the general detection accuracy for visual inspection of items can also be improved. Furthermore, according to the information processing device 20, by adjusting the multimodal prompts based on information about the characteristics and abnormal state of a specific item, the generation model 24B can understand the characteristics and abnormal state of the specific item more accurately, and the accuracy of the output results can be improved.
[0088] (others) While embodiments of the present disclosure have been described in detail above with reference to the attached drawings, the technical scope of the present disclosure is not limited to these examples. It is clear that a person with ordinary skill in the art of the present disclosure may conceive of various modifications or alterations within the scope of the technical idea set forth in the claims, and these modifications or alterations are also understood to fall within the technical scope of the present disclosure.
[0089] Furthermore, the effects described in the above embodiments are descriptive or illustrative, and are not limited to those described in the above embodiments. In other words, the technology relating to this disclosure may produce other effects that would be obvious to a person of ordinary skill in the art of this disclosure from the descriptions in the above embodiments, in addition to or in lieu of the effects described in the above embodiments.
[0090] The processing described in the above embodiment can also be implemented using dedicated hardware circuits. In this case, it may be executed on one piece of hardware or on multiple pieces of hardware.
[0091] In the above embodiment, the input of auxiliary text may be performed in an interactive format between the user and the generation model 24B. This allows the information processing device 20 to flexibly acquire the necessary information by supplementing the characteristics of a specific item in an interactive format with the user, enabling anomaly detection that is in line with the user's intentions.
[0092] In the above embodiment, the verification screen shown in Figure 5, etc., allows the input of supplementary text and reference images based on the user's selection, but the timing of these inputs is not limited. For example, instead of a configuration that allows input at any time the user desires, a configuration may be adopted in which the generation model 24B requests the user to input at least one of the supplementary text and reference images when certain conditions are met.
[0093] For example, the "case where the predetermined conditions are met" described above can be defined as the case where the difference in the confidence level (%) of whether or not an anomaly is present, inferred by the generative model 24B without input of auxiliary text and reference images, is less than or equal to a predetermined value.
[0094] Furthermore, the system may be configured such that, when certain conditions are met, the generation model 24B requests the user to input either supplementary text or a reference image, and when specific conditions are met, the generation model 24B requests the user to input the other supplementary text or reference image.
[0095] For example, the "case where specific conditions are met" described above can be defined as the case where the difference in confidence levels of whether or not an anomaly is inferred by the generative model 24B after inputting either the auxiliary text or the reference image is less than or equal to a predetermined value.
[0096] With the above configuration, the information processing device 20 can deepen its understanding of the generation model 24B by adding specific information regarding the characteristics of the item and whether or not there are abnormalities when the reliability of the output of the generation model 24B is low, thereby enabling more accurate anomaly detection and reducing false detections and missed detections.
[0097] Furthermore, based on the content of the supplementary text entered by the user, the generative model 24B may suggest to the user what supplementary text to input to improve the reliability of the output. For example, when the generative model 24B accepts supplementary text input in an interactive format, if the content of the supplementary text entered by the user is "product name," it may respond with a suggestion such as "Please also enter a specific example of the defect." In addition, when the generative model 24B accepts supplementary text input in an interactive format, it may also make suggestions such as "The product name is foundation, and the defect is uneven color."
[0098] In the above embodiment, the method of representing the heatmap superimposed on the screw image is not particularly limited. For example, multiple colors corresponding to the degree of abnormality may be represented in the heatmap. Alternatively, instead of using a heatmap to indicate information about the abnormality, marks indicating the location and area of the abnormality may be used. Furthermore, the degree of abnormality may be indicated by adding numerical values or text.
[0099] The configuration described in the above embodiment is not limited to the visual inspection of goods, but can be extended to general "anomaly detection" such as hazard prediction (unsafe behavior or illness by workers), foreign object contamination, and detection of suspicious persons.
[0100] In the above embodiment, an example was described in which an image of an item to be inspected is received as input, but the system is not limited to this. For example, instead, an image of the item may be received as input.
[0101] In the above embodiment, the generation model 24B can output not only a binary output of "presence or absence of an abnormality" but also "degree of abnormality." Here, "degree of abnormality" is one element that constitutes "content of the abnormality." "Content of the abnormality" includes, for example, "location of the abnormality," "degree of abnormality," and "nature of the abnormality (type of abnormality)."
[0102] In the above embodiment, the generated model 24B directly provides image data indicating whether or not there is an abnormality in the article. In addition to generating text data, this also includes cases where the generation model 24B outputs the degree of anomaly as numerical data in map format, and internally performs a process to generate image data and text data indicating the presence or absence of anomalies based on that numerical data and the input image.
[0103] In the above embodiment, the CPU 21 may save the output results of the generative model 24B, accept user modifications to the saved output results, and train the generative model 24B using the modified data. This creates a loop for improving the accuracy of the generative model 24B, enabling continuous improvement of anomaly detection performance.
[0104] In the above embodiment, the "characteristics of the article" entered as auxiliary text includes not only information indicating what the article itself is (e.g., foundation), but also information about specific abnormalities to be detected (e.g., type: uneven color, location: on powder, degree: very strong, etc.).
[0105] In the above embodiment, the dialogue format between the user and the generation model 24B when inputting auxiliary text is not particularly limited and may be, for example, a QA format (e.g., What is the product? → Foundation → What specific defect features do you want to detect? → Color unevenness) or a suggestion dialogue format (e.g., Is the product foundation? → Yes). Alternatively, the input format for auxiliary text may not depend on the dialogue format, and the generation model 24B may simply propose the auxiliary text entirely and present it in a format that the user can modify.
[0106] In each of the above embodiments, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).
[0107] Furthermore, the processor operations in each of the above embodiments may not be performed by a single processor, but may also be performed by multiple processors located in physically separate locations working together. Alternatively, some or all of the operations performed by specific multiple processors in each of the above embodiments may be integrated and performed by a single processor. In addition, the order of the processor operations is not limited to the order described in each of the above embodiments, and may be changed as appropriate.
[0108] Furthermore, although the above embodiment describes an embodiment in which the information processing program 24A is pre-stored (installed) in the storage 24, the invention is not limited thereto. The information processing program 24A may be provided in the form of a recording medium such as a CD-ROM (Compact Disk Read Only Memory), DVD-ROM (Digital Versatile Disk Read Only Memory), and USB (Universal Serial Bus) memory. Alternatively, the information processing program 24A may be provided in the form of a download from an external device via a network N. The technology of this disclosure can also be applied to programs and program products.
[0109] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference. [Explanation of Symbols]
[0110] 20 Information Processing Devices 21 CPU (Processor) 24A Information Processing Program 24B Generative Model 26 Display section
Claims
1. Equipped with a processor, The aforementioned processor, The system accepts input of an object image that indicates the object to be detected for anomaly detection. A multimodal prompt, including a specific object image received as input and a predetermined instruction, is input to a generative model capable of outputting whether or not there is an abnormality in the object shown in the object image. The display unit displays whether or not there is an abnormality in the specific object shown in the specific object image, which is the output result of the generation model. If predetermined conditions are met, the generation model prompts the user to input supplementary text that complements the characteristics of the specific object that received the input. The user provides input of the aforementioned supplementary text. A multimodal prompt including the aforementioned specific object image, the predetermined instruction text, and the auxiliary text received as input is input to the generation model. Information processing device.
2. The generation model identifies the type of the specific object from the features of the input image of the specific object, infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific object. The information processing apparatus according to claim 1.
3. If the generation model infers that there is an abnormality in the specific object, it generates at least one of image data and text data indicating the nature of the abnormality. The processor causes the display unit to display at least one of image data and text data indicating the content of the abnormality of the specific object, which is the output result of the generation model. The information processing apparatus according to claim 2.
4. The aforementioned processor, The system accepts input of the aforementioned auxiliary text and a reference image showing at least one of the states in which the specified object is normal and in which it is abnormal. The generative model is input a multimodal prompt that includes the aforementioned specific object image, the predetermined instruction text, and at least one of the auxiliary text and the reference image that has been received as input. The information processing apparatus according to claim 1.
5. The input of the auxiliary text is performed through an interactive format between the user and the generative model. The information processing apparatus according to claim 4.
6. The aforementioned processor, The system accepts input of a definition of abnormality, which indicates the criteria for determining whether or not an abnormality exists in the aforementioned specific object. The information processing apparatus according to claim 1.
7. The aforementioned processor, The system accepts at least one input: the aforementioned auxiliary text, an image showing no abnormality in an object of the same type as the specific object, and an image showing an abnormality in the same object. Based on at least one of the auxiliary text received as input, the image without abnormalities, and the image with abnormalities, the multimodal prompts input to the generation model and at least one of the parameters of the generation model relating to the output content are adjusted. The information processing apparatus according to claim 1.
8. The system accepts input of an object image that indicates the object to be detected for anomaly detection. A multimodal prompt, including a specific object image received as input and a predetermined instruction, is input to a generative model capable of outputting whether or not there is an abnormality in the object shown in the object image. The display unit displays whether or not there is an abnormality in the specific object shown in the specific object image, which is the output result of the generation model. If predetermined conditions are met, the generation model prompts the user to input supplementary text that complements the characteristics of the specific object that received the input. The user provides input of the aforementioned supplementary text. A multimodal prompt including the aforementioned specific object image, the predetermined instruction text, and the auxiliary text received as input is input to the generation model. An information processing method in which a computer performs the processing.
9. The system accepts input of an object image that indicates the object to be detected for anomaly detection. A multimodal prompt, including a specific object image received as input and a predetermined instruction, is input to a generative model capable of outputting whether or not there is an abnormality in the object shown in the object image. The display unit displays whether or not there is an abnormality in the specific object shown in the specific object image, which is the output result of the generation model. If predetermined conditions are met, the generation model prompts the user to input supplementary text that complements the characteristics of the specific object that received the input. The user provides input of the aforementioned supplementary text. A multimodal prompt including the aforementioned specific object image, the predetermined instruction text, and the auxiliary text received as input is input to the generation model. An information processing program that causes a computer to perform a task.