Information processing device, information processing method, and information processing program

JP2026123665AActive Publication Date: 2026-07-30OUEN CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
OUEN CO LTD
Filing Date
2025-01-17
Publication Date
2026-07-30

AI Technical Summary

Benefits of technology

【0024】 以上説明したように、本開示に係る情報処理装置、情報処理方法、及び情報処理プログラムでは、物品が示された物品画像から当該物品の異常の有無を出力可能な生成モデルを用いて外観検査を実施する際に、当該生成モデルによる検査精度を外観検査の実施前にユーザに把握させることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026123665000001_ABST
    Figure 2026123665000001_ABST
Patent Text Reader

Abstract

The present invention provides an information processing device, an information processing method, and an information processing program that allow a user to understand the inspection accuracy of a generative model used to perform a visual inspection using an image of an item that shows the item, and to understand the accuracy of the inspection by the generative model before performing the visual inspection. [Solution] The information processing method executed by the processor of the information processing device receives an input of an item image showing an item to be inspected visually, inputs a multimodal prompt including the specific item image received as input and a predetermined instruction sentence into a generation model capable of outputting whether or not there is an abnormality in the item shown in the item image, and displays the output result of the generation model, which is whether or not there is an abnormality in the specific item shown in the specific item image, on the display unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Patent Document 1 discloses a technique for improving the accuracy of estimating an equipment state or a dangerous state that violates a safety manual based on a field image captured by a camera.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, it is known to perform an appearance inspection of an article in order to guarantee the quality of the product and maintain and improve customer satisfaction. Here, in the appearance inspection by a human, there are problems such as variations in inspection accuracy depending on the skill level and experience of the operator. Therefore, it is required to suppress variations in inspection accuracy in the appearance inspection of articles.

[0005] Therefore, an object of the present disclosure is to provide an information processing apparatus, an information processing method, and an information processing program that can allow a user to grasp the inspection accuracy by a generation model before performing an appearance inspection when performing the appearance inspection using a generation model capable of outputting the presence or absence of an abnormality of an article from an article image in which the article is shown.

Means for Solving the Problems

[0006] The first embodiment of the information processing device includes a processor, which receives an input of an item image showing an item to be inspected visually, inputs a multimodal prompt including the specific item image received as input and a predetermined instruction sentence to a generation model capable of outputting whether or not there is an abnormality in the item shown in the item image, and causes a display unit to display the output result of the generation model, which is whether or not there is an abnormality in the specific item shown in the specific item image.

[0007] In the first embodiment of the information processing device, the processor receives an input of an image of an item that is to be visually inspected. The processor then inputs a multimodal prompt, which includes the received specific image of an item and a predetermined instruction, into a generation model, and causes the display unit to show whether or not there is an abnormality in the specific item shown in the specific image of an item, which is the output result of the generation model. As a result, the information processing device allows the user to understand the inspection accuracy of the generation model before performing the visual inspection by inputting the multimodal prompt into the generation model before performing the visual inspection.

[0008] The second embodiment of the information processing apparatus is the same as the first embodiment, wherein the generation model identifies the type of a specific item from the characteristics of the input specific item image, infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific item.

[0009] In the second embodiment of the information processing device, the generation model identifies a specific type of item from the features of a specific item image input, and infers whether or not there is an abnormality corresponding to that type according to instructions indicated by a predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific item. As a result, the information processing device can automate everything from item type identification to abnormality detection using the generation model, thereby improving the accuracy of visual inspection, increasing work efficiency, and reducing costs.

[0010] The third embodiment of the information processing apparatus is the same as the second embodiment, wherein the generation model, when it infers that there is an abnormality in the specific item, generates at least one of image data and text data indicating the nature of the abnormality, and the processor causes the display unit to display at least one of the image data and text data indicating the nature of the abnormality of the specific item, which is the output result of the generation model.

[0011] In the third embodiment of the information processing device, when the generation model infers that there is an abnormality in a particular item, it generates at least one of image data and text data indicating the nature of the abnormality. The processor then causes the display unit to display at least one of the image data and text data indicating the nature of the abnormality of the particular item, which are the output results of the generation model. As a result, the information processing device displays the nature of the abnormality of the particular item on the display unit as at least one of the image and text, allowing the user to intuitively and concretely grasp the problem area, facilitating quick response and record management.

[0012] The fourth embodiment of the information processing device is an information processing device of any one of the first to third embodiments, wherein the processor receives input of supplementary text that supplements the characteristics of the specific article and a reference image showing at least one of the states in which the predetermined article is normal and in which it is abnormal, and inputs a multimodal prompt to the generation model that includes the image of the specific article, the predetermined instruction sentence, and at least one of the supplementary text and the reference image that was received as input.

[0013] In the fourth embodiment of the information processing device, the processor receives input of supplementary text that supplements the characteristics of a specific article and a reference image showing at least one of the states in which the given article is free from abnormalities and in which it is abnormal. The processor then inputs a multimodal prompt to the generative model, which includes the image of the specific article, a predetermined instruction, and at least one of the received supplementary text and reference image. As a result, the information processing device can deepen the understanding of the generative model by adding specific information regarding the characteristics of the article and whether or not it is abnormal, enabling more accurate anomaly detection and reducing false positives and missed detections.

[0014] The fifth embodiment of the information processing apparatus is the same as the fourth embodiment, wherein the input of the auxiliary text is performed in an interactive format between the user and the generation model.

[0015] In the fifth embodiment of the information processing device, the input of auxiliary text is performed through an interactive format between the user and the generative model. This allows the information processing device to flexibly acquire the necessary information by supplementing the characteristics of specific items through an interactive format with the user, enabling anomaly detection that aligns with the user's intentions.

[0016] The sixth embodiment of the information processing device is one of the first to fifth embodiments, wherein the processor receives input of an abnormality definition indicating the criteria for determining whether or not a particular item is abnormal.

[0017] In the sixth embodiment of the information processing device, the processor accepts input of an abnormality definition that indicates the criteria for determining whether a particular item is abnormal or not. As a result, the information processing device can set flexible judgment criteria according to the user's application by accepting input of an abnormality definition, thereby improving the detection accuracy of the generation model and increasing its adaptability to specific applications.

[0018] The seventh embodiment of the information processing apparatus is any one of the first to sixth embodiments, wherein the processor receives at least one input of supplementary text that supplements the characteristics of the specific article, a good product image showing an article of the same type as the specific article without defects, and a defective product image showing an article of the same type with defects, and adjusts at least one of the parameters of the generation model relating to a multimodal prompt to be input to the generation model and the output content based on the supplementary text, the good product image, and the defective product image received as input.

[0019] In the seventh embodiment of the information processing device, the processor receives at least one input: supplementary text that supplements the characteristics of a specific article, an image of a good product showing that an article of the same type as the specific article is free of defects, and an image of a defective product showing that an article of the same type has defects. Based on the supplementary text, the image of a good product, and the image of a defective product that the processor receives as input, the processor adjusts at least one of the multimodal prompts to be input to the generation model and the parameters of the generation model related to the output content. As a result, the information processing device can optimize the anomaly detection of the generation model for a specific article and specific conditions by adjusting the parameters of the generation model based on information about the characteristics and abnormal state of the specific article. Furthermore, the information processing device can improve the accuracy of the output results by enabling the generation model to understand the characteristics and abnormal state of the specific article more accurately by adjusting the multimodal prompts based on information about the characteristics and abnormal state of the specific article.

[0020] The information processing method of the eighth embodiment involves a computer receiving an input of an item image showing an item to be inspected visually, inputting a multimodal prompt including the specific item image received as input and a predetermined instruction sentence into a generation model capable of outputting whether or not there is an abnormality in the item shown in the item image, and displaying the output result of the generation model, which is whether or not there is an abnormality in the specific item shown in the specific item image, on a display unit.

[0021] In the eighth aspect of the information processing method, the computer performs a process to receive input of an image of an item that is to be visually inspected. The computer also inputs a multimodal prompt, which includes the received specific image of an item and a predetermined instruction, into a generation model, and performs a process to display on the display unit whether or not there is an abnormality in the specific item shown in the specific image of an item, which is the output result of the generation model. As a result, according to the information processing method, by inputting the multimodal prompt into the generation model before performing the visual inspection using the generation model, the user can understand the inspection accuracy of the generation model before performing the visual inspection.

[0022] The information processing program of the ninth embodiment receives an input of an item image showing an item to be inspected visually, inputs a multimodal prompt including the specific item image received as input and a predetermined instruction sentence into a generation model capable of outputting whether or not there is an abnormality in the item shown in the item image, and causes a computer to execute a process that displays on a display unit whether or not there is an abnormality in the specific item shown in the specific item image, which is the output result of the generation model.

[0023] In the ninth aspect of the information processing program, the computer is instructed to perform a process to receive input of an image of an item that is the subject of a visual inspection. The computer then inputs a multimodal prompt, which includes the received specific image of the item and a predetermined instruction, into a generation model, and performs a process to display on the display unit whether or not there is an abnormality in the specific item shown in the specific image of the item, which is the output result of the generation model. As a result, according to the information processing program, by inputting the multimodal prompt into the generation model before performing the visual inspection using the generation model, the user can understand the inspection accuracy of the generation model before the visual inspection is performed. [Effects of the Invention]

[0024] As described above, when performing an appearance inspection using a generation model capable of outputting the presence or absence of an abnormality of an item from an item image showing the item in the information processing apparatus, information processing method, and information processing program according to the present disclosure, the inspection accuracy by the generation model can be grasped by the user before performing the appearance inspection.

Brief Description of Drawings

[0025] [Figure 1] It is a block diagram showing the hardware configuration of the information processing apparatus. [Figure 2] It is a block diagram showing the configuration of the storage of the information processing apparatus. [Figure 3] It is a block diagram schematically showing the functional configuration of the CPU of the information processing apparatus. [Figure 4] It is a flowchart showing the flow of specific processing executed by the information processing apparatus. [Figure 5] It is a first display example displayed on the display unit of the information processing apparatus. [Figure 6] It is a specific example of a reference image. [Figure 7] It is a second display example displayed on the display unit of the information processing apparatus. [Figure 8] It is a third display example displayed on the display unit of the information processing apparatus. [Figure 9] It is a fourth display example displayed on the display unit of the information processing apparatus. [Figure 10] It is a fifth display example displayed on the display unit of the information processing apparatus. [Figure 11] It is a sixth display example displayed on the display unit of the information processing apparatus.

Embodiments for Carrying Out the Invention

[0026] Hereinafter, the information processing apparatus 20 according to the present embodiment will be described. Figure 1 is a block diagram showing the hardware configuration of the information processing device 20. The information processing device 20 may include, as an example, a server computer, a general-purpose computer device such as a PC (Personal Computer), or a mobile device such as a smartphone or tablet.

[0027] As shown in Figure 1, the information processing device 20 includes a CPU (Central Processing Unit) 21, ROM (Read Only Memory) 22, RAM (Random Access Memory) 23, storage 24, input unit 25, display unit 26, and communication unit 27. Each component is connected to the others via a bus 28 so as to be able to communicate with each other. The information processing device 20 is an example of the "information processing device" of this disclosure.

[0028] The CPU 21 is a central processing unit that executes various programs and controls various components. Specifically, the CPU 21 reads a program from the ROM 22 or storage 24 and executes the program using the RAM 23 as a working area. The CPU 21 controls each of the above components and performs various calculations according to the program stored in the ROM 22 or storage 24. The CPU 21 is an example of a "processor" in this disclosure.

[0029] ROM22 stores various programs and data. RAM23 temporarily stores programs or data as a working area.

[0030] Storage 24 consists of storage devices such as HDDs (Hard Disk Drives), SSDs (Solid State Drives), or flash memory, and stores various programs and data.

[0031] The input unit 25 includes, for example, a pointing device such as a mouse, various buttons, a keyboard, a microphone, and a camera, and is used for various types of input.

[0032] The display unit 26 is, for example, a liquid crystal display and displays various information. The display unit 26 may also function as an input unit 25 by employing a touch panel method. The display unit 26 is an example of the "display unit" in this disclosure.

[0033] The communication unit 27 is an interface for communicating with other devices. For such communication, a wired communication standard such as Ethernet® or FDDI, or a wireless communication standard such as 4G, 5G, or Wi-Fi® may be used.

[0034] Next, the configuration of the storage 24 of the information processing device 20 will be described. Figure 2 is a block diagram showing the configuration of the storage 24 of the information processing device 20.

[0035] As shown in Figure 2, the storage 24 stores the information processing program 24A and the generation model 24B.

[0036] The information processing program 24A is a program that causes the CPU 21 to execute various processes described later. When executing the information processing program 24A, the information processing device 20 uses the hardware resources shown in Figure 1 to execute the processes based on the information processing program 24A. The information processing program 24A is an example of an "information processing program" in this disclosure.

[0037] The generative model 24B is a so-called generative AI (Artificial Intelligence). The generative model 24B is a model that can output whether or not there are abnormalities in an item shown in an image of an item that is subject to visual inspection. An example of a component of the generative model 24B is CLIP (Internet Search Engine Provider).<URL: https: / / trail.t.u-tokyo.ac.jp / ja / blog / 22-12-02-clip / > Examples include the following. The generative model 24B is configured as a multimodal generative AI by appropriately using components that can associate images and text, such as CLIP, and image generation such as diffusion models and GANs, and text decoders such as transformers. However, these are merely examples and are not limiting. The generative model 24B is input with a multimodal prompt that includes text data representing text and image data representing images. The generative model 24B performs inference according to the instructions shown by the multimodal prompt and outputs the inference results in data formats such as image data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0038] Next, we will describe the functional configuration of CPU21. Figure 3 is a block diagram showing an example of the functional configuration of CPU21.

[0039] As shown in Figure 3, the CPU 21 includes an input unit 21A, a processing unit 21B, and an output unit 21C.

[0040] The input unit 21A acquires user input received by the information processing device 20. Specifically, the input unit 21A acquires at least one input data from text, audio, images, and video received by the information processing device 20.

[0041] The processing unit 21B performs specific processing using the generation model 24B. Specifically, the processing unit 21B inputs a predetermined multimodal prompt to the generation model 24B and obtains the output result.

[0042] The output unit 21C outputs the output result of the generation model 24B, that is, the result of the specific processing, to the display unit 26. As a result, the result of the specific processing is displayed on the display unit 26.

[0043] Figure 4 is a flowchart showing the flow of a specific process executed by the information processing device 20. The specific process is performed when the CPU 21 reads the information processing program 24A from the storage 24, loads it into the RAM 23, and executes it. The specific process is performed by the CPU 21 functioning as the input unit 21A, processing unit 21B, and output unit 21C. As an example, the specific process is performed when a user executes a predetermined application.

[0044] In step S10 shown in Figure 4, the CPU 21 acquires various information received from the user input. Then, the CPU 21 proceeds to step S11. Here, the various information includes at least an image showing a specific item that is the subject of the visual inspection (hereinafter referred to as the "specific item image"). In the following explanation, "screws" will be used as an example of the item (see Figure 5, etc.), but the type of item is not limited to this. For example, in addition to "screws," various products such as "fabric products," "cosmetics (e.g., foundation)," and "circuit boards" can be used as the item.

[0045] In step S11, the CPU 21 generates a multimodal prompt to output whether or not there is an abnormality in the specific item shown in the specific item image, based on the various information acquired in step S10. Then, the CPU 21 proceeds to step S12.

[0046] As an example, the CPU 21 generates a multimodal prompt like the one below, which includes at least the image of a specific item acquired in step S10 and a predetermined instruction. Note that the instruction shown in the following multimodal prompt is merely an example and is not limited to it.

[0047] "Multimodal prompt" Instructions: Output whether or not the specific item shown in the following image is defective. The output format should be at least one of image data and text data. Specific item image: screwAAA.jpg

[0048] In step S12, the CPU 21 inputs the multimodal prompt generated in step S11 to the generation model 24B and obtains the output result from the generation model 24B. Then, the CPU 21 proceeds to step S13.

[0049] In step S13, the CPU 21 outputs the output result of the generation model 24B in step S12 to the display unit 26. As a result, the display unit 26 displays the result of the specific process. Then, the CPU 21 terminates the specific process.

[0050] Next, an example of the display according to this embodiment will be described. Figure 5 is a first example of the display shown on the display unit 26 of the information processing device 20. Specifically, Figure 5 shows a verification screen for verifying the inspection accuracy by the generation model 24B.

[0051] The verification screen shown in Figure 5 displays GUI (Graphical User Interface) buttons 40-46 and the verification area 50.

[0052] GUI button 40 is a button for selecting a generation model 24B to verify the inspection accuracy of the visual inspection. In this embodiment, multiple generation models 24B with different parameters related to the output content (e.g., generation model 24B-1, generation model 24B-2, etc.) are provided. The user can select the desired generation model 24B by operating the GUI button 40. If the GUI button 40 is not operated, a pre-configured generation model 24B will be selected. Note that at least one model may be selected as the target of the pre-configured generation model 24B, and accuracy verification may be performed on multiple models, or the model with the optimal accuracy and verification result may be automatically selected.

[0053] GUI button 41 is a button for selecting at least one image of a specific item to be inspected for abnormalities. The image of the specific item may also have data attached indicating the location and characteristics of the abnormality. After operating GUI button 41, the user selects the desired image of the specific item by selecting an image from storage 24 or by taking a picture with the camera of the information processing device 20. The verification area 50 shown in Figure 5 displays the screw image 60 selected as the specific item image. If multiple images of specific items are selected, they may be displayed in a list, or they may be viewable sequentially using an image scrolling function. Furthermore, the verification area 50 may also display at least one of the automatically selected images to be judged, for purposes such as re-verifying the image to be judged during actual visual inspection.

[0054] GUI button 42 is a button for setting anomaly detection conditions by the generation model 24B. These detection conditions include a threshold, the nature of the anomaly, and its location. The user can set these detection conditions by operating GUI button 42. As a result, the CPU 21 accepts the detection conditions set by the user as an anomaly definition that indicates the criteria for determining whether or not a particular item is an anomaly. This anomaly definition may be freely entered by the user, or it may be defined by the user modifying a suggestion from the software. Furthermore, if data indicating the correct answer for an anomaly is attached to the image of the specific item, the conditions that result in the highest judgment accuracy may be automatically searched for, for example.

[0055] GUI button 43 is a button for entering supplementary text that complements the characteristics of a specific item. The supplementary text is text that enhances the specificity of the item, such as a specific product name and defect information. The user can enter the supplementary text by operating GUI button 43. The input of the supplementary text may be in the form of the user directly typing into a text box, or it may be in the form of the user modifying suggestions from the software.

[0056] GUI button 44 is a button for inputting a reference image of a good product that shows no abnormalities in a specified item. The specified item may be the same type of item as the specific item, or it may be a different type of item. After operating the GUI button 44, the user selects the desired reference image of a good product by selecting the desired image from the storage 24 or by taking a picture with the camera of the information processing device 20.

[0057] GUI button 45 is a button for inputting a reference defective product image that shows a condition where a specified item has an abnormality. After operating GUI button 45, the user selects the desired reference defective product image by selecting an image from storage 24 or by taking a picture with the camera of the information processing device 20. Furthermore, the user also inputs text indicating the location and content of the abnormality in the selected reference defective product image.

[0058] GUI button 46 is used to verify the inspection accuracy of the visual inspection performed by the generation model 24B. When GUI button 46 is operated, the CPU 21 generates a multimodal prompt based on the settings configured based on the operation of GUI buttons 40 to 45. The CPU 21 then inputs this multimodal prompt into the generation model 24B to output whether or not there is an abnormality in the specific item (e.g., screw image 60) shown in the verification area 50. If multiple specific item images are selected, the contents of multiple images can be viewed using a list display of the presence or absence of abnormalities or an image scrolling function. If data indicating the correct answer for abnormalities is attached to the specific item image, statistical information such as the accuracy of individual images and the overall judgment accuracy may also be displayed. Note that, among GUI buttons 40 to 45, when detection conditions are automatically searched and suggested, inputting abnormality detection conditions in GUI button 42, inputting auxiliary text in GUI button 43, and inputting reference good product images and reference defective product images in GUI buttons 44 and 45 are optional. Even without these inputs, the generation model 24B can output whether or not there is an abnormality in the specific item based on the operation of the GUI button 46.

[0059] Figure 6 is a specific example of a reference image showing a given item in both its normal and abnormal states. As an example, in Figure 6, the given item is a "screw" of the same type as a specific item.

[0060] Figure 6(A) shows screw image 62, which indicates a screw without any abnormalities. Figure 6(B) shows screw image 64, which indicates a screw with an abnormality. For example, if the user operates the GUI button 45 and selects screw image 64 as the reference image of a defective product, they can input text indicating the location and nature of the abnormality, such as "Bent tip of screw."

[0061] Figure 7 shows a second display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 7 shows the state after the GUI button 46 has been operated on the verification screen shown in Figure 5.

[0062] In the verification screen shown in Figure 7, text data 65 indicating whether or not there is an abnormality in the screw shown in the screw image 60 is displayed below the screw image 60 within the verification area 50. Specifically, the text data 65 says "No abnormalities found," indicating that there is no abnormality in the screw shown in the screw image 60.

[0063] Figure 8 shows a third display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 8 shows a state in which a screw image 70, which represents a screw different from the screw image 60, is displayed as a specific item image in the verification area 50.

[0064] Figure 9 shows a fourth display example shown on the display unit 26 of the information processing device 20. Specifically, Figure 9 shows the state after the GUI button 46 has been operated on the verification screen shown in Figure 8.

[0065] In the verification screen shown in Figure 9, text data 75 indicating whether or not there is an abnormality in the screw shown in the screw image 70 is displayed below the screw image 70 within the verification area 50. Specifically, the text data 75 reads, "There is an abnormality. The tip of the screw is bent," indicating that there is an abnormality in the screw shown in the screw image 70 and the nature of that abnormality.

[0066] Here, the display method of the screw image 70 shown in Figure 9 differs from the display method of the screw image 70 shown in Figure 8. Specifically, in the screw image 70 shown in Figure 9, a heat map is superimposed on the screw image to indicate abnormal areas. Note that in Figure 9, for illustrative purposes, the heat map is represented by a pattern of diagonal hatching and dots. In the screw image 70, the area 70A from the head of the screw to near the tip is represented by a first color (diagonal hatching in the figure), and the tip area 70B is represented by a second color different from the first color (dots in the figure). The second color is a predetermined color (e.g., red, yellow) that indicates the area with an abnormality.

[0067] As described above, the verification screen shown in Figure 9 indicates, through the display of the screw image 70 and the text data 75, that there is an abnormality in the screw shown in the screw image 70 and the details of that abnormality.

[0068] Figure 10 shows a fifth display example shown on the display unit 26 of the information processing device 20. As an example, Figure 10 shows an adjustment screen for adjusting at least one of the multimodal prompts input to the generation model 24B and the parameters of the generation model 24B related to the output content. The adjustment screen is displayed when a predetermined operation is performed on the verification screen.

[0069] The adjustment screen shown in Figure 10 displays GUI buttons 80 to 83. GUI button 80 is a button for selecting the generation model 24B to which the multimodal prompt and parameters will be adjusted. The user can select the desired generation model 24B by operating GUI button 80.

[0070] GUI button 81 is a button for selecting the training images to be used for adjusting the multimodal prompts and parameters described above. The training images include images of good products showing no abnormalities in a specific item (e.g., a screw) and an item of the same type (e.g., a screw), and images of defective products showing abnormalities in the same item. After operating GUI button 81, the user selects at least one desired good product image and one desired defective product image by selecting a desired image from storage 24 or by taking a picture with the camera of the information processing device 20. The user may also input text, text data, or structured data such as JSON or YAML that indicates the location and content of the abnormality in the selected defective product image.

[0071] GUI button 82 is a button for entering supplementary text that adds details about the characteristics of a specific item. Users can enter supplementary text by operating GUI button 82.

[0072] GUI button 83 is a button for adjusting at least one of the multimodal prompts and parameters. When GUI button 83 is operated, the CPU 21 adjusts at least one of the multimodal prompts and parameters based on the settings based on the operation of GUI buttons 80 to 82. The adjustment of the multimodal prompts or parameters is performed as appropriate using known techniques.

[0073] Figure 11 shows a sixth display example shown on the display unit 26 of the information processing device 20. As an example, Figure 11 shows the operation screen during actual operation of the visual inspection.

[0074] The operation screen shown in Figure 11 shows GUI buttons 90-93 and a verification area 50. In actual operation, an image of an item captured by a camera (not shown) is displayed in the verification area 50, and the information processing device 20 automatically determines whether or not there is an abnormality each time an image is taken. In the verification area 50 shown in Figure 11, a screw image 100, which represents a screw captured by the camera, is displayed.

[0075] GUI button 90 is for selecting the generation model 24B to be used in the actual operation of visual inspection. By operating this GUI button 90, the user can select the desired generation model 24B. However, the selection is not limited to the user; for example, the model with the highest accuracy in the verification results may be automatically selected.

[0076] GUI button 91 is a button for setting the anomaly detection conditions by the generation model 24B. The user can set these detection conditions by operating GUI button 91.

[0077] GUI button 92 is a button for selecting whether to enable or disable online learning. Users can switch online learning on or off by operating GUI button 92. In this embodiment, online learning refers to a function that adjusts the parameters of the generative model 24B during visual inspection to improve anomaly detection performance. Specifically, this includes methods such as performing learning by assuming the first XX items during visual inspection are good products, or receiving user feedback on good / defective product judgments on the spot and performing learning in real time or retrospectively. With this configuration, in addition to the user explicitly selecting images to train the model at times other than during actual visual inspection, the generative model 24B can continuously improve its accuracy by automatically incorporating feedback.

[0078] GUI button 93 is a button for starting the actual operation of the visual inspection by the generation model 24B. When GUI button 93 is operated, the CPU 21 generates a multimodal prompt based on the settings based on the operation of GUI buttons 90 to 92. The CPU 21 then inputs this multimodal prompt into the generation model 24B and outputs whether or not there is an abnormality in the item shown in the item image (e.g., screw image 100) of the item displayed in the verification area 50.

[0079] In the operation screen shown in Figure 11, text data 105 indicating whether or not there is an abnormality in the screw shown in the screw image 100 is displayed below the screw image 100 within the verification area 50. Specifically, the text data 105 is "There is an abnormality. The tip of the screw is bent," indicating that there is an abnormality in the screw shown in the screw image 100 and the nature of the abnormality. The nature of the abnormality may be written in conjunction with the location of the abnormality, for example, but the display location is not particularly limited. Furthermore, there are no particular restrictions on the nature of the abnormality itself, such as using quantitative or qualitative expressions to describe the degree of the abnormality.

[0080] Furthermore, the display method for screw images 100 that have been determined to have an abnormality is the same as for screw image 70 shown in Figure 9, where a heat map is superimposed on the screw image to indicate the abnormal area. In the screw image 100, the region 100A from the head of the screw to near the tip is represented in the first color, and the region 100B at the tip is represented in the second color.

[0081] Note that the GUI buttons shown in Figure 11 represent only a portion of the operation screen, and other GUI buttons actually exist. For example, above GUI button 90, there is an image input source selection button. This image input source selection button allows you to select the source from which to receive images, such as images saved in a designated folder or images taken directly from the camera.

[0082] As explained above, in the information processing device 20, the CPU 21 receives input of an image of an item that is the subject of the visual inspection. The CPU 21 also inputs a multimodal prompt to the generation model 24B that includes the received image of the specific item and a predetermined instruction. The CPU 21 then displays on the display unit 26 whether or not there is an abnormality in the specific item shown in the image of the specific item, which is the output result of the generation model 24B. Thus, according to the information processing device 20, by inputting the multimodal prompt to the generation model 24B before performing the visual inspection using the generation model 24B, the user can understand the inspection accuracy of the generation model 24B before the visual inspection is performed. Furthermore, according to the information processing device 20, by inputting the multimodal prompt to the generation model 24B during the actual operation of the visual inspection, the user can understand the inspection accuracy of the generation model 24B while the visual inspection is being performed.

[0083] Furthermore, the generation model 24B identifies the type of a specific item from the characteristics of the input specific item image, and infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific item. As a result, the information processing device 20 can automate everything from item type identification to abnormality detection using the generation model 24B, thereby improving the accuracy of visual inspection, increasing work efficiency, and reducing costs.

[0084] Furthermore, if the generation model 24B infers that there is an abnormality in a particular item, it generates at least one of image data and text data indicating the nature of the abnormality. Then, in the information processing device 20, the CPU 21 displays at least one of the image data and text data indicating the nature of the abnormality of the particular item, which are the output results of the generation model 24B, on the display unit 26. As a result, according to the information processing device 20, the nature of the abnormality of the particular item is displayed on the display unit 26 in the form of at least one of the image and text, allowing the user to intuitively and concretely grasp the problem area, making it easier to take prompt action and manage records.

[0085] Furthermore, in the information processing device 20, the CPU 21 receives input of supplementary text that supplements the characteristics of a specific item, and a reference image showing at least one of the states in which the specified item is free from abnormalities and in which it is abnormal. The CPU 21 then inputs a multimodal prompt to the generation model 24B that includes the image of the specific item, a predetermined instruction sentence, and at least one of the received supplementary text and reference image. As a result, the information processing device 20 deepens the understanding of the generation model 24B by adding specific information regarding the characteristics of the item and whether or not it is abnormal, enabling more accurate anomaly detection and reducing false detections and missed detections.

[0086] Furthermore, in the information processing device 20, the CPU 21 accepts input of an abnormality definition that indicates the criteria for determining whether or not a particular item is abnormal. As a result, the information processing device 20 can accept input of an abnormality definition, allowing for flexible judgment criteria to be set according to the user's application, improving the detection accuracy of the generation model 24B and increasing its adaptability to specific applications.

[0087] Furthermore, in the information processing device 20, the CPU 21 receives at least one input: supplementary text that supplements the characteristics of a specific item, an image of a good product showing that an item of the same type as the specific item is free of defects, and an image of a defective product showing that an item of the same type has defects. Based on the supplementary text, the image of a good product, and the image of a defective product that the CPU 21 has received as input, it adjusts at least one of the multimodal prompts to be input to the generation model 24B and the parameters of the generation model 24B related to the output content. Thus, according to the information processing device 20, by adjusting the parameters of the generation model 24B based on information about the characteristics and abnormal state of a specific item, the abnormality detection of the generation model 24B can be optimized for a specific item and specific conditions. In addition, according to the information processing device 20, by optimizing for multiple items, the general detection accuracy for visual inspection of items can also be improved. Furthermore, according to the information processing device 20, by adjusting the multimodal prompts based on information about the characteristics and abnormal state of a specific item, the generation model 24B can understand the characteristics and abnormal state of the specific item more accurately, and the accuracy of the output results can be improved.

[0088] (others) While embodiments of the present disclosure have been described in detail above with reference to the attached drawings, the technical scope of the present disclosure is not limited to these examples. It is clear that a person with ordinary skill in the art of the present disclosure may conceive of various modifications or alterations within the scope of the technical idea set forth in the claims, and these modifications or alterations are also understood to fall within the technical scope of the present disclosure.

[0089] Furthermore, the effects described in the above embodiments are descriptive or illustrative, and are not limited to those described in the above embodiments. In other words, the technology relating to this disclosure may produce other effects that would be obvious to a person of ordinary skill in the art of this disclosure from the descriptions in the above embodiments, in addition to or in lieu of the effects described in the above embodiments.

[0090] The processing described in the above embodiment can also be implemented using dedicated hardware circuits. In this case, it may be executed on one piece of hardware or on multiple pieces of hardware.

[0091] In the above embodiment, the input of auxiliary text may be performed in an interactive format between the user and the generation model 24B. This allows the information processing device 20 to flexibly acquire the necessary information by supplementing the characteristics of a specific item in an interactive format with the user, enabling anomaly detection that is in line with the user's intentions.

[0092] In the above embodiment, the verification screen shown in Figure 5, etc., allows the input of supplementary text and reference images based on the user's selection, but the timing of these inputs is not limited. For example, instead of a configuration that allows input at any time the user desires, a configuration may be adopted in which the generation model 24B requests the user to input at least one of the supplementary text and reference images when certain conditions are met.

[0093] For example, the "case where the predetermined conditions are met" described above can be defined as the case where the difference in the confidence level (%) of whether or not an anomaly is present, inferred by the generative model 24B without input of auxiliary text and reference images, is less than or equal to a predetermined value.

[0094] Furthermore, the system may be configured such that, when certain conditions are met, the generation model 24B requests the user to input either supplementary text or a reference image, and when specific conditions are met, the generation model 24B requests the user to input the other supplementary text or reference image.

[0095] For example, the "case where specific conditions are met" described above can be defined as the case where the difference in confidence levels of whether or not an anomaly is inferred by the generative model 24B after inputting either the auxiliary text or the reference image is less than or equal to a predetermined value.

[0096] With the above configuration, the information processing device 20 can deepen its understanding of the generation model 24B by adding specific information regarding the characteristics of the item and whether or not there are abnormalities when the reliability of the output of the generation model 24B is low, thereby enabling more accurate anomaly detection and reducing false detections and missed detections.

[0097] Furthermore, based on the content of the supplementary text entered by the user, the generative model 24B may suggest to the user what supplementary text to input to improve the reliability of the output. For example, when the generative model 24B accepts supplementary text input in an interactive format, if the content of the supplementary text entered by the user is "product name," it may respond with a suggestion such as "Please also enter a specific example of the defect." In addition, when the generative model 24B accepts supplementary text input in an interactive format, it may also make suggestions such as "The product name is foundation, and the defect is uneven color."

[0098] In the above embodiment, the method of representing the heatmap superimposed on the screw image is not particularly limited. For example, multiple colors corresponding to the degree of abnormality may be represented in the heatmap. Alternatively, instead of using a heatmap to indicate information about the abnormality, marks indicating the location and area of ​​the abnormality may be used. Furthermore, the degree of abnormality may be indicated by adding numerical values ​​or text.

[0099] The configuration described in the above embodiment is not limited to the visual inspection of goods, but can be extended to general "anomaly detection" such as hazard prediction (unsafe behavior or illness by workers), foreign object contamination, and detection of suspicious persons.

[0100] In the above embodiment, an example was described in which an image of an item to be inspected is received as input, but the system is not limited to this. For example, instead, an image of the item may be received as input.

[0101] In the above embodiment, the generation model 24B can output not only a binary output of "presence or absence of an abnormality" but also "degree of abnormality." Here, "degree of abnormality" is one element that constitutes "content of the abnormality." "Content of the abnormality" includes, for example, "location of the abnormality," "degree of abnormality," and "nature of the abnormality (type of abnormality)."

[0102] In the above embodiment, in addition to the case in which the generation model 24B directly generates image data and text data indicating whether or not there is an abnormality in the item, the embodiment also includes a case in which the generation model 24B outputs the degree of abnormality as numerical data in map format, and internally executes a process to generate image data and text data indicating whether or not there is an abnormality based on that numerical data and the input image.

[0103] In the above embodiment, the CPU 21 may save the output results of the generative model 24B, accept user modifications to the saved output results, and train the generative model 24B using the modified data. This creates a loop for improving the accuracy of the generative model 24B, enabling continuous improvement of anomaly detection performance.

[0104] In the above embodiment, the "characteristics of the article" entered as auxiliary text includes not only information indicating what the article itself is (e.g., foundation), but also information about specific abnormalities to be detected (e.g., type: uneven color, location: on powder, degree: very strong, etc.).

[0105] In the above embodiment, the dialogue format between the user and the generation model 24B when inputting auxiliary text is not particularly limited and may be, for example, a QA format (e.g., What is the product? → Foundation → What specific defect features do you want to detect? → Color unevenness) or a suggestion dialogue format (e.g., Is the product foundation? → Yes). Alternatively, the input format for auxiliary text may not depend on the dialogue format, and the generation model 24B may simply propose the auxiliary text entirely and present it in a format that the user can modify.

[0106] In each of the above embodiments, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and dedicated processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.).

[0107] Furthermore, the processor operations in each of the above embodiments may not be performed by a single processor, but may also be performed by multiple processors located in physically separate locations working together. Alternatively, some or all of the operations performed by specific multiple processors in each of the above embodiments may be integrated and performed by a single processor. In addition, the order of the processor operations is not limited to the order described in each of the above embodiments, and may be changed as appropriate.

[0108] Furthermore, although the above embodiment describes an embodiment in which the information processing program 24A is pre-stored (installed) in the storage 24, the invention is not limited thereto. The information processing program 24A may be provided in the form of a recording medium such as a CD-ROM (Compact Disk Read Only Memory), DVD-ROM (Digital Versatile Disk Read Only Memory), and USB (Universal Serial Bus) memory. Alternatively, the information processing program 24A may be provided in the form of a download from an external device via a network N. The technology of this disclosure can also be applied to programs and program products.

[0109] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference. [Explanation of Symbols]

[0110] 20 Information Processing Devices 21 CPU (Processor) 24A Information Processing Program 24B Generative Model 26 Display section

Claims

1. Equipped with a processor, The aforementioned processor, The system accepts input of images of items that are subject to visual inspection. A multimodal prompt, including a specific image of an item received as input and a predetermined instruction, is input to a generation model capable of outputting whether or not there is an abnormality in the item shown in the image. The display unit will show whether or not there is an abnormality in the specific item shown in the specific item image, which is the output result of the generation model. Information processing device.

2. The generation model identifies the type of a specific item from the features of the input image of the specific item, infers whether or not there is an abnormality corresponding to that type according to the instructions indicated by the predetermined instruction statement, and generates at least one of image data and text data indicating whether or not there is an abnormality in the specific item. The information processing apparatus according to claim 1.

3. If the generation model infers that there is an abnormality in the specific item, it generates at least one of image data and text data indicating the nature of the abnormality. The processor causes the display unit to display at least one of image data and text data indicating the content of the abnormality of the specific item, which is the output result of the generation model. The information processing apparatus according to claim 2.

4. The aforementioned processor, The system accepts input of supplementary text that complements the characteristics of the aforementioned specific article, and a reference image showing at least one of the states in which the specified article is normal and in which it is abnormal. The generative model is input a multimodal prompt that includes the image of the specific item, the predetermined instruction text, and at least one of the auxiliary text and the reference image that has been received as input. The information processing apparatus according to claim 1.

5. The input of the aforementioned supplementary text is performed through an interactive process between the user and the generative model. The information processing apparatus according to claim 4.

6. The aforementioned processor, The system accepts input of a definition of abnormality, which indicates the criteria for determining whether or not a particular item is abnormal. The information processing apparatus according to claim 1.

7. The aforementioned processor, The system accepts at least one input: supplementary text that complements the characteristics of the specific item, an image of a good product showing an item of the same type as the specific item without any defects, and an image of a defective product showing an item of the same type with defects. Based on at least one of the auxiliary text, the good product image, and the defective product image received as input, at least one of the multimodal prompts input to the generation model and the parameters of the generation model relating to the output content is adjusted. The information processing apparatus according to claim 1.

8. The system accepts input of images of items that are subject to visual inspection. A multimodal prompt, including a specific image of an item received as input and a predetermined instruction, is input to a generation model capable of outputting whether or not there is an abnormality in the item shown in the image. The display unit will show whether or not there is an abnormality in the specific item shown in the specific item image, which is the output result of the generation model. An information processing method in which a computer performs the processing.

9. The system accepts input of images of items that are subject to visual inspection. A multimodal prompt, including a specific image of an item received as input and a predetermined instruction, is input to a generation model capable of outputting whether or not there is an abnormality in the item shown in the image. The display unit will show whether or not there is an abnormality in the specific item shown in the specific item image, which is the output result of the generation model. An information processing program that causes a computer to perform a task.

Citation Information

Patent Citations

  • State determination device and image analysis device

    JP2022071675A