A method for detecting product defects using a VLM agent and a VLM agent using the same.
The VLM agent addresses the limitations of conventional AOI and AI deep learning by using text prompts and inspection tools for flexible, efficient, and transparent defect detection across diverse products.
Patent Information
- Authority / Receiving Office
- JP Β· JP
- Patent Type
- Patents
- Current Assignee / Owner
- SUPERB AI CO LTD
- Filing Date
- 2025-11-19
- Publication Date
- 2026-07-24
AI Technical Summary
Conventional AOI methods for detecting product defects are sensitive to lighting and camera position changes, require re-setting rules for each product, involve significant training data and computing resources, and lack transparency in defect determination, making them inflexible and resource-intensive.
A VLM agent uses text prompts for defect criteria and inspection tools to sequentially perform multiple inspections, determining defects based on input criteria without additional training, enabling accurate defect detection with precise location and measurement.
The VLM agent allows for defect detection across various products without re-training, understands defect criteria, and requires minimal data, providing accurate and transparent defect assessment.
Smart Images

Figure 0007894668000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to detecting product defects using a VLM (Vision Language Model) agent. More specifically, using a VLM agent, it refers to a method for detecting product defects and a VLM agent using the same, by referring to a text prompt in which clear criteria for product defects are described and the result of inspecting for defects using an inspection tool according to the text prompt.
Background Art
[0002] Generally, in order to detect defects in products produced at a manufacturing site, AOI (Automated Optical Inspection) is used. AOI analyzes a product image captured using a camera to check whether the product meets quality specifications, thereby inspecting for the presence or absence of product defects.
[0003] Such a conventional AOI determines the presence or absence of product defects based on set rules. By comparing a reference image for a pre-set good product with the captured product image, if it deviates from rules such as coordinates, color, size, etc., it is determined as a defect.
[0004] However, while the conventional AOI method has the advantages of being able to quickly determine defects, being effective for simple and repetitive defect detection, and having relatively easy settings, it is sensitive to changes in lighting and the shooting position of the camera. Not only is it necessary to re-set the rules for defect determination for each new product, but there is also a problem that the user has to intervene for rule setting and exception handling.
[0005] Therefore, in recent years, a method of executing AOI by an AI (Artificial Intelligence) deep learning method has been proposed.
[0006] In the AI ββdeep learning method, an AI model that has learned from a large number of normal / defective images independently recognizes the characteristics of defects and patterns to determine whether a product is defective. This method has the advantages of high accuracy, the ability to detect minute and atypical defects, and flexibility in application to various product groups.
[0007] However, the AI ββdeep learning method has the drawback of requiring not only a large amount of training data to train the AI ββmodel, but also a significant amount of time and computing resources to train the AI ββmodel.
[0008] Furthermore, with AI deep learning methods, the AI ββmodel infers the location and measurement values ββof defects from images, which means that the location and measurement values ββof defects may differ from those of the actual product.
[0009] Furthermore, a problem with AI deep learning methods is that it is difficult to understand the criteria the AI ββmodel used to determine that a product is defective.
[0010] In addition, with AI deep learning, if the criteria for determining defects change due to changes in the external environment such as product quality standards and related regulations, it is not only necessary to build new training data to train the AI ββmodel based on the changed defect criteria, but it is also difficult to build a large amount of training data for training the AI ββmodel because it is difficult to easily obtain defect data at the manufacturing site. [Overview of the Initiative] [Problems that the invention aims to solve]
[0011] The purpose of this invention is to solve all of the problems of the prior art described above.
[0012] Another objective of the present invention is to enable the detection of defects in a product without requiring additional learning in response to changes in the product's defect criteria, by having the VLM agent detect defects in the product using the product's defect criteria as input.
[0013] Furthermore, another objective of the present invention is to enable the detection of defects in various product groups without additional training, by having the VLM agent detect defects in products using product defect criteria as input.
[0014] Furthermore, the present invention aims to enable understanding of the criteria on which the VLM agent determined a product to be defective, by having the VLM agent detect defects in the product using product defect criteria as input.
[0015] Furthermore, another objective of the present invention is to enable the detection of product defects using a VLM agent, even with only a small amount of training data, by having the VLM agent detect product defects using product defect criteria as input.
[0016] Furthermore, the present invention aims to enable accurate detection of product defects using the precise location and measurements of defects by having a VLM agent inspect defects using an inspection tool based on input product defect criteria. [Means for solving the problem]
[0017] According to one embodiment of the present invention, in a method for detecting product defects using a VLM (Vision Language Model) agent, (a) when a text prompt relating to product defect judgment criteria and a target image of the product captured by a camera are acquired, the VLM agent refers to the target image to check whether the product has a defect or not, and when at least one defect is confirmed in the target image, it generates a first inspection task to the nth inspection task (where n is an integer of 1 or more) for inspecting the defect based on the defect judgment criteria included in the text prompt; (b) the VLM agent generates the pth inspection task (where p is an integer of 1 or more) which is the first inspection task to the nth inspection task. A method is provided that includes the steps of: (c) obtaining a first to nth inspection result by requesting a pth inspection tool to perform a pth inspection on the defect according to (a) an integer that increases sequentially from 1 to n; having the pth inspection tool perform the pth inspection on the defect; and transmitting the pth inspection result, which is the result of performing the pth inspection; and (d) the VLM agent determining whether the product is defective by referring to the first to nth inspection result and confirming whether the defect falls under the defect determination criteria, and generating a defect determination result for the product.
[0018] In one example, in step (a), the VLM agent inputs the text prompt and the image to be identified into the VLM, and uses the VLM to analyze the text prompt and the image to be identified to generate the first inspection task to the nth inspection task; and in step (c), the VLM agent inputs the first inspection result to the nth inspection result into the VLM, and uses the VLM to refer to the first inspection result to the nth inspection result to generate a result for determining whether or not the product has defects.
[0019] In one example, in step (a), the VLM agent saves the text prompt and the image to be identified as a chat log, updates the chat log by adding the first inspection task to the nth inspection task, thereby generating an updated chat log; in step (c), the VLM agent inputs the updated chat log and the first inspection result to the nth inspection result into the VLM, and uses the VLM to generate the defect detection result by referring to the updated chat log and the first inspection result to the nth inspection result, and adds the first inspection result to the nth inspection result and the defect detection result to the updated chat log.
[0020] In one example, in step (b), when the VLM agent obtains a specific inspection result corresponding to a specific inspection task which is one of the first inspection task to the nth inspection task, it generates an updated specific inspection task which is an updated version of the a updated version of the specific inspection task which is a
[0021] In one example, in step (a), the VLM agent further acquires a vision prompt, which is a sample image indicating whether the product is defective or good, and further refers to the vision prompt to analyze whether the image to be determined has the defect.
[0022] In one example, in step (b) above, the p-inspection result includes a p-inspection result image obtained by applying the state in which the p-inspection was performed to the p-inspection image corresponding to the p-inspection region in which the p-inspection was performed, and p-inspection result text relating to the result value of performing the p-inspection.
[0023] In one example, in step (b), the VLM agent executes one of the following subprocesses: (i) a subprocess that uses the p inspection tool to crop and enlarge an image region corresponding to the p inspection region from the image to be determined to generate the p inspection image, performs the p inspection on the p inspection image, and applies the state after the p inspection has been performed to the p inspection image to generate the p inspection result image; and (ii) a subprocess that uses the p inspection tool to acquire the p inspection image taken by zooming in on the p inspection region of the product through the camera, performs the p inspection on the p inspection image, and applies the state after the p inspection has been performed to the p inspection image to generate the p inspection result image.
[0024] According to one embodiment of the present invention, in a method for detecting product defects using a VLM (Vision Language Model) agent, (a) when a text prompt relating to product defect judgment criteria and a target image of the product captured by a camera are acquired, the VLM agent refers to the target image to check whether the product has a defect or not, and when at least one defect is confirmed in the target image, when p is 1, generates a first inspection task for inspecting the defect based on the defect judgment criteria contained in the text prompt, requests a first inspection tool to perform a first inspection on the defect according to the first inspection task, causes the first inspection tool to perform the first inspection on the defect, and transmits the first inspection result which is the result of performing the first inspection, thereby obtaining a first inspection result; (b) when the VLM agent determines whether p is 2 A method is provided that includes the steps of: (c) when p is an integer increasing from n (where n is an integer of 2 or more), generating a p-th inspection task for inspecting the defect by referring to the defect judgment criterion and the (p-1) inspection result; requesting a p-th inspection tool to perform a p-th inspection on the defect according to the p-th inspection task; having the p-th inspection tool perform the p-th inspection on the defect; and transmitting the p-th inspection result, which is the result of performing the p-th inspection, thereby repeating a subprocess for obtaining the p-th inspection result until p becomes n; and (c) the VLM agent determines whether the product is defective by referring to the first inspection result to the n-th inspection result and confirming whether the defect falls under the defect judgment criterion, and generating a defect determination result for the product.
[0025] In one example, in the step (a), the VLM agent stores the text prompt and the discriminant target image as a first initial chat log, adds the first inspection task to the first initial chat log to generate a first chat log. In the step (b), the VLM agent generates the p-th inspection task by referring to the (p - 1)-th chat log and the (p - 1)-th inspection result, adds the (p - 1)-th inspection result and the p-th inspection task to the (p - 1)-th chat log to generate a p-th chat log. In the step (c), the VLM agent generates a defect presence / absence discrimination result by referring to the (n - 1)-th chat log and the n-th inspection result, and adds the n-th inspection result and the defect presence / absence discrimination result to the (n - 1)-th chat log to generate an n-th chat log.
[0026] In one example, in the step (b), before p reaches n, if the (p - 1)-th inspection result meets the defect determination criteria, the VLM agent interrupts the sub-process. In the step (c), the VLM agent generates a defect presence / absence discrimination result by referring to the inspection results obtained until the sub-process is interrupted.
[0027] According to one embodiment of the present invention, a VLM agent for detecting product defects includes: a memory storing instructions for detecting product defects; and a processor for detecting product defects according to the instructions stored in the memory, wherein the processor (I) when a text prompt relating to product defect criteria and a target image of the product captured by a camera are acquired, it refers to the target image to determine whether the product has a defect, and when at least one defect is confirmed in the target image, it generates a first inspection task to an nth inspection task (where n is an integer of 1 or more) for inspecting the defect based on the defect criteria contained in the text prompt; and (II) the first inspection task (III) A VLM agent is provided that performs a process of obtaining a first inspection result to an nth inspection result by requesting the pth inspection tool to perform the pth inspection on the defect according to the pth inspection task which is the nth inspection task (where p is an integer that increases sequentially from 1 to n), having the pth inspection tool perform the pth inspection on the defect, and transmitting the pth inspection result which is the result of performing the pth inspection; and (III) a process of determining whether the product is defective by referring to the first inspection result to the nth inspection result and confirming whether the defect falls under the defect determination criteria, and generating a defect determination result for the product.
[0028] In one example, in the (I) process, the processor inputs the text prompt and the image to be discriminated into the VLM, and uses the VLM to analyze the text prompt and the image to be discriminated to generate the first inspection task to the nth inspection task. In the (III) process, the processor inputs the first inspection result to the nth inspection result into the VLM, and uses the VLM to refer to the first inspection result to the nth inspection result to generate a discrimination result of whether there is a defect in the product.
[0029] In one example, in the (I) process, the processor saves the text prompt and the image to be discriminated as a chat log, adds the first inspection task to the nth inspection task to the chat log to update the chat log, and generates an updated chat log. In the (III) process, the processor inputs the updated chat log and the first inspection result to the nth inspection result into the VLM, and uses the VLM to refer to the updated chat log and the first inspection result to the nth inspection result to generate a discrimination result of whether there is a defect. The processor adds the first inspection result to the nth inspection result and the discrimination result of whether there is a defect to the updated chat log.
[0030] In one example, in the (II) process, when the processor obtains a specific inspection result corresponding to a specific inspection task which is any one of the first inspection task to the nth inspection task, it generates an updated specific inspection task which is an updated version of the a
[0031] In one example, the processor further acquires a vision prompt, which is a sample image indicating whether the product is defective or good, in the (I) process, and further refers to the vision prompt to analyze whether the image to be determined has the defect.
[0032] In one example, in process (II) above, the p-inspection result includes a p-inspection result image obtained by applying the state in which the p-inspection was performed to the p-inspection image corresponding to the p-inspection region in which the p-inspection was performed, and p-inspection result text relating to the result value of performing the p-inspection.
[0033] In one example, the processor executes one of the following subprocesses in the (II) process: (i) a subprocess that generates a p-inspection image by using the p-inspection tool to crop and enlarge an image region corresponding to the p-inspection region from the image to be determined, performs the p-inspection on the p-inspection image, and generates a p-inspection result image by applying the state in which the p-inspection has been performed to the p-inspection image; and (ii) a subprocess that uses the p-inspection tool to acquire a p-inspection image taken by zooming in on the p-inspection region of the product through the camera, performs the p-inspection on the p-inspection image, and generates a p-inspection result image by applying the state in which the p-inspection has been performed to the p-inspection image.
[0034] According to one embodiment of the present invention, a VLM agent for detecting product defects includes: a memory storing instructions for detecting product defects; and a processor for detecting product defects according to the instructions stored in the memory, wherein the processor (i) when a text prompt relating to product defect criteria and a target image of the product captured by a camera are acquired, it refers to the target image to determine whether the product has a defect; when at least one defect is confirmed in the target image, and p is 1, it generates a first inspection task for inspecting the defect based on the defect criteria contained in the text prompt; requests a first inspection tool to perform a first inspection on the defect according to the first inspection task; causes the first inspection tool to perform the first inspection on the defect; and transmits the first inspection result, which is the result of performing the first inspection. This provides a VLM agent that performs the following steps: (II) a process of obtaining a first inspection result; (II) when p is an integer increasing from 2 to n (where n is an integer greater than or equal to 2), a process of generating a p-th inspection task for inspecting the defect by referring to the defect judgment criterion and the (p-1) inspection result; a process of repeating a subprocess of obtaining the p-th inspection result until p becomes n; and (III) a process of determining whether the product is defective by referring to the first inspection result to the n-th inspection result and confirming whether the defect falls under the defect judgment criterion.
[0035] In one example, the processor, in process (I), saves the text prompt and the image to be identified as a first initial chat log, adds the first inspection task to the first initial chat log to generate a first chat log, in process (II), generates the p inspection task by referring to the (p-1) chat log and the (p-1) inspection result, adds the (p-1) inspection result and the p inspection task to the (p-1) chat log to generate the p chat log, in process (III), generates the defect detection result by referring to the (n-1) chat log and the n inspection result, and adds the n inspection result and the defect detection result to the (n-1) chat log to generate the n chat log.
[0036] In one example, the processor interrupts the subprocess if, in process (II), the inspection result (p-1) satisfies the defect criteria before p reaches n, and in process (III), generates the defect determination result by referring to the inspection results obtained up to the time the subprocess was interrupted. [Effects of the Invention]
[0037] According to the present invention, by having the VLM agent detect product defects using product defect criteria as input, product defects can be detected without requiring additional learning in response to changes in product defect criteria.
[0038] According to the present invention, by having the VLM agent detect defects in products using product defect judgment criteria as input, it becomes possible to detect defects in various product groups without additional training.
[0039] According to the present invention, by having the VLM agent detect defects in a product using product defect judgment criteria as input, it becomes possible to understand what criteria the VLM agent used to determine that a product was defective.
[0040] According to the present invention, by using product defect judgment criteria as input and having the VLM agent detect defects in the product, it becomes possible for the VLM agent to detect product defects even with only a small amount of training data.
[0041] According to the present invention, by having the VLM agent inspect defects using an inspection tool based on the input product defect criteria, it becomes possible to accurately detect product defects using the precise location and measurement values ββof the defects. [Brief explanation of the drawing]
[0042] The following drawings, attached for use in describing embodiments of the present invention, represent only a portion of the embodiments, and a person with ordinary skill in the art to which the present invention pertains (hereinafter referred to as "ordinary art") can obtain other drawings based on these drawings without performing any inventive work.
[0043] [Figure 1] This figure schematically shows a VLM agent for detecting product defects according to one embodiment of the present invention. [Figure 2] This figure schematically illustrates a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Figure 3] This figure illustrates an inspection result image generated by at least one inspection tool in a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Figure 4]This figure illustrates an inspection result image generated by at least one inspection tool in a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Figure 5] This figure illustrates an inspection result image generated by at least one inspection tool in a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Figure 6] This figure illustrates an inspection result image generated by at least one inspection tool in a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Figure 7] This figure illustrates an inspection result image generated by at least one inspection tool in a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Figure 8] This figure illustrates an inspection result image generated by at least one inspection tool in a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Figure 9] This figure illustrates an inspection result image generated by at least one inspection tool in a method for detecting product defects using a VLM agent according to one embodiment of the present invention. [Modes for carrying out the invention]
[0044] The detailed description of the present invention, as described below, refers to the accompanying drawings illustrating specific embodiments in which the present invention may be carried out. These embodiments are described in sufficient detail to enable a person of the ordinary skill to carry out the present invention. It should be understood that the various embodiments of the present invention are different from one another but do not need to be mutually exclusive. For example, certain shapes, structures and characteristics described herein can be realized by modifying one embodiment to another without departing from the spirit and scope of the present invention. It should also be understood that the position or arrangement of individual components within each embodiment can be modified without departing from the spirit and scope of the present invention. Therefore, the detailed description described below should not be taken as restrictive, and the scope of the present invention should be understood to encompass the scope claimed in the claims and all equivalent scopes thereto. In the drawings, similar reference numerals indicate identical or similar components in various aspects.
[0045] In the following, several preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings, so that a person with ordinary skill in the art to which the present invention pertains can easily implement the present invention.
[0046] Figure 1 is a schematic diagram of a VLM agent for detecting product defects according to one embodiment of the present invention. Referring to Figure 1, the VLM agent 1000 may include a memory 1100 that stores instructions for detecting product defects, and a processor 1200 that detects defects in the product according to the instructions stored in the memory 1100.
[0047] Specifically, the VLM agent 1000 may, but is not limited to, achieve desired system performance using a combination of computing devices (e.g., devices that may include computer processors, memory, storage, input and output devices, and other conventional computing device components; electronic communication devices such as routers and switches; and electronic information storage systems such as network-attached storage (NAS) and storage area networks (SANs)) and computer software (i.e., instructions for using computing devices in a specific manner).
[0048] Furthermore, the processor 1200 of the VLM agent 1000 may include hardware configurations such as an MPU (Micro Processing Unit) or CPU (Central Processing Unit), cache memory, and data bus. The computing device may also further include an operating system and software configurations for applications that perform specific purposes.
[0049] However, this does not preclude the case where the VLM agent 1000 includes an integrated processor in which a medium, processor, and memory are integrated for carrying out the present invention.
[0050] On the other hand, the processor 1200 of the VLM agent 1000, in accordance with instructions stored in memory 1100, can obtain a text prompt regarding product defect criteria and an image of the product to be identified captured by a camera from the user terminal 10. The processor 1200 then checks whether the product has defects by referring to the image to be identified. If at least one defect is identified in the image to be identified, the processor 1200 can execute a process to generate at least one inspection task, from the first inspection task to the nth inspection task, for inspecting the defect based on the defect criteria contained in the text prompt. In this case, the processor 1200 of the VLM agent 1000 can also obtain the image to be identified from an inspection tool 20 that captures the product using a camera, instead of obtaining it from the user terminal 10. Next, the processor 1200 of the VLM agent 1000 can perform a process to obtain the first to the nth inspection results by following the instructions stored in memory 1100, requesting the pth inspection tool 20_p to perform the pth inspection on the defect according to the first to the nth inspection task (where p is an integer that increases sequentially from 1 to n), having the pth inspection tool 20_p perform the pth inspection on the defect, and transmitting the pth inspection result, which is the result of performing the pth inspection. The pth inspection tool 20_p, that is, the first inspection tool 20_1 to the nth inspection tool 20_n, may consist of each inspection tool that performs the pth inspection on the defect, i.e., the first to the nth inspection on the defect, or it may consist of a single inspection tool 20 that performs the first to the nth inspection on the defect.Although Figure 1 shows a single inspection tool 20, this may be understood as multiple inspection tools, specifically the p-th inspection tool 20_p, i.e., the first inspection tool 20_1 through the nth inspection tool 20_n. Next, the processor 1200 of the VLM agent 1000 can execute a process to determine whether a product is defective by referring to the first inspection result through the nth inspection result according to the instructions stored in memory 1100 and checking whether the defect meets the defect judgment criteria, thereby generating a defect determination result for the product. This defect determination result can then be provided to the user terminal 10.
[0051] In contrast, the VLM agent 1000's processor 1200, in accordance with instructions stored in memory 1100, obtains a text prompt regarding product defect criteria and an image of the product to be identified captured by the camera. It then refers to the image to be identified to check whether the product has defects. If at least one defect is identified in the image to be identified, and p is 1, it generates a first inspection task to inspect the defect based on the defect criteria contained in the text prompt. In accordance with the first inspection task, it requests the first inspection tool 20_1 to perform a first inspection of the defect. The first inspection tool 20_1 then performs the first inspection of the defect, and the processor 1200 transmits the first inspection result, thereby executing the process of obtaining the first inspection result. Next, the processor 1200 of the VLM agent 1000 can execute a process to obtain the p-th inspection result by referring to the defect judgment criteria and the (p-1) inspection result, when p is an integer increasing from 2 to n, according to the instructions stored in memory 1100, generating a p-th inspection task to inspect the defect, requesting the p-th inspection tool 20_p to perform the p-th inspection on the defect according to the p-th inspection task, having the p-th inspection tool 20_p perform the p-th inspection on the defect, and transmitting the p-th inspection result, which is the result of performing the p-th inspection, thereby repeating the subprocess to obtain the p-th inspection result until p becomes n. Next, the processor 1200 of the VLM agent 1000 can execute a process to determine whether the product is defective by referring to the first inspection result to the n-th inspection result according to the instructions stored in memory 1100, and determining whether the defect meets the defect judgment criteria, thereby generating a product defect determination result.
[0052] A method for detecting product defects using the VLM agent 1000 configured as described above will be described in more detail with reference to FIGS. 1 and 2 as follows.
[0053] First, the VLM agent 1000 can obtain a text prompt regarding the product defect determination criteria and a discriminant image of the product captured by the camera (S100). At this time, the VLM agent 1000 can obtain the text prompt and the discriminant image from the user terminal 10, and the user terminal 10 may be a computing device corresponding to a user who monitors product defects, or a computing device that constitutes a system for automatically detecting product defects at the manufacturing site. Alternatively, the VLM agent 1000 can obtain the text prompt from the user terminal 10 and the discriminant image from the inspection tool 20 that inspects product defects at the manufacturing site.
[0054] As an example, the product defect determination criteria input as the text prompt can be shown as follows. Note that the following defect determination criteria are an example of some defect determination criteria for a PCB (Printed Circuit Board), and the product defect determination criteria input as the text prompt can be set in various ways according to the product.
[0055] <Example of PCB defect determination criteria>
[0056] * The circuit is divided into a critical area and a safe area.
[0057] * In the case of a defect in the critical area, if the physical area of the defect is 30 or more, it is determined as a defect.
[0058] * If the defect crosses 50% or more of the wiring in the critical area, it is determined as a defect.
[0059] * In the case of a defect located in the safe area, if the physical area of ββthe defect is 50 or larger, it will be judged as a defect.
[0060] * Otherwise, the product will be judged as good quality.
[0061] Next, the VLM agent 1000 can refer to the image to be identified and confirm whether or not the product has defects (S200).
[0062] For example, the VLM agent 1000 can input an image to be classified into the VLM, which encodes the image to be classified through a vision encoder to generate a vision embedding, and then performs a learning operation on the vision embedding through an LLM (Large Language Model) to detect defect patterns learned from the image to be classified, thereby confirming whether or not the image to be classified has defects.
[0063] On the other hand, while the VLM agent 1000 checked for defects by referring to the acquired image to be classified as described above, it is also possible to acquire vision prompts, which are sample images indicating whether a product is defective or not. In this case, the vision prompts can be further referred to to analyze whether or not the image to be classified as defective.
[0064] For example, when detecting defects in a product that has not been trained for defect detection through the VLM agent 1000, the user can input a sample image representing a defective product as a vision prompt, and the VLM agent 1000 can analyze whether or not there is a defect in the image to be identified by detecting a defective product pattern corresponding to the vision prompt from the image to be identified. Alternatively, the user can input a sample image representing a good product as a vision prompt, and the VLM agent 1000 can analyze whether or not there is a defect in the image to be identified by detecting a pattern from the image to be identified that is different from the good product pattern corresponding to the vision prompt.
[0065] Next, the VLM agent 1000 checks whether or not the image to be identified has defects. If it confirms that the image does not have defects, it determines that the product is good and can generate and provide a defect determination result for it.
[0066] In contrast, if at least one defect is identified in the image to be judged, the VLM agent 1000 can generate at least one inspection task, from the first to the nth inspection task, for inspecting the defect based on the defect judgment criteria contained in the text prompt (S300).
[0067] In this case, the VLM agent 1000 can input a text prompt and an image to be classified into the VLM, and use the VLM to analyze the text prompt and the image to be classified in order to generate the first inspection task or the nth inspection task.
[0068] As an example, the VLM agent 1000 can input the defect determination criteria corresponding to the text prompt and the discrimination target image into the VLM. Then, the VLM can encode the defect determination criteria through the text encoder to generate a text embedding, encode the discrimination target image through the vision encoder to generate a vision embedding, perform a running operation on the vision embedding and the text embedding through the LLM to detect defects, and generate a first inspection task or an nth inspection task for performing an inspection on the detected defects.
[0069] As an example, the inspection tasks that can be generated according to the above <Example of PCB Defect Determination Criteria> can be shown as follows.
[0070] <Inspection Tasks Based on PCB Defect Determination Criteria Example>
[0071] * The circuit is divided into a critical area and a safe area.
[0072] => First Inspection Task: Measure the area information regarding the critical area and the safe area in the discrimination target image
[0073] => In the case of a defect in the critical area, if the physical area of the defect is 30 or more, it is determined as a defect.
[0074] => Second_1 Inspection Task: Measure the length of the defect
[0075] * If the defect crosses 50% or more of the wiring in the critical area, it is determined as a defect.
[0076] => Third Inspection Task: Measure the ratio of the defect crossing the wiring
[0077] [[ID=δΈεδΊ]] * In the case of a defect located in the safe area, if the physical area of ββthe defect is 50 or larger, it will be judged as a defect.
[0078] => Inspection Task 2_2: Measure the length of the defect
[0079] * Otherwise, the product will be judged as good quality.
[0080] Furthermore, the VLM agent 1000 can save the text prompt and the image to be identified as a chat log, and can generate an updated chat log by adding the first inspection task or the nth inspection task to the chat log and updating the chat log.
[0081] Next, the VLM agent 1000 can obtain the first to nth inspection results (S400) by requesting the pth inspection tool 20_p to perform the pth inspection on the defect according to the pth inspection task, which is the first to nth inspection task, that is, the pth inspection task corresponding to p, an integer that increases sequentially from 1 to n; having the pth inspection tool 20_p perform the pth inspection on the defect; and transmitting the pth inspection result, which is the result of performing the pth inspection.
[0082] In this case, the p-th inspection tool 20_p, that is, the first inspection tool 20_1 to the nth inspection tool 20_n, may consist of each inspection tool that performs the p-th inspection of a defect, that is, the first inspection to the nth inspection of a defect, or it may consist of a single inspection tool 20 that performs the first inspection to the nth inspection of a defect.
[0083] The p-th inspection result can include a p-th inspection result image, which is the p-th inspection image corresponding to the p-th inspection region where the p-th inspection was performed, with the state of the p-th inspection applied, and p-th inspection result text relating to the result value of the p-th inspection.
[0084] In other words, the VLM agent 1000 can generate a p-th inspection image by using the p-th inspection tool 20_p to crop and enlarge the image region corresponding to the p-th inspection region from the image to be discriminated, perform the p-th inspection on the p-th inspection image, and then generate a p-th inspection result image by applying the state after performing the p-th inspection to the p-th inspection image.
[0085] In contrast, the VLM agent 1000 can also use the p-th inspection tool 20_p to acquire a p-th inspection image by zooming in on the p-th inspection area of ββthe product via a camera, perform the p-th inspection on the p-th inspection image, and generate a p-th inspection result image by applying the state after the p-th inspection has been performed to the p-th inspection image.
[0086] Furthermore, when the VLM agent 1000 obtains a specific inspection result corresponding to a specific inspection task, which is one of the first to the nth inspection tasks, it generates an updated specific inspection task by referencing the specific inspection result and updating the specific inspection task. It then requests an updated specific inspection for defects from a specific inspection tool according to the updated specific inspection task and obtains the updated specific inspection result from the specific inspection tool.
[0087] In other words, the VLM agent 1000 can check specific inspection results, and if it finds an error in a specific inspection result, it can update the specific inspection task to correct the error, generate an updated specific inspection task, and then use a specific inspection tool to perform the updated specific inspection on the defect.
[0088] For example, if, after reviewing the results of a specific inspection, it is found that the measurement location for measuring the distance of a defect is incorrect, the measurement location can be corrected, and the distance of the defect can be remeasured according to the corrected measurement location.
[0089] On the other hand, the first to nth inspections for defects can include various types of inspections for measuring defects, and the related inspection tasks and the inspection results are illustrated with reference to Figures 3 to 9 as follows. Figures 3 to 9 illustrate the state in which an inspection is performed on a PCB, and the present invention is not limited thereto, and inspections can be performed in various ways depending on the type of product for which defects are to be detected.
[0090] For example, an inspection tool could be used to preview a portion of the target image a1, namely a region a2.
[0091] At this time, VLM agent 1000, <measure> Overview< / measure> Preview inspection tasks can be generated in this way.
[0092] Referring to Figure 3, the inspection tool can generate preview inspection results, such as an image showing the overview state with the overview region a2 displayed from the image to be identified a1 (as shown in the right diagram of Figure 3), and an enlarged image obtained by cropping and enlarging the overview region a2 from the image to be identified a1 (as shown in the left diagram of Figure 3), and transmit these results to the VLM agent.
[0093] Another example is area information inspection, which uses inspection tools to measure area information.
[0094] At this time, VLM agent 1000, <measure> areainfo< / measure> This allows you to generate domain information inspection tasks.
[0095] Referring to Figure 4, the inspection tool can generate a region information inspection result image in which the critical area b2 and the safe area b3 are separated and displayed in the target image b1. Note that the target image b1 in Figure 4 may be an enlarged image of the overview region a2 in Figure 3.
[0096] Subsequently, the inspection tool generates area information inspection result text, such as AreaInfo(gray: critical area, white: safe area), and can then transmit the area information inspection result, including the area information inspection result image and area information inspection result text, to the VLM agent.
[0097] Another example is a coordinate measurement inspection that uses an inspection tool to measure coordinates, that is, to determine whether the location of a particular coordinate is located in a critical area or a safe area.
[0098] At this time, VLM agent 1000, <measure> points, (500,260)< / measure>A coordinate measurement inspection task can be generated as shown above. <measure> points, (500,260)< / measure> This may also indicate that region information for the (500,260) coordinate is measured in the image to be classified.
[0099] Referring to Figure 5, the inspection tool can generate a coordinate measurement inspection result image that displays the coordinate c2 at (500,260) in the image c1 to be classified.
[0100] Subsequently, the inspection tool generates coordinate measurement inspection result text, such as PointMeasurement(x=500, y=200, area=βcritical areaβ), and then transmits the coordinate measurement inspection result, including the coordinate measurement inspection result image and the coordinate measurement inspection result text, to the VLM agent. Note that PointMeasurement(x=500, y=200, area=βcritical areaβ) may also indicate that coordinate c2 at (500,200) in the image c1 to be classified is located in the critical area.
[0101] Another example would be a rectangle measurement inspection that uses an inspection tool to measure the area of ββa rectangle.
[0102] At this time, VLM agent 1000, <measure> box, (400,200), (600,290)< / measure> A rectangular measurement inspection task can be generated as shown above. <measure> box, (400,200), (600,290)< / measure> This may also indicate that the area of ββthe rectangular box generated by the (400,200) coordinate and the (600,290) coordinate in the image to be classified is measured.
[0103] Referring to Figure 6, the inspection tool can generate a rectangular measurement inspection result image that displays a rectangular box d4 generated by the coordinates (400,200) d2 and (600,290) d3 in the image d1 to be classified.
[0104] Subsequently, the inspection tool generates rectangular measurement inspection result text, such as RectangleMeasurement(width=200, height=90, area=1800, phy_width=2427, phy_height=1092.24, phy_area=265108), and then transmits the rectangular measurement inspection result, including the rectangular measurement inspection result image and the rectangular measurement inspection result text, to the VLM agent. Furthermore, RectangleMeasurement(width=200, height=90, area=1800, phy_width=2427, phy_height=1092.24, phy_area=265108) may indicate that, in the image d1 to be identified, the area of ββthe rectangular box generated by coordinates d2 (400,200) and d3 (600,290) was measured, and the measurement result on the image d1 to be identified is width=200, height=90, area=1800, while the measurement result on the actual product is phy_width=2427, phy_height=1092.24, phy_area=265108. In this case, the measurement result on the actual product may be calculated by referring to the scale ratio between the image to be identified and the actual product.
[0105] Another example might be an angle measurement inspection that uses an inspection tool to measure angles.
[0106] At this time, VLM agent 1000, <measure> angle, (500,400), (400,200), (600,290)< / measure> An angle measurement inspection task can be generated as shown above. <measure> angle, (500,400), (400,200), (600,290)< / measure> This may also indicate that, in the image to be classified, the angle between the line segment formed by the (500,400) coordinates and the (400,200) coordinates and the line segment formed by the (500,400) coordinates and the (600,290) coordinates is measured.
[0107] Referring to Figure 7, the inspection tool can generate an angle measurement inspection result image that displays the line segment e5 formed by coordinates e2 (500,400) and e3 (400,200), and the line segment e6 formed by coordinates e2 (500,400) and e4 (600,290), in the image e1 to be identified.
[0108] Subsequently, the inspection tool generates angle measurement inspection result text, such as AngleMeasurement(angle=68.8), and can then transmit the angle measurement inspection result, including the angle measurement inspection result image and the angle measurement inspection result text, to the VLM agent.
[0109] Another example would be a length measurement inspection using an inspection tool to measure length.
[0110] At this time, VLM agent 1000, <measure> line, (500,250), (650,450)< / measure> A length measurement inspection task can be generated as shown above. Note that line, (500,250), and (650,450) may indicate that the distance between the (500,250) coordinate and the (650,450) coordinate is to be measured in the image to be classified.
[0111] Referring to Figure 8, the inspection tool can generate a distance measurement inspection result image that displays a line f4 connecting coordinates f2 (500,250) and f3 (650,450) in the image f1 to be classified.
[0112] Subsequently, the inspection tool generates distance measurement inspection result text, such as LineMeasurement(length=250.0, phy_length=3034.0), and then transmits the distance measurement inspection result, including the distance measurement inspection result image and the distance measurement inspection result text, to the VLM agent. Note that LineMeasurement(length=250.0, phy_length=3034.0) may indicate that, in the image f1 to be identified, the length of line f4 generated by coordinates (500,250) f2 and (650,450) f3, i.e., the distance between coordinates (500,250) f2 and (650,450) f3, is measured, and the measurement result on the image f1 to be identified is length=250.0, while the measurement result on the actual product is phy_length=3034.0. In this case, the measurement result on the actual product may be calculated by referring to the scale ratio between the image to be identified and the actual product.
[0113] Another example would be a circle measurement inspection that uses an inspection tool to measure circles.
[0114] At this time, VLM agent 1000, <measure> circle, (500,280), (680,290), (550,200)< / measure> A circle measurement inspection task can be generated as shown above. <measure> circle, (500,280), (680,290), (550,200)< / measure> This may also indicate that circles generated by the (500,280) coordinate, (680,290) coordinate, and (550,200) coordinate are measured in the image to be classified.
[0115] Referring to Figure 9, the inspection tool can generate a circle measurement inspection result image that displays the circle g5 generated by coordinates (500,280) g2, (680,290) g3, and (550,200) g4 in the image g1 to be identified, along with the center point g6 of circle g5.
[0116] Subsequently, the inspection tool generates circle measurement inspection result text, such as CircleMeasurement(center_x=590.23, center_y=280.77, radius=90.2, phy_radius=1095.1), and can then transmit the circle measurement inspection result, including the circle measurement inspection result image and the circle measurement inspection result text, to the VLM agent. Furthermore, CircleMeasurement(center_x=590.23, center_y=280.77, radius=90.2, phy_radius=1095.1) may indicate that, in the image g1 to be classified, the circle g5 generated by coordinates (500,280) g2, (680,290) g3, and (550,200) g4 was measured, and the measurement result on the image g1 to be classified shows that the center point coordinates of the circle are (center_x=590.23, center_y=280.77), the radius of the circle is radius=90.2, and the radius of the circle measured on the actual product is phy_radius=1095.1. In this case, the measurement result on the actual product may be the result calculated by referring to the scale ratio between the image to be classified and the actual product.
[0117] Referring again to Figures 1 and 2, the VLM agent 1000 can determine whether a product is defective by referring to the first inspection result or the nth inspection result to confirm whether the defect meets the defect judgment criteria, and generate a product defect determination result (S500).
[0118] In this case, the VLM agent 1000 can input the first inspection result to the nth inspection result into the VLM, and use the VLM to generate a result for determining whether or not the product is defective by referring to the first inspection result to the nth inspection result.
[0119] Furthermore, the VLM agent 1000 can update the chat log by adding the first inspection task or the nth inspection task, input the updated chat log and the first inspection results or the nth inspection results into the VLM, and use the VLM to generate a defect detection result by referring to the updated chat log and the first inspection results or the nth inspection results, and can add the first inspection results or the nth inspection results and the defect detection result to the updated chat log.
[0120] On the other hand, the explanation referring to Figures 1 and 2 described how a first inspection task to the nth inspection task is generated by referring to the defect judgment criteria and the image to be judged, and how the first inspection to the nth inspection is performed using at least one inspection tool. However, contrary to this, it is also possible to detect defects in a product by generating a first inspection task by referring to the defect judgment criteria and the image to be judged, generating a second inspection task by referring to the result of the first inspection, and so on, by generating the next inspection task by referring to the current inspection result, and repeating the operation of performing the inspection. A brief explanation of this is as follows. Note that in the following explanation, explanations of parts that can be easily understood from the explanation referring to Figures 1 and 2 will be omitted.
[0121] First, the VLM agent 1000 can acquire a text prompt regarding the product's defect judgment criteria and an image of the product to be judged, captured by the camera.
[0122] The VLM agent 1000 then refers to the image to be identified to check whether the product has defects. If at least one defect is identified in the image to be identified, it can generate a first inspection task to inspect the defect based on the defect judgment criteria contained in the text prompt.
[0123] Subsequently, the VLM agent 1000 can obtain the first inspection results by requesting the first inspection tool to perform a first inspection of the defect in accordance with the first inspection task, having the first inspection tool perform the first inspection of the defect, and transmitting the first inspection results, which are the result of performing the first inspection.
[0124] In this case, when p is 1, the VLM agent 1000 can save the text prompt and the image to be identified as the first initial chat log, and then add the first inspection task to the first initial chat log to generate the first chat log.
[0125] Next, when p is an integer increasing from 2 to n, the VLM agent 1000 can generate a p-th inspection task to inspect for defects by referring to the defect criteria and the (p-1)th inspection result.
[0126] The VLM agent 1000 then requests the p-th inspection tool to perform the p-th inspection on the defect according to the p-th inspection task, has the p-th inspection tool perform the p-th inspection on the defect, and transmits the p-th inspection result, which is the result of performing the p-th inspection. In this way, the subprocess of obtaining the p-th inspection result can be repeated until p becomes n.
[0127] At this time, the VLM agent 1000 can generate the p-th inspection task by referring to the (p-1)th chat log and the (p-1)th inspection result, and generate the p-th chat log by adding the (p-1)th inspection result and the p-th inspection task to the (p-1)th chat log.
[0128] Next, the VLM agent 1000 can determine whether a product is defective by referring to the first inspection result or the nth inspection result and checking whether the defect meets the defect judgment criteria, thereby generating a product defect determination result.
[0129] At this time, the VLM agent 1000 can generate a defect detection result by referring to the (n-1)th chat log and the nth inspection result, and generate the nth chat log by adding the nth inspection result and the defect detection result to the (n-1)th chat log.
[0130] On the other hand, if the (p-1) inspection result satisfies the defect criteria before p reaches n, the VLM agent 1000 can also interrupt the repeated subprocess and generate a defect determination result by referring to the inspection results obtained up to the point of interruption of the subprocess.
[0131] The embodiments of the present invention described above may be implemented in the form of program instructions that can be executed through various computer components and may be recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., individually or in combination. The program instructions recorded on the computer-readable recording medium may be specially designed and configured for the present invention, or they may be known and available to those skilled in the art in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROMs, RAMs, and flash memory. Examples of program instructions include not only machine code, such as that produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like. The hardware devices may be configured to operate as one or more software modules to perform the processing according to the present invention, and vice versa.
[0132] Although the present invention has been described above with specific details such as concrete components, and with limited embodiments and drawings, these are provided only to aid in a more overall understanding of the invention, and the invention is not limited to the above embodiments. A person with ordinary skill in the art to which the invention pertains can make various modifications and variations from this description.
[0133] Therefore, the concept of the present invention should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the claims described below shall also fall within the scope of the concept of the present invention. [Explanation of Symbols]
[0134] 1000 VLM agents 1100 memory 1200 processors
Claims
1. In a method for detecting product defects using a VLM (Vision Language Model) agent, (a) When a text prompt relating to product defect criteria and an image of the product to be identified captured by a camera are acquired, a VLM agent including a memory storing instructions for detecting product defects and a processor that detects product defects according to the instructions stored in the memory, refers to the image to be identified to determine whether the product has defects, and when at least one defect is identified in the image to be identified, generates a first inspection task to an nth inspection task (where n is an integer of 1 or more) to inspect the defect based on the defect criteria included in the text prompt, (b) The VLM agent obtains the first to nth inspection results by requesting the pth inspection tool to perform the pth inspection on the defect according to the first to nth inspection task, which is the first to nth inspection task (where p is an integer that increases sequentially from 1 to n), having the pth inspection tool perform the pth inspection on the defect, and transmitting the pth inspection result which is the result of performing the pth inspection. (c) The VLM agent determines whether the product is defective by referring to the first inspection result to the nth inspection result and confirming whether the defect falls under the defect determination criteria, and generates a defect determination result for the product. Includes, In step (b) above, The p-inspection result includes a p-inspection result image obtained by applying the state under which the p-inspection was performed to the p-inspection image corresponding to the p-inspection region where the p-inspection was performed, and p-inspection result text relating to the result value of the p-inspection. The VLM agent can perform either of the following subprocesses: (i) a subprocess that uses the p-inspection tool to crop and enlarge an image region corresponding to the p-inspection region from the image to be determined to generate the p-inspection image, performs the p-inspection on the p-inspection image, and applies the state after the p-inspection has been performed to the p-inspection image to generate the p-inspection result image; or (ii) a subprocess that uses the p-inspection tool to acquire the p-inspection image taken by zooming in on the p-inspection region of the product through the camera, performs the p-inspection on the p-inspection image, and applies the state after the p-inspection has been performed to the p-inspection image to generate the p-inspection result image.
2. In step (a) above, The VLM agent inputs the text prompt and the image to be identified into the VLM, and uses the VLM to analyze the text prompt and the image to be identified to generate the first inspection task to the nth inspection task. In step (c) above, The method according to claim 1, wherein the VLM agent inputs the first inspection result to the nth inspection result to the VLM, and uses the VLM to refer to the first inspection result to the nth inspection result to generate a result for determining whether or not the product has defects.
3. In step (a) above, The VLM agent saves the text prompt and the image to be identified as a chat log, and updates the chat log by adding the first inspection task to the nth inspection task, thereby generating an updated chat log. In step (c) above, The method according to claim 2, wherein the VLM agent inputs the updated chat log and the first inspection result to the nth inspection result into the VLM, uses the VLM to generate a defect detection result by referring to the updated chat log and the first inspection result to the nth inspection result, and adds the first inspection result to the nth inspection result and the defect detection result to the updated chat log.
4. In step (b) above, The method according to claim 1, wherein the VLM agent, upon obtaining a specific inspection result obtained by performing a specific inspection corresponding to a specific inspection task which is any one of the first inspection task to the nth inspection task, checks whether there is an error in the specific inspection result, and if there is an error, generates an updated specific inspection task which is an updated version of the specific inspection task in order to correct the error, requests an updated specific inspection for the defect from a specific inspection tool according to the updated specific inspection task, and obtains an updated specific inspection result from the specific inspection tool.
5. In step (a) above, The method according to claim 1, wherein the VLM agent further acquires a vision prompt which is a sample image indicating a defective or good product, and further analyzes whether the image to be determined has the defect by referring to the vision prompt.
6. In a method for detecting product defects using a VLM (Vision Language Model) agent, (a) When a text prompt relating to product defect criteria and an image of the product to be identified taken from a camera are acquired, a VLM agent including a memory storing instructions for detecting product defects and a processor that detects product defects according to the instructions stored in the memory, refers to the image to be identified to determine whether the product has defects, and when at least one defect is identified in the image to be identified, if p is 1, generates a first inspection task to inspect the defect based on the defect criteria included in the text prompt, requests a first inspection tool to perform a first inspection on the defect according to the first inspection task, causes the first inspection tool to perform the first inspection on the defect, and transmits the first inspection result which is the result of performing the first inspection, thereby acquiring a first inspection result; (b) The VLM agent, when p is an integer increasing from 2 to n (where n is an integer greater than or equal to 2), generates a p-th inspection task for inspecting the defect by referring to the defect judgment criterion and the (p-1) inspection result, requests the p-th inspection tool to perform the p-th inspection on the defect according to the p-th inspection task, causes the p-th inspection tool to perform the p-th inspection on the defect, and transmits the p-th inspection result which is the result of performing the p-th inspection, thereby repeating the subprocess for obtaining the p-th inspection result until p becomes n. (c) The VLM agent determines whether the product is defective by referring to the first inspection result to the nth inspection result and confirming whether the defect falls under the defect determination criteria, and generates a defect determination result for the product. Includes, In step (a) above, The first inspection result includes a first inspection result image obtained by applying the state in which the first inspection was performed to a first inspection image corresponding to the first inspection region in which the first inspection was performed, and first inspection result text relating to the result value of performing the first inspection. The VLM agent executes either one of the following subprocesses: (i) a subprocess that uses the first inspection tool to crop and enlarge an image region corresponding to the first inspection region from the image to be determined to generate the first inspection image, performs the first inspection on the first inspection image, and applies the state after the first inspection has been performed to the first inspection image to generate the first inspection result image; or (ii) a subprocess that uses the first inspection tool to acquire the first inspection image by zooming in on the first inspection region of the product through the camera, performs the first inspection on the first inspection image, and applies the state after the first inspection has been performed to the first inspection image to generate the first inspection result image. In step (b) above, The p-inspection result includes a p-inspection result image obtained by applying the state under which the p-inspection was performed to the p-inspection image corresponding to the p-inspection region where the p-inspection was performed, and p-inspection result text relating to the result value of the p-inspection. The VLM agent can perform either of the following subprocesses: (i) a subprocess that uses the p-inspection tool to crop and enlarge an image region corresponding to the p-inspection region from the image to be determined to generate the p-inspection image, performs the p-inspection on the p-inspection image, and applies the state after the p-inspection has been performed to the p-inspection image to generate the p-inspection result image; or (ii) a subprocess that uses the p-inspection tool to acquire the p-inspection image taken by zooming in on the p-inspection region of the product through the camera, performs the p-inspection on the p-inspection image, and applies the state after the p-inspection has been performed to the p-inspection image to generate the p-inspection result image.
7. In step (a) above, The VLM agent saves the text prompt and the image to be identified as a first initial chat log, and adds the first inspection task to the first initial chat log to generate a first chat log. In step (b) above, The VLM agent generates the p-inspection task by referring to the (p-1) chat log and the (p-1) inspection result, and generates the p-th chat log by adding the (p-1) inspection result and the p-inspection task to the (p-1) chat log. In step (c) above, The method according to claim 6, wherein the VLM agent generates a defect detection result by referring to the (n-1)th chat log and the nth inspection result, and generates the nth chat log by adding the nth inspection result and the defect detection result to the (n-1)th chat log.
8. In step (b) above, The VLM agent interrupts the subprocess if the (p-1) inspection result satisfies the failure criteria before p reaches n. In step (c) above, The method according to claim 6, wherein the VLM agent generates the defect determination result by referring to the inspection results obtained until the subprocess is interrupted.
9. In a VLM agent that detects product defects, A memory containing instructions for detecting product defects, A processor that detects defects in the product according to the instructions stored in the memory, Includes, The processor, in accordance with the instructions, (I) when it obtains a text prompt relating to the defect judgment criteria for the product and an image of the product to be judged taken from the camera, it checks whether the product has a defect by referring to the image to be judged, and when at least one defect is confirmed in the image to be judged, it generates a first inspection task to the nth inspection task (where n is an integer of 1 or more) for inspecting the defect based on the defect judgment criteria included in the text prompt, and (II) the pth inspection task which is the first inspection task to the nth inspection task (where p is an integer of 1 or more) (III) A process to obtain a first inspection result to an nth inspection result by requesting the pth inspection tool to perform the pth inspection on the defect according to an integer that increases sequentially up to n, having the pth inspection tool perform the pth inspection on the defect, and transmitting the pth inspection result which is the result of performing the pth inspection, and (III) a process to determine whether the product is defective by referring to the first inspection result to the nth inspection result and confirming whether the defect falls under the defect determination criteria, and generating a defect determination result for the product, In the above process (II), The p-inspection result includes a p-inspection result image obtained by applying the state under which the p-inspection was performed to the p-inspection image corresponding to the p-inspection region where the p-inspection was performed, and p-inspection result text relating to the result value of the p-inspection. The aforementioned processor, In the process of (II) described above, a VLM agent executes one of the following subprocesses: (i) a subprocess that uses the p-inspection tool to crop and enlarge an image region corresponding to the p-inspection region from the image to be determined to generate the p-inspection image, performs the p-inspection on the p-inspection image, and generates the p-inspection result image by applying the state in which the p-inspection has been performed to the p-inspection image; and (ii) a subprocess that uses the p-inspection tool to acquire the p-inspection image taken by zooming in on the p-inspection region of the product through the camera, performs the p-inspection on the p-inspection image, and generates the p-inspection result image by applying the state in which the p-inspection has been performed to the p-inspection image.
10. The aforementioned processor, In the process described in (I) above, the text prompt and the image to be identified are input to the VLM, and the VLM is used to analyze the text prompt and the image to be identified in order to generate the first inspection task to the nth inspection task. The VLM agent according to claim 9, wherein in the process of (III), the first inspection result to the nth inspection result is input to the VLM, and the VLM is used to generate a result for determining whether or not the product is defective by referring to the first inspection result to the nth inspection result.
11. The aforementioned processor, In the process described in (I) above, the text prompt and the image to be identified are saved as a chat log, and the chat log is updated by adding the first inspection task to the nth inspection task to the chat log, thereby generating an updated chat log. The VLM agent according to claim 10, wherein in the process of (III), the updated chat log and the first inspection result to the nth inspection result are input to the VLM, the VLM is used to generate the defect detection result by referring to the updated chat log and the first inspection result to the nth inspection result, and the first inspection result to the nth inspection result and the defect detection result are added to the updated chat log.
12. The aforementioned processor, In the process of (II) above, when a specific inspection result is obtained in which a specific inspection corresponding to a specific inspection task which is any one of the first inspection task to the nth inspection task is performed, the VLM agent according to claim 9 checks whether there is an error in the specific inspection result, and if there is an error, generates an updated specific inspection task which is an updated version of the specific inspection task in order to correct the error, requests an updated specific inspection for the defect from a specific inspection tool according to the updated specific inspection task, and obtains an updated specific inspection result from the specific inspection tool.
13. The aforementioned processor, The VLM agent according to claim 9, wherein in the process of (I) above, a vision prompt is obtained which is a sample image indicating a defective or good product, and the VLM agent further analyzes whether the image to be determined has the defect by referring to the vision prompt.
14. In a VLM agent that detects product defects, A memory containing instructions for detecting product defects, A processor that detects defects in the product according to the instructions stored in the memory, Includes, The processor, in accordance with the instructions, (I) when a text prompt relating to the defect judgment criteria for the product and a target image of the product taken from the camera are acquired, it refers to the target image to check whether the product has a defect, and when at least one defect is confirmed in the target image, when p is 1, it generates a first inspection task for inspecting the defect based on the defect judgment criteria included in the text prompt, requests a first inspection tool to perform a first inspection on the defect according to the first inspection task, has the first inspection tool perform the first inspection on the defect, and transmits the first inspection result which is the result of performing the first inspection, thereby obtaining a first inspection result, (II) when p is 2 to n ( (III) When n is an increasing integer (where n is an integer of 2 or more), the process of obtaining the p-th inspection result is repeated until p becomes n, by referring to the defect judgment criterion and the (p-1) inspection result, generating a p-th inspection task for inspecting the defect, requesting the p-th inspection tool to perform the p-th inspection on the defect according to the p-th inspection task, having the p-th inspection tool perform the p-th inspection on the defect, and transmitting the p-th inspection result which is the result of performing the p-th inspection, and (III) determining whether the product is defective by referring to the first inspection result to the n-th inspection result and determining whether the product is defective, and generating a defect determination result for the product. In the above process (I), The first inspection result includes a first inspection result image obtained by applying the state in which the first inspection was performed to a first inspection image corresponding to the first inspection region in which the first inspection was performed, and first inspection result text relating to the result value of performing the first inspection. The aforementioned processor, In the process described in (I), one of the following subprocesses is executed: (i) a subprocess that uses the first inspection tool to crop and enlarge an image region corresponding to the first inspection region from the image to be determined to generate the first inspection image, performs the first inspection on the first inspection image, and generates the first inspection result image by applying the state after the first inspection has been performed to the first inspection image; and (ii) a subprocess that uses the first inspection tool to acquire the first inspection image by zooming in on the first inspection region of the product through the camera, performs the first inspection on the first inspection image, and generates the first inspection result image by applying the state after the first inspection has been performed to the first inspection image. In the above process (II), The p-inspection result includes a p-inspection result image obtained by applying the state under which the p-inspection was performed to the p-inspection image corresponding to the p-inspection region where the p-inspection was performed, and p-inspection result text relating to the result value of the p-inspection. The aforementioned processor, In the process of (II) described above, a VLM agent executes one of the following subprocesses: (i) a subprocess that uses the p-inspection tool to crop and enlarge an image region corresponding to the p-inspection region from the image to be determined to generate the p-inspection image, performs the p-inspection on the p-inspection image, and generates the p-inspection result image by applying the state in which the p-inspection has been performed to the p-inspection image; and (ii) a subprocess that uses the p-inspection tool to acquire the p-inspection image taken by zooming in on the p-inspection region of the product through the camera, performs the p-inspection on the p-inspection image, and generates the p-inspection result image by applying the state in which the p-inspection has been performed to the p-inspection image.
15. The aforementioned processor, In the process described in (I) above, the text prompt and the image to be identified are saved as a first initial chat log, and the first inspection task is added to the first initial chat log to generate a first chat log. In the process described in (II), the p-inspection task is generated by referring to the (p-1) chat log and the (p-1) inspection result, and the p-inspection task is added to the (p-1) chat log to generate the p-th chat log. The VLM agent according to claim 14, wherein in the (III) process, the (n-1) chat log and the n inspection results are referenced to generate the defect detection result, and the n inspection results and the defect detection result are added to the (n-1) chat log to generate the n chat log.
16. The aforementioned processor, In the process (II) described above, if the inspection result (p-1) satisfies the defect criteria before p reaches n, the sub-process is interrupted. The VLM agent according to claim 14, wherein in the (III) process, the VLM agent generates the defect determination result by referring to the inspection results obtained until the subprocess is interrupted.