Method for generating training data for training an agent model based on a vision-language model, and a data generation apparatus using the same.

JP7911805B1Active Publication Date: 2026-08-27SUPERB AI CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025248360
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2025-11-05
Filing Date
2025-12-15
Publication Date
2026-08-27
Estimated Expiration
2045-12-15

AI Technical Summary

Benefits of technology

【0034】 本発明は、製品に対する不良検査基準および製品に対応するモデリング画像が取得されることにより、ディフェクトジェネレーターを通じて、不良検査基準を参照してディフェクトマスクを生成し、モデリング画像に適用される所定の傾きおよび回転に関するレンダリング情報を生成し、不良検査基準、ディフェクトマスクおよびレンダリング情報を含むディフェクトパラメーターを参照して、モデリング画像に基づく判別対象画像を生成するように支援することにより、ビジョン言語モデルに基づくエージェントモデルの学習に必要な不良画像を確保する効果がある。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007911805000001_ABST
    Figure 0007911805000001_ABST
Patent Text Reader

Abstract

This invention provides a method for generating training data that secures defective images necessary for training an agent model based on a vision-language model, and a data generation apparatus using the same. [Solution] The method acquires defect inspection criteria for a product and a modeling image corresponding to the product, generates a defect mask by referencing the defect inspection criteria through a defect generator, generates rendering information regarding predetermined tilt and rotation applied to the modeling image, assists in generating a discriminant image based on the modeling image by referencing the defect inspection criteria and defect parameters, secures the defect images necessary for training an agent model based on a vision language model, and generates training data for training the agent model based on a vision language model based on a document generator based on a large-scale language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating learning data for learning an agent model based on a vision language model and a data generation device using the same. More specifically, when a defect inspection standard for a product and a modeling image corresponding to the product are acquired, a discrimination target image and chat log data for learning the agent model are generated, and learning data is generated by merging the discrimination target image and the chat log data. The present invention also relates to a data generation device using the same method.

Background Art

[0002] Generally, in order to detect defects in products produced at a manufacturing site, AOI (Automated Optical Inspection) is used. AOI analyzes a product image captured using a camera to check whether the product meets quality specifications, thereby inspecting for the presence or absence of product defects.

[0003] Such a conventional AOI determines the presence or absence of product defects based on set rules. It compares a reference image for non-defective products set in advance with the captured product image and determines a defect when it deviates from rules such as coordinates, color, and size.

[0004] However, while the conventional AOI method has the advantages of being able to quickly determine defects, being effective for simple and repetitive defect detection, and having relatively easy settings, it is sensitive to changes in lighting and the shooting position of the camera. Not only is it necessary to reset the rules for defect determination for each new product, but there is also a problem that the user must intervene for rule setting and exception handling.

[0005] Therefore, in recent years, a method of executing AOI by an AI (Artificial Intelligence) deep learning method has been proposed. <00000In the AI ​​deep learning method, an AI model that has learned from a large number of normal and defective images independently recognizes the characteristics of defects and patterns to determine whether or not a product is defective. This method has the advantages of high accuracy, the ability to detect minute and atypical defects, and flexibility in application to various product groups.

[0007] However, the AI ​​deep learning method has the drawback of requiring not only a large amount of training data to train the AI ​​model, but also a significant amount of time and computing resources to train the AI ​​model.

[0008] Specifically, securing a large amount of training data for training an AI model requires a large quantity of both normal and defective images. While the quantity of image data of products acquired at the manufacturing site may be sufficient, the occurrence rate of defective products becomes extremely low as manufacturing sites become more sophisticated. Therefore, securing a large quantity of defective images is extremely difficult. The process of classifying whether an item is normal or defective based on defect inspection criteria must be performed on the available large amount of training data, but this is difficult to do with limited computing resources.

[0009] Furthermore, while vision-language models are primarily used for AI models that perform AOI (Autonomous Object Inspection), they have the drawback of being poor at recognizing small objects in images due to their characteristics. When training an AI model using training data that includes defective images with extremely small defects, the AI ​​model cannot determine what defect criteria to use to classify extremely defective images as defects, resulting in poor learning performance.

[0010] Therefore, there is a need for improvement measures to solve the above problems. [Overview of the Initiative] [Problems that the invention aims to solve]

[0011] The purpose of this invention is to solve all of the problems mentioned above.

[0012] Furthermore, the present invention also aims to secure defective images necessary for training an agent model based on a Vision Language Model (VLM) by acquiring defective inspection criteria for a product and a modeling image corresponding to the product, generating a defect mask by referencing the defective inspection criteria through a defect generator, generating rendering information regarding predetermined tilt and rotation applied to the modeling image, and assisting in generating a discriminant image based on the modeling image by referencing defect parameters including the defective inspection criteria, defect mask, and rendering information.

[0013] Furthermore, the present invention also aims to generate training data for training an agent model based on a vision language model by generating a description of an image to be classified using defect parameters through a document generator based on a large language model (LLM), and then generating a tool usage plan, tool usage request information, tool usage simulation result information, and defect detection-related information for the image to be classified through interaction between the agent large language model and the tool simulation large language model included in the document generator, and generating chat log data including defect inspection criteria, the description of the image to be classified, the tool usage plan information, the tool usage request information, the tool usage simulation result information, and defect detection-related information for the image to be classified, and merging the image to be classified generated through the defect generator into the chat log data. [Means for solving the problem]

[0014] According to one embodiment of the present invention, in a method for generating training data for training an agent model based on a vision language model, (a) when a defect inspection criterion for a product and a modeling image corresponding to the product are acquired, a data generation device inputs the defect inspection criterion and the modeling image to a defect generator, and uses the defect generator to generate a defect mask (the defect mask includes information on the type and location of defects applied to the modeling image) by referring to the defect inspection criterion, and generates rendering information including predetermined tilt and rotation information applied to the modeling image, and acquires defect parameters including the defect inspection criterion, the defect mask and the rendering information; (b) a subprocess in which the data generation device (i) uses the defect generator to generate a discriminant image based on the modeling image by referring to the defect parameters, and (ii) uses the defect parameters A subprocess is performed to input into a document generator, use the document generator to generate the target image description by referring to the defect parameters, generate tool usage plan information by referring to the defect inspection criteria and the target image description, generate tool usage request information for detecting the defect by referring to the tool usage plan information, generate tool usage simulation result information by simulating the tool usage result based on the tool usage request information in the target image description, generate information related to the determination of whether or not there is a defect in the target image by referring to the defect inspection criteria and the tool usage simulation result information, and generate chat log data including the defect inspection criteria, the target image description, the tool usage plan information, the tool usage request information, the tool usage simulation result information, and the information related to the determination of whether or not there is a defect in the target image;Furthermore, a method is provided which includes the step of (c) the data generation device inputting the image to be discriminated and the chat log data into a Merger, and using the Merger to generate training data by merging the image to be discriminated with the chat log data.

[0015] In one example, the document generator includes an agent large-scale language model and a tool simulation large-scale language model, and in step (b), the data generation device uses the agent large-scale language model to generate first tool usage plan information to nth tool usage plan information (where n is an integer of 1 or more) by referring to the defect inspection criteria and the image description to be determined, generates pth tool usage request information for detecting the defect with the pth tool based on the pth tool usage plan information (where p is an integer that increases sequentially from 1 to n) corresponding to the first tool usage plan information to nth tool usage plan information, and inputs the image description to be determined and the pth tool usage request information into the tool simulation large-scale language model, thereby enabling the tool A large-scale simulation language model is used to generate p-tool usage simulation result information, which is the result of simulating the use of the p-tool based on the p-tool usage request information in the discriminant image description; the p-tool usage simulation result information is input to the agent large-scale language model; the agent large-scale language model is used to generate information related to the determination of whether or not there is a defect in the discriminant image by referring to the defect inspection criteria and the p-tool usage simulation result information; and the chat log data is generated, which includes the defect inspection criteria, the discriminant image description, the p-tool usage plan information, the p-tool usage request information, the p-tool usage simulation result information, and the information related to the determination of whether or not there is a defect in the discriminant image.

[0016] In one example, in step (b), when the data generation device obtains specific tool usage simulation result information corresponding to a specific tool usage plan which is at least a part of the first tool usage plan to the n tool usage plans from the tool simulation large language model, the data generation device inputs the specific tool usage simulation result information into the agent large language model using the tool simulation large language model, the agent large language model uses the specific tool usage simulation result information to generate specific tool usage plan update information which updates the specific tool usage plan, the data generation device generates specific tool usage request update information for detecting the defect using a specific tool based on the specific tool usage plan update information, and inputs the specific tool usage request update information into the tool simulation large language model. The large-scale language model for tool simulation is configured to generate specific tool usage simulation result update information, which simulates the use of the specific tool based on the specific tool usage request update information in the image description to be determined; the specific tool usage simulation result update information is input to the large-scale language model for agent use; the large-scale language model for agent use is configured to generate update information related to the determination of whether or not there is a defect in the image to be determined, by referring to the defect inspection criteria and the specific tool usage simulation result update information; and the chat log data is configured to include the defect inspection criteria, the image description to be determined, the specific tool usage plan update information, the specific tool usage request update information, the specific tool usage simulation result update information, and the update information related to the determination of whether or not there is a defect in the image to be determined.

[0017] In one example, the document generator includes an agent large-scale language model and a tool simulation large-scale language model, and in step (b), the data generation device (i) uses the agent large-scale language model to generate first tool usage plan information by referring to the defect inspection criteria and the image description to be identified when p is 1, generates first tool usage request information for detecting the defect with the first tool based on the first tool usage plan information, inputs the image description to be identified and the first tool usage request information into the tool simulation large-scale language model, and uses the tool simulation large-scale language model to generate first tool usage simulation result information by simulating the use of the first tool based on the first tool usage request information in the image description to be identified, and the first tool usage simulation result (ii) A subprocess that inputs information into the agent large-scale language model; and (ii) a subprocess that uses the agent large-scale language model to generate p-th tool usage plan information by referring to the defect inspection criteria and the (p-1) tool usage simulation result information when p is an integer increasing from 2 to n (where n is an integer greater than or equal to 2), generates p-th tool usage request information for detecting the defect using the p-th tool based on the p-th tool usage plan information, inputs the p-th tool usage request information into the tool simulation large-scale language model, uses the tool simulation large-scale language model to generate p-th tool usage simulation result information simulating the use of the p-th tool in the discriminant image description based on the p-th tool usage request information, and inputs the p-th tool usage simulation result information into the agent large-scale language model;This process is repeated until p becomes n, and the agent's large-scale language model is used to generate information related to the determination of whether or not a defect exists in the target image, by referring to the defect inspection criteria and the first tool usage simulation result information or the nth tool usage simulation result information. Chat log data is then generated that includes the defect inspection criteria, the target image description, the first tool usage plan information or the nth tool usage plan information, the first tool usage request information or the nth tool usage request information, the first tool usage simulation result information or the nth tool usage simulation result information, and the information related to the determination of whether or not a defect exists in the target image.

[0018] In one example, in step (b), the data generation device generates a text template for describing the image to be identified using the document generator, extracts important keywords from the defect parameters, and then applies the important keywords to the text template to generate the image description to be identified.

[0019] In one example, in step (c), the data generation device generates the training data by using the merger to change the image description to be identified in the chat log data to the image to be identified.

[0020] In one example, in step (c), the data generation device generates the training data by using the merger to refer to the tool usage simulation result information contained in the chat log data, generating a tool usage result image corresponding to the tool usage simulation result information, and changing the tool usage simulation result information to the tool usage result image.

[0021] In one example, in step (a) above, the data generation device uses the defect generator to apply the rendering information to the modeling image to generate a rendering image, and inputs the rendering image and the defect mask into an image synthesis model (the image synthesis model is included in the defect generator) to generate the image to be discriminated through the image synthesis model.

[0022] In one example, in step (b), the tool usage simulation result information includes at least a portion of the measured values ​​obtained by using the tool, the overview image description, and the image description in which the use of the tool was simulated.

[0023] In one example, the procedure further includes (d) the data generation device fine-tuning an agent model based on a vision language model through learning using the training data.

[0024] According to one embodiment of the present invention, a data generation device for generating training data for training an agent model based on a vision language model includes at least one memory for storing instructions; and at least one processor configured to execute the instructions, wherein the processor (i) when a defect inspection criterion for a product and a modeling image corresponding to the product are acquired, inputs the defect inspection criterion and the modeling image to a defect generator, and uses the defect generator to generate a defect mask (the defect mask includes information on the type and location of defects applied to the modeling image) by referring to the defect inspection criterion, generates rendering information including predetermined tilt and rotation information applied to the modeling image, and acquires defect parameters including the defect inspection criterion, the defect mask, and the rendering information;(II)(i) A subprocess that uses the defect generator to generate a discriminant image based on the modeling image by referring to the defect parameters, and (ii) inputting the defect parameters into a document generator, using the document generator to generate a discriminant image description by referring to the defect parameters, generating tool usage plan information by referring to the defect inspection criteria and the discriminant image description, generating tool usage request information for detecting the defect by referring to the tool usage plan information, and simulating the tool usage result based on the tool usage request information in the discriminant image description, A data generation device is provided that performs a subprocess to generate tool usage simulation result information, generate information related to the determination of whether or not a defect exists for the target image by referring to the defect inspection criteria and the tool usage simulation result information, and generate chat log data including the defect inspection criteria, the target image description, the tool usage plan information, the tool usage request information, the tool usage simulation result information, and the information related to the determination of whether or not a defect exists for the target image; and (III) a process to input the target image and the chat log data into a Merger, and use the Merger to generate training data by merging the chat log data with the target image.

[0025] In one example, the document generator includes an agent large-scale language model and a tool simulation large-scale language model, and the processor generates first tool usage plan information to nth tool usage plan information (where n is an integer of 1 or more) using the agent large-scale language model, referring to the defect inspection criteria and the image description to be determined, and generates pth tool usage request information for detecting the defect using the pth tool based on the pth tool usage plan information (where p is an integer that increases sequentially from 1 to n) corresponding to the first tool usage plan information to nth tool usage plan information, and inputs the image description to be determined and the pth tool usage request information into the tool simulation large-scale language model, and the tool simulation A large-scale simulation language model is used to generate p-tool usage simulation result information, which is the result of simulating the use of the p-tool based on the p-tool usage request information in the image description to be determined; the p-tool usage simulation result information is input to the agent large-scale language model; the agent large-scale language model is used to generate information related to the determination of whether or not there is a defect in the image to be determined, by referring to the defect inspection criteria and the p-tool usage simulation result information; and chat log data is generated, which includes the defect inspection criteria, the image description to be determined, the p-tool usage plan information, the p-tool usage request information, the p-tool usage simulation result information, and the information related to the determination of whether or not there is a defect in the image to be determined.

[0026] In one example, when the processor obtains, in the (II) process, specific tool usage simulation result information corresponding to a specific tool usage plan that is at least part of the first tool usage plan to the nth tool usage plan from the tool simulation large language model, the tool simulation large language model is used to input the specific tool usage simulation result information into the agent large language model, and the agent large language model is used to refer to the specific tool usage simulation result information to generate specific tool usage plan update information obtained by updating the specific tool usage plan. Based on the specific tool usage plan update information, specific tool usage requirement update information for detecting the defect by a specific tool is generated, the specific tool usage requirement update information is input into the tool simulation large language model, and the tool simulation large language model is used to generate specific tool usage simulation result update information obtained by simulating the use of the specific tool based on the specific tool usage requirement update information in the discrimination target image description. The specific tool usage simulation result update information is input into the agent large language model, and the agent large language model is used to refer to the defect inspection criteria and the specific tool usage simulation result update information to generate defect presence / absence discrimination related update information for the discrimination target image. The chat log data including the defect inspection criteria, the discrimination target image description, the specific tool usage plan update information, the specific tool usage requirement update information, the specific tool usage simulation result update information, and the defect presence / absence discrimination related update information for the discrimination target image is generated.

[0027] In one example, the document generator includes an agent large-scale language model and a tool simulation large-scale language model, and the processor, in the (II) process, (i) uses the agent large-scale language model to generate first tool usage plan information by referring to the defect inspection criteria and the image description to be identified when p is 1, generates first tool usage request information for detecting the defect with the first tool based on the first tool usage plan information, inputs the image description to be identified and the first tool usage request information into the tool simulation large-scale language model, and uses the tool simulation large-scale language model to generate first tool usage simulation result information by simulating the use of the first tool based on the first tool usage request information in the image description to be identified, and the first tool usage simulation result (ii) A subprocess that inputs information into the agent large-scale language model; and (ii) a subprocess that uses the agent large-scale language model to generate p-th tool usage plan information by referring to the defect inspection criteria and the (p-1) tool usage simulation result information when p is an integer increasing from 2 to n (where n is an integer greater than or equal to 2), generates p-th tool usage request information for detecting the defect using the p-th tool based on the p-th tool usage plan information, inputs the p-th tool usage request information into the tool simulation large-scale language model, uses the tool simulation large-scale language model to generate p-th tool usage simulation result information simulating the use of the p-th tool in the discriminant image description based on the p-th tool usage request information, and inputs the p-th tool usage simulation result information into the agent large-scale language model;Repeat until p becomes n, and use the agent large language model to generate the defect presence / absence discrimination related information for the image to be discriminated by referring to the defect inspection criteria and the first tool usage simulation result information to the nth tool usage simulation result information, and generate chat log data including the defect inspection criteria, the description of the image to be discriminated, the first tool usage plan information to the nth tool usage plan information, the first tool usage requirement information to the nth tool usage requirement information, the first tool usage simulation result information to the nth tool usage simulation result information, and the defect presence / absence discrimination related information for the image to be discriminated.;

[0028] In one example, in the (II) process, the processor uses the document generator to generate a text template for explaining the image to be discriminated, extracts important keywords from the defect parameters, and then applies the important keywords to the text template to generate the description of the image to be discriminated.

[0029] In one example, in the (III) process, the processor uses the merger to change the description of the image to be discriminated in the chat log data to the image to be discriminated, thereby generating the learning data.

[0030] In one example, in the (III) process, the processor uses the merger to generate a tool usage result image corresponding to the tool usage simulation result information by referring to the tool usage simulation result information included in the chat log data, and changes the tool usage simulation result information to the tool usage result image, thereby generating the learning data.

[0031] In one example, the processor, in process (I), uses the defect generator to apply the rendering information to the modeling image to generate a rendered image, inputs the rendered image and the defect mask into an image synthesis model (the image synthesis model is included in the defect generator), and generates the image to be discriminated through the image synthesis model.

[0032] In one example, in process (II) above, the tool usage simulation result information includes at least a portion of the measured values ​​obtained by using the tool, the overview image description, and the image description in which the use of the tool was simulated.

[0033] In one example, (IV) the processor further performs a process of fine-tuning an agent model based on a vision language model through learning using the training data. [Effects of the Invention]

[0034] The present invention has the effect of securing defective images necessary for training an agent model based on a vision language model by acquiring defect inspection criteria for a product and a modeling image corresponding to the product, generating a defect mask by referencing the defect inspection criteria through a defect generator, generating rendering information regarding predetermined tilt and rotation applied to the modeling image, and assisting in generating a discriminable image based on the modeling image by referencing defect parameters including the defect inspection criteria, defect mask, and rendering information.

[0035] Furthermore, the present invention generates a description of the image to be discriminated using defect parameters through a document generator based on a large-scale language model, and through interaction between the agent large-scale language model and the tool simulation large-scale language model included in the document generator, generates a tool usage plan, tool usage request information, tool usage simulation result information, and information related to the determination of whether or not there are defects in the image to be discriminated. It also generates chat log data including defect inspection criteria, the description of the image to be discriminated, the tool usage plan information, the tool usage request information, the tool usage simulation result information, and information related to the determination of whether or not there are defects in the image to be discriminated, and merges the image to be discriminated generated through the defect generator with the chat log data, thereby generating training data for training an agent model based on a vision language model. [Brief explanation of the drawing]

[0036] The following drawings, attached for use in describing embodiments of the present invention, represent only a portion of the embodiments, and a person with ordinary skill in the art to which the present invention pertains (hereinafter referred to as "ordinary art") can obtain other drawings based on these drawings without performing any inventive work.

[0037] [Figure 1] This figure schematically shows a data generation device that generates training data for training an agent model based on a vision language model according to one embodiment of the present invention. [Figure 2] This figure schematically shows a flowchart for generating training data for training an agent model based on a vision language model according to one embodiment of the present invention. [Figure 3] This figure schematically shows examples of a modeling image, a defect mask, and an image to be discriminated according to one embodiment of the present invention. [Modes for carrying out the invention]

[0038] The detailed description of the present invention, as described below, will refer to the accompanying drawings illustrating specific embodiments in which the present invention may be carried out, in order to clarify the object, technical solution, and advantages of the present invention. These embodiments will be described in sufficient detail so that a person of the ordinary skill can carry out the present invention.

[0039] Furthermore, throughout the detailed description and claims of the present invention, the word “including” and its variations are not intended to exclude other technical features, additions, components, or steps. To an ordinary person, some of the other purposes, advantages, and characteristics of the present invention will become apparent from this specification and some from the practice of the present invention. The following examples and drawings are provided as illustrative examples and are not intended to limit the present invention.

[0040] Furthermore, the present invention encompasses all possible combinations of the embodiments shown herein. It should be understood that while the various embodiments of the present invention differ from one another, they do not necessarily have to be mutually exclusive. For example, certain shapes, structures, and characteristics described herein can be realized in other embodiments in relation to one embodiment without departing from the spirit and scope of the invention. It should also be understood that the position or arrangement of individual components within each disclosed embodiment can be modified without departing from the spirit and scope of the invention. Therefore, the detailed descriptions set forth below should not be taken as restrictive, and the scope of the present invention is limited only by the appended claims, along with all equivalent scopes claimed by those claims, provided they are adequately described. Similar reference numerals in the drawings refer to identical or similar functions across various aspects.

[0041] The titles or abstracts of the inventions provided herein are provided for convenience only and are not intended to limit or imply any limitation of the scope or meaning of these embodiments.

[0042] Furthermore, even if each component is listed singly below, this does not rule out the possibility of multiple components.

[0043] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings, so that persons with ordinary skill in the art to which the present invention pertains can easily implement the present invention.

[0044] Figure 1 is a schematic diagram showing a data generation device that generates training data for training an agent model based on a vision language model according to one embodiment of the present invention.

[0045] Referring to Figure 1, the data generation device 1000 may include a memory 1100 that stores instructions for generating training data for training an agent model based on a vision language model, and a processor 1200 that generates training data for training an agent model based on a vision language model in response to the instructions stored in the memory 1100. In this case, the data generation device 1000 may include various computing devices such as servers, PCs (personal computers), laptops, workstations, tablets, mobile computers, PDAs / EDAs, mobile phones, smartphones, and IoT devices.

[0046] Specifically, the data generation device 1000 may typically achieve desired system performance using a combination of computing equipment (e.g., equipment that may include computer processors, memory, storage, input and output devices, and other conventional computing equipment components; electronic communication equipment such as routers and switches; and electronic information storage systems such as network-attached storage (NAS) and storage area networks (SANs)) and computer software (i.e., instructions for using the computing equipment in a particular manner).

[0047] Furthermore, the processor 1200 of the data generation device 1000 may include hardware components such as an MPU (Micro Processing Unit) or CPU (Central Processing Unit), cache memory, and a data bus. The computing device may also further include an operating system and software components for applications that perform specific purposes.

[0048] However, this does not exclude the case in which the data generation device 1000 includes an integrated processor in which a medium, processor, and memory for carrying out the present invention are integrated.

[0049] On the other hand, when generating training data for training an agent model based on a vision language model through the data generation device 1000, the system may include a user terminal 100 that transmits the defect inspection criteria for the product and the corresponding modeling image to the data generation device 1000, and a training data generation module 200 that generates training data by referring to the defect inspection criteria and the modeling image. However, although Figure 1 shows the training data generation module 200 as a separate configuration separate from the data generation device 1000, the training data generation module 200 may also be included in the data generation device 1000. Note that the modeling image may be generated by the user in 2D or 3D form using a CAD-based modeling tool such as SolidWorks, Inventor, CATIA, NX, or AutoCAD, but is not limited to this.

[0050] Furthermore, the training data generation module 200 may include a defect generator 210, a document generator 220, and a merger 230, the configurations of which are as follows.

[0051] First, when the defect generator 210 obtains the defect inspection criteria and modeling image from the data generation device 1000, it refers to the defect inspection criteria to generate a defect mask containing the type and location information of defects applied to the modeling image, generates rendering information containing predetermined tilt and rotation information applied to the modeling image, and generates an image to be identified by referring to the defect parameters, which include the defect inspection criteria, defect mask, and rendering information. At this time, the defect mask may be a product segmentation mask containing information about defects that can be confirmed in products produced in an actual manufacturing site, but a specific example will be described later in Figure 3.

[0052] Specifically, the defect generator 210 applies rendering information to the modeling image to generate a rendering image in which the modeling image is transformed into a viewing point that facilitates defect detection of the actual product (i.e., the same or similar viewing point as the direction in which a camera photographs the actual product to inspect it in the manufacturing plant). The defect generator 210 then inputs the rendering image and the defect mask into the image synthesis model 211 included in the defect generator 210, and through the image synthesis model 211, it can generate a discriminant image in which the defect mask is applied to the rendering image. In this case, a diffusion model may be used as the image synthesis model 211, and the process of generating the discriminant image may correspond to the process of securing defect images necessary for training an agent model based on a vision language model.

[0053] Next, when the document generator 220 obtains the defect parameters generated by the defect generator 210, it can refer to the defect parameters and generate a description of the image to be classified, which is the description corresponding to the image to be classified. Generally, when comparing vision language models and large-scale language models, large-scale language models are advantageous in terms of cost and performance. Therefore, the goal is to obtain various data included in the process of determining whether or not a product is defective using a large-scale language model (i.e., data necessary for training the agent model) and use these as part of the training data for the agent model. In other words, since large-scale language models cannot process image data, and the defect generator 210 generates the image to be classified by referring to the defect parameters, the goal is to use the document generator 220 based on the large-scale language model to generate a text template to describe the image to be classified, extract important keywords from the defect parameters, and then apply the important keywords to the text template to generate a description of the image to be classified, and to use the description of the image to be classified in place of the image to be classified in the document generator 220.

[0054] For example, an example of a discriminant image description generated by the document generator 220 is as follows. The document generator 220 may extract information such as which area of ​​the product the defect is located in, the size, shape, and color of the defect as important keywords, and then apply these to a template.

[0055] "The top surface of the screw head is barely visible. This is a defect where the screw threads are stripped. The defect is located at (210,100) to (250,150) in the image, with an actual length of 2 cm and a width of 0.23 cm. The defect is formed in a direction from the lower left to the upper right of the image."

[0056] Furthermore, the document generator 220 includes an agent large-scale language model 221 and a tool simulation large-scale language model 222. After the image description to be classified is generated, the interaction between the agent large-scale language model 221 and the tool simulation large-scale language model 222 executes a process to generate the data necessary for the agent model based on the vision language model to learn. This enables the generation of chat log data including defect inspection criteria, the image description to be classified, and data generated by the interaction between the agent large-scale language model 221 and the tool simulation large-scale language model 222. The interaction between the agent large-scale language model 221 and the tool simulation large-scale language model 222 will be described in detail in Figure 2.

[0057] Finally, when the merger 230 receives the image to be classified generated by the defect generator 210 and the chat log data generated by the document generator 220 as input, it can generate training data that can be used to train an agent model based on a vision language model by merging the image to be classified with the chat log data. Specifically, while the training of an agent model based on a vision language model can use both text data and image data, the chat log data only contains text data for the image to be classified. Therefore, by merging the image to be classified with the chat log data, the merger attempts to generate training data that includes both image data and text data.

[0058] The method for generating training data for training an agent model through a system including the data generation device 1000 configured in this way will be described in detail below with reference to Figures 1 and 2.

[0059] Figure 2 is a schematic diagram showing a flowchart for generating training data for training an agent model based on a vision language model according to one embodiment of the present invention.

[0060] Referring to Figure 2, first, the defect inspection criteria for the product and the modeling image corresponding to the product can be obtained (S210). As described above in Figure 1, this data can be obtained by the data generation device 1000 from the user terminal 100. As an example, the present invention describes a "screw" product, but it can be applied not only to PCB substrates and semiconductor wafers, but to any product manufactured on the factory floor. And, as explained in Figure 1, the modeling image may be a CAD image generated using a 2D or 3D modeling tool, but is not limited to this.

[0061] <Example of criteria for determining defects>

[0062] *In some cases, the threads at the end of a screw may be damaged, but if the width and length of the damage exceed 5 mm, it will be considered defective.

[0063] *The screw head allows for defects of up to 2cm x 2cm. However, if the area is 4cm... 2 If the dimensions are greater than or equal to the above, or if the length of one side exceeds 2 cm, it will be considered defective.

[0064] *In other parts, a screw thread with a length of 1 mm or more that is damaged will be considered defective.

[0065] *Items other than those listed above will be judged as good quality.

[0066] Next, the data generation device 1000 can use the defect generator 210 to generate a defect mask and rendering information and obtain defect parameters (S220).

[0067] As an example, the data generation device 1000, using a defect generator 210, can generate a defect mask containing information on the type and location of defects applied to a modeling image by referring to a defect inspection criterion, and generate rendering information containing information on predetermined tilt and rotation applied to the modeling image, thereby obtaining defect parameters including the defect inspection criterion, defect mask, and rendering information. In this case, the process of obtaining the defect mask and rendering information included in the defect parameters will be explained in more detail with reference to Figure 3.

[0068] Figure 3 is a schematic diagram showing examples of a modeling image, a defect mask, and an image to be discriminated according to one embodiment of the present invention.

[0069] Referring to Figure 3(a), an example of a modeling image generated by a CAD-based modeling tool is shown, and it can be confirmed that a modeling image 300 of a screw, including a screw head 310, a screw body 320, and screw threads 330, is shown. In this case, the defect generator 210 may generate rendering information corresponding to a predetermined tilt angle and rotation angle that is the same as the viewpoint from which the actual agent model views the product, and apply the rendering information to the modeling image 300 to generate a rendering image. However, the rendering information may also include, but is not limited to, information to distinguish the product by part, such as the screw head 310, the center of the screw body 321, the end of the screw body 322, and the screw threads 330, using color, etc.

[0070] Furthermore, referring to Figure 3(b), the defect generator 210 can refer to the defect inspection criteria and generate a defect mask 400 that includes the defect 410 to be applied to the modeling image. For example, the defect generator 210 may refer to the defect inspection criteria that "a screw thread is considered defective if the length of the damaged thread is 1 mm or more" and generate a defect 410 in which the screw thread is damaged for 1 mm or more.

[0071] Referring again to Figure 2, the data generation device 1000 can execute a subprocess that uses the defect generator 210 to generate an image to be identified by referring to the defect parameters (S231); and uses the defect generator 210 to input the defect parameters to the document generator 220, and uses the document generator 220 to generate an image to be identified description, a tool usage plan for detecting defects in the image to be identified description, and tool usage request information (S232_1); generates a tool usage simulation result that simulates the tool usage result in the image to be identified description based on the tool usage request information, and information related to the determination of whether or not there are defects in the image to be identified (S232_2); and generates chat log data including defect inspection criteria, image to be identified description, tool usage plan information, tool usage request information, tool usage simulation result information, and information related to the determination of whether or not there are defects in the image to be identified (S232_3). In this case, processes S231 may be executed first, followed by processes S232_1 through S232_3, but processes S231 and S232_1 may be executed simultaneously, and processes S232_1 through S232_3 may be executed before process S231. However, for the sake of explanation, process S231 will be explained first, followed by processes S232_1 through S232_3.

[0072] First, in process S231, the data generation device 1000 can use the defect generator 210 to refer to defect parameters and generate a discriminant image based on the modeling image. Referring again to Figure 3, the data generation device 1000 can use the defect generator 210 to input the modeling image 300 corresponding to Figure 3(a) and the defect mask 400 corresponding to Figure 3(b) into the image synthesis model 211 based on the diffusion model included in the defect generator 210, and through the image synthesis model 211 corresponding to Figure 3(c), generate a discriminant image 500 that corresponds to a defective image for a product to which the defect mask 400 has been applied to the modeling image 300.

[0073] Next, referring again to Figure 2, in process S232_1, the data generation device 1000 can input defect parameters to the document generator 220, and the document generator 220 can refer to the defect parameters to generate a discriminant image description, refer to the defect inspection criteria and the discriminant image description to generate tool usage plan information, and refer to the tool usage plan information to generate tool usage request information for detecting defects. At this time, the document generator 220 may include an agent large-scale language model 221 and a tool simulation large-scale language model 222, as explained in Figure 1, and of these, the agent large-scale language model 221 may generate the tool usage plan information and the tool usage request information. At this time, the method by which the agent large-scale language model 221 generates the tool usage plan information and the tool usage request information can be broadly divided into two categories. However, the process for generating the discriminant image description is as explained in the configuration of the document generator 220 in Figure 1, so redundant explanations will be omitted.

[0074] As an example, the data generation device 1000 may use the agent large-scale language model 221 to generate first tool usage plan information to nth tool usage plan information by referring to the defect inspection criteria and the image description to be judged, and generate pth tool usage request information for detecting defects using the pth tool based on the pth tool usage plan information (where p is an integer that increases sequentially from 1 to n) corresponding to the first to nth tool usage plan information. In other words, this may correspond to an embodiment in which interaction takes place between the agent large-scale language model 221 and the tool simulation large-scale language model 222 after the agent large-scale language model 221 has generated first tool usage plan information to nth tool usage plan information and the corresponding first tool usage request information to nth tool usage request information. This is because, once defect inspection criteria are determined at the manufacturing site, a series of processes for detecting defects are usually determined and then executed sequentially to detect defects. Accordingly, it is thought that interaction between the agent large-scale language model 221 and the tool simulation large-scale language model 222 takes place after the overall tool usage plan information has been generated in advance.

[0075] In this case, the agent large-scale language model 221 has information about the functions of the tools (magnification, length measurement, size measurement, angle measurement, triangle area measurement, quadrilateral area measurement, etc.) that the agent model based on the vision language model can use to detect defects. For example, if p is 2, the agent large-scale language model 221 may, but is not limited to, refer to the defect inspection criteria and the image description to be classified and determine the first tool usage plan information to be "magnification" and the second tool usage plan information to be "length measurement". Furthermore, according to the above example, the agent large-scale language model 221 can generate the first tool usage request information by referring to the first tool usage plan information, "magnification," and specifying the area to be magnified in the image description to be classified using coordinates or other methods. It can also generate the second tool usage request information by referring to the second tool usage plan information, "length measurement," and specifying the section to be measured in the image description to be classified (or the simulated result data based on the first tool usage request information) using coordinates or other methods.

[0076] As another example, the data generation device 1000 may use the agent large-scale language model 221 to generate first tool usage plan information when p is 1, by referring to the defect inspection criteria and the image description to be determined, and generate first tool usage request information for detecting the defect using the first tool based on the first tool usage plan information. That is, this may correspond to an embodiment in which, with the agent large-scale language model 221 having generated only the first tool usage plan information and the first tool usage request information, the process of generating the nth tool usage plan information in the same way as the process of generating the second tool usage plan information based on the results of the interaction between the agent large-scale language model 221 and the tool simulation large-scale language model 222 based on the first tool usage plan information is repeated until the nth tool usage plan information is generated (i.e., until p becomes n). This may slightly increase the computational load on the agent large-scale language model 221, as it generates the next tool usage plan information and tool usage request information by reflecting the output results of the tool simulation large-scale language model 222. However, this is thought to help the agent large-scale language model 221 generate the tool usage plan information and tool usage request information more accurately.

[0077] As an example, if p is 2, the agent large-scale language model 221, after referring to the defect inspection criteria and the image description to be judged, determines the first tool usage plan information to be "enlargement," and then, referring to the first tool usage plan information, "enlargement," it can generate the first tool usage request information by explicitly indicating the area to be enlarged in the image description to be judged using coordinates or other methods, and then input the first tool usage request information to the tool simulation large-scale language model 222. Subsequently, the agent large-scale language model 221 receives the output data generated by the tool simulation large-scale language model 222, and, referring to the defect inspection criteria and the output data generated by the tool simulation large-scale language model 222, determines the second tool usage plan information suitable for detecting defects to be "circle size measurement," and then, referring to "circle size measurement," it can generate the second tool usage request information by explicitly indicating the position to be measured in the image description to be judged (or the result data simulated based on the first tool usage request information) using coordinates or other methods. However, it is not limited to this.

[0078] Furthermore, in process S232_2, the data generation device 1000 generates tool usage simulation result information by simulating the results of tool usage based on tool usage request information in the image description to be determined using the document generator 220, and generates information related to the determination of whether or not there are defects in the image to be determined by referring to the defect inspection criteria and the tool usage simulation result information. In this case, the tool usage simulation result information may include at least a portion of the measured values ​​obtained by using the tool, the overview image description, and the image description in which the use of the tool has been simulated, and the measured values ​​obtained by using the tool may include measured values ​​corresponding to defects.

[0079] As an example, the data generation device 1000 can, using the agent large-scale language model 221, generate first tool usage request information to the nth tool usage request information according to the first embodiment in process S232_1, and then input the discriminant image description and the pth tool usage request information (where p is an integer that increases sequentially from 1 to n) into the tool simulation large-scale language model 222. The data generation device 1000 can then, using the tool simulation large-scale language model 222, generate the pth tool usage simulation result information, which is the result of simulating the use of the pth tool based on the pth tool usage request information in the discriminant image description. Furthermore, the data generation device 1000 can, using the tool simulation large-scale language model 222, input the pth tool usage simulation result information into the agent large-scale language model 221, and the agent large-scale language model 221 can, by referring to the defect inspection criteria and the pth tool usage simulation result information, generate defect detection related information for the discriminant image.

[0080] As an example, if p is 2, the data generation device 1000 uses the tool simulation large-scale language model 222 to reference the coordinate values ​​corresponding to the area to be enlarged based on the first tool usage request information, and simulates the result of enlarging that area in the discriminant image description. This generates first tool usage simulation result information such as "The defect appears large in the enlarged image. The defect is a straight line and is black. The endpoints of the line are (210,100) and (250,150)," and the first tool usage simulation result information can be input to the agent large-scale language model 221.

[0081] Furthermore, in this embodiment, since the second tool usage simulation result information has been generated previously, the data generation device 1000 can, after confirming the first tool usage simulation result information with the agent large-scale language model 221, input the second tool usage request information, including coordinate values ​​for measuring the length of the defect, into the tool simulation large-scale language model 222 if there are no abnormalities in the second tool usage request information. The tool simulation large-scale language model 222 then uses the second tool usage request information to reference the coordinate values ​​for which the length is to be measured and simulates the result of measuring the length in the discriminant image description (or the first tool usage simulation result information), thereby generating second tool usage simulation result information such as "The length of the straight line is 6.4 mm," and inputting the second tool usage simulation result information into the agent large-scale language model 221.

[0082] The data generation device 1000 then uses the agent large-scale language model 221 to refer to the defect inspection criteria, the first tool usage simulation result information, and the second tool usage simulation result information to confirm that a straight line (i.e., a defect) is formed 6.4 mm above the screw thread. This allows the device to determine whether the image to be judged is defective or not based on the defect inspection criteria, which state that "if the length of the damaged screw thread is 1 mm or more, it is considered defective." Finally, it can generate defect determination-related information for the image to be judged that contains content corresponding to a defect.

[0083] On the other hand, when the data generation device 1000 obtains specific tool usage simulation result information corresponding to a specific tool usage plan, which is at least a part of the first to the nth tool usage plan, from the tool simulation large-scale language model 222, it can use the tool simulation large-scale language model 222 to input the specific tool usage simulation result information into the agent large-scale language model 221, and the agent large-scale language model 221 can use the specific tool usage simulation result information to generate specific tool usage plan update information that updates the specific tool usage plan.

[0084] For example, when n is 3 or greater, if the data generation device 1000 obtains third tool usage simulation result information corresponding to a specific tool usage simulation result information from the first to the nth tool usage plan among the first to nth tool usage plans from the tool simulation large-scale language model 222, it can use the tool simulation large-scale language model 222 to input the third tool usage simulation result information to the agent large-scale language model 221. In this case, the agent large-scale language model 221 may refer to the third tool usage simulation result information and determine that it is difficult to detect defects based on the third tool usage plan information, and that it is necessary to update the third tool usage plan information.

[0085] The data generation device 1000 can update the third tool usage plan information using the agent large-scale language model 221 to generate third tool usage plan update information, which is specific tool usage plan update information. Based on the third tool usage plan update information, it can generate third tool usage request update information, which is specific tool usage request update information for detecting defects using a specific tool, the third tool, and input the third tool usage request update information into the tool simulation large-scale language model 222. At this time, the agent large-scale language model 221 can update the third tool usage plan information by, but is not limited to, using other tools that can measure size or length from the domain expansion tool, or by changing the target coordinate values.

[0086] Subsequently, the data generation device 1000 may use the tool simulation large-scale language model 222 to generate third tool usage simulation result update information, which is specific tool usage simulation result update information that simulates the use of the third tool, based on the third tool usage request update information in the image description to be discriminated (or third tool usage simulation result information). The data generation device 1000 may then input the third tool usage simulation result update information to the agent large-scale language model 221, which may then use the agent large-scale language model 221 to generate defect detection related update information for the image to be discriminated, by referring to the defect inspection criteria and the third tool usage simulation result update information.

[0087] As another example, according to the second embodiment in process S232_1, the agent large-scale language model 221 generates first tool usage plan information (e.g., magnification), and based on the first tool usage plan information, generates first tool usage request information (e.g., explicitly indicating the area to be magnified in the image description to be discriminated using coordinates, etc.). Then, the data generation device 1000 can input the image description to be discriminated and the first tool usage request information into the tool simulation large-scale language model 222 using the agent large-scale language model 221 when p is 1. The data generation device 1000 can then use the tool simulation large-scale language model 222 to generate first tool usage simulation result information, which simulates the use of the first tool based on the first tool usage request information in the image description to be discriminated, and input the first tool usage simulation result information into the agent large-scale language model 221. In this case, if the agent large-scale language model 221 can determine whether or not there is a defect in the target image by referring to the defect inspection criteria and the first tool simulation results, the data generation device 1000 can also generate information related to the determination of whether or not there is a defect in the target image by referring to the defect inspection criteria and the first tool simulation results, without generating second tool usage plan information using the agent large-scale language model 221.

[0088] Subsequently, the data generation device 1000, using the agent large-scale language model 221, can generate p-th tool usage plan information by referring to the defect inspection criteria and the (p-1) tool usage simulation result information when p is an integer increasing from 2 to n (where n is an integer greater than or equal to 2), generate p-th tool usage request information for detecting defects using the p-th tool based on the p-th tool usage plan information, input the p-th tool usage request information into the tool simulation large-scale language model 222, and use the tool simulation large-scale language model 222 to generate p-th tool usage simulation result information, which simulates the use of the p-th tool in the image description to be discriminated, based on the p-th tool usage request information, and input the p-th tool usage simulation result information into the agent large-scale language model 221. Then, the data generation device 1000, using the agent large-scale language model 221, can generate defect detection information for the image to be discriminated by referring to the defect inspection criteria, the first tool usage simulation result information or the nth tool usage simulation result information, while repeating this process until p becomes n.

[0089] As an example, if p is 2, the data generation device 1000 can use the agent large-scale language model 221 to refer to the first tool usage simulation result information (for example, the result of simulating the magnification of a predetermined area of ​​the image to be discriminated) obtained from the defect inspection criteria and the tool simulation large-scale language model 222 to confirm that a defect has been detected in the magnified area, and generate second tool usage plan information for measuring the size or length of the confirmed defect. Furthermore, the data generation device 1000 can use the agent large-scale language model 221 to measure the size or length of the defect based on the second tool usage plan information, thereby specifying concrete coordinate values ​​for detecting the defect, and generate second tool usage request information, which can then be input into the tool simulation large-scale language model 222.

[0090] Subsequently, the data generation device 1000 uses the agent large-scale language model 221 to simulate the measurement of the size or length of the defect using coordinate values ​​based on the second tool usage request information in the image description to be discriminated (or the simulation result using the first tool), thereby generating second tool usage simulation result information such as "the length of the straight line is 6.4 mm," and inputting the second tool usage simulation result information into the agent large-scale language model 221. Then, the data generation device 1000 uses the agent large-scale language model 221 to refer to the defect inspection criteria, the first tool usage simulation result information, and the second tool usage simulation result information to confirm that the straight line (i.e., the defect) is formed 6.4 mm above the screw thread. This allows the data generation device 1000 to determine whether the image to be discriminated is defective or not based on the defect inspection criteria, which state that "if the length of the damaged screw thread is 1 mm or more, it is considered defective." Finally, it can generate defect determination-related information for images to be discriminated that contain content corresponding to defects.

[0091] Furthermore, the data generation device 1000 can use the document generator 220 to generate chat log data (S232_3) using data used or generated during the execution of processes S232_1 and S232_2. Specifically, it can generate chat log data that includes defect inspection criteria, a description of the image to be judged, tool usage plan information, tool usage request information, tool usage simulation result information, and information on whether or not the image to be judged is defective.

[0092] For example, the data generation device 1000 can use a document generator 220 to generate chat log data that includes defect inspection criteria, a description of the image to be judged, p-th tool usage plan information which is a first tool usage plan information to the nth tool usage plan information, p-th tool usage request information generated based on the p-th tool usage plan information, p-th tool usage simulation result information, and information related to the determination of whether or not there is a defect in the image to be judged.

[0093] On the other hand, the data generation device 1000 can also generate chat log data using the document generator 220 when a specific tool usage plan is updated by specific tool usage simulation result information corresponding to a specific tool usage plan, which is at least a part of the first tool usage plan information to the nth tool usage plan information, by including defect inspection criteria, a description of the image to be judged, updated information on the specific tool usage plan, updated information on the specific tool usage request, updated information on the specific tool usage simulation result, and updated information related to the determination of whether or not there is a defect in the image to be judged.

[0094] Furthermore, once the execution of processes S231 and S232_1 to S232_2 is completed, the data generation device 1000 can input the image to be discriminated and the chat log data into the merger 230, and use the merger 230 to generate training data (S240) by merging the image to be discriminated with the chat log data.

[0095] Specifically, the data generation device 1000 can generate training data by using the merger 230 to change the image description to be classified in the chat log data to the image to be classified. In other words, in processes S232_1 to S232_3, in order to generate chat log data, which is text data for training an agent model based on a vision language model, while minimizing the use of computing resources, an image description to be classified is generated to replace the image to be classified, and the chat log data is generated by including data obtained through interaction between two large-scale language models. However, since image data can also be used in training the agent model, the training data is generated by including the image to be classified generated by the defect generator 210 instead of the image description to be classified.

[0096] In addition, the data generation device 1000 can also generate training data by using the merger 230 to refer to the tool usage simulation result information contained in the chat log data, generating tool usage result images corresponding to the tool usage simulation result information, and converting the tool usage simulation result information into tool usage result images. However, the data generation device 1000 can also generate training data by directly generating tool usage result images corresponding to the tool usage simulation result information, inputting the tool usage result images into the merger 230, and using the merger 230 to convert the tool usage simulation result information into tool usage result images, but is not limited to this.

[0097] Thus, once training data for training the agent model based on the vision-language model is generated, the data generation device 1000 can fine-tune the agent model based on the vision-language model through learning using the training data.

[0098] Furthermore, once the agent model is fine-tuned, the user can input text prompts corresponding to the defect inspection criteria and test images to be identified (which can also include example images of the product) obtained at the manufacturing site through the user terminal. As a result of the agent model interacting with various tools for defect detection, it can determine whether the product corresponding to the image is defective or not.

[0099] Furthermore, the embodiments of the present invention described above can be implemented in the form of program instructions executable through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., individually or in combination. The program instructions recorded on the computer-readable recording medium may be specially designed and configured for the present invention, or they may be publicly known and available to the average technician in the field of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions, such as ROMs, RAMs, and flash memory. Examples of program instructions include not only machine code, such as that produced by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like. The hardware device may be configured to operate as one or more software modules to perform the processing according to the present invention, and vice versa.

[0100] Although the present invention has been described above with specific details such as concrete components, and with limited embodiments and drawings, these are provided only to aid in a more general understanding of the invention, and the invention is not limited to the above embodiments. A person with ordinary skill in the art to which the invention pertains can make various modifications and variations from this description.

[0101] Therefore, the concept of the present invention should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the claims described below shall also fall within the scope of the concept of the present invention.

Claims

1. In a method for generating training data for training an agent model based on a vision language model, (a) When a defect inspection criterion for a product and a model image corresponding to the product are acquired, the data generation device inputs the defect inspection criterion and the model image to a defect generator, and uses the defect generator to generate a defect mask (the defect mask includes information on the type and location of defects applied to the model image) by referring to the defect inspection criterion, and generates rendering information including predetermined tilt and rotation information applied to the model image, and acquires defect parameters including the defect inspection criterion, the defect mask and the rendering information, (b) The data generation device performs the following steps: (i) a subprocess in which the defect generator generates a target image for discrimination based on the modeling image by referring to the defect parameters using the defect generator; (ii) inputting the defect parameters into a document generator based on a large-scale language model, and using the document generator generating a target image description by referring to the defect parameters, generating tool usage plan information by referring to the defect inspection criteria and the target image description, generating tool usage request information for detecting the defect by referring to the tool usage plan information, generating tool usage simulation result information by simulating the tool usage result based on the tool usage request information in the target image description for discrimination, generating information related to the determination of whether or not there is a defect in the target image by referring to the defect inspection criteria and the tool usage simulation result information, and generating chat log data including the defect inspection criteria, the target image description for discrimination, the tool usage plan information, the tool usage request information, the tool usage simulation result information, and the information related to the determination of whether or not there is a defect in the target image; (c) The data generation device inputs the image to be discriminated and the chat log data into a Merger, and uses the Merger to generate training data by merging the image to be discriminated with the chat log data, Includes, In step (b) above, A method to generate a description of an image to be identified by having the data generation device generate a text template for describing the image to be identified using the document generator, extracting important keywords from the defect parameters, and then applying the important keywords to the text template.

2. The aforementioned document generator includes an agent-based large-scale language model and a tool-based simulation large-scale language model. In step (b) above, The data generation device generates first tool usage plan information to nth tool usage plan information (where n is an integer of 1 or more) by referring to the defect inspection criteria and the image description to be determined using the agent large-scale language model, generates pth tool usage request information for detecting the defect using the pth tool based on the pth tool usage plan information (where p is an integer that increases sequentially from 1 to n) corresponding to the first to nth tool usage plan information, and inputs the image description to be determined and the pth tool usage request information into the tool simulation large-scale language model. The large-scale language model for tool simulation is used to generate p-tool usage simulation result information, which is the result of simulating the use of the p-tool based on the p-tool usage request information in the discriminant image description, and the p-tool usage simulation result information is input to the agent large-scale language model, and the agent large-scale language model is used to generate information related to the determination of the presence or absence of defects in the discriminant image by referring to the defect inspection criteria and the p-tool usage simulation result information, and the defect inspection criteria, the discriminant image description, the p-tool The method according to claim 1, wherein the chat log data is generated including p-tool usage plan information, p-tool usage request information, p-tool usage simulation result information, and information related to the determination of whether or not there is a defect in the target image.

3. In step (b) above, When the data generation device obtains specific tool usage simulation result information corresponding to a specific tool usage plan, which is at least a part of the first tool usage plan to the n tool usage plans, from the tool simulation large-scale language model, it inputs the specific tool usage simulation result information into the agent large-scale language model, and the agent large-scale language model, by referring to the specific tool usage simulation result information, generates specific tool usage plan update information that updates the specific tool usage plan, generates specific tool usage request update information for detecting the defect using a specific tool based on the specific tool usage plan update information, and inputs the specific tool usage request update information into the tool simulation large-scale language model, and the tool simulation The method according to claim 2, wherein a large-scale language model is used to generate specific tool usage simulation result update information, which simulates the use of the specific tool in the image description to be determined, based on the specific tool usage request update information; the specific tool usage simulation result update information is input to the large-scale language model; the large-scale language model is used to generate update information related to the determination of whether or not there is a defect in the image to be determined, by referring to the defect inspection criteria and the specific tool usage simulation result update information; and the chat log data is generated, which includes the defect inspection criteria, the image description to be determined, the specific tool usage plan update information, the specific tool usage request update information, the specific tool usage simulation result update information, and the update information related to the determination of whether or not there is a defect in the image to be determined.

4. The aforementioned document generator includes an agent-based large-scale language model and a tool-based simulation large-scale language model. In step (b) above, A subprocess in which the data generation device (i) uses the agent large-scale language model to generate first tool usage plan information when p is 1, by referring to the defect inspection criteria and the image description to be determined; generates first tool usage request information for detecting the defect using the first tool based on the first tool usage plan information; inputs the image description to be determined and the first tool usage request information into the tool simulation large-scale language model; uses the tool simulation large-scale language model to generate first tool usage simulation result information, which simulates the use of the first tool based on the first tool usage request information in the image description to be determined; and inputs the first tool usage simulation result information into the agent large-scale language model. ; and (ii) using the agent large-scale language model, when p is an integer increasing from 2 to n (where n is an integer greater than or equal to 2), generate p-th tool usage plan information by referring to the defect inspection criteria and the (p-1) tool usage simulation result information; generate p-th tool usage request information for detecting the defect using the p-th tool based on the p-th tool usage plan information; input the p-th tool usage request information into the tool simulation large-scale language model; use the tool simulation large-scale language model to generate p-th tool usage simulation result information simulating the use of the p-th tool in the discriminant image description based on the p-th tool usage request information; and input the p-th tool usage simulation result information into the agent large-scale language model. The method according to claim 1, wherein the subprocess is repeated until p becomes n, and the agent large-scale language model is used to generate information related to the determination of whether or not a defect exists for the target image, by referring to the defect inspection criteria and the first tool usage simulation result information or the nth tool usage simulation result information, and chat log data is generated which includes the defect inspection criteria, the target image description, the first tool usage plan information or the nth tool usage plan information, the first tool usage request information or the nth tool usage request information, the first tool usage simulation result information or the nth tool usage simulation result information, and the information related to the determination of whether or not a defect exists for the target image.

5. In step (c) above, The method according to claim 1, wherein the data generation device generates the training data by using the merger to change the discriminant image description in the chat log data to the discriminant image.

6. In step (c) above, The method according to claim 5, wherein the data generation device uses the merger to reference the tool usage simulation result information contained in the chat log data to generate a tool usage result image corresponding to the tool usage simulation result information, and generates the training data by changing the tool usage simulation result information to the tool usage result image.

7. In step (a) above, The method according to claim 1, wherein the data generation device uses the defect generator to apply the rendering information to the modeling image to generate a rendering image, and inputs the rendering image and the defect mask into an image synthesis model (the image synthesis model is included in the defect generator) to generate the image to be discriminated through the image synthesis model.

8. In step (b) above, The method according to claim 1, wherein the tool usage simulation result information includes at least a portion of the measured values ​​obtained by using the tool, the overview image description, and the image description in which the use of the tool was simulated.

9. (d) A step in which the data generation device fine-tunes an agent model based on a vision language model through learning using the training data; The method according to claim 1, further comprising:

10. In a data generation device that generates training data for training an agent model based on a vision language model, At least one memory to store instructions, The system includes at least one processor configured to execute the aforementioned instructions, and in doing so, (i) A process in which the processor, once a defect inspection criterion for a product and a modeling image corresponding to the product are acquired, inputs the defect inspection criterion and the modeling image into a defect generator, and uses the defect generator to generate a defect mask (the defect mask includes information on the type and location of defects applied to the modeling image) by referring to the defect inspection criterion, generates rendering information including predetermined tilt and rotation information applied to the modeling image, and acquires defect parameters including the defect inspection criterion, the defect mask, and the rendering information; (ii) A subprocess in which the defect generator generates a discriminant image based on the modeling image by referring to the defect parameters, and (ii) inputs the defect parameters into a document generator based on a large-scale language model, and uses the document generator A process that uses a language to generate a description of an image to be identified by referring to the defect parameters; generates tool usage plan information by referring to the defect inspection criteria and the description of the image to be identified; generates tool usage request information for detecting the defect by referring to the tool usage plan information; generates tool usage simulation result information by simulating the tool usage result based on the tool usage request information in the description of the image to be identified; generates information related to the determination of whether or not there is a defect in the image to be identified by referring to the defect inspection criteria and the tool usage simulation result information; and generates chat log data including the defect inspection criteria, the description of the image to be identified, the tool usage plan information, the tool usage request information, the tool usage simulation result information, and the information related to the determination of whether or not there is a defect in the image to be identified;(III) A process of inputting the image to be discriminated and the chat log data into a Merger, and using the Merger to generate training data by merging the image to be discriminated with the chat log data; The aforementioned processor, In the above process (II), A data generation device that generates a text template for describing the image to be identified using the document generator, extracts important keywords from the defect parameters, and then applies the important keywords to the text template to generate the image description to be identified.

11. The aforementioned document generator includes an agent-based large-scale language model and a tool-based simulation large-scale language model. The aforementioned processor, In the above process (II), The agent's large-scale language model is used to generate first tool usage plan information to nth tool usage plan information (where n is an integer of 1 or more) by referring to the defect inspection criteria and the image description to be determined, and based on the pth tool usage plan information (where p is an integer that increases sequentially from 1 to n) corresponding to the first tool usage plan information to nth tool usage plan information, the pth tool usage request information for detecting the defect using the pth tool is generated, and the image description to be determined and the pth tool usage request information are input to the tool simulation large-scale language model, and the tool simulation large-scale language model is used to input the pth tool usage request information in the image description to be determined The data generation device according to claim 10, which generates p-tool usage simulation result information, which is the result of simulating the use of the p-tool based on the report; inputs the p-tool usage simulation result information into the agent large-scale language model; uses the agent large-scale language model to generate information related to the determination of whether or not there is a defect in the target image, by referring to the defect inspection criteria and the p-tool usage simulation result information; and generates chat log data including the defect inspection criteria, the target image description, the p-tool usage plan information, the p-tool usage request information, the p-tool usage simulation result information, and the information related to the determination of whether or not there is a defect in the target image.

12. The aforementioned processor, In the above process (II), When specific tool usage simulation result information corresponding to a specific tool usage plan, which is at least a part of the first tool usage plan to the nth tool usage plan, is obtained from the tool simulation large-scale language model, the tool simulation large-scale language model is used to input the specific tool usage simulation result information into the agent large-scale language model, the agent large-scale language model is used to refer to the specific tool usage simulation result information to generate specific tool usage plan update information that updates the specific tool usage plan, the agent large-scale language model is used to refer to the specific tool usage simulation result information to generate specific tool usage plan update information for detecting the defect using a specific tool, and the specific tool usage request update information is used to input the tool simulation large-scale language model, the tool simulation large-scale language model A data generation device according to claim 11, comprising: a language model, which generates specific tool usage simulation result update information that simulates the use of a specific tool based on the specific tool usage request update information in the image description to be determined; inputting the specific tool usage simulation result update information into the agent large language model; using the agent large language model, which generates update information related to the determination of whether or not there is a defect in the image to be determined, by referring to the defect inspection criteria and the specific tool usage simulation result update information; and generating chat log data including the defect inspection criteria, the image description to be determined, the specific tool usage plan update information, the specific tool usage request update information, the specific tool usage simulation result update information, and the update information related to the determination of whether or not there is a defect in the image to be determined.

13. The aforementioned document generator includes an agent-based large-scale language model and a tool-based simulation large-scale language model. The aforementioned processor, In the above process (II), (i) A subprocess that, using the agent large language model, generates first tool usage plan information by referring to the defect inspection criteria and the image description to be determined when p is 1, generates first tool usage request information for detecting the defect with the first tool based on the first tool usage plan information, inputs the image description to be determined and the first tool usage request information into the tool simulation large language model, uses the tool simulation large language model to generate first tool usage simulation result information by simulating the use of the first tool based on the first tool usage request information in the image description to be determined, and inputs the first tool usage simulation result information into the agent large language model; and (ii) A subprocess that, using the agent large language model, generates first tool usage simulation result information by referring to the defect inspection criteria and the image description to be determined when p is 1 (where n is an integer of 2 or more) A subprocess that, when p is an increasing integer, generates p-th tool usage plan information by referring to the defect inspection criteria and the (p-1) tool usage simulation result information, generates p-th tool usage request information for detecting the defect by the p-th tool based on the p-th tool usage plan information, inputs the p-th tool usage request information into the tool simulation large language model, uses the tool simulation large language model to generate p-th tool usage simulation result information simulating the use of the p-th tool based on the p-th tool usage request information in the discriminant image description, and inputs the p-th tool usage simulation result information into the agent large language model; repeats this until p becomes n, and uses the agent large language model to generate p-th tool usage simulation result information by referring to the defect inspection criteria and the first tool usage simulation result information or the n-th tool usage The data generation device according to claim 10, which generates information related to the determination of whether or not a defect exists for the target image by referring to simulation result information, and generates chat log data including the defect inspection criteria, the description of the target image, the first tool usage plan information or the nth tool usage plan information, the first tool usage request information or the nth tool usage request information, the first tool usage simulation result information or the nth tool usage simulation result information, and the information related to the determination of whether or not a defect exists for the target image.

14. The aforementioned processor, In the above (III) process, The data generation device according to claim 10, wherein the merger is used to change the image description to be identified in the chat log data to the image to be identified, thereby generating the training data.

15. The aforementioned processor, In the above (III) process, The data generation device according to claim 14, wherein the merger is used to generate a tool usage result image corresponding to the tool usage simulation result information by referring to the tool usage simulation result information contained in the chat log data, and the learning data is generated by changing the tool usage simulation result information to the tool usage result image.

16. The aforementioned processor, In the above process (I), The data generation device according to claim 10, wherein the defect generator is used to apply the rendering information to the modeling image to generate a rendered image, and the rendered image and the defect mask are input to an image synthesis model (the image synthesis model is included in the defect generator) to generate the image to be determined through the image synthesis model.

17. In the above process (II), The data generation apparatus according to claim 10, wherein the tool usage simulation result information includes at least a portion of the measured values ​​obtained by using the tool, the overview image description, and the image description in which the use of the tool was simulated.

18. (IV) The data generation apparatus according to claim 10, further comprising the process of the processor fine-tuning an agent model based on a vision language model through learning using the training data.

Citation Information

Patent Citations

  • Generation of Agentic Trajectories for Training Artificial Intelligence Agents to Automate Multimodal Interface Task Workflows

    US20250299098A1

  • Open-vocabulary object detection based on frozen vision and language models

    WO2024006340A1

  • Synthetic data generation for training visual language models

    WO2025087788A1