Method for generating training data to be used for training machine learning model and training data generating device using the same

KR103012795B1Active Publication Date: 2026-09-02SUPERB AI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
KR1020250118613
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-09-02
Estimated Expiration
2045-08-25

Smart Images

  • Figure 112025097072961-PAT00002_ABST
    Figure 112025097072961-PAT00002_ABST
Patent Text Reader

Abstract

The present invention relates to a method for generating training data for training a machine learning model, comprising: (a) when an original image is acquired, a training data generating device performs a sub-process of performing captioning on the original image to generate an image caption for the original image, extracting at least one noun phrase from the image caption, and performing open vocabulary object detection on the original image by referencing the at least one noun phrase to generate at least one first pseudo-label including at least one first category name and a first bounding box; and a sub-process of extracting at least one proposal corresponding to at least one object from the original image, generating at least one region description corresponding to the at least one proposal, and generating at least one second pseudo-label including at least one second category name and a second bounding box. and (b) the training data generating device generates a combined water label by filtering the at least one first water label and the at least one second water label according to preset filtering conditions, and generates training data by annotating the original image with the combined water label; the method is to include the step of generating training data.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a method for generating training data for training a machine learning model and a training data generation device using the same. More specifically, the invention relates to a method for generating training data through auto-labeling based on a combination of a Vision Language Model (VLM) and a Vision Foundation Model (VFM), and a training data generation device using the same. Background Technology

[0002] Generally, labeling methods for generating training data include manual labeling, where a labeler manually labels images, and auto-labeling, where a trained machine learning model automatically labels images. Recently, as the performance of machine learning models has improved, auto-labeling methods are widely used for generating training data.

[0003] The process of generating training data using the auto-labeling method involves training object detection models, such as the Vision Language Model (VLM) or Vision Foundation Model (VFM), using a small amount of manually labeled, high-quality training data. Subsequently, pseudo-labels or auto-labels are generated by predicting objects on unlabeled datasets using the trained models—that is, by automatically assigning information such as object location and class to unlabeled images using the trained models. These pseudo-labels or auto-labels are then used to annotate the unlabeled datasets, thereby creating the training dataset. Finally, the training dataset generated by auto-labeling is utilized again to train the object detection model, and if necessary, the quality of the labels is improved using quality validation or misslabel detection algorithms. By repeating this process, the object detection model performing auto-labeling evolves autonomously as it is applied to increasingly larger amounts of data and diverse environments.

[0004] However, in such conventional auto-labeling methods, if the initial manually labeled training data is insufficient or of low quality, the first auto-labeled result becomes inaccurate. Furthermore, false positives and false negatives accumulate in the object detection model during subsequent iterative training, and incorrect pseudo-labels or auto-labels may be repeatedly included in the training data. In this case, there is a risk that the object detection model will learn incorrect patterns. Additionally, while mislabels or noise are prone to being mixed in during large-scale auto-labeling, it is difficult for humans to verify them individually. Moreover, in specialized domains such as industrial sites, additional quality control is required due to the many details and rare events that general VLM or VFM methods miss.

[0005] Therefore, the applicant intends to propose a new method for constructing a large-scale, high-quality training data set. The problem to be solved

[0006] The present invention aims to solve all the problems of the aforementioned prior art.

[0007] In addition, the present invention has another objective of enabling the generation of large-scale and high-quality training data through a combination of a vision language model and a vision foundation model.

[0008] In addition, another objective of the present invention is to prevent the accumulation of incorrect labels through filtering and verification of the outputs of the vision language model and the vision foundation model.

[0009] In addition, another objective of the present invention is to enable the rapid construction of high-quality, large-scale multi-domain data sets without manual work, thereby reducing the cost and time required for data construction in actual industrial and research settings. means of solving the problem

[0010] A representative configuration of the present invention for achieving the above objective is as follows.

[0011] According to one embodiment of the present invention, a method for generating training data for training a machine learning model comprises: (a) when an original image is acquired, a training data generating device performs a subprocess of performing captioning on the original image to generate an image caption for the original image, extracting at least one noun phrase from the image caption, and performing open vocabulary object detection on the original image by referencing the at least one noun phrase to generate at least one first pseudo label including at least one first category name and a first bounding box; and a subprocess of extracting at least one proposal corresponding to at least one object from the original image, and generating at least one region description corresponding to the at least one proposal to generate at least one second pseudo label including at least one second category name and a second bounding box. and (b) the training data generating device generates a combined water label by filtering the at least one first water label and the at least one second water label according to preset filtering conditions, and generates training data by annotating the original image with the combined water label; a method is provided.

[0012] In the above embodiment, in step (b), the training data generating device may calculate image-text alignment scores, which are alignment scores between a cropped image corresponding to a bounding box and a category name, for each of the at least one first pseudonym label and the at least one second pseudonym label, and filter the at least one first pseudonym label and the at least one second pseudonym label by at least one filtering condition among a first filtering condition that removes a pseudonym label corresponding to an image-text alignment score that is less than or equal to a preset threshold score among the image-text alignment scores, a second filtering condition that removes a pseudonym label for each of the at least one first pseudonym label and the at least one second pseudonym label in which the size of the bounding box is less than or equal to a preset threshold size, and a third filtering condition that removes a pseudonym label for each of the at least one first pseudonym label and the at least one second pseudonym label in which the IOU between bounding boxes is greater than or equal to a preset threshold IOU.

[0013] In the above embodiment, in step (b), the training data generation device may compute the image-text alignment scores through any one of the following models under the first filtering condition: CLIP (Contrastive Language-Image Pre-training) model, SigLIP (Sigmoid Loss for Language Image Pre-training)-2 model, GLIP (Grounded Language-Image Pre-training) model, and BLIP (Bootstrapping Language-Image Pre-training)-2 model.

[0014] In the above embodiment, in step (b), the training data generating device filters a capital label corresponding to either a specific first bounding box and a specific second bounding box, wherein the IOU is greater than or equal to the preset threshold IOU according to the third filtering condition, by referring to a specific first category name corresponding to the specific first bounding box and a specific second category name corresponding to the specific second bounding box, the device may include the integrated capital label corresponding to any one specific bounding box that matches the task of a machine learning model—the machine learning model is a model to be learned using the training data—and remove another specific capital label corresponding to another specific bounding box.

[0015] In the above embodiment, in step (a), the training data generation device may perform a sub-process of generating the first pseudo-label by extracting the image caption for the original image through a first Vision Language Model, extracting the at least one noun phrase from the image caption through a natural language processing model, and inputting the at least one noun phrase and the original image into an open vocabulary object detection model to enable the open vocabulary object detection model to detect the at least one object corresponding to the at least one noun phrase in the original image.

[0016] In the above embodiment, in step (a), the training data generation device may perform a sub-process of generating the second pseudo-label by inputting the original image into an object proposal model so that the object proposal model extracts the at least one proposal corresponding to the at least one object from the original image, and input the at least one proposal and the original image into a second vision language model so that the second vision language model generates the at least one region description for the at least one proposal.

[0017] According to one embodiment of the present invention, a training data generating device for generating training data for training a machine learning model comprises: a memory storing instructions for generating training data for training a machine learning model; and a processor for generating training data for training a machine learning model according to the instructions stored in the memory. A training data generation device is provided, comprising: (I) a subprocess of, when an original image is acquired, performing captioning on the original image to generate an image caption for the original image, extracting at least one noun phrase from the image caption, and performing open vocabulary object detection on the original image by referencing the at least one noun phrase to generate at least one first pseudo label including at least one first category name and a first bounding box; and a subprocess of, when the original image is acquired, performing at least one proposal corresponding to at least one object, and performing at least one region description corresponding to the at least one proposal to generate at least one second pseudo label including at least one second category name and a second bounding box; and (II) a process of, when the at least one first pseudo label and the at least one second pseudo label are filtered according to preset filtering conditions to generate an integrated pseudo label, and when the integrated pseudo label is annotated on the original image, a training data generation device is provided.

[0018] In the above embodiment, the processor may, in the process (II), calculate image-text alignment scores, which are alignment scores between a cropped image corresponding to a bounding box and a category name, for each of the at least one first pseudonym and the at least one second pseudonym, and filter the at least one first pseudonym and the at least one second pseudonym by at least one filtering condition among the first filtering condition for removing a pseudonym corresponding to an image-text alignment score that is less than or equal to a preset threshold score among the image-text alignment scores, a second filtering condition for removing a pseudonym for each of the at least one first pseudonym and the at least one second pseudonym for which the size of the bounding box is less than or equal to a preset threshold size, and a third filtering condition for removing a pseudonym for each of the at least one first pseudonym and the at least one second pseudonym for which the IOU between bounding boxes is greater than or equal to a preset threshold IOU.

[0019] In the above embodiment, the processor may compute the image-text alignment scores through any one of the following models in the process (II), in the first filtering condition: CLIP (Contrastive Language-Image Pre-training) model, SigLIP (Sigmoid Loss for Language Image Pre-training)-2 model, GLIP (Grounded Language-Image Pre-training) model, and BLIP (Bootstrapping Language-Image Pre-training)-2 model.

[0020] In the above embodiment, the processor, in the process (II), filters a number label corresponding to either a specific first bounding box and a specific second bounding box in which the IOU is greater than or equal to the preset threshold IOU according to the third filtering condition, by referring to a specific first category name corresponding to the specific first bounding box and a specific second category name corresponding to the specific second bounding box, can include the integrated number label corresponding to any one specific bounding box that matches the task of a machine learning model—the machine learning model is a model to be learned using the training data—and remove another specific number label corresponding to another specific bounding box.

[0021] In the above embodiment, the processor, in performing a sub-process of generating the first pseudo-label in the process (I), extracts the image caption for the original image through a first Vision Language Model, extracts the at least one noun phrase from the image caption through a natural language processing model, and inputs the at least one noun phrase and the original image into an open vocabulary object detection model so that the open vocabulary object detection model detects the at least one object corresponding to the at least one noun phrase in the original image.

[0022] In the above embodiment, the processor, in performing a sub-process of generating the second pseudo-label in the process (I), inputs the original image into an object proposal model so that the object proposal model extracts the at least one proposal corresponding to the at least one object from the original image, and inputs the at least one proposal and the original image into a second vision language model so that the second vision language model generates the at least one region description for the at least one proposal.

[0023] In addition, a computer-readable recording medium for recording a computer program for executing the method of the present invention is further provided. Effects of the invention

[0024] According to the present invention, large-scale and high-quality training data can be generated by combining a vision language model and a vision foundation model.

[0025] According to the present invention, the accumulation of incorrect labels can be prevented through filtering and verification of the outputs of the vision language model and the vision foundation model.

[0026] According to the present invention, high-quality, large-scale multi-domain data sets can be rapidly constructed without manual work, thereby reducing the cost and time of data construction in actual industrial and research sites. Brief explanation of the drawing

[0027] The drawings attached below for use in describing embodiments of the present invention are merely some of the embodiments of the present invention, and other drawings can be obtained based on these drawings without inventive work by a person skilled in the art to which the present invention pertains (hereinafter "person skilled in the art"). FIG. 1 schematically illustrates a training data generation device for generating training data for training a machine learning model according to an embodiment of the present invention, and FIG. 2 schematically illustrates a method for generating training data for training a machine learning model according to an embodiment of the present invention, and FIG. 3 exemplarily illustrates a first water level label and a second water level label generated on an original image in a method for generating training data for training a machine learning model according to an embodiment of the present invention, and FIG. 4 schematically illustrates a state in which a first filtering is performed using an image-text alignment score on the first pseudo-label and the second pseudo-label of FIG. 3 in a method for generating training data for training a machine learning model according to an embodiment of the present invention. FIG. 5 schematically illustrates a state in which a second filtering is performed using the bounding box size at the first pseudonym label and the second pseudonym label of FIG. 3 in a method for generating training data for training a machine learning model according to an embodiment of the present invention. FIG. 6 schematically illustrates a state in which a third filtering is performed using IOUs on the first pseudonym label and the second pseudonym label of FIG. 3 in a method for generating training data for training a machine learning model according to an embodiment of the present invention, and FIG. 7 schematically illustrates an integrated water label obtained by performing first to third filtering on the first water label and the second water label of FIG. 3 in a method for generating training data for training a machine learning model according to an embodiment of the present invention. Specific details for implementing the invention

[0028] The following detailed description of the invention refers to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. It should be understood that various embodiments of the invention are different but need not be mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be modified from one embodiment to another without departing from the spirit and scope of the invention. It should also be understood that the location or arrangement of individual components within each embodiment may be modified without departing from the spirit and scope of the invention. Accordingly, the following detailed description is not intended to be limited in meaning, and the scope of the invention should be understood to encompass the scope claimed by the claims and all equivalents thereof. Similar reference numerals in the drawings indicate identical or similar components across various aspects.

[0029] Hereinafter, in order to enable a person skilled in the art to easily practice the present invention, various preferred embodiments of the present invention will be described in detail with reference to the attached drawings.

[0030] FIG. 1 schematically illustrates a training data generating device for generating training data for training a machine learning model according to an embodiment of the present invention. Referring to FIG. 1, the training data generating device (1000) may include a memory (1100) in which instructions for generating training data for training a machine learning model are stored, and a processor (1200) that generates training data for training a deep learning model according to the instructions stored in the memory (1100).

[0031] Specifically, the learning data generating device (1000) may achieve desired system performance by utilizing a combination of a computing device (e.g., a device that may include components of a computer processor, memory, storage, input device and output device, and other conventional computing devices; an electronic communication device such as a router, switch, etc.; an electronic information storage system such as a Network Attached Storage (NAS) and a Storage Area Network (SAN)) and computer software (i.e., instructions that cause the computing device to function in a specific way), but is not limited thereto.

[0032] Additionally, the processor (1200) of the learning data generation device (1000) may include hardware configurations such as an MPU (Micro Processing Unit) or CPU (Central Processing Unit), cache memory, and data bus. Additionally, the computing device may further include software configurations such as an operating system and an application for a specific purpose.

[0033] However, this does not exclude the case where the learning data generation device (1000) includes an integrated processor in which a medium, a processor, and a memory are integrated for implementing the present invention.

[0034] Meanwhile, the processor (1200) of the learning data generation device (1000) may perform a process according to instructions stored in memory (1100), wherein when an original image is acquired, it performs captioning on the original image to generate an image caption for the original image, extracts at least one noun phrase from the image caption, and performs open vocabulary object detection on the original image by referencing at least one noun phrase to generate at least one first pseudo label including at least one first category name and a first bounding box, and performs a sub-process to extract at least one proposal corresponding to at least one object from the original image, generates at least one region description corresponding to at least one proposal, and generates at least one second pseudo label including at least one second category name and a second bounding box. And, the processor (1200) of the training data generation device (1000) can perform a process of generating training data by filtering at least one first pseudonym label and at least one second pseudonym label according to pre-set filtering conditions according to instructions stored in memory (1100), generating an integrated pseudonym label, and annotating the original image with the integrated pseudonym label.

[0035] The method of generating training data for training a machine learning model in the training data generating device (1000) configured in this way is explained in more detail with reference to FIG. 2 as follows.

[0036] First, the training data generation device (1000) can acquire the original image required for training data generation (S100).

[0037] For example, the original image may be obtained from an unlabeled dataset that stores unlabeled images collected for the generation of training data. In this case, the unlabeled dataset may be stored in a storage device or cloud storage connected to the training data generation device (1000), but the present invention is not limited thereto and may be stored in various storage environments capable of recording data.

[0038] And, the learning data generation device (1000) can generate an image caption for the original image by performing captioning on the original image (S211).

[0039] For example, a training data generation device (1000) can extract an image caption for an original image through a first vision language model. At this time, the image caption can be assumed to be, for example, “there is a lot of man playing an acoustic guitar,” which is used to derive category names of “man” and “acoustic guitar” corresponding to the solid line box in FIG. 3. The first vision language model may include CLIP (Contrastive Language-Image Pre-training), BLIP (Bootstrapping Language-Image Pre-training), InternVL (large-scale vision-language foundation model), etc., but the present invention is not limited thereto and may include various multimodal models that perform image captioning.

[0040] Afterwards, the learning data generation device (1000) can extract at least one noun phrase from the image caption (S212).

[0041] For example, a learning data generation device (1000) can extract at least one noun phrase from an image caption through a natural language processing model, and the extracted noun phrase may be a coarse-grained noun phrase. At this time, the extracted noun phrase may be, for example, “man,” “acoustic guitar,” etc., corresponding to the solid line box of FIG. 3. The natural language processing model may include spaCy, WordNet, etc., but the present invention is not limited thereto and may include various models for processing natural language. The noun phrase may include nouns representing objects, object attributes such as color, shape, and material.

[0042] And, the learning data generation device (1000) may perform a sub-process of generating at least one first pseudo label (S214) including at least one first category name and a first bounding box by performing open vocabulary object detection in the original image by referencing at least one noun phrase (S213). At this time, the first category name and the first bounding box may be, for example, texts and solid boxes corresponding to the solid boxes of FIG. 3.

[0043] For example, a training data generation device (1000) inputs at least one noun phrase and an original image into an open vocabulary object detection model to enable the open vocabulary object detection model to detect at least one object corresponding to at least one noun phrase in the original image, that is, to predict the location of at least one object corresponding to at least one noun phrase, for example, a bounding box, and accordingly, can generate at least one first pseudo-label including at least one first category name which is at least one noun phrase and at least one first bounding box corresponding thereto. At this time, the open vocabulary object detection model may include at least some of the Grounding-DINO model, zero-shot model, etc., but the present invention is not limited thereto and may include various models that perform object detection in an open vocabulary manner.

[0044] Additionally, the training data generation device (1000) that has acquired the original image (S100) can extract at least one proposal corresponding to at least one object from the original image (S221).

[0045] For example, a training data generation device (1000) may input an original image into an object proposal model to enable the object proposal model to extract at least one proposal corresponding to at least one object in the original image. At this time, the proposal may be, for example, the dotted box of FIG. 3. In addition, the object proposal model may include a Universal Proposal Network (UPN), and the UPN may extract coarse-grained proposals corresponding to instance-level objects and fine-grained proposals corresponding to part-level objects.

[0046] And, the learning data generation device (1000) can perform a sub-process of generating at least one second pseudo-label (S223) including at least one second category name and a second bounding box by generating at least one region description corresponding to at least one proposal (S222). At this time, the region description may be, for example, “old man”, “brown acoustic guitar”, etc., corresponding to the dotted boxes of FIG. 3, and may be a description of an object on the region corresponding to the proposal.

[0047] For example, a training data generation device (1000) inputs at least one proposal and an original image into a second vision language model to cause the second vision language model to generate at least one region description for at least one proposal, thereby generating at least one second pseudo-label including at least one bounding box which is at least one proposal and at least one second category name which is the corresponding region description. At this time, the second vision language model may include a ChatRex model. And, the region description output from ChatRex may be a fine-grained phrase, which is a more detailed and contextual sentence for each proposal.

[0048] When at least one first waterway label and at least one second waterway label for an original image are generated by the method described above, the training data generation device (1000) can generate an integrated waterway label (S300) by filtering at least one first waterway label and at least one second waterway label according to preset filtering conditions.

[0049] At this time, the learning data generation device (1000) calculates image-text alignment scores, which are alignment scores between a crop image corresponding to a bounding box and a category name, for each of at least one first pseudo-label and at least one second pseudo-label, for example, alignment scores regarding how well a category name describes an object on a crop image, and can filter at least one first pseudo-label and at least one second pseudo-label by at least one filtering condition among the following: a first filtering condition for removing a pseudo-label corresponding to an image-text alignment score below a preset threshold score among the image-text alignment scores; a second filtering condition for removing a pseudo-label in which the size of the bounding box is below a preset threshold size for each of at least one first pseudo-label and at least one second pseudo-label; and a third filtering condition for removing a pseudo-label in which the IOU between bounding boxes is above a preset threshold IOU for each of at least one first pseudo-label and at least one second pseudo-label.

[0050] For example, with reference to FIGS. 3 to 7, the process of a learning data generation device (1000) generating an integrated water label by filtering at least one first water label and at least one second water label according to a first filtering condition to a third filtering condition is described as follows. For reference, in FIGS. 3 to 7, a solid line box may represent a first water label, and a dotted line box may represent a second water label.

[0051] As shown in FIG. 3, with the original image auto-labeled to generate first pseudo-labels and second pseudo-labels, as shown in FIG. 4, the training data generation device (1000) can calculate image-text alignment scores for each of the first pseudo-labels and second pseudo-labels through any one of the CLIP (Contrastive Language-Image Pre-training) model, SigLIP (Sigmoid Loss for Language Image Pre-training)-2 model, GLIP (Grounded Language-Image Pre-training) model, and BLIP (Bootstrapping Language-Image Pre-training)-2 model according to the first filtering condition, and can remove pseudo-labels of “partially visible car” (11), “small wooden table” (12), “black sandals” (13) and “black and white sneakers” (14) that are confirmed to have image-text alignment scores below a set threshold score.

[0052] And, as shown in FIG. 5, the learning data generation device (1000) can check the size of the bounding boxes of each of the first and second water labels according to the second filtering condition, and can remove the water label of “mouth” (21) that is confirmed to have a bounding box size less than or equal to the threshold size.

[0053] Subsequently, as shown in FIG. 6, the learning data generation device (1000) can identify pairs of first and second pseudo-labels of “acoustic guitar” (31) and “brown acoustic guitar” (31') and “man” (32) and “older man” (32') that overlap each other according to a third filtering condition, i.e., the IOU of the bounding box is greater than or equal to a set threshold IOU, and by removing the pseudo-labels of “brown acoustic guitar” (31') and “older man” (32'), it can generate an integrated pseudo-label that filters at least one first pseudo-label and at least one second pseudo-label as shown in FIG. 7.

[0054] At this time, the learning data generation device (100) filters a number label corresponding to either a specific first bounding box and a specific second bounding box, wherein the IOU is greater than or equal to a set threshold IOU according to a third filtering condition, by referring to a specific first category name corresponding to a specific first bounding box and a specific second category name corresponding to a specific second bounding box, the device may include a number label corresponding to one specific bounding box that matches the task of the machine learning model to be learned, and remove another number label corresponding to another specific bounding box.

[0055] That is, as explained with reference to FIG. 6, in the pairs of first pseudolabels and second pseudolabels that overlap each other, such as “acoustic guitar” (31) and “brown acoustic guitar” (31’) and “man” (32) and “older man” (32’), if the machine learning model to be learned simply needs the class of the object, the second pseudolabels “brown acoustic guitar” (31’) and “older man” (32’) can be removed and the first pseudolabels “acoustic guitar” (31) and “man” (32) can be included in the combined pseudolabel, and if the machine learning model to be learned needs the attribute of the object, the first pseudolabels “acoustic guitar” (31) and “man” (32) can be removed and the second pseudolabels “brown acoustic guitar” (31’) and “older man” (32’) can be included in the combined pseudolabel.

[0056] Again, referring to FIG. 2, the training data generation device (1000) can generate training data (S400) by annotating the original image with integrated water labels.

[0057] Subsequently, the models used to generate the training data can be retrained using the generated training data, and the process of auto-labeling newly acquired unlabeled datasets can be repeated.

[0058] The embodiments according to the present invention described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware device may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.

[0059] Although the present invention has been described above with specific details such as specific components, limited embodiments, and drawings, this is provided only to aid in a more comprehensive understanding of the invention, and the invention is not limited to the above embodiments, and a person skilled in the art to which the invention belongs can make various modifications and variations from this description.

[0060] Accordingly, the scope of the present invention should not be limited to the embodiments described above, and all modifications equivalent to or equivalent to the claims set forth below, as well as the claims described below, shall be considered to fall within the scope of the concept of the present invention. Explanation of the symbols

[0061] 1000: Training data generation device, 1100: Memory 1200: Processor

Claims

Claim 1 A method for generating training data for training an object detection model, comprising: (a) when an original image is acquired, a training data generating device performing captioning on the original image to generate an image caption for the original image, extracting at least one noun phrase from the image caption—the at least one noun phrase is a coarse-grained noun phrase—and performing open vocabulary object detection on the original image by referencing the at least one noun phrase to generate at least one first pseudo label including at least one first category name and a first bounding box; and a subprocess of extracting at least one proposal corresponding to at least one object in the original image, generating at least one region description corresponding to the at least one proposal—the at least one region description is a fine-grained phrase—to generate at least one second pseudo label including at least one second category name and a second bounding box; and (b) the training data generating device generates a combined capital label by filtering the at least one first capital label and the at least one second capital label according to preset filtering conditions, and generates training data by annotating the original image with the combined capital label;A method comprising, wherein in step (a), the training data generating device performs a sub-process of generating the first pseudo-label by extracting the image caption for the original image through a first Vision Language Model, extracting the at least one noun phrase from the image caption through a natural language processing model, and inputting the at least one noun phrase and the original image into an open vocabulary object detection model to cause the open vocabulary object detection model to detect the at least one object corresponding to the at least one noun phrase in the original image. Claim 2 In claim 1, in step (b), the training data generating device calculates image-text alignment scores, which are alignment scores between a cropped image corresponding to a bounding box and a category name, for each of the at least one first pseudonym label and the at least one second pseudonym label, and filters the at least one first pseudonym label and the at least one second pseudonym label by at least one filtering condition among the at least one filtering condition for removing a pseudonym label corresponding to an image-text alignment score of less than or equal to a preset threshold score among the image-text alignment scores, a second filtering condition for removing a pseudonym label for each of the at least one first pseudonym label and the at least one second pseudonym label in which the size of the bounding box is less than or equal to a preset threshold size, and a third filtering condition for removing a pseudonym label for each of the at least one first pseudonym label and the at least one second pseudonym label in which the IOU between bounding boxes is greater than or equal to a preset threshold IOU. Claim 3 In paragraph 2, in step (b), the training data generating device calculates the image-text alignment scores through any one of the following models under the first filtering condition: CLIP (Contrastive Language-Image Pre-training) model, SigLIP (Sigmoid Loss for Language Image Pre-training)-2 model, GLIP (Grounded Language-Image Pre-training) model, and BLIP (Bootstrapping Language-Image Pre-training)-2 model. Claim 4 In paragraph 2, in step (b), the training data generating device filters a capital label corresponding to either a specific first bounding box and a specific second bounding box, wherein the IOU is greater than or equal to the preset threshold IOU according to the third filtering condition, by referring to a specific first category name corresponding to the specific first bounding box and a specific second category name corresponding to the specific second bounding box, the method includes the integrated capital label corresponding to any one specific bounding box that matches the task of a machine learning model—the machine learning model is a model to be learned using the training data—and removes another specific capital label corresponding to another specific bounding box. Claim 5 delete Claim 6 In claim 1, in step (a), the training data generating device performs a sub-process of generating the second pseudo-label by inputting the original image into an object proposal model so that the object proposal model extracts the at least one proposal corresponding to the at least one object from the original image, and inputting the at least one proposal and the original image into a second vision language model so that the second vision language model generates the at least one region description for the at least one proposal. Claim 7 A training data generating device for generating training data for training an object detection model, comprising: a memory in which instructions for generating training data for training an object detection model are stored; and a processor that generates training data for training the object detection model according to the instructions stored in the memory; wherein the processor comprises: (I) a subprocess that, when an original image is acquired, performs captioning on the original image to generate an image caption for the original image, extracts at least one noun phrase from the image caption—the at least one noun phrase is a coarse-grained noun phrase—and performs open vocabulary object detection on the original image by referencing the at least one noun phrase to generate at least one first pseudo label including at least one first category name and a first bounding box; and a subprocess that extracts at least one proposal corresponding to at least one object in the original image, and generates at least one region description corresponding to the at least one proposal—the at least one region description is a fine-grained phrase—including at least one second category name and a second bounding box. A process for performing a sub-process of generating at least one second pseudo-label, and (II) a process for generating an integrated pseudo-label by filtering the at least one first pseudo-label and the at least one second pseudo-label according to preset filtering conditions, and generating training data by annotating the original image with the integrated pseudo-label, wherein in the process (I), the processor extracts the image caption for the original image through a first vision language model while performing the sub-process of generating the first pseudo-label.A learning data generation device that extracts at least one noun phrase from the image caption through a natural language processing model, inputs the at least one noun phrase and the original image into an open vocabulary object detection model, and causes the open vocabulary object detection model to detect at least one object corresponding to the at least one noun phrase in the original image. Claim 8 In claim 7, the processor, in the process (II), calculates image-text alignment scores, which are alignment scores between a cropped image corresponding to a bounding box and a category name for each of the at least one first pseudonym and the at least one second pseudonym, and filters the at least one first pseudonym and the at least one second pseudonym by at least one filtering condition among the image-text alignment scores, which removes a pseudonym corresponding to an image-text alignment score below a preset threshold score, a second filtering condition for each of the at least one first pseudonym and the at least one second pseudonym, wherein the size of the bounding box is below a preset threshold size, and a third filtering condition for each of the at least one first pseudonym and the at least one second pseudonym, wherein the IOU between bounding boxes is above a preset threshold IOU. Claim 9 In claim 8, the processor is a training data generation device that, in the process (II), calculates the image-text alignment scores through any one of the CLIP (Contrastive Language-Image Pre-training) model, SigLIP (Sigmoid Loss for Language Image Pre-training)-2 model, GLIP (Grounded Language-Image Pre-training) model, and BLIP (Bootstrapping Language-Image Pre-training)-2 model under the first filtering condition. Claim 10 In claim 8, the processor, in the process (II), filters a number label corresponding to either a specific first bounding box and a specific second bounding box in which the IOU is greater than or equal to the preset threshold IOU according to the third filtering condition, and by referring to a specific first category name corresponding to the specific first bounding box and a specific second category name corresponding to the specific second bounding box, a learning data generation device that includes the integrated number label corresponding to any one specific bounding box that matches the task of a machine learning model—the machine learning model is a model to be learned using the training data—and removes another specific number label corresponding to another specific bounding box. Claim 11 delete Claim 12 A learning data generation device according to claim 7, wherein, in the process (I), the processor performs a sub-process of generating the second pseudo-label by inputting the original image into an object proposal model so that the object proposal model extracts the at least one proposal corresponding to the at least one object from the original image, and inputs the at least one proposal and the original image into a second vision language model so that the second vision language model generates the at least one region description for the at least one proposal.

Citation Information

Patent Citations

  • Association analysis method for output image from generative model and input text prompt and analysis apparatus

    KR1020250075209A

  • Apparatus for generating a dataset for an image generation AI model based on caption data describing images

    KR1020250078296A