Article inspection apparatus and method therefor

By combining natural language processing and image processing technologies, the problem of robots' inability to accurately identify items has been solved, enabling efficient and accurate item inspection and delivery.

CN120898128APending Publication Date: 2025-11-04LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380095688.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-15
Filing Date
2023-12-04
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Robots cannot proactively identify delivery targets like humans, resulting in low efficiency, especially in accurately determining whether items meet constraints.

Method used

An object inspection device and method based on natural language constraints and images is proposed. Using an image processor, a text processor, and an inspector, a pre-trained dataset is used to determine whether an object image meets the constraints.

Benefits of technology

It improves the accuracy and efficiency of item inspection, and can clearly display or output items that do not meet the constraints and the reasons therefor, thereby improving the efficiency of item delivery and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120898128A_ABST
    Figure CN120898128A_ABST
Patent Text Reader

Abstract

The present invention provides an article inspection apparatus, which may include: an image processor that performs image processing on an input article image; a text processor that performs text processing for the input constraint conditions for article inspection; and a checker which determines whether the input article image satisfies the constraint condition by using the constraint condition and a pre-trained data set of a reference article image related to the constraint condition or information related to the pre-trained data set.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to an article inspection device and a method for the device based on a constraint matter in natural language, in more detail. BACKGROUND

[0002] Due to the aging of the population, the avoidance phenomenon of job seekers, the cost burden of employers, and the like, robots have begun to replace jobs for simple physical labor. When visiting business sites such as restaurants, a robot that delivers food to an orderer (customer) can often be seen. Such a delivery robot is expected to be applied to more fields, and is currently used outside of restaurants.

[0003] Robots are expected to perform work that has been performed by human laborers in the past, but the problem is that robots cannot actively recognize a delivery target article like a person. In addition, if a robot cannot recognize an instruction or a command of an instruction giver like a person, a step for inputting the instruction or the command will be increased, and thus it is possible to cause a low use efficiency of the robot.

[0004] Therefore, the present application proposes a scheme in which a robot can efficiently and accurately determine whether a constraint condition of an instruction giver for an article, such as whether an order fruit type is loaded in a tray loaded with a plurality of fruits when the tray is delivered to a customer, the number of related fruits, and the like, is satisfied. SUMMARY

[0005] Problems to be Solved

[0006] The present application proposes an article inspection device and a method for the device.

[0007] In more detail, a device, a method, and the like that perform article inspection by a combination of a constraint condition based on natural language and an image are proposed.

[0008] The problems to be solved in the present application are not limited to the above-described problems to be solved, and other problems not mentioned can be clearly understood by a person having ordinary knowledge in the technical field to which the present application pertains from the following description.

[0009] Technical Solution to the Problems

[0010] The present application proposes an article inspection device, and the article inspection device can include an image processor that performs image processing for an input article image, a text processor that performs text processing for an input constraint condition for article inspection, and an inspector that determines whether the input article image satisfies the constraint condition using the constraint condition and a pre-trained dataset of a reference article image related to the constraint condition or information related to the pre-trained dataset.

[0011] The present application proposes an article inspection method, which is performed by an apparatus for performing article inspection, the article inspection method including: a step of performing image processing for an input article image and performing text processing for an input constraint condition for article inspection; and a step of determining whether the input article image satisfies the constraint condition using the constraint condition and a pre-trained dataset of a reference article image related to the constraint condition or information related to the pre-trained dataset.

[0012] The solution to the above problem is only a part of the embodiments of the present application, and various embodiments reflecting the technical features of the present application can be derived and understood by those having ordinary knowledge in the art based on the above detailed description of the present application.

[0013] Effects of Invention

[0014] The present application has the following effects.

[0015] According to the present application, the accuracy and efficiency of article inspection can be improved.

[0016] According to the present application, the result of article inspection can be displayed or output for individual articles and individual conditions that do not satisfy the constraint condition, thereby improving the efficiency of article inspection or article distribution and delivery.

[0017] The effects of the present application are not limited to the above-mentioned effects, and other effects not mentioned can be explicitly understood by those having ordinary knowledge in the art to which the present application belongs from the detailed description of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0018] To help understand the present application, the accompanying drawings included as a part of the detailed description provide illustrations about the embodiments of the present application, which explain the technical idea of the present application together with the detailed description.

[0019] Figure 1 An architecture of article inspection according to the prior art is illustrated.

[0020] Figure 2 A reference image obtained by photographing an article and an article image compared therewith are illustrated.

[0021] Figure 3 An architecture of article inspection according to the present application is illustrated.

[0022] Figure 4 An architecture of article inspection according to the present application is illustrated.

[0023] Figure 5 An architecture of article inspection according to the present application is illustrated.

[0024] Figure 6 A detailed architecture of the article inspection according to the present application is illustrated.

[0025] Figure 7 Results of performance improvement according to the present application are shown.

[0026] Figure 8 A detailed architecture of the article inspection according to the present application is illustrated.

[0027] Figure 9 A flowchart of the article inspection method according to the present application is illustrated. DETAILED DESCRIPTION

[0028] Hereinafter, embodiments disclosed in the present specification will be described in detail with reference to the accompanying drawings, but the same or similar constituent elements are designated by the same reference numerals regardless of the drawings and repetitive explanations will be omitted. The suffixes "module" and "part" used in the following description are assigned or mixed to facilitate writing of the specification, and do not have meanings or roles mutually distinguished by themselves. In addition, when explaining the embodiments disclosed in the present specification, in the case where it is judged that specific explanation of related known technology can confuse the gist of the embodiments disclosed in the present specification, detailed explanation thereof will be omitted. In addition, it should be understood that the drawings are only for the purpose of easily understanding the embodiments disclosed in the present specification, and the technical idea disclosed in the present specification should not be limited by the drawings, including all modifications, equivalents, and even alternatives included in the concept and technical scope of the present application.

[0029] Numerical terms such as first, second, and the like can be used to describe various constituents, but the above-described constituents are not limited by the above-described terms. The above-described terms are used only for the purpose of distinguishing one constituent from other constituents.

[0030] In the case where a certain constituent is mentioned to be "connected" or "coupled" to other constituents, although the constituent can be directly connected or coupled to the other constituent, it should be understood that there can be other constituents "existing" between the constituents. In contrast, in the case where a certain constituent is mentioned to be "directly connected" or "directly coupled" to other constituents, it should be understood that there are no other constituents between the constituents.

[0031] Unless the context clearly indicates otherwise, the singular expression should include the plural expression.

[0032] In the present application, it should be understood that the terms such as "comprise" or "have" are intended to designate the existence of characteristics, numbers, steps, actions, constituent elements, parts, or combinations thereof described in the specification and are not intended to preclude the possibility of adding one or more other characteristics, numbers, steps, actions, constituent elements, parts, or combinations thereof.

[0033] Figure 1 An architecture of visual anomaly detection AD according to the prior art is illustrated. A visual embedding of an image is acquired with a pre-trained image encoder and decoded to acquire a log-likelihood so as to detect whether an anomaly exists in the image.

[0034] Due to significant progress in deep learning, research on visual anomaly detection AD work in recent years has achieved important development by practical deployment in each business block such as detecting various defects on the surface of a product, such as scratches, dents, and scratches. However, in reality, in the visual anomaly detection work, not only the surface defects of the object need to be detected, but also high-dimensional anomalies such as confirming the number of items in the package or confirming the position of the assembled constituent elements need to be detected.

[0035] Therefore, it is expected that the visual anomaly detection AD in the real world can only improve its utilization rate by dealing with two types of anomalies. As one type, there is a structural anomaly caused by physical damage to the surface of an existing product, and as the other type, there is a logical anomaly that violates the logical constraint condition designated by a human.

[0036] Referring to Figure 2 , examples regarding logical anomalies are illustrated. Figure 2 (a) of FIG. illustrates a state in which fruits and cereals, bananas, almonds, etc. are correctly placed on a tray, Figure 2 (b) of FIG. illustrates a state in which a logical anomaly occurs in which fruits and cereals, etc. are interchanged with each other, Figure 2 (c) of FIG. illustrates a state in which a logical anomaly occurs in which a part of the constituent products is missing. The image corresponding to (a) of FIG. Figure 2 may be referred to as an item reference image, and the constraint condition corresponding thereto is "breakfast always includes two oranges and one peach, which must be located on the left side of the tray or box. In addition, a mixture of cereals, banana slices, and almonds should be located on the right side of the tray or box, in which the cereals should be located on the upper side and fill half, and the mixture of banana slices and almonds should be located on the lower side and fill the remaining half." In the following description, the same constraint condition is described as a basis. The above constraint condition is only an example, and other constraint conditions can be applied.

[0037] The two types of anomalies mentioned above often coexist in a single instance, and different approaches are required for detection from each other. Structural anomalies require identification of localized defects, and for logical anomalies, understanding of logical constraints reflecting the global visual context of the product is required. Most of the previous studies on visual anomaly detection AD have focused on identifying localized structural defects in product images, and thus a new approach for handling both structural and logical anomalies is required.

[0038] Hereinafter, visual anomaly detection AD will be described as "article inspection". Article inspection refers to judging whether a constraint item related to an article, which is prepared by a user (administrator), is found in an article image. That is, the article inspection which the present invention mainly handles targets logical anomaly detection, but does not exclude structural anomaly detection.

[0039] Figure 3 A first mode of article inspection according to the present invention is illustrated.

[0040] The unit 100 for inputting an article image receives an article image and transmits it to the text-based object detector 300. The unit 200 for inputting a constraint condition receives a constraint condition composed of natural language voice or text information and transmits it to the text-based object detector 300 and the inspector 400. The text-based object detector 300 performs object detection based on the input constraint condition, and identifies individual articles from the article image according to the result thereof.

[0041] The text-based object detector 300 is capable of matching the identified individual articles with the text information related to the identified individual articles within the constraint condition, and converting it into another form of information. The converted information can be stored in a semantic space 310. For example, the text-based object detector 300 is capable of vectorizing the result of matching the individual article information with the text information, and storing the resultant product in a vector space 310.

[0042] The inspector 400 judges whether the article image complies with the input constraint condition. If the article image complies with the input constraint condition, the inspector 400 is capable of outputting the result thereof using an audiovisual outputter such as a display or a speaker. In addition, if the article image does not comply with the input constraint condition, the inspector 400 is capable of outputting the result thereof using an audiovisual outputter such as a display or a speaker. If the article image does not comply with the input constraint condition, the inspector 400 is capable of outputting the non-compliant individual article and its basis, reason, etc. using the audiovisual outputter.

[0043] Accordingly, the individual articles can be clearly judged within the constraint condition and the basis and reason of the judgment can be confirmed.

[0044] On the other hand, in the manner of the article inspection illustrated in Figure 3 the above-described manner can be pre-trained before performing the article inspection. That is, a plurality of article reference images can be input into the article image inputter 100, and a single constraint condition corresponding thereto can be input so that the inspector 400 trains the "correct answer". In the pre-training step, the inspector 400 can average the matching results of the plurality of article reference images and the text of the constraint condition, and set the averaged information or value as reference information. Thereafter, when the article image and the constraint condition are actually input in order to perform the article inspection, the inspector 400 can compare the converted matching information obtained as a result of the text-based object detection action using the relevant article image and the constraint condition with the reference information, and determine whether the article inspection has failed or not depending on whether the difference therebetween deviates from a certain range set in advance. That is, if the difference between the two information deviates from the certain range set in advance, it is determined that the article inspection has failed; if the difference between the two information is within the certain range set in advance, it is determined that the article inspection has succeeded.

[0045] Figure 4 A second manner of article inspection according to the present application is illustrated.

[0046] The second manner employs CLIP (Contrastive Language-Image Pre-training), which, when simply explained, trains combinations of N*N (image, text) pairs made of N (image, text) pairs. CLIP distinguishes N 2 N incorrect pairs while matching N correct (image, text) pairs. To this end, CLIP is basically composed of a dual-encoder structure composed of an image encoder and a text encoder.

[0047] The unit 100 for inputting the article image receives the article image and transmits it to the image encoder 110. The image encoder 110 performs visual embedding on the article image or obtains a result product of the visual embedding from the article image. The result product of the embedding can be expressed as an arrangement or a vector of numbers of a certain size, and in the present specification, performing embedding or obtaining the corresponding result can be referred to as "obtaining embedding".

[0048] The unit 200 for inputting the constraint condition receives the constraint condition composed of natural language speech or text information and transmits it to the text encoder 210. The text encoder 210 obtains a text embedding from the constraint condition.

[0049] The acquired visual embeddings and acquired text embeddings can be aligned with each other based on similarity. Such alignment is called semantic alignment. Although not illustrated, the alignment of visual and text embeddings can be performed using a controller or inspector 400 that controls the image encoder and / or text encoder. To do this, the controller or inspector 400 is able to acquire the similarity between the visual and text embeddings.

[0050] Inspector 400 can determine whether an item image meets the constraints by comparing a previously acquired similarity score with a reference threshold. That is, if the acquired similarity score exceeds the reference threshold, the item image meets the constraints; otherwise, the item image does not meet the constraints.

[0051] If the image of an item meets the constraints, the inspector 400 can output the result using an audiovisual output device such as a display or speaker. Conversely, if the image of an item does not meet the constraints, the inspector 400 can output the result using an audiovisual output device such as a display or speaker. If the image of an item does not meet the input constraints, the inspector 400 can output the individual non-compliant items along with their basis and reasons using an audiovisual output device.

[0052] Therefore, it is possible to make a clear judgment on individual items within the constraints and to confirm the basis and reasons for the judgment.

[0053] On the other hand, Figure 4 In the item inspection method illustrated in the diagram, the method can be pre-trained before performing the inspection. Specifically, multiple item reference images are input into the item image input device 100, along with their corresponding single constraints, thereby training the inspector 400 to "correctly identify the correct answer." During the pre-training step, the inspector can obtain visual and text embeddings using each of the multiple item reference images and the constraints, and calculate the similarity between the visual and text embeddings, training it until the calculated similarity exceeds a reference threshold. This process allows the setting of the reference threshold. Subsequently, when item images and constraints are actually input for item inspection, the inspector 400 can obtain the visual and text embeddings obtained using the relevant item images and constraints, calculate their similarity, and compare the calculated similarity with the reference threshold to determine whether the item inspection was successful.

[0054] Figure 5 The illustration depicts a third method of article inspection according to the present invention. The third method combines the previously described first method with the second method.

[0055] Unit 100, used for inputting object images, receives the object images and transmits them to text-based object detector 300. Unit 200, used for inputting constraints, receives constraints consisting of natural language speech or text information and transmits them to text-based object detector 300. Text-based object detector 300 performs object detection based on the input constraints and identifies individual objects from the object images based on the results.

[0056] The text-based object detector 300 can match identified individual items with text information related to those items within constraints, and convert this information into another form. The converted information can be stored in the semantic space 310. Furthermore, the text-based object detector 300 or inspector 400 can obtain the visual embedding of each object from the converted information stored in the semantic space. The obtained visual embedding can be averaged.

[0057] The unit 100 for inputting an object image transmits the input object image to the image encoder 110. The image encoder 110 obtains a visual embedding from the object image. Additionally, the unit 200 for inputting constraints transmits the input constraints, consisting of natural language speech or text information, to the text encoder 210. The text encoder 210 obtains a text embedding from the constraints.

[0058] The average visual embedding of each object, the acquired visual embedding, and the acquired text embedding can be aligned with each other based on similarity. Such alignment is called semantic alignment. Although not illustrated, this alignment can be performed using a controller or inspector 400 that controls the image encoder and / or text encoder. For this purpose, the controller or inspector 400 can acquire the similarity between the average visual embedding and the text embedding of each object, as well as the similarity between the visual embedding and the text embedding.

[0059] Inspector 400 can determine whether an item image meets the constraints by comparing a previously acquired similarity score with a reference threshold. That is, if the acquired similarity score exceeds the reference threshold, the item image meets the constraints; otherwise, the item image does not meet the constraints.

[0060] If the image of an item meets the constraints, the inspector 400 can output the result using an audiovisual output device such as a display or speaker. Conversely, if the image of an item does not meet the constraints, the inspector 400 can output the result using an audiovisual output device such as a display or speaker. If the image of an item does not meet the input constraints, the inspector 400 can output the individual non-compliant items along with their basis and reasons using an audiovisual output device.

[0061] Therefore, it is possible to make a clear judgment on individual items within the constraints and to confirm the basis and reasons for the judgment.

[0062] On the other hand, Figure 5 In the item inspection method illustrated in the diagram, the method can be pre-trained before performing the item inspection. That is, multiple item reference images can be input into the item image input device 100, along with their corresponding single constraints, to train the inspector 400 to "correctly identify the correct answer." In the pre-training step, the inspector can use each of the multiple item reference images and constraints to obtain the average visual embedding, visual embedding, and text embedding of each object, and calculate the similarity between the average visual embedding and the text embedding, as well as the similarity between the visual embedding and the text embedding, training it to ensure that the calculated similarity exceeds a reference threshold. This process allows setting the reference threshold. Subsequently, when item images and constraints are actually input for item inspection, the inspector 400 can obtain the average visual embedding, visual embedding, and text embedding of each object obtained using the relevant item images and constraints, calculate their similarity, and compare the calculated similarity with the reference threshold to determine whether the item inspection was successful.

[0063] Figure 6 The detailed architecture of the article inspection according to the present invention is illustrated.

[0064] This invention proposes a visual anomaly detection method capable of detecting both structural and logical anomalies in the aforementioned items. Therefore, inspired by the scenarios in which visual inspections are performed, the textual descriptions of logical constraints are used as prior knowledge for visual anomaly detection. To this end, it is necessary to integrate the textual descriptions of logical constraints and apply a multimodal representation model to the visual anomaly detection model.

[0065] Therefore, such as Figure 6 As illustrated in the diagram, CLIP (Contrastive Language-Image Pre-training) is employed, which uses a dual visual and text encoder trained to align image-text pairs. Furthermore, a CLIP stream is proposed that utilizes not only the visual spatial representation of image data but also representations in the semantic space based on logical constraints for visual anomaly detection.

[0066] exist Figure 6 The diagram illustrates three variations of the CLIP flow. Figure 6 (a) is equivalent to CLIP stream - visualization. Figure 6 (b) is equivalent to CLIP stream-text. Figure 6(c) is equivalent to CLIP stream distillation. All three variants use fine-tuned CLIP to leverage more accurate cross-modal representations, and image encoding uses a pre-trained ImageNet model.

[0067] For semantic encoding, CLIP stream-visualization and CLIP stream-text use the CLIP visual encoder and text encoder, respectively, employing both the ImageNet encoder and the CLIP encoder during training and detection. In the case of CLIP stream-distillation, knowledge distillation is performed using an additional mapping module that concatenates the visual encoding from the image encoder with the semantic encoding from the CLIP visual encoder. Unlike CLIP stream-visualization and CLIP stream-text, CLIP stream-distillation eliminates the CLIP encoder, using only the visual encoder during inference.

[0068] -Preparation

[0069] 1. Normalizing flow for anomaly detection

[0070] Normalized flow is a deep generative model that explicitly models the target distribution p*(y) using the bijective transformation g⁻¹: Y -> Z in the latent space Z of the Gaussian distribution pz(z). Using the variable transformation formula, z = g -1 The probability density function of the model for input y∈Y is as follows.

[0071] [Formula 1]

[0072]

[0073] For visual anomaly detection, y is collectively selected as image features from a pre-trained feature extractor. The normalized flow g... -1 It is executed as a mapping function from the high-dimensional feature space to the latent space.

[0074] Then, g -1 The parameters are optimized by minimizing a negative log-likelihood loss.

[0075] [Formula 2]

[0076]

[0077] Here, It is the Jacobian determinant of the parameters θ of the flow. Then, from the estimated parameterized density...

[0078]

[0079] Calculate the likelihood of image features. During detection, the relevant model assigns low probabilities to ideal inputs, identifying them as difficult to sample from the normal training set.

[0080] 2. Multimodal contrastive learning

[0081] Recent research on visual-language representation models has proposed a multimodal contrastive learning approach to align images and text descriptions in a visual-language joint embedding space, providing a new paradigm for visual models. Specifically, CLIP employs an image encoder I:x→R... d A text encoder T:t→R d The architecture is a simple dual encoder. Joint visual-linguistic embeddings of CLIP, pre-trained on a dataset of 400 million image-text pairs, effectively link rich object concepts across different modalities, enabling visual embeddings to be used as semantic text embeddings. CLIP allows text-based constraints to be used directly as prior information, providing an incentive to leverage multimodal representations for visual anomaly detection that utilizes logical constraints.

[0082] -CLIP stream

[0083] 1. Fine-tuning CLIP using text logic constraints

[0084] To improve the similarity between visual and text embeddings and textual logical constraints for product images in non-private industries, some parameters of CLIP need to be fine-tuned using their target dataset. The last block of the CLIP visual encoder and text encoder, along with all layers except the projection layer used to optimize the MVTec LOCO AD, are frozen. Since there is a single definition of logical constraints in each category, category-based visual-language contrastive learning proposed by UniCL (Unified Contrastive Learning) is used to provide articles for each category's logical constraints as general sample descriptions. Figure 7 The correlation between similarity based on fine-tuned normal image-text constraints and AD performance is shown. It is confirmed that fine-tuning performance affects not only logical anomalies but also structural anomalies.

[0085] 2. CLIP Flow - Visual and Text

[0086] The overall framework diagram of CLIP flow-vision is shown in Figure 8To utilize logical constraints as input, the output features of the CLIP visual encoder are used as conditional semantic encodings for the normalized flow AD model. The existing normalized flow model is extended by providing a joint representation of multi-scale features from both fine-tuned CLIP and pre-trained ImageNet to a reversible transformation module. Furthermore, since image embeddings and text embeddings are semantically aligned, a CLIP flow-text model is constructed using a CLIP text encoder instead of the CLIP visual encoder to directly inject text into the AD model.

[0087] Reference Figure 8 Multi-scale visual and semantic embeddings generated in the pre-trained ImageNet encoder and the fine-tuned CLIP visual encoder are passed to the streaming decoder model by forming a joint visual-semantic representation. After training, well-defined structural and logical anomalies for each category, present in a single model, are detected using log-likelihood anomaly scores. The dashed arrows from the fine-tuned CLIP text encoder to the semantic space represent the neighborhood of normal image-text constraint pairs whose embeddings are located within the semantic space after the fine-tuning process.

[0088] 3. CLIP Flow Distillation

[0089] CLIP Stream Distillation aims to improve the feedforward process of both CLIP and ImageNet encoders during inference. To achieve this goal, a lightweight version of CLIP Stream Vision is proposed, along with a semantic mapping module that utilizes knowledge distillation techniques.

[0090] As mentioned above, by using a multimodal model for detecting visual anomalies, it is expected that improved detection results can be obtained in the areas of logical and structural anomalies.

[0091] Figure 9 A flowchart illustrating the article inspection method according to the present invention is shown. Figure 9 The method can be executed using an item inspection device or a processor or controller of the item inspection device, and the name of the item inspection device does not limit the scope of the invention. Hereinafter, the method will be executed using an item inspection device. Figure 9 The method will be explained in the following way.

[0092] The item inspection device is capable of performing image processing on the input item image (S910).

[0093] The item inspection device is capable of performing text processing on the input constraints for item inspection (S920).

[0094] The order of image processing (S910) of the input item image and text processing (S920) of the input constraints for item inspection does not limit the scope of the invention. That is, the order of S910 and S920 is the same as that in... Figure 9 The order shown in the diagram is different; text processing (S920) can be performed simultaneously or before image processing (S910).

[0095] The item inspection device can determine whether the input item image satisfies the input constraints using the input constraints and a pre-trained dataset of reference item images associated with the input constraints (S930). Alternatively, the item inspection device can determine whether the input item image satisfies the input constraints using information associated with the aforementioned pre-trained dataset (S930).

[0096] The pre-trained dataset can include data used to train input-constrained object detection on multiple reference object images.

[0097] Information associated with the pre-training dataset may include threshold information regarding the similarity between images and constraints or textual information related to constraints, obtained from the pre-training dataset. More specifically, information associated with the pre-training dataset may include threshold information regarding the similarity between visual embeddings obtained from input object images and textual embeddings obtained from input constraints. Additionally, information associated with the pre-training dataset may include: the similarity between the average vector of the visual embedding of each object obtained as a constraint-based object detection result for the input object image and the textual embedding obtained from the constraints, as well as threshold information regarding the similarity between the visual embeddings obtained from the input object image and the textual embeddings obtained from the constraints.

[0098] Alternatively, information associated with the pre-trained dataset may include a threshold for the log-likelihood anomaly score of the visual-semantic representation data, which combines visual embeddings obtained from the input object images with semantic embeddings obtained from the input object images.

[0099] The item inspection device can detect objects from an input item image and match the detected objects with the text information of the constraints associated with the detected objects.

[0100] When the similarity between the visual embedding obtained from the input item image and the text embedding obtained from the above constraints exceeds a preset threshold, the item inspection device can determine that the input item image satisfies the above constraints.

[0101] When the similarity between the average vector of the visual embedding of each object obtained as the object detection result based on the input constraints for the input object image and the text embedding obtained from the constraints, and the similarity between the visual embedding obtained from the input object image and the text embedding obtained from the constraints, exceed a preset threshold, the object inspection device determines that the input object image satisfies the constraints.

[0102] When the log-likelihood anomaly score of the visual-semantic representation data, which combines the visual embedding obtained from the input item image with the semantic embedding obtained from the input item image, exceeds a preset threshold, the item inspection device can determine that the input item image does not meet the above constraints.

[0103] When the input image of an item does not meet the input constraints, the item inspection device can output information related to the unmet constraints to an audiovisual output device such as a display or speaker.

[0104] When Figure 2 The examples shown in the diagram illustrate whether or not the constraints are met. As shown in the table below, in the case of an input item image that does not meet the constraints, the item inspection device can output the individual constraints that are not met to the audiovisual output device.

[0105] [Table 1]

[0106]

[0107] Furthermore, as another aspect of the present invention, the previously described solutions or inventive actions can also be provided as code or computer-readable storage media or computer program products that can be implemented, carried out, or executed using a "computer" (including a comprehensive concept such as a system on chip (SoC) or (micro) processor, etc.). The scope of the present invention can be extended to the aforementioned code or computer-readable storage media or computer program products that store or contain the aforementioned code.

[0108] Detailed description of preferred embodiments of the invention disclosed above is provided to enable those skilled in the art to implement and practice the invention. While the invention has been described above with reference to preferred embodiments, those skilled in the art will understand that various modifications and alterations can be made to the invention as set forth in the following claims. Therefore, the invention is not limited to the embodiments shown herein, but is to be given the maximum scope consistent with the principles and novel features disclosed herein.

Claims

1. An item inspection device, wherein, include: An image processor performs image processing on an input image of an object; A text processor that performs text processing on the input constraints used for item inspection; as well as The checker uses the constraints and a pre-trained dataset of reference object images associated with the constraints, or information associated with the pre-trained dataset, to determine whether the input object image satisfies the constraints.

2. The article inspection device according to claim 1, wherein, It includes a text-based object detector that detects objects from the input image of an object and matches the detected objects with textual information of the constraints associated with the detected objects.

3. The article inspection device according to claim 1, wherein, The pre-trained dataset contains data trained on object detection based on the constraints for a plurality of reference object images.

4. The article inspection device according to claim 1, wherein, The pre-trained dataset includes: Data that is semantically aligned based on the similarity between i) the visual embeddings obtained from the plurality of reference object images and ii) the text embeddings obtained from the constraints.

5. The article inspection device according to claim 1, wherein, When the similarity between the visual embedding obtained from the input item image and the text embedding obtained from the constraint exceeds a preset threshold, the checker determines that the input item image satisfies the constraint.

6. The article inspection device according to claim 1, wherein, The pre-trained dataset includes: Data for semantic alignment is based on the similarity between i) the average vector of the visual embedding of each object obtained from object detection results based on the constraints for a plurality of reference object images, ii) the visual embeddings obtained from the plurality of reference object images, and iii) the text embeddings obtained from the constraints.

7. The article inspection device according to claim 1, wherein, The inspector determines that the input item image satisfies the constraints when the similarity between the average vector of the visual embedding of each object obtained as an object detection result based on the constraints for the input item image, the visual embedding obtained from the input item image, and the text embedding obtained from the constraints exceeds a preset threshold.

8. The article inspection device according to claim 1, wherein, The pre-trained dataset contains visual-semantic representation data obtained by combining visual embeddings obtained from a plurality of reference object images with semantic embeddings obtained from the plurality of reference object images; The semantic embedding is adjusted based on the text embedding obtained from the constraints.

9. The article inspection device according to claim 1, wherein, When the log-likelihood anomaly score of the visual-semantic representation data obtained by combining the visual embedding obtained from the input item image with the semantic embedding obtained from the input item image exceeds a preset threshold, the checker determines that the input item image does not meet the constraint condition.

10. The article inspection device according to claim 1, wherein, The system includes an output device that outputs information related to the unmet constraints when the item image does not meet the constraints.

11. A method for inspecting articles, wherein, The article inspection method is performed using a device for performing article inspection, and the article inspection method includes: Perform image processing on the input image of the item, and perform text processing on the input constraints for item inspection; and The step of determining whether the input item image satisfies the constraints using the constraints and a pre-trained dataset of reference item images associated with the constraints, or information associated with the pre-trained dataset.

12. The article inspection method according to claim 11, wherein, The method includes the steps of detecting objects from the image of the item and matching the detected objects with text information of the constraints associated with the detected objects.

13. The article inspection method according to claim 11, wherein, The pre-trained dataset contains data trained on object detection based on the constraints for a plurality of reference object images.

14. The article inspection method according to claim 11, wherein, The pre-trained dataset includes: Data is semantically aligned based on the similarity between i) visual embeddings obtained from a plurality of reference object images and ii) text embeddings obtained from the constraints.

15. The article inspection method according to claim 11, wherein, The method includes the step of determining that the input item image satisfies the constraint condition when the similarity between the visual embedding obtained from the input item image and the text embedding obtained from the constraint condition exceeds a preset threshold.

16. The article inspection method according to claim 11, wherein, The pre-trained dataset includes: Data for semantic alignment is based on the similarity between i) the average vector of the visual embedding of each object obtained from object detection results based on the constraints for a plurality of reference object images, ii) the visual embeddings obtained from the plurality of reference object images, and iii) the text embeddings obtained from the constraints.

17. The article inspection method according to claim 11, wherein, include: The step of determining that the input item image satisfies the constraints is when the similarity between the average vector of the visual embedding of each object obtained as the object detection result based on the constraints for the input item image, the visual embedding obtained from the input item image, and the text embedding obtained from the constraints exceeds a preset threshold.

18. The article inspection method according to claim 11, wherein, The pre-trained dataset contains visual-semantic representation data obtained by combining visual embeddings obtained from a plurality of reference object images with semantic embeddings obtained from the plurality of reference object images; The semantic embedding is adjusted based on the text embedding obtained from the constraints.

19. The article inspection method according to claim 11, wherein, include: The step of determining that the input item image does not meet the constraint condition is when the log-likelihood anomaly score of the visual-semantic representation data obtained by combining the visual embedding obtained from the input item image with the semantic embedding obtained from the input item image exceeds a preset threshold.

20. The article inspection method according to claim 11, wherein, include: The step of outputting information related to the unmet constraints when the item image does not meet the constraints.

21. A non-volatile computer-readable medium, wherein, The computer program is stored in which the method of any one of claims 11 to 20 is performed.