Visual inspection method and rationale-generative estimation method using large vision–language model

The proposed method leverages a large-scale visual language model to perform efficient visual inspections across various products by incorporating pre-learning and additional learning steps, addressing the limitations of current models in learning specialized knowledge and reducing the need for extensive retraining.

WO2025115537A1PCT designated stage expired Publication Date: 2025-06-05NAT UNIV CORP TOKAI NAT HIGHER EDUCATION & RES SYST

Patent Information

Application Number
PCT/JP2024/039341
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-15
Filing Date
2024-11-06
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current large-scale visual language models are insufficient in learning specialized knowledge and inspection standards for visual inspection of various products, requiring extensive training data and retraining for each product.

Method used

A method using a large-scale visual language model for visual inspection that includes a pre-learning step with context-aware input, an additional learning step with images and text data of good and defective products, and an inspection step that outputs inspection results with grounds for judgment.

Benefits of technology

Enables efficient visual inspection of a wide variety of products with less learning, reduces the burden of creating extensive training datasets, and improves the explainability of large-scale visual language models by generating grounds for judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039341_05062025_PF_FP_ABST
    Figure JP2024039341_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides an article visual inspection method and a rationale-generative estimation method using a large vision–language model. The article visual inspection method comprises: a pretraining step of performing context-sensitive pretraining by inputting information containing images and sentences into a large vision–language model; an additional training step of performing additional training by inputting training information on a good-quality group and training information on a poor-quality group of a plurality of articles into the large vision–language model; and an inspection step of inputting a visual appearance image for inspection into the large vision–language model and causing the large vision–language model to output an inspection result. The article visual inspection method is characterized in that the training information on a good-quality group includes visual appearance images of good-quality articles and text information about good-quality articles, the training information on a poor-quality group includes visual appearance images of poor-quality articles and text information about poor-quality articles, and the information to be inspected includes an image to be inspected and text information requesting an inspection result. The estimation method comprises: a pretraining step of performing pretraining by inputting a plurality of datasets into a large vision–language model, each dataset containing a single image paired with a sentence containing a specified subject's preference information regarding the image; and a result outputting step of inputting an image of an object to be estimated into the large vision–language model and causing the large vision–language model to output an estimation result regarding individual preference and a rationale. The specified subject's preference information includes a judgment of whether or not the subject prefers the single image and text information indicating a rationale for the judgment.
Need to check novelty before this filing date? Find Prior Art

Description

Visual inspection method and decision-making basis generation estimation method using large-scale visual language models

[0001] The present invention relates to a data processing method using a large-scale visual language model, and more particularly to a method for visually inspecting an article and presenting the basis for the inspection results by training a large-scale visual language model using both image data and text data.

[0002] In recent years, large-scale language models capable of document generation and machine translation have been developed by using large amounts of text data as input data and performing deep learning on neural networks with parameters in the hundreds of millions. Furthermore, multimodal models, which are neural networks of a similar scale to large-scale language models but can handle multiple different types of data, have been developed. One example of a multimodal model is a large-scale visual language model that can handle image data and text data simultaneously. Current large-scale visual language models can output data that takes into account the correlations and interactions between image data and text data, enabling them to classify, identify, and estimate patterns contained in the data.

[0003] Specific examples of large-scale visual language models are published in Non-Patent Documents 1, 2, 3, 4, and 5. These large-scale visual language models utilize knowledge from large-scale language models to map visual features to a language space, enabling visual language tasks such as captioning and visual question answering. The model disclosed in Non-Patent Document 2 learns image-language alignment by converting the output of an image encoder into a language space using a linear layer and inputting the image embedding vector and the prompt embedding vector into a large-scale language model. Non-Patent Document 4 proposes a querying transformer that outputs related information in response to a query in order to learn image-language alignment, enabling the combination of a pre-trained image encoder and a large-scale language model.

[0004] One area where deep learning models are expected to be applied is product visual inspection. To date, in order to automate visual inspection using deep learning, it has been necessary to collect data specific to a single product to be inspected and perform training. In other words, in order to automate the visual inspection of an untrained product, it has been necessary to retrain the model each time.

[0005] Furthermore, conventional deep learning models used in visual inspection are often trained using image data of non-defective products. Non-Patent Document 6 discloses PaDiM, which performs pass / fail judgment using feature maps of multiple resolutions of non-defective product images obtained from a pre-trained model. Like conventional deep learning models, PaDiM requires the collection of training samples and model training for each product to be inspected. The same model cannot be applied to a new product.

[0006] In recent years, methods have been developed for applying deep learning to multimodal models using input data that combines images and language to visual inspection. The learning model known as SAA, disclosed in Non-Patent Document 7, is a deep learning model that combines SAM (Non-Patent Document 8) and GroundingDINO (Non-Patent Document 9), and outputs segmentation masks and failure modes for defect locations in language using zero-shot processing. The SAA disclosed in Non-Patent Document 7 requires hyperparameter adjustment for each product. The AnomalyGPT disclosed in Non-Patent Document 10 can identify defect locations by training a decoder from images of good products, pseudo-faulty products, and the corresponding language. However, AnomalyGPT requires training on a specific dataset. The GPT-4V disclosed in Non-Patent Document 11 is said to be capable of determining pass / fail to a certain extent for multiple products in the visual inspection image dataset disclosed in Non-Patent Document 12 (Non-Patent Document 13).

[0007] In conventional visual inspections by inspectors, inspectors are given inspection standards and limit samples for pass / fail products and perform the inspection according to these standards. Applying a learning method similar to the conventional visual inspections performed by inspectors may enable visual inspections of different inspection targets with less learning effort. An example of a model to which a learning method similar to the inspector's skill acquisition method of providing inspection standards can be applied is in-context learning, as disclosed in Non-Patent Document 14. In-context learning is a model that learns from a small number of given examples without updating the model parameters and performs inference on unknown data. Non-Patent Document 15 discloses "Otter" as a large-scale visual language model capable of in-context learning. "Otter" is a large-scale visual language model that has undergone in-context learning and additional learning using the multimodal model "Open Flamingo" disclosed in Non-Patent Document 16, with the "Instruction Tuning" format dataset disclosed in Non-Patent Documents 17 and 18. By performing these additional learnings, the large-scale visual language model is able to execute even unknown tasks when examples are given. However, the information learned by conventional large-scale visual language models is mostly general and easily accessible, and the learning of specialized knowledge and information is lacking compared to the learning of general topics. For this reason, it is currently difficult to judge the quality of visual inspections according to specific criteria.

[0008] One of the expected uses of large-scale visual language models is to present the reasons for test results. For example, as a technology for estimating a subject's preferences, Non-Patent Document 21 discloses an attempt to concretely express a user's vague preferences and values ​​by having a language model conduct a conversation with the subject. If the general-purpose knowledge accumulated by a large-scale visual language model can be used to estimate the preferences of individual subjects and present the basis for that judgment, it may contribute to product development and improved user satisfaction with existing products.

[0009] Zhilian Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, Furu Wei, "Kosmos-2: Grounding Multimodal Large Language Models to the World", arXiv:2306.14824, November 20, 2023. Haotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae Lee, "Visual Instruction Tuning", arXiv:2304.08485, 2023. Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, Mohamed Elhoseiny, "MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models", arXiv:2304.10592, 2023. Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi, "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models", arXiv:2301.12597, 2023. Peng Xu, Wenqi Shao, Kaipeng Zhang, Peng Gao, Shuo Liu, Meng Lei, Fanqing Meng, Siyuan Huang, Yu Qiao, Ping Luo, "LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models", arXiv:2306.09265, 2023. Thomas Defard, Aleksandr Setkov, Angelique Loesch, and Romaric Audigier, "PaDiM: a Patch Distribution Modeling Framework for Anomaly Detection and Localization,In "International Conference on Pattern Recognition" (ICPR), pages 475 - 489, Springer, 2021. Yunkang Cao, Xiaohao Xu, Chen Sun, Yuqi Cheng, Zongwei Du, Liang Gao, Weiming Shen, "Segment Any Anomaly without Training via Hybrid Prompt Regularization", arXiv:2305.10724, 2023. Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan - Yen Lo, Piotr Dollar, Ross Girshick, "Segment Anything", arXiv:2304.02643, 2023. Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang, "Grounding DINO: Marrying DINO with Grounded Pre - Training for Open - Set Object Detection", arXiv:2303.05499, 2023. Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen, Ming Tang, Jinqiao Wang, "AnomalyGPT: Detecting Industrial Anomalies using Large Vision - Language Models", arXiv:2308.15366, 2023. "OpenAI: GPT - 4V(vision) System Card", https: / / cdn.openai.com / papers / GPTV_System_Card.pdf Paul Bergmann, Michael Fauser, David Sattlegger,Carsten Steger, "MVTec AD - A Comprehensive Real - World Dataset for Unsupervised Anomaly Detection", in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9584 - 9592, 2019; Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung - Ching Lin, Zicheng Liu, Lijuan Wang, "The Dawn of LMMs: Preliminary Explorations with GPT - 4V(ision)", arXiv:2309.17421, 2023; Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Pratul Dhariwal, Arvin Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hess, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, Dario Amodei, "Language Models are Few - Shot Learners, Advances in neural information processing systems" (NeurIPS), Vol. 33, pp. 1877 - 1901, 2020; Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang,Jingkang Yang, Ziwei Liu, "Otter: A Multi-Modal Model with In-Context Instruction Tuning", arXiv:2305.03726, 2023 Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, Ludwig Schmidt, "OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models", arXiv:2308.01390, 2023 Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, Quoc V. Le, "Fine-tuned Language Models Are Zero-Shot Learners", arXiv:2109.01652, 2022 Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, Ziwei Liu, "MIMIC-IT: Multi-Modal In-Context Instruction Tuning", arXiv:2306.05425, 2023 Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger,“Learning Transferable Visual Models From Natural Language Supervision” by Ilya Sutskever In International conference on machine learning (PMLR) pp. 8748-8763, 2021 Edited by The MosaicML NLP Team “Introducing MPT-7B: A New Standard for Open-Source,” Commercially Usable LLMs” https: / / www. mosaicml. com / blog / mpt-7bBelinda Z. "Eliciting Human Preferences with Language Models" by Li, arXiv:2310.11589, 2023 Nils Reimers, Iryna Gurevych, “Sentence-bert: Sentence embeddings using Siamese bert-networks” arXiv:1908.10084, 2019,

[0010] There is a demand for efficient visual inspection of various products using large-scale visual language models. However, currently available large-scale visual language models are insufficiently trained with the specialized knowledge and inspection standards corresponding to the products to be inspected. Therefore, in order to have a large-scale visual language model perform actual visual inspections, it is necessary to collect a large number of training samples for each product to be inspected and train the model.

[0011] The present invention has been made in consideration of the current situation, and aims to solve the problem of providing a method for visual inspection of objects using a large-scale visual language model, which enables visual inspection of a wide variety of products with less learning than conventional methods.

[0012] Furthermore, there have been few attempts to generate judgment rationales using large-scale visual language models. For example, if it were possible to generate language to represent the rationale for a test subject's judgment on "preferences," one of their evaluations, it would be possible to improve the explainability of large-scale visual language models. Estimating personal preferences using large-scale visual language models requires creating and training a separate dataset for each subject. However, traditionally, the amount of dataset required for training large-scale visual language models is enormous, which poses a problem: creating the datasets is time-consuming and expensive. Furthermore, the process of creating the subject's dataset places a heavy burden on the subject.

[0013] One of the objectives of the present invention is to provide an estimation method that can generate estimation results and the basis for the judgment in language. To this end, this specification specifically considers a method for estimating "preferences," which are one type of personal evaluation, and simultaneously presenting the basis for that estimation.

[0014] The present invention provides a method for performing visual inspection of an article using a large-scale visual language model. The method for performing visual inspection of an article of the present invention comprises a pre-learning step of inputting information consisting of images and text into the large-scale visual language model to perform pre-learning taking context into consideration, an additional learning step of inputting learning information for a group of good products and learning information for a group of defective products for a plurality of articles into the large-scale visual language model to perform additional learning, and an inspection step of inputting appearance images of an inspection target into the large-scale visual language model and outputting inspection results. The method for visual inspection of an article of the present invention is characterized in that the learning information for the group of good products includes appearance images of the good products and text information about the good products, the learning information for the group of defective products includes appearance images of the defective products and text information about the defective products, and the information about the inspection target includes an image of the inspection target and text information requesting the inspection results.

[0015] In the visual inspection method of the present invention, the text information on a non-defective product preferably includes a fixed question asking whether the product is non-defective or defective, and a fixed answer indicating that the product is non-defective.

[0016] The visual inspection method of the present invention is characterized in that the text information of a defective product includes a standard question asking whether the product is good or defective, a defect explanation including the name of the defect classification, and a standard answer indicating that the product is defective.

[0017] Furthermore, the present invention uses a large-scale visual language model to present the basis for the judgment of the inference result. As a specific example, a method for estimating "preferences," which are personal evaluations, includes a pre-learning step and a result output step. In the pre-learning step, a large-scale visual language model is input with multiple sets of data, each set consisting of an image and a sentence containing preference information of a specific target person regarding that image, for pre-learning. In the result output step, an image of the inference target is input to the large-scale visual language model, and the inference result regarding the individual's preference and the basis for the judgment are output. The method of the present invention is characterized in that the preference information of the specific target person includes a judgment of whether or not the image is a preference, and text information indicating the basis for the judgment.

[0018] In this example, preference information of a specific target person is preferably collected using the evaluation grid method (registered trademark).

[0019] The large-scale visual language model used in the visual inspection method of the present invention can perform inspections without collecting training samples for different products and training the model after an additional training process that strengthens specialized knowledge. As a result, visual inspection can be performed on other products that have not undergone training, providing a more general-purpose visual inspection method that can handle a wide variety of products.

[0020] The method for inspecting the appearance of an article of the present invention can reduce the amount of processing required for learning compared to conventional methods when inspecting the appearance of a wide variety of articles, thereby enabling efficient inspection of the appearance of articles.

[0021] The visual inspection method for articles of the present invention converts the learning information for the group of good products, the learning information for the group of defective products, and the information on the inspection target into image and text information, respectively, so that the inspection result of whether the product is good or defective can be clearly output in text.

[0022] The estimation method of the present invention can generate language that represents the basis for judgment on evaluation results. The basis for judgment obtained from the results of estimation performed on an item can be used for product development and the improvement of existing items. It also makes it possible to clarify evaluation criteria when conducting sensory testing of items. The basis for judgment on preferences output by using a large-scale visual language model can be used to improve the explainability of the large-scale visual language model itself.

[0023] The preference information of specific subjects used for learning in the estimation method of the present invention can be collected using the Evaluation Grid Method (registered trademark). The method of the present invention can be performed more efficiently than conventional methods, thereby reducing the burden on specific subjects who cooperate in collecting preference information.

[0024] FIG. 1 is a diagram showing an example of input information input to a large-scale visual language model for the visual inspection of an article and output information as a result. FIG. 2 is a diagram schematically showing the structure of a large-scale visual language model used in an embodiment of a method for visually inspecting an article. FIG. 3 is a diagram showing an example of an image input in an additional learning process for the visual inspection of an article. FIG. 4 is a drawing-substitute photograph showing an example of an image in which it is difficult to distinguish between a good and a defective product. FIG. 5 is a diagram showing an example of text information input together with an image to be inspected and an example of text information output as an inspection result. FIG. 6 is a drawing-substitute photograph showing an example of an image used in training the method for visually inspecting an article. FIG. 7 is a diagram schematically showing the structure of a large-scale visual language model used in an embodiment of an estimation method. FIG. 8 is a flowchart showing a method for creating a pre-training dataset in an embodiment of the estimation method.

[0025] Hereinafter, one embodiment of the visual inspection method for an article of the present invention will be described with reference to the drawings. Note that this embodiment is one example of realizing the invention and is not intended to limit the scope of the claims.

[0026] The large-scale visual language model in this invention is a computer program that can simultaneously handle image and text data. The large-scale visual language model is trained using data consisting of a large number of images and their associated text, and in the training process learns numerous patterns and relationships, which enables it to appropriately analyze new data and problems.

[0027] In this embodiment, in addition to the conventional pre-training process, an additional training process using unique data is performed on a typical large-scale visual language model to strengthen specialized knowledge related to visual inspection. In-context learning is also performed on the large-scale visual language model. This results in a general-purpose visual inspection model capable of visually inspecting products that have not undergone training. Furthermore, by standardizing the output format of the large-scale visual language model, quantitative evaluation of visual inspection using the large-scale visual language model is possible. Figure 1 shows an example of input information input to the large-scale visual language model in the additional training process and the resulting output information.

[0028] In this embodiment, "Otter," the configuration of which is disclosed in Non-Patent Document 15, is used as the base model of the large-scale visual language model. An overview of the structure of the large-scale visual language model used in this embodiment is shown in Figure 2. The large-scale visual language model is composed of an image encoder for extracting image features, a language model for generating language, and a Perceiver Resampler that resamples data to connect the image encoder and the language model. The image encoder of the large-scale visual language model used in this embodiment is CLIP ViT-L / 14, the structure of which is disclosed in Non-Patent Document 19. The language model is MosaicML Pretrained Transformer 7B, the structure of which is disclosed in Non-Patent Document 20. During additional training, the language model and image encoder are frozen according to the large-scale visual language model, and only the parameters of the Perceiver Resampler, the cross-attention layer inserted in the language model, and the input / output embedding of the language model are updated. The total number of learning parameters is approximately 1.3B (1.03 billion).

[0029] The image and text data used in the additional learning process of this embodiment will be described below. The additional learning uses images of various non-defective and defective products collected from the web in order to enhance the specialized knowledge of visual inspection for the large-scale visual language model.

[0030] The collection of images used for additional learning consists of the following four steps: (Step 1) Using the AI ​​model "GPT-4 (OpenAI)," the product name and text information for inspection keywords are generated to collect images of good and defective products. The inspection text information generated in this embodiment was "good," "chipped," "cracked," "scratched," "cloudy," "stained," and "faded." (Step 2) The generated text information for keywords is expanded to eight languages. (Step 3) Images are collected by performing an image search on the Internet using the expanded text information for keywords. (Step 4) Images that match the keywords and scraped images are visually selected and have no restrictions on data use. Images of a variety of products are collected through steps 1 to 4. Each product image is selected to include a good product image and up to five defective mode images.

[0031] When training a large-scale visual language model to learn visual inspection criteria, it can be difficult to distinguish between good and defective products from appearance images alone. An example of an image where it is difficult to determine whether a product is good or defective is shown in FIG. 4. In this embodiment, in-context learning is used to improve the accuracy of visual inspection while also increasing versatility and enabling the inspection of a larger number of items. Specifically, an appearance image of a good product and an appearance image of a defective product are each combined with explanatory text that serves as the judgment criteria as text information, and these are used as input data for additional learning. An example of an image used for additional learning is shown in FIG. 3.

[0032] To be used in the additional learning process, images combined with text information are divided into example images and test images. Then, for each test image, one good product image and one defective product image of the same product are selected from the example images, creating a set of three images: an example good product image, an example defective product image, and a test image. A maximum of 25 such sets are created for each test image, and the data is expanded. However, if the test image is a defective product, the example defective product image is an image of the same product with the same defective text information added from the example images.

[0033] Furthermore, in the additional learning process, the format of the text information input to and output from the large-scale visual language model is unified, enabling quantitative evaluation of the pass / fail judgment that is the output of the visual inspection results. The contents of the new data set used to unify the input and output formats are explained below with reference to Figures 2 and 5.

[0034] The text information set in this embodiment as a question for each image is "This is {product name}. Does it have a defect called {defect example}?". The text information in the case of an English version is "This is an image of {product name}. Does this {product name} have any defects such as {defect example}?" as shown in FIG. 5. In this way, the question includes text information that requests an answer.

[0035] In the case of a non-defective image, the text information of the response sentence is "No. This {product name} does not have any {defect example}. It is a non-defective product." In the case of an English version, the text information is "No. This {product name} does not have any defects such as {defect example}, so it is non-defective."

[0036] In the case of a defective product image, the text information of the response sentence is "Yes. This {product name} has {example of defect}. It is a defective product." In the case of the English version, the text information is "Yes. This {product name} has some {defect name}, so it is defective."

[0037] By performing additional training using text information that unifies the format of the input question and the format of the output answer, the large-scale visual language model is able to clearly determine whether a product is good or bad.

[0038] In the additional training process, additional training of a large-scale visual language model is first performed using a dataset of good products and a dataset of defective products with a unified output format. First, a first example image and a corresponding question are input, and an answer to the first example image is predicted. Next, a question and answer corresponding to the first example image and a question corresponding to the second example image are input, and an answer to the second example image is predicted. Finally, two example images and their corresponding questions and answers, and a test image and a corresponding question are input, and an answer to the test image is predicted. Loss is calculated using cross-entropy from these three predicted answers and the correct answer.

[0039] A large-scale visual language model that has undergone the above additional learning process can be processed as a general-purpose appearance inspection model that judges the quality of inspection images input in the inspection process.

[0040] Next, an embodiment of a method for estimating individual preferences, as one of the decision-basis generation type estimation methods of the present invention, will be described with reference to the drawings.

[0041] As with the previously described method for visually inspecting an object, the large-scale visual language model of this embodiment uses "Otter," which is disclosed in Non-Patent Document 15. "Otter" is trained on the multimodal model "Open Flamingo" using both in-context learning and an "Instruction Tuning" format dataset, and learns general information through these training methods. Furthermore, by performing additional pre-training using the dataset of this embodiment described below, "Otter" becomes able to estimate individual preferences and generate the basis for the judgments that led to the estimation.

[0042] A method for creating a dataset used in the individual preference estimation method of this embodiment will be described. A dataset is created for each individual subject, i.e., for a specific subject. In this embodiment, the evaluation grid method is applied to create the dataset, and objects are classified and judgment criteria for the specific subject are extracted based on the preferences of the specific subject.

[0043] Figure 8 is a flowchart showing a method for creating a dataset for pre-learning. In the first step (S1) of creating a dataset, several dozen different images of an object whose preferences are to be estimated are prepared. The prepared images are then presented to a specific subject, who is asked to arrange the images in order of preference. Once the arrangement is complete, the images are classified into two groups: a preferred group and an unpreferred group, based on the specific subject's own preferences.

[0044] In the second step (S2) of creating a dataset, the specific subjects are interviewed about the reasons for the classification based on preferences in the first step. The results of the interview are recorded as the basis for judgment based on preferences. By performing the second step (S2) of creating a dataset, the basis for judgment of preferences recognized by the specific subjects themselves can be extracted.

[0045] In the third step (S3) of creating a dataset, several hundred images of objects in the same field as the images used in the first and second steps are prepared. The prepared images are presented one by one to the specific subject, and the specific subject is asked to judge for each image whether or not they like it, along with the reasons for their judgment, which are then recorded. The preference judgment, or whether or not the image is liked, is made by categorizing the image as "yes" if it is liked and "no" if it is not liked. The reasons for the judgment are freely presented in the form of words or sentences. The preference judgment and the reasons for the judgment are used as the specific subject's preference information. The preference information of the specific subject is recorded as text information in one-to-one correspondence with the images. By performing the third step (S3) in a way that allows the specific subject to appropriately review the images used in the second step, it is possible to reduce fluctuations in the basis for the judgment.

[0046] The dataset necessary for additional pre-training of a large-scale visual language model is completed through the first step (S1) to the third step (S3) of creating the dataset described above.

[0047] FIG. 7 shows a schematic diagram of the structure of a large-scale visual language model used in the individual preference estimation method of this embodiment.

[0048] In the pre-training process, the text formats of the questions input to the large-scale visual language model and the estimated results and reasons for judgment about individual preferences that are output by the large-scale visual language model as answers to the questions are unified.

[0049] The format of the question to be input to the large-scale visual language model was "Is this image your preferred {object}? Please answer first with 'yes' or 'no' and then give the reason for your decision." The format of the question in the English version was specified as "Is this image the preferred {object}? Please answer first with 'yes' or 'no' and then give the reason for your decision." In this way, the question contains text information that requests an answer.

[0050] The text format of the answer sentences output by the large-scale visual language model in response to questions was specified as "Yes (or) no, because...". The format of the answer sentences for the English version was specified as "Yes / No, Because...".

[0051] For the learning, the cross entropy error between the correct text and the output text was used as the error function.

[0052] When an image of an object to be estimated in the same field as the image that was pre-trained is input to the large-scale visual language model that has undergone the above-described pre-training, an estimation result about the individual's preferences and the reason for the judgment are output. The large-scale visual language model can output an answer in the text format specified in the pre-training.

[0053] An example of visual inspection of an article based on the method of the embodiment will be described below. In this example, additional learning was performed using images of 37 products collected from the Internet.

[0054] The images of 37 products used for additional training included images of good products and up to five defective mode images. A total of 4,693 images were collected and divided into training data and validation data in an 8:2 ratio. Of the 3,738 training data, 1,834 were used as example images and 1,904 were used as inspection images. The images were expanded to a set of 43,943 images as training data. Input and output question and answer sentences were prepared for each product name and defect name.

[0055] As image data for the items to be visually inspected, MVTec-AD (MVTec Software GmbH), a visual inspection image dataset, was used. From the test data for each item, one image of a good product and one example image of a defective product for each defect mode were selected, and all remaining test data were used as inspection images. The inspection image was changed from the example image to perform the visual inspection. However, if the inspection image was a defective product, an example image of a defective product for the same defect mode was used. The accuracy rate was used as an evaluation index for inspection accuracy. The accuracy rate here was calculated as the percentage of cases where the large-scale visual language model answered "Yes" or "No" to the questions in the inspection image and the answers were correct.

[0056] As a comparative example, a similar appearance inspection was also performed on a large-scale visual language model that had not undergone an additional learning process. As a result, none of the large-scale visual language models that had not undergone an additional learning process were able to answer "Yes" or "No," making quantitative evaluation impossible. For example, the large-scale visual language model that had not undergone an additional learning process output the following answer in response to an input image in which "color abnormality" of a "carpet" was a defect mode: "This is a close-up image of a carpet, but it does not provide enough information to determine if the carpet has any specific defects."

[0057] In contrast, the large-scale visual language model that underwent the additional learning process of the embodiment output answers in a consistent output format. Table 1 shows the accuracy of the appearance inspection results of the large-scale visual language model of the embodiment. Table 1 shows not only the simple judgment of good and bad products, but also the average value calculated for the output rate of correct answers for multiple bad items.

[0058] Table 2 shows, as an example, the results of determining whether a carpet is good or bad for each defect category.

[0059] As shown in Table 1, the large-scale visual language model that underwent the additional training process, particularly for "carpet," "leather," and "wood," achieved highly accurate pass / fail judgments. The image data collected from the Internet for "carpet," "leather," and "wood" used in the additional training process included images similar to those in MVTec-AD. This likely contributed to the more accurate pass / fail judgments. For example, the images of "leather" obtained from an Internet website for use in the additional training process included 26 good images and 52 bad images. Although additional training was not performed using the "leather" images in MVTec-AD, a highly accurate pass / fail judgment was achieved with an accuracy rate of 0.99. (If additional training were performed using images in MVTec-AD, the same or similar images would be used throughout the entire process of pre-training, additional training, and example creation, so a higher accuracy rate would naturally be expected.)

[0060] Furthermore, although images of foreign matter contamination such as "metal contamination" or "loose or contaminated threads" in "carpets," which are unknown defect modes, were not present in the additional training data, a pass / fail judgment could be made, confirming the versatility of the large-scale visual language model of this example.

[0061] On the other hand, for items other than "carpet," "leather," and "wood," the large-scale visual language model output answers biased toward either "Yes" or "No," sometimes failing to adequately determine pass / fail. Furthermore, among the evaluated items, "bottle" and "tile" had low pass / fail accuracy, even though they were items present in the additional training data. This is thought to be because, as shown in Figure 6, the images of "bottle" and "tile" used as input data in the additional training process were significantly different from the images of "bottle" and "tile" taken by MVTec-AD during the actual visual inspection.

[0062] Thus, in order to improve the accuracy of visual inspection through the additional learning process, it is important to input examples of good and defective images of the items to be visually inspected. To confirm the effectiveness of the additional learning process, we compared the visual inspection results for "carpet," "lattice," "leather," and "wood" when good and defective images corresponding to the items to be evaluated were used for additional learning, and when no items corresponding to additional learning were used. For the examples, one good image and one defective image for each defect category were used.

[0063] The comparative evaluation results for "Carpet" are shown in Table 3. The comparative evaluation results for "Grid" are shown in Table 4. The evaluation results for "Leather" are shown in Table 5. The evaluation results for "Wood" are shown in Table 6.

[0064] As shown in Tables 3 to 6, the large-scale visual language model of this example correctly output the results of the visual inspection when examples of good and bad products were input as examples of the items to be visually inspected. Note that the accuracy rate is not shown for the cases with and without examples of good products because the calculation methods are different and a simple comparison is not possible.

[0065] As explained above, it has been confirmed that the large-scale visual language model of this embodiment, by performing a unique additional learning process, can efficiently perform visual inspection with higher accuracy than conventional methods with less learning.

[0066] An example will be described below in which, based on the method of the embodiment, "preferences," which are one of the personal evaluations, are estimated, the basis for the estimation is generated, and the accuracy and quality of the output results are evaluated.

[0067] In this example, the object for which preference estimation is performed is "fried egg." Also, "Otter" is used as the large-scale visual language model.

[0068] In this study, five individuals were selected as specific subjects for which preferences were to be estimated. The same image of a fried egg was used for all specific subjects.

[0069] A pre-training dataset was created according to the flowchart in FIG. 8. In a first step (S1), 30 images of fried eggs were prepared and classified based on the preferences of a specific subject. In a second step (S2), judgments of whether or not each of the 30 images of fried eggs was a preference and the reasons for that judgment were collected and recorded as text data. In a third step (S3), 200 images of fried eggs different from those in the second step (S2) were prepared, and preference judgments of whether or not each fried egg was a preference and the reasons for that judgment were collected from the specific subject. The collected preference judgments and reasons for that judgment were associated with each image as preference information, thereby obtaining a dataset consisting of images and text information.

[0070] The first step (S1) to the third step (S3) were performed for each of the five specific subjects, and 200 sets of data consisting of images and text information were obtained for each specific subject.

[0071] Of the 200 sets of images and preference information, 120 sets were used in the pre-training process of the large-scale visual language model. 30 sets of images and preference information were used for validation. The remaining 50 images that were not used in pre-training and validation were input into the large-scale visual language model to generate preference estimation results and judgment rationales.

[0072] The quality and accuracy of our method were evaluated by comparing the similarity between the preference information associated with the image and the output generated by the large-scale visual language model, assuming it to be the correct answer.

[0073] Because individual preferences are not limited to a single correct answer, evaluating the quality and accuracy of large-scale visual language models is more difficult than in the past. Therefore, to evaluate the accuracy of the estimation results, in addition to the BLEU (Bilingual Evaluation Understudy) score and Rouge-L (Recall-Oriented Understudy for Gisting Evaluation), which are conventionally known document similarity evaluation indices, we performed evaluation using Sentence BERT, as disclosed in Non-Patent Document 22, and human evaluation. Sentence BERT is a natural language processing model that can compare the content of two sentences with high accuracy.

[0074] In the human processing, if the decision-making basis output by the large-scale visual language model contained elements of the correct decision-making basis, even if the preference estimation result was incorrect, a standard was set in which an evaluation score was given based on the proportion of elements of the correct decision-making basis included.If all of the correct decision-making basis elements were included, the score was given as 1, and if none were included, the score was given as 0.

[0075] The accuracy rate of the preference estimation results using the large-scale visual language model, and the quality and accuracy of the judgment basis were evaluated using four methods: BLEU score, ROUGE-L, Sentence BERT, and human evaluation. The results are shown in Table 7 below.

[0076] The evaluation results confirmed that the system was able to generally accurately estimate preference judgments, i.e., whether a specific target person is considered desirable or not. Furthermore, the similarity score based on Sentence BERT for the basis of judgment was extremely high, at around 0.9 for all respondents, confirming that the pre-learning process had generated high-quality basis for judgment. This was supported by the evaluation scores of human evaluations ranging from 0.386 to 0.555, confirming that the system was able to generate correct basis elements for judgment in just under 40% to just over 50% of cases.

[0077] From the above, it has become clear that the method of estimating individual preferences for an estimation target using a large-scale visual language model in this embodiment can generate individual preferences and judgment grounds with high accuracy by pre-training using a set of 120 pairs.

[0078] By applying the visual inspection method for an article of the present invention, it is possible to efficiently distinguish between good and bad articles using a large-scale visual language model.

[0079] By applying the decision-reason-generating estimation method of the present invention, it is possible to estimate an individual's opinion about an object using a large-scale visual language model and improve the product to make it more popular.

[0080] Furthermore, by applying the judgment basis generation type estimation method of the present invention, sensory inspection of items, which has traditionally been performed by workers, can be performed using a large-scale visual language model. By classifying good products as "yes," which indicates a judgment that they are desirable, and bad products as "no," which indicates a judgment that they are undesirable, and creating a data set by associating the judgment basis, it becomes possible to estimate whether an object is good or bad.

Claims

1. A method for performing visual inspection of an object using a large-scale visual language model, comprising: a pre-learning step in which information consisting of images and text is input into the large-scale visual language model to perform pre-learning that takes context into consideration; an additional learning step in which learning information of a group of good products and learning information of a group of defective products for a plurality of objects are input into the large-scale visual language model to perform additional learning; and an inspection step in which an appearance image of an inspection target is input into the large-scale visual language model and inspection results are output, wherein the learning information of the group of good products includes appearance images of good products and text information of the good products, the learning information of the group of defective products includes appearance images of defective products and text information of the defective products, and the information of the inspection target includes an image of the inspection target and text information requesting an inspection result.

2. The appearance inspection method according to claim 1, wherein the text information on the non-defective product includes a standard question asking whether the product is non-defective or defective, and a standard answer indicating that the product is non-defective.

3. The appearance inspection method according to claim 1, characterized in that the text information of the defective product includes a standard question asking whether the product is good or defective, a defect explanation including the name of the defect classification, and a standard answer indicating that the product is defective.

4. A method for estimating an evaluation of an object to be estimated and the reason for the evaluation using a large-scale visual and language model, comprising: a pre-learning step of inputting multiple sets of data, each set being a single image and a sentence including a specific subject's evaluation information regarding the image, into the large-scale visual and language model to perform pre-learning; and a result output step of inputting an image of the object to be estimated into the large-scale visual and language model to output an estimation result regarding the individual's evaluation and the basis for the judgment, characterized in that the evaluation information of the specific subject includes a judgment as to whether the single image is a preference, and text information indicating the basis for the judgment.

5. The estimation method according to claim 4, characterized in that the evaluation information of the specific subjects is collected using the Evaluation Grid Method (registered trademark).

Citation Information

Patent Citations

  • Cross-modal processing for vision and language

    CN115017911A

  • Inspection device, inspection method, and inspection program

    JP2022056389A

  • How to determine damage to vehicle parts

    JP2023524608A

  • Multimodal few-shot learning with frozen language models

    WO2022258666A1

Cited By

  • Information processing system, information processing method, and information processing program

    JP7843097B1

  • Information processing device, information processing method, and program

    JP7892871B1

  • A method for detecting product defects using a VLM agent and a VLM agent using the same.

    JP7894668B1