Inspection system and inspection method

The inspection system improves visual inspection efficiency and accuracy by generating natural language questions about food item placement and number, addressing the inefficiencies of existing methods.

JP2025136431APending Publication Date: 2025-09-19PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Patent Information

Application Number
JP2024035007
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2025-09-19

Smart Images

  • Figure 2025136431000001_ABST
    Figure 2025136431000001_ABST
Patent Text Reader

Abstract

To contribute to a provision of an inspection system and an inspection method capable of improving the efficiency of an appearance inspection.SOLUTION: An inspection system comprises: a question generating section that generates a first question based on reference information indicative of a reference on the number and arrangement of each of a plurality of articles which a reference article has; a question answering section that generates a first answer to the first question about an image of an object-of-inspection article by using a learnt model configured to answer a question of natural language about a content of the image; and a determining section that determines whether or not the object-of-inspection article satisfies a reference indicated by the reference information based on the first answer.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an inspection system and an inspection method. [Background technology]

[0002] When serving meals in school or company cafeterias, or when producing boxed lunches in factories, visual inspections such as food presentation inspections to check whether necessary food items have been placed in predetermined locations, and food omission inspections to check whether necessary food items have been left out, are being considered (see Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-114462 Summary of the Invention [Problem to be solved by the invention]

[0004] However, there is room for improvement in the efficiency of visual inspections, such as inspections of food presentation and inspections for missing food items. Generally, instructions for workers who prepare boxed lunches are written in natural language, and deviations in the number and placement of items are allowed to a certain extent within the scope of interpretation of the instructions. However, the configuration of Patent Document 1 performs visual inspections based on whether the positions and numbers of ingredients fall within preset ranges, which can lead to overly strict judgments regarding deviations in the number and placement. As a result, there is a risk of frequent errors occurring during visual inspections, reducing the efficiency of visual inspections.

[0005] Non-limiting examples of the present disclosure contribute to providing an inspection system and an inspection method that can improve the efficiency of visual inspection. [Means for solving the problem]

[0006] An inspection system according to one embodiment of the present disclosure includes a question generation unit that generates a first question written in natural language and inquiring about the likelihood of the number and / or placement of items in a reference product having a plurality of items, based on reference information indicating criteria for the number and placement of each of the items; a question answering unit that generates a first answer to the first question generated by the question generation unit about an image of an item to be inspected, using a trained model configured to answer questions in natural language about the content of the image; and a determination unit that determines whether the item to be inspected satisfies the criteria indicated by the reference information, based on the first answer.

[0007] In an inspection method according to one embodiment of the present disclosure, an inspection system including at least one information processing device generates a first question written in natural language and asking about the likelihood of the number and / or placement of items based on reference information indicating criteria for the number and placement of each item in a reference item having a plurality of items, generates a first answer to the first question about an image of an item to be inspected using a trained model configured to answer natural language questions about the content of the image, and determines whether the item to be inspected satisfies the criteria indicated by the reference information based on the first answer.

[0008] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]

[0009] According to an embodiment of the present disclosure, it is possible to make the visual inspection more efficient.

[0010] Further advantages and benefits of an embodiment of the present disclosure will become apparent from the specification and drawings. Such advantages and / or benefits may be provided by some of the embodiments and features described in the specification and drawings, respectively, but not necessarily all of them may be provided to obtain one or more identical features. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram showing a configuration example of an inspection system 1 according to an embodiment of the present disclosure. [Figure 2] A diagram showing a first example of the procedure for generating questions. [Figure 3] A diagram showing a second example of how questions are generated. [Figure 4] Diagram showing an example of how to generate multiple questions [Figure 5A] An example of generating questions from multiple sample images. [Figure 5B] FIG. 5B shows a first example of a process for performing the question generation of FIG. 5A. [Figure 5C] FIG. 5B illustrates a second example of a process for performing the question generation of FIG. 5A. [Figure 6] A diagram showing an example of a flow for outputting answers to questions and judgment results based on the answers. [Figure 7A] A diagram showing an example of generating an answer to a question [Figure 7B] A diagram showing a second example of generating an answer to a question. [Figure 8] An example of appearance accuracy judgment [Figure 9A] FIG. 10 shows an example in which it is determined that there is no error. [Figure 9B] FIG. 10 is a diagram showing a first example in which it is determined that an error may exist; [Figure 9C] FIG. 2 shows a second example in which it is determined that an error may exist. [Figure 9D] FIG. 10 is a diagram showing an example in which it is determined that an error exists. [Figure 10A] FIG. 10 is a diagram showing a first example of a method for calculating reliability. [Figure 10B] FIG. 10 is a diagram showing a second example of a method for calculating reliability. [Figure 10C] FIG. 10 is a diagram showing a third example of a method for calculating reliability. [Figure 11] Figure showing the first example of the UI [Figure 12] Figure showing a second example of the UI [Figure 13] Figure showing the third example of the UI [Figure 14] Figure showing the fourth example of the UI [Figure 15] A diagram showing the hardware configuration of a computer that implements the functions of each device through a program. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings as appropriate. However, more detailed explanation than necessary may be omitted. For example, detailed explanation of already well-known matters or redundant explanation of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following explanation and to facilitate understanding by those skilled in the art.

[0013] The accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter described in the claims.

[0014] (One embodiment) When serving meals in school or company cafeterias or when producing boxed lunches in factories, visual inspections such as food presentation inspections to check whether necessary food items have been placed in predetermined locations and food omission inspections to check whether necessary food items have been left out are being considered.

[0015] For example, in visual inspection, it is considered to use image processing to match an image of a sample that serves as a reference having a correct appearance with an image of the object to be inspected in the visual inspection. In matching between images, a comparison of the feature values ​​of each image and / or a comparison of the results of object recognition performed on each image are performed.

[0016] However, since image matching does not allow for a qualitative and / or quantitative definition of errors, it may be difficult to distinguish between small differences between the image of the sample and the image of the inspection object and errors that actually occur in the inspection object. In such cases, for example, a difference in the position of an object that is acceptable between the image of the sample and the image of the inspection object may be judged as an error, which may result in a deterioration in inspection accuracy or a decrease in the efficiency of the visual inspection due to frequent errors.

[0017] Furthermore, for example, in visual inspection, a method is being considered in which the contents of the object to be inspected (for example, the contents of a lunch box) are individually sensed to detect the number and relative positions of the contents of the lunch box, thereby determining whether there has been a mistake.

[0018] However, this method of sensing the contents of a lunch box cannot accurately detect correlations from the meta-information obtained by sensing. For example, if the contents of a lunch box containing spaghetti and a hamburger are sensed, the positional relationship between the spaghetti and the hamburger in the image can be obtained as meta-information. However, the only correlation that can be detected from this meta-information is that the area where the spaghetti is detected and the area where the hamburger is detected are different in the planar direction of the image. In other words, this method cannot distinguish between a case where the hamburger is placed next to the spaghetti and a case where the hamburger is placed on top of the spaghetti. Furthermore, because the logic for determining errors corresponding to differences between the test object and the sample must be manually created, there is a risk of errors being overlooked. Furthermore, even if an error is detected, it is difficult to identify which step in the process caused the error and what improvements should be made.

[0019] Therefore, in the present embodiment described below, a method is described in which questions regarding the number and arrangement of items contained in a sample that serves as a visual reference for visual inspection are generated using language processing, answers to the generated questions are generated for the inspection object of the visual inspection, and the presence or absence of visual errors is determined based on the generated answers. This method makes it possible to detect differences between the inspection object and the sample that should be considered errors, defined in natural language, and therefore makes it possible to distinguish between minor differences between the sample and the inspection object and errors that have occurred in the inspection object, thereby reducing misjudgments and improving the accuracy and efficiency of the visual inspection. Furthermore, questions defined by language processing can automate the judgment of errors corresponding to differences between the inspection object and the sample, thereby reducing or eliminating omissions and errors.

[0020] The following description will be given taking an example of a case where a visual inspection of a lunch box is performed. In this case, the sample for the visual inspection corresponds to a lunch box in which a predetermined number (or a predetermined amount) of food items, etc. are arranged in predetermined locations on the lunch box tray. The inspection target for the visual inspection corresponds to a lunch box produced in a factory or the like by an employee or a robot arranging a plurality of food items, etc. on the lunch box tray. The visual inspection checks whether the plurality of food items, etc. are arranged in predetermined locations on the lunch box tray in the predetermined number (or a predetermined amount) of the inspection target.

[0021] <System configuration example> Fig. 1 is a diagram showing an example of the configuration of an inspection system 1 according to the present embodiment. The inspection system 1 shown in Fig. 1 includes an imaging unit 11, a visual inspection unit 12, a manual database (hereinafter referred to as "manual DB") 13, and a display unit 14.

[0022] The imaging unit 11 captures an image including the inspection object and outputs the captured image to the appearance inspection unit 12. The imaging unit 11 may, for example, capture a still image of the inspection object or a moving image of the inspection object.

[0023] The appearance inspection unit 12 inspects the appearance of the inspection object. For example, the appearance inspection unit 12 determines whether the appearance of the inspection object matches the appearance of a sample and / or the degree of match between the appearance of the inspection object and the appearance of a sample. The appearance inspection unit 12 outputs the determination result to the display unit 14.

[0024] The appearance inspection unit 12 may be configured by, for example, an information processing device (for example, a PC (personal computer) or the like).

[0025] The manual DB 13 stores a manual regarding the appearance of an inspection target. For example, when a person (hereinafter referred to as an operator) performs work to improve the appearance of an inspection target, the manual describes the method, procedure, rules, etc. of the work to improve the appearance in a natural language text format (or document format). When the operator performs work to improve the appearance of the inspection target according to the manual, the appearance of the inspection target matches the appearance of the sample.

[0026] The display unit 14 displays the result of the appearance inspection based on the judgment result output from the appearance inspection unit 12. The display unit 14 is an example of a UI. Display examples on the display unit 14 will be described later.

[0027] The appearance inspection unit 12 includes a question generation unit 121, a question answering unit 122, and an error determination unit 123.

[0028] The question generation unit 121 generates a question from, for example, a manual document. Alternatively, the question generation unit 121 generates a question from a sample image. The question generation unit 121 also generates an expected answer to the generated question. The expected answer corresponds to an answer to the question according to the manual or an answer according to the sample image. After generating a question for one sample, the question generation unit 121 may update (e.g., add or delete) the generated question.

[0029] An example of question generation will be described later.

[0030] The question answering unit 122 answers the question generated by the question generating unit 121. For example, the question answering unit 122 has a VQA (Visual Question Answering) module. The VQA module is a module that receives input of an image and a question written in natural language and outputs an answer to the question. The VQA module can be constructed, for example, by learning pairs of images and questions through machine learning. Note that the details of how to construct a VQA module are known techniques and will not be described in detail here. The question answering unit 122 obtains an answer from the VQA module by inputting an image of the inspection target and the generated question to the VQA module. Also, for example, the question answering unit 122 may perform object detection processing on the image of the inspection target and generate an answer to a question regarding the presence or absence of an object and the positional relationship between multiple objects based on the detected object and the positional relationship between the object. Note that specific examples of methods for answering questions will be described later.

[0031] The error determination unit 123 determines whether the appearance is correct or incorrect based on the actual answer to the question. The error determination unit 123 determines whether there is an error based on whether the actual answer to the question is inconsistent with the expected answer. The error determination unit 123 outputs the determination result, including the presence or absence of an error, to the display unit 14. Specific examples of determining whether the appearance is correct or incorrect will be described later.

[0032] The inspection system 1 described above may be realized by a configuration including at least one information processing device (e.g., a personal computer). For example, the appearance inspection unit 12 is included in at least one information processing device. Alternatively, the inspection system 1 may have a configuration in which the question generation unit 121 of the appearance inspection unit 12 is included in a first information processing device, and the question answering unit 122 and the error determination unit 123 are included in a second information processing device, or the question answering unit 122 and the error determination unit 123 may be included in separate information processing devices.

[0033] As described above, the inspection system according to the present embodiment executes a process of generating questions, a process of generating answers to the generated questions, and a process of making a determination based on the answers. Each of these processes will be described below.

[0034] <Question generation example> An example of generating a question by the question generator 121 will be described below. In the following, the inspection object is a grilled meat bento, and the appearance inspection of the inspection object corresponds to an inspection of whether the food items, etc., served in the grilled meat bento are arranged in predetermined locations on the tray of the bento, in a predetermined number (or a predetermined amount).

[0035] Figure 2 is a diagram showing a first example of a procedure for generating questions. Figure 2 illustrates the flow when questions are generated from a manual. Note that Figure 2 shows only four procedures and seven questions.

[0036] The manual in Figure 2 is a document in which the procedures for arranging Yakiniku bento are written in natural language. When employees arrange food items according to the manual, error-free bento are produced. The manual is an example of standard information that indicates the standards for the number and arrangement of each food item in a sample Yakiniku bento. An item that is completed by arranging multiple items in the number and positions according to the standard information is also called a "standard product." The Yakiniku bento described above is an example of a standard product.

[0037] The question generation unit 121 divides the manual procedures and processes them using Large Language Models (LLMs). As a result of the LLM processing, multiple questions and expected answers to the questions are generated. Specifically, by inputting a prompt to the LLM, such as "The given text is part of the manual. Please generate one or more questions and answers to confirm whether this manual has been followed correctly," multiple questions and expected answers are generated based on the given text. Note that the accuracy of the expected answers to the questions may be confirmed visually by the user, for example. Furthermore, even if there is a risk of inaccurate expected answers being included, the accuracy of the visual inspection can be guaranteed as long as most of the expected answers are accurate. Therefore, the user may limit confirmation to some questions, or may not confirm at all. FIG. 2 shows the generated multiple questions, expected answers, and manual procedures corresponding to the questions, in association with each other.

[0038] In the present embodiment, generating questions may include generating questions and generating expected answers to the questions. The generated questions may be referred to as a question list or a question group hereinafter.

[0039] The expected answer is the answer that follows the manual. For example, the expected answer "YES" to the question "Are yakiniku bento trays being used?" follows step 1.

[0040] Fig. 3 is a diagram showing a second example of a procedure for generating questions. Fig. 3 illustrates a flow when questions are generated from a sample image. Seven questions are selected and shown in Fig. 3.

[0041] The sample image in FIG. 3 is an image of a sample of a yakiniku bento box. The sample image is an example of reference information indicating standards regarding the number and arrangement of each food item in the sample of a yakiniku bento box. The question generation unit 121 performs image captioning on the sample image. Image captioning uses the image as input to generate a caption text that describes events occurring in the image, the behavior of people, etc. For example, image captioning recognizes each element in the image using an image recognition model, and then generates a caption using a language model based on the image's features and / or identified labels. In the example of FIG. 3, the question generation unit 121 performs image captioning on the sample image and converts the sample image into natural language text that expresses the features of the sample image (e.g., appearance features). The question generation unit 121 processes the converted text using LLM. As a result of the LLM processing, multiple questions and expected answers to the questions are generated.

[0042] The question generator 121 may generate questions based on a manual as shown in Fig. 2, or may generate questions based on a sample image as shown in Fig. 3. Note that the question generator 121 may generate questions based on both the manual and the sample image, or may generate questions based on information different from the manual and the sample image.

[0043] <Variation to generate questions for false positive check> The question generator 121 may generate a question for checking for misjudgment for a generated question. A question for checking for misjudgment for a certain question X is a question that asks about the matter asked in question X in a different way or from a different perspective than question X. A question for checking for misjudgment may be considered as a question for double-checking. For example, the question generator 121 may again instruct the LLM to generate a question for checking for misjudgment.

[0044] Fig. 4 shows an example of a method for generating multiple questions. Fig. 4 shows an example in which multiple questions for checking for incorrect judgment are generated from one question A, "Is the dried daikon radish to the left of the rice?"

[0045] For example, the question generator 121 generates a redundant question corresponding to a superordinate concept or a subordinate concept of a certain matter (for example, a certain question).

[0046] In the example shown in Figure 4, question B, "Is kiriboshi daikon radish available?" is generated as a superordinate question of question A, "Is kiriboshi daikon radish to the left of the rice?". For example, as shown in Figure 4, question A is input to the LLM, and an instruction to "generate a confirmation statement as to whether the target object exists" is input to the LLM as a prompt, thereby generating question B, which corresponds to a confirmation statement as to whether the "kiriboshi daikon radish" that is the target of question A exists.

[0047] Here, if the answer to question A is YES, it is assumed that the answer to question B is also YES. By including both question A and question B in the question list, it is possible to detect errors. For example, if the answer to question A, "Is the kiriboshi daikon to the left of the rice?" is YES, and the answer to question B, "Is there kiriboshi daikon?" is NO, there is a contradiction between the two answers, so it can be determined that an error has occurred or that the reliability of the answers to questions A and B is low.

[0048] In this way, by generating a question with a higher or lower concept for a certain matter (for example, a certain question), errors in the target matter can be detected with high accuracy, or the reliability of the target matter can be calculated accurately.

[0049] For example, the question generator 121 generates a question that expresses a certain matter (for example, a certain question) in a different form (for example, from a different point of view or a different viewpoint).

[0050] In the example shown in Figure 4, question A, "Is the dried daikon radish to the left of the rice?", is expressed in a different form to generate question C, "Is the rice to the right of the dried daikon radish?". Illustratively, as shown in Figure 4, question A is input to the LLM, and an instruction is input to the LLM as a prompt, saying, "Please convert it into a sentence that is equivalent in meaning but has a different positional relationship," and question C, which is equivalent in meaning to question A but has a different positional relationship, is generated.

[0051] For example, if the answers to questions A and C are both YES, it is determined that no error exists, and if the answers to questions A and C are both NO, it is determined that an error exists. Also, for example, if one of the answers to questions A and C is YES and the other is NO, it is determined that the reliability of the answers to questions A and C is low.

[0052] In this way, by generating multiple questions that express a certain matter (e.g., a certain question) in different ways (e.g., viewpoints or perspectives), errors in the target matter can be detected with high accuracy, or the reliability of the target matter can be calculated accurately.

[0053] For example, the question generation unit 121 generates closed questions that require a response from multiple options, such as a binary choice of "YES / NO" or a three-way choice of "A / B / C," and open questions that require a response to the question in sentences and / or words.

[0054] For example, the question generation unit 121 generates an open-ended question E, "How many fried shrimp are there?" based on a closed-ended question D, "Are there two fried shrimp?", which is generated based on the item "Place two fried shrimp" described in the manual. For example, if the answer to the closed-ended question D is YES and the answer to the open-ended question E is "two," it is determined that there is no contradiction between the two answers and therefore no error. For example, if the answer to the closed-ended question D is YES and the answer to the open-ended question E is "three," it is determined that there is a contradiction between the two answers and therefore an error exists or the reliability of the two answers is low. Note that the absence of a contradiction between the two answers may be equivalent to the two answers having the same meaning, and the existence of a contradiction between the two answers may be equivalent to the two answers not having the same meaning.

[0055] By generating multiple questions with different perspectives on the same subject, the same subject can be checked multiple times, making it possible to detect errors in the subject with high accuracy, or to accurately calculate the reliability of the subject.

[0056] The question generation unit 121 may generate a question based on a sample image of a target for which a question is to be generated (hereinafter referred to as the question generation target) and an image of a product different from the question generation target. For example, the product different from the question generation target may be at least one of a product similar to the question generation target, a product that is easily mistaken for the question generation target, a product manufactured in the same place as the question generation target, or a product manufactured by the same person who manufactures the question generation target.

[0057] The question generator 121 may extract differences between a sample image for which a question is to be generated and an image of a finished product that is different from the sample image for which a question is to be generated, and generate a question for confirming the extracted differences.

[0058] 5A is a diagram showing an example of generating a question from multiple images. Fig. 5A shows an example of generating a question when the target of the question generation is a pork cutlet curry bento and the product different from the target of the question generation is a chicken cutlet curry bento.

[0059] In the example of FIG. 5A , the question generator 121 extracts the difference between a sample image of a pork cutlet curry bento and an image of a finished chicken cutlet curry bento. For example, the question generator 121 may extract the difference between two images by performing object detection on the images, or by inputting the two images into a VQA module. In the example of FIG. 5A , the difference extracted is whether or not tartar sauce is on top of the cutlet. Then, to confirm this extracted difference, the question generator 121 generates a question such as “Is there tartar sauce on top of the cutlet?” and determines that the expected answer to the generated question is “NO.” Note that when extracting the difference using a VQA module, the question “Is there tartar sauce on top of the cutlet?” can be automatically generated by, for example, preparing a prompt in advance such as “Please generate a question that asks about the difference output by the VQA module” and querying the LLM after obtaining the output from the VQA module. Also, when extracting differences by performing object detection on images, the question "Is there tartar sauce on top of the cutlet?" can be automatically generated by preparing a prompt in advance, for example, "Please generate a question that asks about the differences obtained by object detection," and then querying the LLM after the differences are obtained by object detection.

[0060] If the question generation target is a chicken cutlet curry bento and the item different from the question generation target is a pork cutlet curry bento, the question "Is there tartar sauce on top of the cutlet?" may be generated to confirm the difference, i.e., whether or not there is tartar sauce on top of the cutlet, and the expected answer to the generated question may be determined to be "YES." In this way, questions about each of the items in the two images may be generated based on the difference between the two images.

[0061] Fig. 5B is a diagram showing a first example of the process for generating questions in Fig. 5A. Similar to Fig. 5A, Fig. 5B shows an example of generating questions when the question generation target is a pork cutlet curry bento and the item different from the question generation target is a chicken cutlet curry bento. In Fig. 5B, object detection is performed on the images to extract the differences between the two images.

[0062] In FIG. 5B, foods corresponding to curry roux, cutlet, and pickles are detected in the images of the pork cutlet curry bento and the chicken cutlet curry bento. Also, in the example of FIG. 5B, an object corresponding to tartar sauce is detected in the image of the chicken cutlet curry bento, but not in the image of the pork cutlet curry bento. The object corresponding to tartar sauce, which is detected only in the image of the chicken cutlet curry bento, is extracted as the difference between the two images. Then, to confirm this extracted difference, the question generator 121 generates a question such as "Is there tartar sauce on top of the cutlet?" and determines that the expected answer to the generated question is "NO."

[0063] Figure 5C is a diagram showing a second example of the process for generating questions in Figure 5A. Similar to Figure 5A, Figure 5C shows an example of generating questions when the target of question generation is a pork cutlet curry bento and the item different from the target of question generation is a chicken cutlet curry bento. In Figure 5C, two images and the query "What is the difference between the two images?" are input to the VQA module, and the differences between the two images are extracted.

[0064] 5C, the VQA module outputs a difference indicating that there is no tartar sauce. To confirm this difference, the question generator 121 generates a question such as "Is there tartar sauce on the cutlet?" and determines that the expected answer to the generated question is "NO."

[0065] <Question and answer flow> FIG. 6 is a diagram showing an example of a flow for outputting answers to questions and determination results based on the answers.

[0066] The question and answering section 122 acquires an image (still image or video) from the imaging section 11 (ST01). For example, when the question and answering section 122 acquires a video, it may convert it into a still image at a specific frame rate.

[0067] The question answering unit 122 generates answers to each question based on the acquired image and the set of questions generated by the question generation unit 121 (ST02). The question answering unit 122 answers with a binary value of YES or NO to questions that require a "YES or NO" answer, answers the number of applicable items to questions that require a number answer, and answers the positional relationship of applicable items to questions that require a positional relationship of two or more items. An example of answer generation will be described later.

[0068] The question answering unit 122 calculates the reliability (ST03). For example, the reliability may be a score value output by a model (for example, an answer generation model described later). The score is a value indicating how accurate the VQA module evaluates the answer generated by the VQA module to be. Alternatively, the question answering unit 122 may calculate the reliability by referring to the results of a judgment model such as an object detector. Furthermore, when double-checking questions are generated, the question answering unit 122 may calculate the reliability based on whether or not there is a contradiction between the answers to the double-checking questions. Furthermore, the reliability may be calculated using results obtained by an image recognition method. For example, the image recognition method may be a method that matches the features of multiple images or that uses an object detector. An example of calculating the reliability will be described later.

[0069] Note that the reliability does not have to be calculated. For example, if the reliability is not used in error determination section 123 (described later), the reliability does not have to be calculated. In this case, ST03 may be omitted.

[0070] The error determination unit 123 determines whether the appearance is correct or incorrect (ST04). For example, the error determination unit 123 determines whether the appearance is correct or incorrect (whether there is an error in the appearance) based on the answer to the question. Alternatively, the error determination unit 123 may determine whether the appearance is correct or incorrect (whether there is an error in the appearance) based on the answer to the question and the reliability of the answer.

[0071] For example, if all the actual answers to each question in a question group about a certain test object are the same as the expected answers, it is determined that there is no error in the appearance of the test object.Also, if at least N actual answers (N is an integer greater than or equal to 1) to each question in a question group about a certain test object are not the same as the expected answers, it is determined that there is an error in the appearance of the test object.

[0072] Note that the reliability of each answer may be used in this determination. For example, if all actual answers to a question are the same as the expected answers and the reliability of all actual answers is equal to or greater than a threshold, it is determined that there is no error in the appearance of the test object. For example, if all actual answers to a question are the same as the expected answers but the reliability of at least some (e.g., at least one) of the actual answers is less than a threshold, it is determined that there may be an error in the appearance of the test object. For example, if some (e.g., at least one) of the actual answers to a question are different from the expected answers, the remaining actual answers are the same as the expected answers, and the reliability of the partial actual answers is less than a threshold, it is determined that there may be an error in the appearance of the test object. For example, if some (e.g., at least one) of the actual answers to a question are different from the expected answers, the remaining actual answers are the same as the expected answers, and the reliability of the partial actual answers is equal to or greater than a threshold, it is determined that there is an error in the appearance of the test object.

[0073] Note that the determination of the accuracy of the appearance is not limited to the above-mentioned examples. For example, the determination of whether there is a match between the actual answer and the expected answer and / or the determination using the reliability are not limited to the above-mentioned examples. For example, the user of the inspection system 1 may adjust which answers to questions are used for the determination, the reliability of the answers to which questions are used for the determination, etc. Note that examples of the determination of the accuracy of the appearance will be described later.

[0074] The error determination unit 123 generates a message (ST05). The generated message may include at least the result of the appearance accuracy determination, an image of the test object, and a combination of the question and answer used in the appearance accuracy determination. The message is output to the display unit 14. The display unit 14 performs display based on the message. An example of the display on the display unit 14 will be described later.

[0075] <Answer generation example> The question answering unit 122 generates answers to the respective questions based on the acquired images and the set of questions generated by the question generating unit 121.

[0076] 7A is a diagram showing an example of generating answers to questions, in which an answer is given to each question in a group of questions based on an image of a grilled meat lunch box to be inspected.

[0077] As shown in FIG. 7A, the answer generation unit 122 inputs images of the yakiniku bento box to be inspected and a set of questions into a VQA (Visual Question Answering) module, and obtains answers to questions about the inspection target and the reliability of the answers. In the example of FIG. 7A, the reliability is expressed as a number ranging from 0 to 100, with "100" indicating the highest reliability. The VQA module is an example of a trained model configured to answer questions in natural language about the content of the image.

[0078] Note that instead of using a VQA module, object detection may be used to generate answers.

[0079] 7B is a diagram showing a second example of generating an answer to a question. In FIG. 7B, an actual answer to the question is generated by performing object detection on an image of a grilled meat bento box to be inspected.

[0080] In FIG. 7B, rice, grilled meat, boiled greens, and dried strips of daikon radish are detected. The answer generation unit 122 generates an answer to the question based on the detection result. For example, the answer generation unit 122 generates an answer of "YES" to the question "Is rice served?" based on the detection result that rice has been detected. Furthermore, for example, the answer generation unit 122 generates an answer of "YES" to the question "Is grilled meat served to the left of the rice?" based on the detection result that rice and grilled meat have been detected and that the grilled meat is located to the left of the rice.

[0081] Note that the answer to the question may be generated by combining the VQA module shown in FIG. 7A with the object detection module shown in FIG. 7B.

[0082] For example, object detection may be performed on the portion of the lunch box where side dishes are present (e.g., the portion from the center to the left of the image of the yakiniku bento in FIG. 7B). An answer as to whether the arrangement of food items in the lunch box, including the side dishes, is correct is then generated using a VQA module. For objects that are more accurately detected by learning individually, such as the presence or absence of side dishes, object detection improves detection accuracy, thereby improving the accuracy (e.g., reliability) of the answer to the question. Furthermore, using the results of object detection and the results of the VQA module enables double-checking, improving the accuracy of error detection.

[0083] <Example of appearance accuracy judgment> The error determining unit 123 uses the answer to the question generated by the answer generating unit 122 to determine whether the appearance is correct or incorrect.

[0084] For example, if an actual answer to a certain question matches an expected answer, the actual answer is determined to be the expected answer. If an actual answer to a certain question does not match the expected answer, the actual answer is determined to be an answer that differs from the expected answer. Note that the actual answer matching the expected answer may correspond to the actual answer not contradicting the expected answer, and the actual answer not matching the expected answer may correspond to the actual answer contradicting the expected answer.

[0085] Fig. 8 is a diagram showing an example of determining whether an appearance is correct or incorrect. Fig. 8 shows, in a table format, a question, an expected answer to the question, an actual answer for test object α, and an actual answer for test object β. Note that the actual answer for test object α is the actual answer to the question when the question is asked about test object α.

[0086] In the example shown in Figure 8, all of the actual answers for test object α are as expected, so test object α is determined to have no errors. Also, in the example shown in Figure 8, among the actual answers for test object β, the answer to the question "Is the dried daikon radish to the left of the rice?" is an answer that differs from the expected answer. In this way, if there is even one answer that differs from the expected answer, test object β is determined to have an error.

[0087] 8 shows an example in which the determination is made based on the actual answer, but the reliability of the actual answer may be calculated and used to determine whether the appearance is correct. In this case, the error determination unit 123 determines whether the appearance is correct or incorrect using the answer to the question generated by the answer generation unit 122 and the reliability of the answer. The calculation of the reliability will be described later.

[0088] For example, if the actual answer to a question matches the expected answer, the actual answer is determined to be the expected answer. If the actual answer to a question does not match the expected answer, the actual answer is determined to be a different answer from the expected answer.

[0089] For example, if an actual answer to a certain question matches an expected answer and the reliability of the actual answer is equal to or greater than a first threshold, the actual answer is determined to be as expected and have a high reliability. For example, if an actual answer to a certain question matches an expected answer and the reliability of the actual answer is less than a first threshold, the actual answer is determined to be as expected and have a low reliability. For example, if an actual answer to a certain question does not match an expected answer and the reliability of the actual answer is less than a first threshold, the actual answer is determined to be different from the expected answer and have a low reliability. For example, if an actual answer to a certain question does not match an expected answer and the reliability of the actual answer is equal to or greater than a second threshold, the actual answer is determined to be different from the expected answer and have a high reliability. Note that the second threshold may be smaller than the first threshold or may be equal to or greater than the first threshold.

[0090] For example, the answers to each of the question groups may be classified into the four categories mentioned above, namely, “as expected and highly reliable,” “as expected and low reliability,” “unexpected and low reliability,” and “unexpected and highly reliable.” Then, the error determination unit 123 may determine the presence or absence of an error based on the number of answers included in the four categories.

[0091] For example, if the number of answers in the category "Unlikely expected and highly reliable" is L or more (for example, L is an integer equal to or greater than 1), it is determined that an error exists in the test object. In this case, the determination result includes information indicating the existence of an error (for example, information indicating "NG").

[0092] For example, even if the number of answers in the category "Unlikely expected and highly reliable" is less than L, if the number of answers in the category "Unlikely expected and low reliability" is M or more (M is an integer greater than or equal to 1), it is determined that there may be an error in the test object. In this case, the determination result includes information indicating a warning because there may be an error.

[0093] For example, if the number of answers in the category "unexpected and highly reliable" is less than L, and the number of answers in the category "unexpected and low reliability" is less than M, but the number of answers in the category "as expected and low reliability" is N or more (N is an integer equal to or greater than 1), it is determined that there may be an error in the test object. In this case, the determination result includes information indicating a warning because there may be an error.

[0094] For example, if the number of answers in the category "unexpected and highly reliable" is less than L, the number of answers in the category "unexpected and low reliability" is less than M, and the number of answers in the category "as expected and low reliability" is less than N, it is determined that there is no error in the test object. In this case, the determination result includes information indicating that there is no error (for example, information indicating "OK").

[0095] Below, an example of the determination made by the error determination unit 123 when the first threshold and the second threshold are 80 and L = M = N = 1 will be described. Note that by adjusting the values ​​of L, M, and N according to user input, etc., it is possible to adjust the degree of error to be tolerated and the category of error to be tolerated.

[0096] 9A is a diagram showing an example in which it is determined that there is no error, in which a question, an expected answer to the question, an actual answer to the question, and the reliability of the actual answer are shown in a table format.

[0097] In the example of Fig. 9A, the actual answers to the questions each match the expected answers, and the reliability of the actual answers is 80 or higher. In other words, all of the actual answers to the questions are classified into the category of "as expected and highly reliable." In this case, the error determination unit 123 determines that there is no error in the test target ("OK" in the example of Fig. 9A).

[0098] 9B is a diagram showing a first example in which it is determined that an error may exist. Similar to FIG. 9A, FIG. 9B shows a question, a possible answer to the question, an actual answer to the question, and the reliability of the actual answer in a table format.

[0099] In the example of FIG. 9B, the actual answers to the questions each match the expected answers, but the reliability of the actual answer to the question "Is the dried daikon radish to the left of the rice?" is less than 80. In other words, the actual answer to the question "Is the dried daikon radish to the left of the rice?" is classified into the category "as expected and low reliability", and the actual answers to the remaining questions are classified into the category "as expected and high reliability". In this case, since the number of answers in the category "as expected and low reliability" is 1, the error determination unit 123 determines that there is a possibility that an error exists in the test target ("Warning" in the example of FIG. 9B).

[0100] 9C is a diagram showing a second example in which it is determined that an error may exist. Similar to FIG. 9A, FIG. 9C shows a question, a possible answer to the question, an actual answer to the question, and the reliability of the actual answer in a table format.

[0101] In the example of FIG. 9C, the actual answer to the question "Is the dried daikon radish to the left of the rice?" does not match the expected answer, and the reliability of the actual answer is less than 80. In other words, the actual answer to the question "Is the dried daikon radish to the left of the rice?" is classified into the category "different from expected and low reliability", and the actual answers to the remaining questions are classified into the category "as expected and high reliability". In this case, since the number of answers in the category "different from expected and low reliability" is 1, the error determination unit 123 determines that there is a possibility that an error exists in the test target ("Warning" in the example of FIG. 9C).

[0102] 9D is a diagram showing an example in which it is determined that an error exists. Similar to FIG. 9A, FIG. 9D shows a question, an expected answer to the question, an actual answer to the question, and the reliability of the actual answer in a table format.

[0103] In the example of FIG. 9D, the actual answer to the question "Is the dried daikon radish to the left of the rice?" does not match the expected answer, and the reliability of the actual answer is 80 or more. In other words, the actual answer to the question "Is the dried daikon radish to the left of the rice?" is classified into the category "different from expected and highly reliable", and the actual answers to the remaining questions are classified into the category "as expected and highly reliable". In this case, since the number of answers in the category "different from expected and highly reliable" is 1, the error determination unit 123 determines that an error exists in the test target ("NG" in the example of FIG. 9D).

[0104] 8 and 9A to 9D, the error determination unit 123 determines whether or not an error exists in the test object based on the answers to questions about the test object, and therefore, if an error exists in the test object, the cause of the error can be identified (or estimated) based on the questions for which the expected answer does not match the actual answer. This allows feedback to be provided to the user (for example, an employee of a lunch box manufacturing factory) in a form that allows the user to understand the cause of the error.

[0105] As shown in Figures 9A to 9D, by judging the appearance using the reliability, it is possible to perform a more accurate appearance inspection. In particular, even if the expected answer and the actual answer match, if the reliability of the actual answer is low, a warning can be issued, which prevents errors from being overlooked. Furthermore, even if the expected answer and the actual answer do not match, if the reliability of the actual answer is low, a warning can be issued, which prevents erroneous judgments in which an error is detected even when there is no error.

[0106] For example, inspection targets with a NG result can be manually or automatically excluded. On the other hand, for inspection targets with a Warning result, employees can visually check for errors and manually adjust them as necessary.

[0107] <Example of reliability calculation> The question answering unit 122 may calculate the reliability. For example, the reliability may be calculated using a score value output by an answer generation model. The answer generation model is included in the question answering unit 122 and has a function of generating an answer to a question. The answer generation model may include, for example, answer generation using the VQA module shown in FIG. 7A or answer generation using object detection shown in FIG. 7B.

[0108] FIG. 10A is a diagram showing a first example of a method for calculating reliability. FIG. 10A shows an example in which an answer generation model generates an answer to a question and also calculates a score. In the example shown in FIG. 10A, an image of a certain object and the question "What is this?" are input to the answer generation model. In response to the input question, the answer generation model outputs an answer stating that candidates for the object in the image are an egg, a golf ball, and a carrot, and that the reliability of each candidate is 60%, 30%, and 10%, respectively. In this case, the answer generation unit 122 determines that the candidate with the highest reliability is the object in the image, generates an answer that is "egg" to the question "What is this?", and determines that the reliability of this answer is 60.

[0109] Although the above describes an example in which the answer generation model generates an answer and a confidence level, the present disclosure is not limited to this. The answer generation model may generate an answer, and an object identification model separate from the answer generation model may output a score, and the output score may be used as the confidence level.

[0110] FIG. 10B is a diagram showing a second example of a method for calculating reliability. FIG. 10B is an example in which an answer generation model generates an answer to a question while a score is calculated by an object recognition model. In the example shown in FIG. 10B, an image of a certain object and a question "What is this?" are input into the answer generation model. The answer generation model generates an answer that the object in the image is an egg for the input question. Then, the answer that the object in the image is an egg and the image of the object taken are input into the object recognition model. The object recognition model determines that the score indicating the similarity between the object and an egg is 80%. In this case, the answer generation unit 122 generates an answer "egg" for the question "What is this?", and determines that the reliability of the answer is 80.

[0111] When a plurality of questions are generated for the same matter, the reliability of the matter may be calculated based on the answers to the plurality of questions.

[0112] FIG. 10C is a diagram showing a third example of a method for calculating reliability. FIG. 10C shows five questions Q1 to Q5, assumed answers to the five questions, and actual answers. In FIG. 10C, for one question Q1 "Is a fried shrimp placed on the right side of the hamburger?", questions Q2 to Q5 are generated as non-contradictory questions. In the case of FIG. 10C, since the number of actual answers that match the assumed answer is 3 and the number of actual answers that do not match the assumed answer is 2, the ratio of the actual answers that match the assumed answer is 60%. In this case, the answer generation unit 122 may determine that the reliability that the answer to question Q1 is "YES" is 60%.

[0113] <Example of UI (user interface)> Next, a display example of the display unit 14, which is an example of a UI, will be described. The display unit 14 acquires a message including a determination result from the error determination unit 123 of the visual inspection unit 12, and performs display based on the message. The message includes information on the determination result indicating whether or not an error exists in a certain inspection object (for example, information indicating any of "OK," "NG," and "Warning"). The message may also include information on the reason for the determination result. The information on the reason for the determination result includes information such as a question where the expected answer differed from the actual answer, a manual procedure associated with the question, etc. The message may also include information on the inspection object (for example, an image of the inspection object, identification information for identifying the inspection object).

[0114] Fig. 11 is a diagram showing a first example of the UI. The example in Fig. 11 is a display for an inspection object that has been determined to be error-free by the error determination unit 123. In Fig. 11, an image of the photographed inspection object is displayed, along with text information indicating that the inspection object is error-free ("OK" in the example in Fig. 11).

[0115] Fig. 12 is a diagram showing a second example of the UI. The example in Fig. 12 is a display of an inspection object that has been determined to have an error by the error determination unit 123. In Fig. 12, a photographed image of the inspection object is displayed, along with text information indicating that the inspection object has an error ("NG" in the example in Fig. 12). Also, in Fig. 12, text information indicating the error, the assumed reason for the error, and text information indicating how to correct the error are displayed. The text information indicating the assumed reason for the error and how to correct the error are included in the message acquired by the display unit 14.

[0116] Fig. 13 is a diagram showing a third example of the UI. The example in Fig. 13 is a display for an inspection object that has been determined by the error determination unit 123 to possibly contain an error. In Fig. 13, a photographed image of the inspection object is displayed, along with text information indicating that the inspection object may contain an error ("Warning" in the example in Fig. 13). Also, in Fig. 13, text information urging the user to check the error, text information indicating the assumed reason for the error, and text information indicating how to correct the error are displayed. The text information urging the user to check the error, text information indicating the assumed reason for the error, and text information indicating how to correct the error are included in the message acquired by the display unit 14.

[0117] In addition to the display for checking whether there are any errors, the UI (for example, the display unit 14) may also display the process of determining whether there are any errors. The process of determining whether there are any errors may include, for example, information indicating up to which step in the manual it was determined that there were no errors, or information indicating which answer to which question was different from the expected answer.

[0118] FIG. 14 is a diagram illustrating a fourth example of the UI. The example in FIG. 14 is a display for an inspection object determined by the error determination unit 123 to contain an error. FIG. 14 includes a first display and a second display. The first display in FIG. 14 displays an image of the inspection object captured and text information indicating that the inspection object contains an error ("NG" in the example in FIG. 14). The first display in FIG. 14 also displays text information indicating the error, a presumed reason for the error, and text information indicating how to correct the error. The first display in FIG. 14 also displays an "Analyze" button. When a user (e.g., an employee of a lunch box manufacturing factory) operates an operation unit (e.g., a touch panel, a mouse, etc.) connected to the display unit 14 to press the "Analyze" button, the display unit 14, which has acquired operation information indicating that the "Analyze" button has been pressed, switches from the first display to the second display. The second display in FIG. 14 displays a question posed about the inspection object, an expected answer to the question, an actual answer, and the reliability of the actual answer. By checking the second display, the user can understand the cause of the error.

[0119] 11 to 14 are merely examples, and the present disclosure is not limited to these. For example, instead of displaying the determination result as text information, the determination result may be shown to the user by other means such as sound or the illumination of a warning light or the like.

[0120] As described above, in the inspection system 1 according to this embodiment, the question generator 121 generates a question written in natural language, asking about the accuracy of the number and / or arrangement of items, based on a manual and / or sample image indicating criteria for the number and arrangement of each food item in a sample of a yakiniku bento box containing multiple foods. The question answerer 122 then generates an answer to the question generated by the question generator 121 about the image of the inspection target, using a trained model (e.g., a VQA module) configured to answer questions in natural language about the content of the image. The error determiner 123 determines, based on the answer, whether the inspection target satisfies the criteria indicated by the manual and / or sample image.

[0121] According to the inspection system 1 of the above-described embodiment, differences between the inspection object and the sample that should be considered errors, as defined in natural language, can be detected, and therefore it is possible to distinguish between small differences between the sample and the inspection object and errors that have occurred in the inspection object, reducing misjudgments and improving the accuracy and efficiency of the visual inspection. Furthermore, questions defined by language processing can automate the judgment of errors corresponding to differences between the inspection object and the sample, reducing or eliminating omissions and errors, improving the accuracy and efficiency of the visual inspection.

[0122] In the present embodiment, an example has been given in which the reference product to be inspected is a bento lunch (e.g., a yakiniku bento lunch), but the present disclosure is not limited to this. For example, the present disclosure may be applied to inspecting the appearance, such as the number and arrangement, of multiple different types of items in an item in which multiple different types of items are placed in a single outer box (e.g., a juice assortment, a sweets assortment, etc.). Also, for example, in a manufacturing site, when assembling a single product from multiple parts, necessary parts are placed on a tray before assembly to prevent mistakes such as forgetting to install a screw. In such a case, the present disclosure may be applied to inspecting the number of parts placed on the tray by a worker.

[0123] Furthermore, in this embodiment, both the number and arrangement of items in the standard product to be inspected are confirmed, but it is also possible to configure the system to check only one of them. For example, in the case of cooking sets that are produced in a central kitchen and then plated at each restaurant, or in meals provided to hospitalized patients, accuracy in the number and amount of food is required, but accuracy in arrangement may not be required. Furthermore, if the color, etc., of the items when they are arranged is important, accuracy in the number of items may not be required.

[0124] Furthermore, the notation "... section" in the above-described embodiments may be replaced with other notations such as "... circuitry," "... assembly," "... device," "... unit," or "... module."

[0125] The embodiments of the present disclosure have been described above in detail with reference to the drawings, but the functions of the above-described product management system 1 can be realized by a computer program.

[0126] 15 is a diagram showing the hardware configuration of a computer that realizes the functions of each device by a program. This computer 1100 includes an input device 1101 such as a keyboard, a mouse, or a touchpad, an output device 1102 such as a display or a speaker, a central processing unit (CPU) 1103, a graphics processing unit (GPU) 1104, a read only memory (ROM) 1105, a random access memory (RAM) 1106, a storage device 1107 such as a hard disk drive or a solid state drive (SSD), a reading device 1108 that reads information from a recording medium such as a digital versatile disk read only memory (DVD-ROM) or a universal serial bus (USB) memory, and a transmitting / receiving device 1109 that communicates via a network, and each unit is connected by a bus 1110.

[0127] The reading device 1108 then reads the program for realizing the functions of each of the above-mentioned devices from a recording medium on which the program is recorded, and stores the program in the storage device 1107. Alternatively, the transmitting / receiving device 1109 communicates with a server device connected to the network, and stores the program for realizing the functions of each of the above-mentioned devices downloaded from the server device in the storage device 1107.

[0128] The CPU 1103 then copies the program stored in the storage device 1107 to the RAM 1106, and sequentially reads out and executes instructions contained in the program from the RAM 1106, thereby realizing the functions of the above-mentioned devices.

[0129] The present disclosure can be realized in software, hardware, or software in conjunction with hardware.

[0130] Each functional block used in the description of the above embodiments may be partially or entirely realized as an LSI, which is an integrated circuit, and each process described in the above embodiments may be partially or entirely controlled by a single LSI or a combination of LSIs. The LSI may be composed of individual chips, or may be composed of a single chip that includes some or all of the functional blocks. The LSI may have data input and output. Depending on the degree of integration, the LSI may be called an IC, system LSI, super LSI, or ultra LSI.

[0131] The integrated circuit method is not limited to LSI, but may be realized by a dedicated circuit, a general-purpose processor, or a dedicated processor. Also, a field programmable gate array (FPGA) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells within the LSI, may be used. The present disclosure may be realized as digital processing or analog processing.

[0132] Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology or other derivative technologies, it is natural that such technology may be used to integrate functional blocks. The application of biotechnology, etc. is also a possibility.

[0133] The present disclosure may be implemented in any type of apparatus, device, or system (collectively referred to as a communications apparatus) that has a communications function. The communications apparatus may include a wireless transceiver and processing / control circuitry. The wireless transceiver may include a receiver and a transmitter, or both functions. The wireless transceiver (transmitter and receiver) may include a radio frequency (RF) module and one or more antennas. The RF module may include an amplifier, an RF modulator / demodulator, or the like. Non-limiting examples of communication devices include telephones (e.g., cell phones, smartphones), tablets, personal computers (PCs) (e.g., laptops, desktops, notebooks), cameras (e.g., digital still / video cameras), digital players (e.g., digital audio / video players), wearable devices (e.g., wearable cameras, smartwatches, tracking devices), game consoles, digital book readers, telehealth / telemedicine devices, communication-enabled vehicles or mobile transportation (e.g., cars, airplanes, ships), and combinations of the above devices.

[0134] Communications equipment is not limited to portable or mobile equipment, but also includes non-portable or fixed equipment, devices, and systems of any kind, such as smart home devices (such as appliances, lighting equipment, smart meters or metering devices, control panels, etc.), vending machines, and any other "things" that may exist on an IoT (Internet of Things) network.

[0135] Furthermore, in recent years, in the field of IoT (Internet of Things) technology, CPS (Cyber ​​Physical Systems) has been attracting attention as a new concept that creates new added value by linking information between physical space and cyberspace. This CPS concept can also be adopted in the above-mentioned embodiments.

[0136] That is, as a basic configuration of a CPS, for example, an edge server located in physical space and a cloud server located in cyberspace can be connected via a network, and processing can be distributed and performed by processors installed on both servers. Here, it is preferable that each piece of processing data generated on the edge server or cloud server is generated on a standardized platform, and the use of such a standardized platform can improve the efficiency of building a system that includes a variety of sensor groups and IoT application software.

[0137] In the above-described embodiments, for example, the edge server may be located in a store and perform product recognition processing and product misrecognition risk assessment processing. The cloud server may perform model learning using data received from the edge server via a network. Alternatively, for example, the edge server may be located in a store and perform product recognition processing, and the cloud server may perform product misrecognition risk assessment processing using data received from the edge server via a network.

[0138] Communications include data communications via cellular systems, wireless LAN systems, communications satellite systems, etc., as well as data communications via combinations of these.

[0139] A communications apparatus also includes devices such as controllers and sensors connected or coupled to a communications device that performs the communications functions described in this disclosure, such as controllers and sensors that generate control and data signals used by the communications device to perform the communications functions of the communications apparatus.

[0140] The communication apparatus also includes infrastructure facilities, such as base stations, access points, and any other apparatus, device, or system that communicates with or controls the various apparatuses listed above, but are not limited to these.

[0141] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications or alterations within the scope of the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure. Furthermore, the components of the above-described embodiments may be combined in any manner without departing from the spirit of the disclosure.

[0142] Although specific examples of the present disclosure have been described in detail above, these are merely examples and do not limit the scope of the claims. The technology described in the claims includes various modifications and alterations of the specific examples exemplified above.

[0143] An inspection system according to one embodiment of the present disclosure includes a question generation unit that generates a first question written in natural language and inquiring about the likelihood of the number and / or placement of items in a reference product having a plurality of items, based on reference information indicating criteria for the number and placement of each of the items; a question answering unit that generates a first answer to the first question generated by the question generation unit about an image of an item to be inspected, using a trained model configured to answer questions in natural language about the content of the image; and a determination unit that determines whether the item to be inspected satisfies the criteria indicated by the reference information, based on the first answer.

[0144] In one embodiment of the present disclosure, the question generator generates a plurality of second questions from one of the criteria based on the criteria information, and the first questions include the second questions.

[0145] In one embodiment of the present disclosure, the second question includes a third question and a fourth question that corresponds to a superordinate or subordinate concept of the third question.

[0146] In one embodiment of the present disclosure, the second question includes a third question and a fourth question that expresses in a different language the likelihood of the number and / or location of the items asked about by the third question.

[0147] In one embodiment of the present disclosure, the second question includes a closed question asking whether the number and / or placement of items matches the criteria and an open question requesting an answer about the number and / or placement of the items.

[0148] In one embodiment of the present disclosure, the judgment unit judges that the product to be inspected does not meet the criteria if at least one of the second answers to the multiple second questions contradicts at least one other of the second answers.

[0149] In an embodiment of the present disclosure, the determination unit determines that the inspection target item does not satisfy the standard when the first answers to one or more of the first questions contradict the standard.

[0150] In one embodiment of the present disclosure, when the first answers to one or more of the first questions contradict the criteria, the determination unit displays criteria information used to generate the first questions corresponding to the first answers that contradict the criteria.

[0151] In one embodiment of the present disclosure, if the first answers to one or more of the first questions contradict the criteria, the judgment unit displays a judgment result that the inspected item does not satisfy the criteria and at least one pair of the first questions and the first answers.

[0152] In one embodiment of the present disclosure, the question generation unit generates the first question by summarizing the reference information, which is written in natural language and indicates a manufacturing process of the reference product, using a large-scale language model.

[0153] In one embodiment of the present disclosure, the reference information is an image of the reference product, and the question generation unit generates the first question by summarizing a sentence generated from the image of the reference product by image captioning using a large-scale language model.

[0154] In an inspection method according to one embodiment of the present disclosure, an inspection system including at least one information processing device generates a first question written in natural language and asking about the likelihood of the number and / or placement of items based on reference information indicating criteria for the number and placement of each item in a reference item having a plurality of items, generates a first answer to the first question about an image of an item to be inspected using a trained model configured to answer natural language questions about the content of the image, and determines whether the item to be inspected satisfies the criteria indicated by the reference information based on the first answer. [Industrial Applicability]

[0155] An embodiment of the present disclosure is useful in an inspection system that performs visual inspection. [Explanation of symbols]

[0156] 1. Inspection system 11 Imaging unit 12 Visual Inspection Department 13 Manual DB 14 Display section 121 Question generation part 122 Question and answer section 123 Error detection unit

Claims

1. a question generator that generates a first question, written in natural language, inquiring about the likelihood of the number and / or the location of the items, based on reference information indicating criteria for the number and location of each of the items in a reference item having a plurality of items; a question answering unit that generates a first answer to the first question generated by the question generating unit about the image of the product to be inspected, using a trained model configured to answer natural language questions about the content of the image; a determination unit that determines whether the inspection target product satisfies the standard indicated by the standard information based on the first response; An inspection system comprising:

2. the question generator generates a plurality of second questions from one of the criteria based on the criteria information; the first question includes the second question; The inspection system of claim 1 .

3. The second question includes a third question and a fourth question that corresponds to a superordinate or subordinate concept of the third question. The inspection system of claim 2 .

4. the second question includes a third question and a fourth question expressing in another language the likelihood of the number and / or the location of the items asked by the third question; The inspection system of claim 2 .

5. the second questions include a closed question asking whether the number and / or location of the items matches the criteria, and an open question requesting an answer about the number and / or location of the items; The inspection system of claim 2 .

6. the determination unit determines that the inspection target product does not satisfy the criteria when at least one of the second answers to the plurality of second questions contradicts at least one other of the second answers. The inspection system of claim 2 .

7. the determination unit determines that the inspection target item does not satisfy the standard when the first answers to one or more of the first questions are inconsistent with the standard; The inspection system of claim 1 .

8. When the first answers to one or more of the first questions contradict the criteria, the determination unit displays criteria information used to generate the first questions corresponding to the first answers that contradict the criteria. The inspection system of claim 7 .

9. When the first answers to one or more of the first questions are inconsistent with the standard, the determination unit displays a determination result that the inspection target product does not satisfy the standard and at least one pair of the first questions and the first answers. The inspection system of claim 7 .

10. the question generation unit generates the first question by summarizing the reference information, which is written in a natural language and indicates a manufacturing process of the reference product, using a large-scale language model. The inspection system of claim 1 .

11. the reference information is an image of the reference product, the question generation unit generates the first question by summarizing, using a large-scale language model, a sentence generated by image captioning from the image of the reference product; The inspection system of claim 1 .

12. An inspection system including at least one information processing device, generating a first question written in natural language and asking about the likelihood of the number and / or the location of the items based on reference information indicating criteria regarding the number and location of each of the items in a reference item having a plurality of items; generating a first answer to the first question about the image of the item being inspected using a trained model configured to answer natural language questions about the content of the image; determining whether the inspection target product satisfies the standard indicated by the standard information based on the first response; Testing method.

Citation Information

Patent Citations

  • Check device and check method

    JP2022114462A

Cited By

  • Information processing device, information processing method, and information processing program

    JP7784017B1

  • Information processing system, information processing method, and information processing program

    JP7843097B1