The present invention provides an article
visual inspection method and a rationale-generative
estimation method using a large vision–
language model. The article
visual inspection method comprises: a pretraining step of performing context-sensitive pretraining by inputting information containing images and sentences into a large vision–
language model; an additional training step of performing additional training by inputting training information on a good-quality group and training information on a poor-quality group of a plurality of articles into the large vision–
language model; and an inspection step of inputting a
visual appearance image for inspection into the large vision–language model and causing the large vision–language model to output an inspection result. The article
visual inspection method is characterized in that the training information on a good-quality group includes
visual appearance images of good-quality articles and text information about good-quality articles, the training information on a poor-quality group includes
visual appearance images of poor-quality articles and text information about poor-quality articles, and the information to be inspected includes an image to be inspected and text information requesting an inspection result. The
estimation method comprises: a pretraining step of performing pretraining by inputting a plurality of datasets into a large vision–language model, each dataset containing a
single image paired with a
sentence containing a specified subject's preference information regarding the image; and a result outputting step of inputting an image of an object to be estimated into the large vision–language model and causing the large vision–language model to output an
estimation result regarding individual preference and a rationale. The specified subject's preference information includes a judgment of whether or not the subject prefers the
single image and text information indicating a rationale for the judgment.