Image evaluation method and image evaluation system

The image evaluation method and system address the challenge of diverse visual features in machine learning by using a multimodal base model to vectorize and evaluate images based on linguistic and non-linguistic features, enhancing recognition accuracy through targeted data selection and distribution adjustment.

WO2026110440A1PCT designated stage Publication Date: 2026-05-28ASTEMO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ASTEMO LTD
Filing Date
2025-08-26
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing machine learning models for image recognition face challenges in accurately handling diverse visual features of objects within the same category, such as vehicles, due to variations in make, model, color, angle, and occlusions, leading to suboptimal training data selection and recognition accuracy issues.

Method used

An image evaluation method and system that utilizes a multimodal base model to vectorize language and image queries, calculate linguistic and non-linguistic feature indices, and evaluate images based on orthogonal vectors to index objects by their linguistic and non-linguistic features, enabling precise data selection for retraining.

Benefits of technology

Enhances the recognition accuracy of machine learning models by providing a means to finely adjust the distribution of training data, ensuring optimal data balance and improving model performance through targeted data selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025029898_28052026_PF_FP_ABST
    Figure JP2025029898_28052026_PF_FP_ABST
Patent Text Reader

Abstract

This image evaluation method is executed by a computer and comprises: a vectorization step for vectorizing a language query and a recognition target image into a language vector and an image vector using a multimodal base model for language and images; a first evaluation index calculation step for calculating a first evaluation index on the basis of similarity between the language vector and the image vector; an orthogonal vector extraction step for extracting an orthogonal vector having a starting point on the language vector and directed toward the image vector; a second evaluation index calculation step for calculating a second evaluation index on the basis of similarity between a predetermined reference vector, which is on a plane including the orthogonal vector and orthogonal to the language vector, and the orthogonal vector; and an evaluation step for evaluating the recognition target image using the first evaluation index and the second evaluation index.
Need to check novelty before this filing date? Find Prior Art

Description

Image evaluation methods, image evaluation systems

[0001] This invention relates to an image evaluation method and an image evaluation system.

[0002] Machine learning models for image recognition achieve high performance by training them with large amounts of labeled data. However, even objects within the same category can have diverse visual features. For example, a vehicle's image features will differ depending on its make, model, body color, shooting angle and size, illumination, and the presence or absence of occluding areas. Such differences in image features affect the recognition accuracy of machine learning models and also affect the difficulty of training. Therefore, in order to build a highly versatile model that is not affected by differences in image features, it is necessary to prepare optimal training data. Optimal training data generally refers to a dataset that includes target images with diverse features and contains a sufficient amount of data for objects that are difficult to recognize. However, classifying the features of target images and determining the quantitative balance is not easy.

[0003] To solve this problem, a method called active learning has been proposed. Active learning is a method in which the model selects the training data itself and requests humans to label it, allowing for effective learning even with a small amount of data. However, existing active learning methods have the problem that the criteria are general, such as simply selecting image data with low recognition performance, and cannot be said to be optimally selected based on the features and performance of the images.

[0004] Patent Document 1 describes a technique for discriminating the attributes of an object to be recognized by utilizing a multimodal base model of language and images. However, it only remains at discriminating general attributes such as the type, shape, and color of fruits, fish, meat, etc., and there is no description of classifying the differences in visual features within the same attribute. Non-Patent Document 1 describes a technique for clustering from feature quantities that associate an image of an object to be recognized with a multimodal base model of language and images, and assigns linguistic meaning, correct labels for object recognition, and recognition performance. Thereby, a method for extracting a small amount of data included in scenes and learning data that are likely to cause recognition errors is described.

[0005] Japanese Patent Application Laid-Open No. 2024-091177

[0006] Sabri Eyuboglu, et al. 7, "DOMINO: DISCOVERING SYSTEMATIC ERRORS WITH CROSS-MODAL EMBEDDINGS", [online], March 24, 2022, Cornell University, [searched on October 28, 2024], Internet <URL:https: / / arxiv.org / abs / 2203.14960>

[0007] In the techniques described in Patent Document 1 and Non-Patent Document 1, images are classified by using the association with linguistic features, which is useful for grasping the tendency of recognition performance and learning data distribution. However, since the object recognition model does not recognize based on linguistic features, the correlation with the distribution of learning data and recognition performance cannot be fully explained only by classification based on features that can be expressed in language. Also, it is difficult to perfectly express the features of an image in language, and it is insufficient as a classification method for images that may include features difficult to verbalize and abstract concepts. Also, even if there is a necessary and sufficient classification method, in order to obtain a higher learning effect, it is necessary to finely adjust the distribution of images included in the dataset used for learning, and more detailed classification and analysis than the generally assumed number of hierarchies and clusters are required. That is, the inventions described in Patent Document 1 and Non-Patent Document 1 cannot provide a means for indexing the object to be recognized included in an image by its linguistic meaning and non-linguistic features.

[0008] A first aspect of the present invention is an image evaluation method performed by a computer, comprising: a vectorization step of vectorizing a language query and an image to be recognized into language vectors and image vectors using a multimodal base model of language and images; a first evaluation index calculation step of calculating a first evaluation index based on the similarity between the language vectors and the image vectors; an orthogonal vector extraction step of extracting orthogonal vectors having an origin on the language vectors and pointing toward the image vectors; a second evaluation index calculation step of calculating a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vectors and is orthogonal to the language vectors and the orthogonal vectors; and an evaluation step of evaluating the image to be recognized using the first evaluation index and the second evaluation index. A second aspect of the present invention provides an image evaluation system comprising: a vectorization unit that vectorizes a language query and a target image into language vectors and image vectors using a multimodal base model of language and images; a first evaluation index calculation unit that calculates a first evaluation index based on the similarity between the language vectors and the image vectors; an orthogonal vector extraction unit that extracts orthogonal vectors having an origin on the language vectors and pointing toward the image vectors; a second evaluation index calculation unit that calculates a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vectors and is orthogonal to the language vectors and the orthogonal vectors; and an evaluation unit that evaluates the target image using the first evaluation index and the second evaluation index.

[0009] According to the present invention, a means can be provided to index objects to be recognized in an image based on their linguistic meaning and non-linguistic features. Details, configuration, and effects of the above means will be revealed by the following embodiments.

[0010] Diagram of the configuration of the image analysis device that performs the image analysis method Hardware configuration of the image analysis device Diagram of the configuration of the image classification unit and image language model Flowchart showing the processing of the image classification unit Diagram of the configuration of the non-verbal feature index extraction unit Diagram showing the processing of the non-verbal feature index extraction unit Flowchart showing the processing of the non-verbal feature index extraction unit Diagram showing an example of a learning count map Diagram showing an example of a performance map Flowchart showing the learning count map creation process performed by the performance analysis unit Flowchart showing the performance map creation process performed by the performance analysis unit Flowchart showing the count evaluation process by the classification evaluation unit Flowchart showing the performance evaluation process by the classification evaluation unit Flowchart showing the processing of the image selection unit Diagram showing the configuration of the image analysis device from a different perspective Diagram of the configuration of the image analysis device in the second embodiment Diagram of the configuration of the image classification unit and image language model in the second embodiment Diagram showing an example of arranging the configuration on a cloud environment Diagram showing an example of arranging the configuration on a vehicle Diagram showing a second example of arranging the configuration on a cloud environment

[0011] —First Embodiment— The first embodiment of the image analysis method will be described below with reference to Figures 1 to 15. In this embodiment, the "processed image" is an image containing one or more objects. The "recognition target image" is an image containing one object included in the processed image. Since the processed image may contain only one object, the processed image and the recognition target image may be the same. Examples of objects include automobiles, people, traffic lights, road signs, etc.

[0012] Figure 1 is a diagram showing the configuration of an image analysis device 1 that performs an image analysis method. The image analysis device 1 comprises a storage device 4 and a processing unit 5. The storage device 4 includes an object recognition model 41, an image language model 42, a trained image recording unit 43, a verification image recording unit 44, and a training image recording unit 45. The processing unit 5 comprises an image extraction unit 51, an image classification unit 52, a performance analysis unit 53, a classification evaluation unit 54, and an image selection unit 55.

[0013] The object recognition model 41 is a pre-trained model that detects objects in an image. The input to the object recognition model 41 is the processed image, and the output of the object recognition model 41 is the predicted label. A predicted label is a combination of the object's name and data indicating the object's position in the image, such as a bounding box. As mentioned above, in this embodiment, the processed image contains one or more objects. Therefore, the predicted label may contain two or more combinations of object names and bounding boxes. The trained image recording unit 43, the verification image recording unit 44, and the training image recording unit 45 each store a combination of the processed image and the correct label. The correct label is data indicating the object's name and its position in the image, and is, so to speak, a pre-created predicted label. The correct label can also be called a "pre-created correct answer." The difference between the correct label and the predicted label is whether it was created in advance or created by the object recognition model 41. The processed image and correct label stored in the trained image recording unit 43 were used to train the object recognition model 41. The processed images and correct labels stored in the training image recording unit 45 will be used for future retraining of the object recognition model 41. The training image recording unit 45 stores the processed images and correct labels through a process described later.

[0014] The image extraction unit 51 extracts the image to be recognized from the processed image using the predicted label output by the object recognition model 41, or the correct label contained in the trained image recording unit 43 or the verification image recording unit 44. When using the object recognition model 41, the image extraction unit 51 first inputs the processed image to the object recognition model 41 to obtain the predicted label. Then, the image extraction unit 51 obtains the image to be recognized by cropping the image from the processed image input to the object recognition model 41 based on the bounding box contained in the predicted label. When using the trained image recording unit 43 or the verification image recording unit 44, the image extraction unit 51 targets the processed image combined with the correct label and obtains the image to be recognized by cropping the image based on the bounding box contained in the correct label. The processing of the image classification unit 52 and the performance analysis unit 53 will be described later.

[0015] The classification and evaluation unit 54 performs either a quantity evaluation process using the learning image count map 161, which will be described later, or a performance evaluation process using the performance map 162, which will be described later. Whether the classification and evaluation unit 54 performs the quantity evaluation process or the performance evaluation process may be specified by, for example, operator 90, by someone other than operator 90, determined in advance by some criteria, or determined randomly. The image selection unit 55 selects a set of images for retraining from the learning image recording unit 45 that matches the selection criteria acquired by the classification and evaluation unit 54.

[0016] Figure 2 is a hardware configuration diagram of the image analysis device 1. The image analysis device 1 comprises a CPU 501, which is a central processing unit; a ROM 502, which is a read-only storage device; a RAM 503, which is a read-write storage device; an input / output device 504, which is a user interface; and a communication device 505. The CPU 501 loads the program stored in the ROM 502 into the RAM 503 and executes it, thereby realizing various calculations in the processing unit 5. The image analysis device 1 may be implemented using an FPGA (Field Programmable Gate Array), which is a rewritable logic circuit, or an ASIC (Application Specific Integrated Circuit), which is an integrated circuit for a specific application, instead of the combination of CPU 501, ROM 502, and RAM 503. Alternatively, the image analysis device 1 may be implemented using a different combination of configurations, for example, a combination of CPU 501, ROM 502, RAM 503 and FPGA.

[0017] In Figure 2, for convenience, the image analysis device 1 is shown as being composed of a single hardware device, but the image analysis device 1 may be composed of multiple hardware devices. In this case, the hardware devices may be installed adjacent to each other, or they may be connected via a local area network or the internet. The input / output device 504 is, for example, a keyboard that receives input from the operator 90. The communication device 505 enables the image analysis device 1 to communicate with other components. The communication device 505 is, for example, a communication interface. Note that the input / output device 504 and the communication device 505 may be configured as a single unit, and input from the operator 90 may be received via communication.

[0018] Figure 3 is a diagram showing the configuration of the image classification unit 52 and the image language model 42. The image classification unit 52 comprises a query generation unit 521, a similarity calculation unit 522, and a non-verbal feature index extraction unit 523. The image language model 42 comprises a language encoder 421 and an image encoder 422. The image language model 42 is a multimodal base model of language and images. The language encoder 421 converts the input language query 83 into a language vector 84. The image encoder 422 converts the input recognition target image 81 into an image vector 82. The language encoder 421 and the image encoder 422 are pre-trained so that the language vector 84 and the image vector 82 match when the linguistic meaning of the language query 83 and the recognition target image 81 match. The language vector 84 and the image vector 82 are vector data represented in the same vector space.

[0019] Operator 90 specifies the object 80 to the image classification unit 52. The object 80 is data that identifies an object and is synonymous with the name of the object included in the correct label. The object 80 is, for example, the string "automobile" or a predetermined code that indicates "automobile". The image classification unit 52 outputs the specified object 80 to the image extraction unit 51 and obtains the recognition target image 81 from the image extraction unit 51. If the object 80 is "automobile", the recognition target image 81 is an image of various automobiles extracted from the processed image. In Figure 3, it is shown as if there is only one recognition target image 81, but in reality there are often multiple recognition target images 81.

[0020] The query generation unit 521 of the image classification unit 52 generates a language query 83 using the object 80 specified by the operator 90 and inputs it to the language encoder 421 of the image language model 42. The query generation unit 521 may use the object 80 itself as the language query 83, or it may generate the language query 83 using a template (not shown). The language query 83 may be, for example, "car" or "a photo of a car". It is preferable that the language query 83 be input in the language language encoder 421 supports, and the query generation unit 521 may have a dictionary, or the object 80 may be described in a specific language. One language query 83 is generated for each object 80.

[0021] The language encoder 421 generates a language vector 84 corresponding to the language query 83 and outputs it to the image classification unit 52. Strictly speaking, the language query 83 is not directly input to the language encoder 421; preprocessing as described below is required. However, the preprocessing described here is just one example, and other methods may be used. Here, we will describe an example in which tokenization, ID conversion, padding, embedding, and positional encoding are performed as preprocessing. First, in tokenization, the input text is divided into words and subwords. For example, the sentence "A cute dog playing on the grass" is divided into "A", "cute", "dog", "playing", "on", "the", "grass". Next, each token is converted into a number, i.e., an ID. This is necessary for the model to treat the text as a number. For example, "dog" is converted to ID 1234, "playing" to ID 5678, etc., based on a predefined dictionary.

[0022] Then, to maintain a consistent text length, specific IDs are added to short texts for padding. This enables batch processing. In the next embedding step, each token ID is converted into a high-dimensional vector. The embedding is a vector representation that contains the semantic information of the tokens. Next, a positional encoding is added to preserve the order information of the tokens. This allows the model to understand the order of the tokens. After these preprocessing steps, these vectors are input to the language encoder 421, and a vector representing the meaning of the text is output.

[0023] The image classification unit 52 outputs the image to be recognized 81 to the image encoder 422 of the image language model 42. The image encoder 422 generates an image vector 82 corresponding to the input image to be recognized 81 and outputs it to the image classification unit 52. The image classification unit 52 inputs the language vector 84 output by the language encoder 421 to the similarity calculation unit 522 and the non-linguistic feature index extraction unit 523, and inputs the image vector 82 output by the image encoder 422 to the similarity calculation unit 522 and the non-linguistic feature index extraction unit 523. In other words, the image vector 82 and the language vector 84 are input to the similarity calculation unit 522 and the non-linguistic feature index extraction unit 523, respectively. The similarity calculation unit 522 calculates a linguistic feature index 88 using the image vector 82 and the language vector 84. The non-linguistic feature index extraction unit 523 calculates a non-linguistic feature index 89 using the image vector 82 and the language vector 84. The processing of the nonverbal feature index extraction unit 523 will be explained in detail later with reference to Figures 5 to 7. In the following, the verbal feature index 88 may be referred to as the "first index," and the nonverbal feature index 89 may be referred to as the "second index."

[0024] The similarity calculation unit 522 calculates the similarity between two input vectors, namely the image vector 82 and the language vector 84. For example, the similarity calculation unit 522 calculates the cosine similarity between the two vectors, that is, the cosine value of the angle between the two vectors. Specifically, the similarity calculation unit 522 calculates the cosine similarity by dividing the dot product of the two vectors by the product of the magnitudes of the two vectors. The similarity calculation unit 522 outputs the calculated similarity as a linguistic feature index 88. As mentioned above, there may be multiple recognition target images 81 for a single object 80. In that case, the similarity calculation unit 522 receives two or more image vectors 82 and one language vector 84 as input. In this case, the similarity calculation unit 522 calculates the similarity between the language vector 84 and each image vector 82 and outputs them as a linguistic feature index 88.

[0025] Figure 4 is a flowchart showing the processing of the image classification unit 52. First, in step S301, the image classification unit 52 receives an object 80, for example, the string "automobile," from the operator 90. In the following step S302, the image classification unit 52 has a query generation unit 521 that generates a language query 83 based on the object 80. In the following step S303, the image classification unit 52 outputs the language query 83 generated in step S302 to the language encoder 421 and obtains a language vector 84 from the language encoder 421. In the following step S304, the image classification unit 52 uses an image extraction unit 51 to obtain a recognition target image 81. In the following step S305, the image classification unit 52 inputs the recognition target image 81 obtained in step S304 to the image encoder 422 and obtains an image vector 82 from the image encoder 422. In the following step S306, the image classification unit 52 has a similarity calculation unit 522 that calculates the similarity between the image vector 82 and the language vector 84 as a linguistic feature index 88. In the following step S307, the image classification unit 52 has a non-linguistic feature index extraction unit 523 that uses the image vector 82 and the language vector 84 to calculate a non-linguistic feature index 89, and the process shown in Figure 4 is completed. The details of step S307 will be explained later with reference to Figures 5 and 7.

[0026] Figure 5 shows the correlation of data calculated by the non-verbal feature index extraction unit 523. The non-verbal feature index extraction unit 523 receives an image vector 82 and a language vector 84 as input. The non-verbal feature index extraction unit 523 generates a reference vector 153 that is orthogonal to the language vector 84. The non-verbal feature index extraction unit 523 also calculates an inter-vector angle 151, which is the angle between the language vector 84 and the image vector 82. Furthermore, the non-verbal feature index extraction unit 523 calculates a non-verbal feature vector 154 using the image vector 82, the language vector 84, and the inter-vector angle 151. Finally, the non-verbal feature index extraction unit 523 calculates the similarity between the reference vector 153 and the non-verbal feature vector 154 as a non-verbal feature index 89.

[0027] Figure 6 is a vector correlation diagram, and Figure 7 is a diagram showing the processing of the non-verbal feature index extraction unit 523. The non-verbal feature index extraction unit 523 receives a language vector 84 output by the language encoder 421 and an image vector 82 output by the image encoder 422 as inputs. In step S311, the non-verbal feature index extraction unit 523 first determines an arbitrary reference vector 153 orthogonal to the language vector 84. In the following step S312, the non-verbal feature index extraction unit 523 calculates the vector-to-vector angle 151 between the image vector 82 and the language vector 84.

[0028] In the following step S313, the non-verbal feature index extraction unit 523 uses the language vector 84, the image vector 82, and the inter-vector angle 151 to extract an orthogonal vector as the non-verbal feature vector 154 that has its origin on the language vector 84 and points toward the endpoint of the image vector 82. As shown in Figure 6, once the language vector 84, the image vector 82, and the inter-vector angle 151 are determined, the non-verbal feature vector 154 that satisfies the conditions is uniquely determined. Also, as shown in the lower part of Figure 6, the reference vector 153 and the non-verbal feature vector 154 lie on the same plane.

[0029] In the following step S314, the non-verbal feature index extraction unit 523 calculates the similarity between the reference vector 153 and the non-verbal feature vector 154. For example, the non-verbal feature index extraction unit 523 calculates cosine similarity by calculating the cosine value of the angle between the two vectors. The similarity obtained in this way is output as a non-verbal feature index 89. In the following step S315, it is determined whether all image vectors 82 have been processed. If it is determined that all image vectors 82 have been processed, the process shown in Figure 7 is terminated. If it is determined that there are unprocessed image vectors 82, the process returns to step S312.

[0030] Furthermore, as can be seen from the fact that the return destination is step S312, the reference vector 153 is not changed. In other words, even when one language vector 84 and multiple image vectors 82 are input to the non-verbal feature index extraction unit 523, the non-verbal feature index extraction unit 523 determines only one reference vector 153 and uses that reference vector 153 to calculate all non-verbal feature indices 89.

[0031] The linguistic feature index 88 is an index that expresses the similarity to a linguistic query 83 that has linguistic meaning, and can be said to be an index that reflects linguistic features. On the other hand, the same value of the linguistic feature index 88 can be extracted from image vectors 82 that have different vector components. This is because, when a linguistic vector 84 is defined, there can be multiple image vectors 82 that have the same inter-vector angle 151. In other words, even though the components of the image vector 82 that express the features as an image are different, the closeness to the linguistic meaning represented by the linguistic query 83 is treated as equivalent. This is equivalent to, for example, the existence of multiple images that are judged to be semantically equivalent to the word "car". Since there are various images of cars, it is natural to treat them all as images of cars linguistically, but this classification is insufficient for purposes such as analyzing the performance of the object recognition model 41 and linking it to the distribution of training data.

[0032] The non-verbal feature index 89 is calculated by separating the image vector 82 into components orthogonal to the language vector 84 and comparing them with a reference vector 153 that has equivalent features. This makes it possible to treat image vectors 82 having different vector components that show the same value of the language feature index 88 as different image vectors 82. Here, the components orthogonal to the language vector 84 can be considered as the components that least reflect the linguistic features represented by the language query 83, and are therefore called non-verbal features.

[0033] The performance analysis unit 53 uses the linguistic feature index 88 and non-linguistic feature index 89 calculated by the image classification unit 52 to analyze the correlation between the data distribution of the recognition target image 81 and the recognition performance when using the object recognition model 41. Specifically, the performance analysis unit 53 defines an evaluation space 160 with the linguistic feature index 88 on one axis and the non-linguistic feature index 89 on the other axis, and maps the data distribution or index of multiple recognition target images 81. In the evaluation space 160, recognition target images 81 with similar linguistic and non-linguistic features can be mapped to similar spaces, so the similarity of multiple recognition target images 81 can be quantitatively determined.

[0034] Figures 8 and 9 show examples of evaluations performed by the performance analysis unit 53, with Figure 8 being the training image map 161 and Figure 9 being the performance map 162. In this embodiment, the three-dimensional space where the evaluation values ​​created by the performance analysis unit 53 are shown, as in Figures 8 and 9, is called the evaluation space 160. In the evaluation space 160, the linguistic feature index 88 is set on the vertical axis and the non-linguistic feature index 89 is set on the horizontal axis, and the height axis can be one of two types, as will be explained below. In this embodiment, the plane in the evaluation space 160 is divided into predetermined regions, and each region is called a "grid". For example, if the linguistic feature index 88 is divided into 10 and the non-linguistic feature index 89 is also divided into 10, 100 grids are defined.

[0035] Figure 8 shows the distribution of the number of training images, i.e., the number of recognition target images 81 extracted from the trained image recording unit 43, on the height axis. Figure 9 shows the performance index, i.e., the recognition rate when recognition is performed using the object recognition model 41, on the height axis. In Figure 8, the more images there are, the whiter they are, and the fewer images there are, the blacker they are. In Figure 8, the distribution of multiple recognition target images 81 is represented by a smooth 3D curved surface, and it can be seen that there are white areas where images are concentrated, i.e., peaks and surrounding areas with a relatively small number of images.

[0036] On the other hand, Figure 9 shows an example of the recognition rate of the object recognition model 41, with higher recognition rates indicated in white and lower recognition rates in black. In Figure 8, the white areas where images are concentrated are represented as black valleys in Figure 9, indicating that the recognition rate in these areas is extremely low. In other words, despite the large number of images used for training, these are areas where recognition is difficult for the object recognition model 41, and therefore areas with high data value for retraining.

[0037] On the other hand, it can be seen that there are regions where the recognition rate is high in areas with a relatively small number of trained image distributions. Such regions are easy for the object recognition model 41 to recognize, and can be judged as regions where high recognition performance can be expected with a small amount of training data. In other words, these are regions where the data value for retraining is not very high. Through such visualization and analysis, it becomes possible to evaluate each region of the evaluation space 160 and the images classified into those regions.

[0038] The training image map 161 maps each reference recognition target image 81A into the evaluation space 160 based on the linguistic and non-linguistic features of the images. Images mapped to the same region in the evaluation space 160 have similar linguistic and non-linguistic features. Therefore, the evaluation space 160 can also be viewed as a grid space divided into a certain range width by linguistic feature index 88 and non-linguistic feature index 89.

[0039] Figure 10 is a flowchart showing the process of creating a training image map executed by the performance analysis unit 53. First, in step S320, the performance analysis unit 53 initializes each grid of the evaluation space 160. For example, if 100 grids are defined, this process assigns zero to an integer array having 100 elements. In the following step S321, all processed images and correct labels are extracted from the trained image recording unit 43 and the process proceeds to step S322. The combinations of processed images and correct labels extracted in step S321 are used in order in step S322, as will be described later. In step S322, the performance analysis unit 53 outputs one set of unprocessed processed images and correct labels to the image extraction unit 51, and extracts all recognition target images 81 from those processed images based on the correct labels, and the process proceeds to step S323. The recognition target images 81 extracted in step S322 are used in order in step S323, as will be described later.

[0040] In step S323, the performance analysis unit 53 uses the unprocessed recognition target image 81 extracted in step S322 to cause the image classification unit 52 to calculate two indices, namely a linguistic feature index 88 and a non-linguistic feature index 89. The processing of the image classification unit 52 at this time is generally as described with reference to Figure 4, but is modified as follows. That is, the image classification unit 52 uses the name of an object included in one of the correct labels output from the performance analysis unit 53 instead of the object 80 received from the operator 90 in step S301. Also, instead of executing step S304, it is treated as if the aforementioned unprocessed recognition target image 81 has been extracted.

[0041] In the following step S324, the performance analysis unit 53 maps the recognition target image 81 to the evaluation space 160 and proceeds to step S325. Specifically, the performance analysis unit 53 increases the grid value corresponding to the combination of the two indicators calculated in step S303 by 1. In the following step S325, the performance analysis unit 53 determines whether or not all of the recognition target images 81 extracted in step S322 have been processed. If the performance analysis unit 53 determines that all of the recognition target images 81 extracted in step S322 have been processed, it proceeds to step S326; if it determines that there are unprocessed recognition target images 81, it returns to step S323.

[0042] In step S326, the performance analysis unit 53 determines whether all the processing images extracted in step S321 have been processed. If the performance analysis unit 53 determines that all processing images have been processed, it terminates the process shown in Figure 10. If it determines that there are unprocessed processing images, it returns to step S322. The learning image count map 161 can be created by counting the number of images corresponding to each grid, so data creation is completed by repeating the voting process in step S324 for all processing images for all recognition target images 81. In addition to the process shown in Figure 10, the performance analysis unit 53 may also create three-dimensional graph data as shown in Figure 8.

[0043] Figure 11 is a flowchart showing the performance map creation process executed by the performance analysis unit 53. First, in step S330, the performance analysis unit 53 initializes each grid in the evaluation space 160. For example, if 100 grids are defined, this process assigns zero to all elements of a multidimensional array having 100 elements, each element capable of storing any number of integers or decimals. In the following step S331, all processed images and correct labels are extracted from the verification image recording unit 44, and the process proceeds to step S332. The combinations of processed images and correct labels extracted in step S331 are used in order in step S332, as will be described later.

[0044] In step S332, the performance analysis unit 53 outputs an unprocessed set of processed images and the correct label to the image extraction unit 51, and extracts all recognition target images 81 using the bounding box included in the correct label from the processed images, and then proceeds to step S333. However, in the performance sheet map creation process, the recognition target image 81 extracted using the bounding box included in the correct label is called the reference recognition target image 81A. When a plurality of reference recognition target images 81A are extracted in step S332, they are used in step S333 in order as described later. The processed images and the correct label output to the image extraction unit 51 in this step are called "specific processed images" and "specific correct labels". These specific processed images and specific correct labels are not changed until a positive determination is made in step S337 described later.

[0045] In step S333, the performance analysis unit 53 uses the unprocessed reference recognition target image 81A among the reference recognition target images 81A extracted in step S332 to cause the image classification unit 52 to calculate two indicators, namely the linguistic feature indicator 88 and the non-linguistic feature indicator 89. The processing of the image classification unit 52 at this time is generally as described with reference to FIG. 4, but is modified as follows. That is, the image classification unit 52 uses the name of the object included in the specific correct label instead of the object 80 received from the operator 90 in step S301. Also, instead of executing step S304, it is handled as if the aforementioned unprocessed reference recognition target image 81A has been extracted. In the subsequent step S334, the performance analysis unit 53 extracts the recognition target image 81 using the object recognition model 41. However, the recognition target image 81 extracted in this step is called the evaluation recognition target image 81B. Specifically, the performance analysis unit 53 inputs the name of the object included in the specific correct label and the specific processed image into the object recognition model 41 to obtain the evaluation recognition target image 81B.

[0046] In the subsequent step S335, the performance analysis unit 53 compares and evaluates the reference recognition target image 81A and the evaluation recognition target image 81B. This evaluation may be the recognition rate indicating the success or failure of recognition, or the IoU (Intersection over Union) representing the degree of overlap between two regions of the reference recognition target image 81A and the evaluation recognition target image 81B, and is not limited to a specific performance metric. In the subsequent step S336, the performance analysis unit 53 maps the evaluation result in step S335 to the evaluation space 160. The mapping in this step means storing the evaluation result in step S335 in the multi-dimensional array initialized in step S330. In step S324 of FIG. 10, the value of the corresponding grid was increased by 1, but this step is different in that the evaluation result is saved.

[0047] In the subsequent step S337, the performance analysis unit 53 determines whether all of the recognition target images 81 extracted in step S332 have been processed. If the performance analysis unit 53 determines that all of the recognition target images 81 extracted in step S332 have been processed, it proceeds to step S338. If it determines that there are unprocessed recognition target images 81, it returns to step S333. In step S338, the performance analysis unit 53 determines whether all of the processed images extracted in step S331 have been processed. If the performance analysis unit 53 determines that all of the processed images have been processed, it proceeds to step S339. If it determines that there are unprocessed processed images, it returns to step S332.

[0048] In step S339, the performance analysis unit 53 aggregates the evaluation results mapped to the evaluation space 160 and ends the process shown in FIG. 11. Specifically, the performance analysis unit 53 calculates the average value of the evaluation results for each grid in the evaluation space 160. Since the respective evaluation results are saved in step S336, the average value can be calculated in this step. In addition to the process shown in FIG. 11, the performance analysis unit 53 may create three-dimensional graph data as shown in FIG. 9.

[0049] Figure 12 is a flowchart showing the image count evaluation process by the classification evaluation unit 54. First, in step S341, the classification evaluation unit 54 calculates the total number of recognition target images 81 mapped to the learning image count map 161. In the following step S342, the classification evaluation unit 54 extracts, in other words, counts, the number of recognition target images 81 in each grid of the learning image count map 161. In the following step S343, the classification evaluation unit 54 calculates the ratio of the extracted number of images to the total number for each grid. In the following step S344, the classification evaluation unit 54 evaluates the data sufficiency of each grid in the evaluation space 160 according to the evaluation criteria. The evaluation criteria may be whether the number of images is 1 or more, or whether the calculated ratio is greater than or equal to a predetermined value.

[0050] In the following step S345, the classification and evaluation unit 54 labels each grid in the evaluation space 160 according to the evaluation results of step S344. Specifically, the classification and evaluation unit 54 assigns a label indicating that the fewer the number of recognition target images 81 included in the grid, the higher the value of the images in that grid. The label may be a binary value indicating high or low value, or it may be an identification value of 3 or more indicating high value, for example, a number from 1 to 10. In the following step S346, the classification and evaluation unit 54 sets the required amount of data for each grid in the evaluation space 160 based on the labeling results and terminates the process shown in Figure 12. For example, based on the labeling results, the classification and evaluation unit 54 can determine data selection criteria for each grid so that the number of recognition target images 81 is reduced in grids where the number of recognition target images 81 is already excessive, and so that the number of recognition target images 81 is increased in areas where there is a shortage of recognition target images 81. As a result, it is possible to provide guidance for a suitable data balance in the entire evaluation space 160.

[0051] Figure 13 is a flowchart of the performance evaluation process performed by the classification and evaluation unit 54. However, some of the processes shown in Figure 13 are similar to those in Figure 12. In step S351, the classification and evaluation unit 54 extracts the performance mapped to each grid of the performance map 162. This step S351 corresponds to steps S341 to S343 in Figure 12. In the following step S354, the classification and evaluation unit 54 evaluates the data sufficiency of each grid in the evaluation space 160 according to the evaluation criteria, similar to step S344 in Figure 12. The evaluation criteria used in this step are different from those used in step S344 in Figure 12. In the following step S355, the classification and evaluation unit 54 labels the data value of each grid in the evaluation space 160 according to the evaluation results, similar to step S345 in Figure 12. In the following step S356, the classification and evaluation unit 54 sets the required amount of data for each grid in the evaluation space 160 based on the labeling results, similar to step S346 in Figure 12, and terminates the process shown in Figure 13.

[0052] Figure 14 is a flowchart showing the processing of the image selection unit 55. First, in step S360, the image selection unit 55 extracts all the processing images from the training image recording unit 45 and proceeds to step S361. The processing images extracted in this step are used sequentially in steps S361 to S364, as will be described later. In step S361, the image selection unit 55 uses the object recognition model 41 to extract all the images to be recognized 81 from the processing images and proceeds to step S362. However, if the processing images have correct labels, the images to be recognized 81 may be extracted using the bounding boxes of the correct labels.

[0053] In step S362, the image selection unit 55 selects one of the unprocessed recognition target images 81 extracted in step S361 and has the image classification unit 52 calculate two indices, namely a linguistic feature index 88 and a non-linguistic feature index 89, before proceeding to step S363. If the recognition target image 81 was extracted in step S361 using the object recognition model 41, the name of the object included in the predicted label output by the object recognition model 41 is set as the target object 80. If the recognition target image 81 was extracted in step S361 using the correct label assigned to the processed image, the name of the object included in that correct label is set as the target object 80. In step S363, the image selection unit 55 determines whether all the recognition target images 81 extracted in step S361 have been processed. If the image selection unit 55 determines that all the recognition target images 81 have been processed, it proceeds to step S364; if it determines that there are unprocessed recognition target images 81, it returns to step S362.

[0054] In step S364, the image selection unit 55 aggregates the distribution of the number of recognition target images 81 included in the processing images and proceeds to step S365. Specifically, the image selection unit 55 aggregates the number of images included in each grid of the evaluation space 160. In the following step S365, the image selection unit 55 determines whether all the processing images extracted in step S360 have been processed. If the image selection unit 55 determines that all processing images have been processed, it proceeds to step S366; if it determines that there are unprocessed processing images, it returns to step S361.

[0055] In step S366, the image selection unit 55 selects a predetermined number of processed images from the processed images extracted in step S360. This number may be a predetermined number, or it may be a random value determined each time step S366 is executed. In the following step S367, the image selection unit 55 sums up the distribution of the number of recognition target images 81 included in the selected processed images. Since the number of recognition target images 81 included in each grid was previously tallied for each processed image in step S364, in this step it is only necessary to tally the number of images selected in step S366.

[0056] In the following step S368, the image selection unit 55 determines whether the aggregation result in step S367 matches the selection criteria. If the image selection unit 55 determines that it matches the selection criteria, it proceeds to step S369; otherwise, it returns to step S366. In step S369, the image selection unit 55 adopts the processed images selected in step S366 as a dataset for retraining and terminates the process shown in Figure 14. This makes it possible to obtain a dataset suitable for retraining based on the linguistic and non-linguistic features of the images and the performance of the object recognition model 41.

[0057] Figure 15 is a diagram that organizes some of the components of the image analysis device 1 from a different perspective than Figures 1 and 3. The image analysis device 1 comprises a vectorization unit 71, a first evaluation index calculation unit 72, an orthogonal vector extraction unit 73, a second evaluation index calculation unit 74, and an evaluation unit 75. The vectorization unit 71 uses an image language model 42, which is a multimodal base model of language and images, to vectorize the language query 83 and the recognition target image 81 into a language vector 84 and an image vector 82. The vectorization unit 71 is a query generation unit 521 that outputs the language query 83 to a language encoder 421, and an image classification unit 52 that outputs the recognition target image 81 to an image encoder 422.

[0058] The first evaluation index calculation unit 72 calculates a first evaluation index, i.e., a linguistic feature index 88, based on the similarity between the language vector 84 and the image vector 82. The first evaluation index calculation unit 72 is the similarity calculation unit 522. The orthogonal vector extraction unit 73 extracts an orthogonal vector, i.e., a non-linguistic feature vector 154, which has its origin on the language vector 84 and points toward the image vector 82. The orthogonal vector extraction unit 73 corresponds to a part of the function of the non-linguistic feature index extraction unit 523, and specifically corresponds to the processing in step S313 shown in Figure 7. The second evaluation index calculation unit 74 calculates a second evaluation index, i.e., a non-linguistic feature index 89, based on the similarity between the non-linguistic feature vector 154 and a predetermined reference vector 153 on a plane that includes the non-linguistic feature vector 154 and is orthogonal to the language vector 84. The second evaluation index calculation unit 74 corresponds to a part of the function of the non-linguistic feature index extraction unit 523, and specifically corresponds to the processing in step S314 shown in Figure 7. The evaluation unit 75 evaluates the recognition target image 81 using linguistic feature index 88 and non-linguistic feature index 89. The evaluation unit 75 combines the functions of both the performance analysis unit 53 and the classification evaluation unit 54.

[0059] According to the first embodiment described above, the following effects can be obtained. (1) The image evaluation method performed by the computer-based image analysis device 1 includes a vectorization step (S303, S305 in Figure 4) in which a language query 83 and a recognition target image 81 are vectorized into a language vector 84 and an image vector 82 using an image language model 42 which is a multimodal base model of language and image, a first evaluation index calculation step (S306 in Figure 4) in which a linguistic feature index 88 is calculated based on the similarity between the language vector 84 and the image vector 82, and an orthogonal vector having its origin on the language vector 84 and pointing toward the image vector 82, i.e. The process includes an orthogonal vector extraction step (S313 in Figure 7) for extracting non-verbal feature vectors 154, a second evaluation index calculation step (S314 in Figure 7) for calculating a non-verbal feature index 89 based on the similarity between the non-verbal feature vectors 154 and a predetermined reference vector 153 on a plane that includes the non-verbal feature vectors 154 and is orthogonal to the language vectors 84, and the non-verbal feature index 154, and an evaluation step (training image map creation process, performance map creation process, image count evaluation process, performance evaluation process) for evaluating the recognition target image 81 using the language feature index 88 and the non-verbal feature index 89. Therefore, recognition target objects contained in an image can be indexed by their language meaning and non-verbal features. Specifically, it is possible to calculate the numerical values ​​necessary for creating the training image map 161 shown in Figure 8 and the performance map 162 shown in Figure 9.

[0060] (2) The images to be recognized 81 are trained images used to train the object recognition model 41. In the training image count map creation process and the count evaluation process, the images to be recognized 81 are mapped onto an evaluation space 160 with linguistic feature index 88 and non-linguistic feature index 89 as axes (S324 in Figure 10), and the mapping results of multiple images to be recognized 81 are accumulated on the evaluation space 160 as a training data distribution (S324 in Figure 10) to evaluate the data distribution in each space. The training image count map 161 is a visualization of this data distribution.

[0061] (3) Based on the data distribution of the trained images, the data proportion included in each region within the defined evaluation space is extracted (S343 in Figure 12), the data sufficiency within the evaluation space is evaluated from the data proportion extraction results (S344), and each region is labeled according to the evaluation criteria (S345). Therefore, the value of each grid in the evaluation space 160, or more precisely, the value of the image corresponding to each grid, can be labeled based on the number of trained images.

[0062] (4) In the evaluation step, the verification images are input to the object recognition model 41, and the output of the object recognition model 41 is compared with the correct labels to calculate the performance of the object recognition model 41 as an evaluation value (S334, S335 in Figure 11). The evaluation value is mapped onto an evaluation space 160 with the first and second evaluation indicators as axes (S336), and the mapping results of multiple verification images are integrated as a performance distribution on the evaluation space 160 to evaluate the performance characteristics of each space (S339).

[0063] (5) Based on the performance distribution of the verification images, the performance included in each region within the defined evaluation space is extracted (S351 in Figure 13), the performance achievement status of each region within the evaluation space is evaluated from the performance extraction results (S354), and each region is labeled according to the evaluation criteria (S355).

[0064] (6) As shown in Figure 14, the image selection unit 55 extracts images for the object recognition model 41 to retrain based on the labeling results of each region within the defined evaluation space.

[0065] (7) The image language model 42 includes a language encoder 421 and an image encoder 422. The image encoder 422 receives the entire processing image, or an extracted image cut out from the processing image using a bounding box created in advance, as the image to be recognized 81.

[0066] (Modification 1) In the first embodiment described above, either a count evaluation process to evaluate the number of training images map 161 or a performance evaluation process to evaluate the performance map 162 was performed. However, the number of training images map 161 and the performance map 162 may be analyzed in combination to determine the data selection criteria. The performance evaluation process using the performance map 162 reveals the performance achievement status of each grid in the evaluation space 160. In addition, the count evaluation process using the number of training images map 161 reveals the distribution of trained images in each grid. Therefore, data selection criteria can be determined for each grid based on the distribution of the number of training images map 161. For example, data selection criteria can be determined for each grid such that the proportion of areas where performance has been achieved is less than the proportion of known trained images, and the proportion of areas where performance has not been achieved is higher. This makes it possible to determine the sufficiency of each grid based on the features of the trained images and the performance of the object recognition model 41, and to provide guidance for a suitable data balance in the entire evaluation space 160.

[0067] (Modification 2) In the first embodiment described above, only one indicator was calculated in the performance evaluation process. However, multiple indicators may be calculated and multiple performance maps 162 may be created. In the performance evaluation process described in the first embodiment, a suitable data balance within the evaluation space 160 can be shown for a given performance indicator. Therefore, a suitable data balance can also be calculated for different performance indicators and target values. For example, a suitable data balance can be calculated based on two performance indicators: recognition rate and IoU. When using multiple performance guidelines, a suitable data balance guideline that takes multiple performance indicators into account can be shown by methods such as summing, averaging, or weighted averaging the data balance of each guideline.

[0068] (Modification 3) In the first embodiment described above, the classification and evaluation unit 54 was capable of performing both a number evaluation process to evaluate the learning number map 161 and a performance evaluation process to evaluate the performance map 162. However, the classification and evaluation unit 54 may only be capable of performing either the number evaluation process or the performance evaluation process.

[0069] (Modification 4) In the first embodiment described above, the image selection unit 55 selected an image dataset for retraining from the training image recording unit 45, but the retraining of the object recognition model 41 was not performed. However, the processing unit 5 may further include a retraining unit that performs the retraining of the object recognition model 41. Once the retraining of the object recognition model 41 is completed by the retraining unit, the image dataset selected by the image selection unit 55 is stored in the trained image recording unit 43. This makes it possible to generate a training image count map 161 and a performance map 162 that are different from those before retraining. By reviewing the optimal data balance using the newly generated training image count map 161 and performance map 162, the data balance can be dynamically adjusted.

[0070] —Second Embodiment— A second embodiment of the image evaluation method will be described with reference to Figures 16 to 17. In the following description, the same reference numerals are used for components that are the same as in the first embodiment, and the differences will be mainly explained. Points that are not specifically explained are the same as in the first embodiment. This embodiment differs from the first embodiment mainly in that it generates various language queries 83 and identifies language queries 83 in which the training image count map 161 and performance map 162 show significant bias.

[0071] Figure 16 is a configuration diagram of the image analysis device 1A in the second embodiment. The difference from the first embodiment is that the processing unit 5 has a deviation language query identification unit 56. The configuration and processing of the image classification unit 52A also differ slightly from the first embodiment, as will be described later. The image classification unit 52A generates a plurality of language queries 83, as will be described later. In this embodiment, the processing of the non-language feature index extraction unit 523 shown in Figure 5, which includes the calculation of linguistic feature indexes 88 and the calculation of non-language feature vectors 154, also called orthogonal vectors, and the calculation of non-language feature indexes 89 are performed for each generated language query 83.

[0072] The deviation language query identification unit 56 compares multiple training sheet count maps 161 or multiple performance maps 162 corresponding to multiple language queries 83 to identify the language query 83 with the largest deviation in the data distribution in the training sheet count map 161 or the largest deviation in the performance features in the performance map 162. The deviation in the data distribution in the training sheet count map 161 is, for example, the difference between the minimum and maximum number of sheets in each grid. The deviation in the performance features in the performance map 162 is, for example, the difference between the minimum and maximum recognition rate in each grid.

[0073] Figure 17 is a diagram showing the configuration of the image classification unit 52 and the image language model 42 in the second embodiment. The difference from Figure 3 in the first embodiment is that the query generation unit 521 reads a template 521A. The template 521A contains multiple templates. By using this template 521A, the query generation unit 521 generates language queries 83 such as "car", "A photo of a car", "An image of a car from any angle", "A photo of a car at night", and "An image of a car driving in the dark" when the object 80 is "car".

[0074] By changing the language query 83 in this way, the linguistic feature index 88 changes, and as a result, the non-linguistic feature index 89 extracted using the language vector 84 extracted from the language query 83 also changes. In other words, the distribution of the number of training images map 161 and the performance map 162 also changes. To put it another way, by preparing multiple language queries 83, it is possible to generate diverse number of training images map 161 and performance map 162. Therefore, by searching for the language query 83 in which the desired performance difference is most pronounced, it becomes possible to accurately analyze the conditions under which the object recognition model 41 can easily recognize objects or the conditions under which it has difficulty recognizing objects. It is also possible to analyze the conditions under which the bias of the trained images is most pronounced. The conditions obtained in this way become useful information when constructing a dataset for retraining.

[0075] According to the second embodiment described above, the following effects can be obtained. (8) Includes a language query generation step of generating multiple language queries 83 by combining the input object 80 with a plurality of templates 521A. The processing of the non-linguistic feature index extraction unit 523 shown in Figure 5, which includes the calculation of a linguistic feature index 88 and the calculation of a non-linguistic feature vector 154 also called an orthogonal vector, and the calculation of a non-linguistic feature index 89 are performed for each generated language query 83. Includes a deviation language query identification process that identifies the language query 83 with the largest deviation in the data distribution in the training count map 161 or the deviation in the performance features in the performance map 162. Therefore, the conditions under which the object recognition model 41 can easily recognize or the conditions under which it cannot recognize can be analyzed with high accuracy.

[0076] The technologies described in each of the embodiments and modifications above can be used for a variety of applications. For example, this technology can be used as an evaluation and retraining technology for an object recognition model 41 for autonomous driving of automobiles. In autonomous driving of automobiles, an object recognition model 41 for recognizing objects is deployed on the vehicle body. Generally, an object recognition model 41 is not a one-time deployment; it is often configured to be replaceable when defects, insufficient performance, or the need for functional updates arise. For example, it is conceivable to accumulate data collected from automobiles in a cloud environment, update the object recognition model 41 based on that data, and redeploy the updated object recognition model 41 to the vehicle.

[0077] One concept known as a connected car involves a vehicle that is constantly connected to the internet, collecting data from sensors and cameras, as well as operational data, and distributing necessary information and software, including an object recognition model 41, to the vehicle. Using such a system, as shown in Figure 18, the collected images 116 from the vehicle 20 can be stored in the learning image recording unit 45, which stores candidate images for retraining on the cloud environment 9, and can be used for retraining using this technology. The retrained object recognition model 41 is then distributed back to the vehicle 20, and the updated new object recognition model 41A is used in the vehicle 20.

[0078] Alternatively, as shown in Figure 19, the functions of this technology can be deployed on the vehicle 20. In this case, an image acquisition unit 6, such as a camera, is mounted on the vehicle 20A, and the collected images 116 obtained by the image acquisition unit 6 are recorded in the learning image recording unit 45. With this configuration, the object recognition model 41 can be updated without connecting to the cloud environment 9. Furthermore, as shown in Figure 20, it is also possible to deploy all or part of the functions included in the processing unit 5 on the vehicle 20B and perform the analysis of the collected images 116 on the vehicle 20B. In this case, by sharing the data selection criteria 117 determined by the classification evaluation processing unit 124 from the cloud environment 9 with the vehicle 20B, only images with high data value can be selected from the collected images 116 and sent to the cloud environment 9. This reduces the amount of data of collected images 116 sent to the cloud environment 9, and allows for the efficient collection and storage of only data with high retraining value.

[0079] A similar mechanism can be applied to applications using the object recognition model 41 in factories, commercial facilities, buildings, train stations, etc. This technology can be applied to applications where the object recognition model 41 is retrained using images collected during the operation of the application, and the updated object recognition model 41A is reused.

[0080] In the embodiments and modifications described above, the configuration of the functional blocks is merely an example. Several functional configurations shown as separate functional blocks may be integrated, or a configuration represented in one functional block diagram may be divided into two or more functions. Furthermore, some of the functions of each functional block may be provided by other functional blocks.

[0081] In the embodiments and modifications described above, the program is stored in ROM 502, but the program may also be stored in a non-volatile memory device. Furthermore, the image analysis device 1 may be equipped with an input / output interface (not shown), and the program may be read from another device via the input / output interface and a medium available to the image analysis device 1 when necessary. Here, "medium" refers to, for example, a storage medium detachable from the input / output interface, or a communication medium, i.e., a wired, wireless, or optical network, or a carrier wave or digital signal propagating through such a network. Also, some or all of the functions realized by the program may be realized by hardware circuits or FPGAs.

[0082] The embodiments and modifications described above may be combined in any way. Although various embodiments and modifications have been described above, the present invention is not limited to these. Other embodiments that can be conceivable within the scope of the technical idea of ​​the present invention are also included within the scope of the present invention.

[0083] 1: Image analysis device 41: Object recognition model 42: Image language model 52: Image classification unit 53: Performance analysis unit 54: Classification evaluation unit 55: Image selection unit 56: Deviation language query identification unit 80: Target object 81: Recognition target image 82: Image vector 83: Language query 84: Language vector 88: Language feature index 89: Non-language feature index 124: Classification evaluation processing unit 151: Inter-vector angle 153: Reference vector 154: Non-language feature vector 160: Evaluation space 161: Training image count map 162: Performance map 421: Language encoder 422: Image encoder 521: Query generation unit 521A: Template 522: Similarity calculation unit 523: Non-language feature index extraction unit

Claims

1. A computer-based image evaluation method comprising: a vectorization step of vectorizing a language query and an image to be recognized into language vectors and image vectors using a multimodal base model of language and images; a first evaluation index calculation step of calculating a first evaluation index based on the similarity between the language vectors and the image vectors; an orthogonal vector extraction step of extracting orthogonal vectors having an origin on the language vectors and pointing toward the image vectors; a second evaluation index calculation step of calculating a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vectors and is orthogonal to the language vectors and the orthogonal vectors; and an evaluation step of evaluating the image to be recognized using the first evaluation index and the second evaluation index.

2. An image evaluation method according to claim 1, wherein the image to be recognized is a trained image used to train an object recognition model, and in the evaluation step, the image to be recognized is mapped onto an evaluation space with the first evaluation index and the second evaluation index as axes, and the data distribution of each space is evaluated by accumulating the mapping results of a plurality of the images to be recognized as a training data distribution on the evaluation space.

3. An image evaluation method according to claim 2, wherein the evaluation step further comprises: calculating the proportion of data contained in each region within a defined evaluation space based on the data distribution of the trained image; evaluating the data sufficiency within the evaluation space using the calculated proportion; and labeling each region according to the evaluation criteria.

4. An image evaluation method according to claim 1, wherein the image to be recognized is a verification image for verifying the performance of an object recognition model, the evaluation step involves inputting the verification image into the object recognition model, comparing the output of the object recognition model with a pre-created correct answer to calculate the performance of the object recognition model as an evaluation value, mapping the evaluation value onto an evaluation space with the first evaluation index and the second evaluation index as axes, and integrating the mapping results of a plurality of verification images as a performance distribution on the evaluation space to explain the performance characteristics of each space.

5. An image evaluation method according to claim 4, wherein the evaluation step further includes: extracting the performance included in each region within the evaluation space based on the performance distribution of the verification image; evaluating the performance achievement status of each region using the evaluation value; and labeling each region according to the evaluation criteria.

6. An image evaluation method according to claim 3 or claim 5, further comprising a training data extraction step of extracting images for retraining the object recognition model based on the labeling results of each of the regions.

7. An image evaluation method according to claim 1, wherein the multimodal base model comprises an image encoder and a language encoder, and in the vectorization step, the image encoder is input as the image to be recognized the entire processed image, or an extracted image cut out using a bounding box previously created from the processed image.

8. An image evaluation method according to claim 2, further comprising a language query generation step of generating a plurality of language queries by combining an input object with a plurality of templates, wherein the first evaluation index calculation step, the orthogonal vector extraction step, the second evaluation index calculation step, and the evaluation step are performed for each language query, and further comprising a deviation language query identification step of identifying the language query that has the largest deviation in the data distribution in the evaluation step.

9. An image evaluation method according to claim 4, further comprising a language query generation step of generating a plurality of language queries by combining an input object with a plurality of templates, wherein the first evaluation index calculation step, the orthogonal vector extraction step, the second evaluation index calculation step, and the evaluation step are performed for each language query, and further comprising a deviation language query identification step of identifying the language query that has the largest deviation of the performance features in the evaluation step.

10. An image evaluation system comprising: a vectorization unit that vectorizes language queries and images to be recognized into language vectors and image vectors using a multimodal platform model of language and images; a first evaluation index calculation unit that calculates a first evaluation index based on the similarity between the language vectors and the image vectors; an orthogonal vector extraction unit that extracts orthogonal vectors having an origin on the language vectors and pointing toward the image vectors; a second evaluation index calculation unit that calculates a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vectors and is orthogonal to the language vectors and the orthogonal vectors; and an evaluation unit that evaluates the images to be recognized using the first evaluation index and the second evaluation index.

11. An image evaluation system according to claim 10, wherein the vectorization unit, the first evaluation index calculation unit, the orthogonal vector extraction unit, the second evaluation index calculation unit, and the evaluation unit are all mounted on a vehicle or in a cloud environment.

12. An image evaluation system according to claim 10, wherein the image to be recognized is a verification image for verifying the performance of an object recognition model, the evaluation unit inputs the verification image to the object recognition model, compares the output of the object recognition model with a pre-created correct answer to calculate the performance of the object recognition model as an evaluation value, maps the evaluation value onto an evaluation space with the first evaluation index and the second evaluation index as axes, and integrates the mapping results of a plurality of verification images as a performance distribution on the evaluation space to show the performance characteristics of each space, the evaluation unit further extracts the performance included in each region within the evaluation space based on the performance distribution of the verification image, evaluates the performance achievement status of each region using the evaluation value, labels each region according to the evaluation criteria, and transmits the evaluation criteria to the vehicle, the vectorization unit, the first evaluation index calculation unit, the orthogonal vector extraction unit, the second evaluation index calculation unit, and the evaluation unit are installed in a cloud environment, and the vehicle includes an image collection unit that transmits images that satisfy the evaluation criteria received from the cloud environment to the cloud environment.