Image evaluation methods, image evaluation systems

The image evaluation method uses a multimodal base model to vectorize language and image queries, enabling precise classification and analysis of image features, addressing the challenge of varying object features and improving recognition model training through optimal data selection.

JP2026088960APending Publication Date: 2026-05-29ASTEMO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ASTEMO LTD
Filing Date
2024-11-19
Publication Date
2026-05-29

Smart Images

  • Figure 2026088960000001_ABST
    Figure 2026088960000001_ABST
Patent Text Reader

Abstract

This invention provides an image evaluation method and system that indexes objects to be recognized in an image based on their linguistic meaning and non-linguistic features. [Solution] The image analysis device 1, which is a computer that performs an image evaluation method, comprises: a vectorization unit 71 that vectorizes language queries and images to be recognized into language vectors and image vectors using a multimodal base model of language and images; a first evaluation index calculation unit 72 that calculates a first evaluation index based on the similarity between the language vectors and image vectors; an orthogonal vector extraction unit 73 that extracts orthogonal vectors that have an origin on the language vectors and point toward the image vectors; a second evaluation index calculation unit 74 that calculates a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes orthogonal vectors and is orthogonal to the language vectors and the orthogonal vectors; and an evaluation unit 75 that evaluates the images to be recognized using the first evaluation index and the second evaluation index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image evaluation method and an image evaluation system.

Background Art

[0002] In a machine learning model for image recognition, high performance is achieved by training the model using a large amount of labeled data. On the other hand, even for objects of the same category, their visual features vary widely. For example, in the case of vehicles, the features of vehicle images differ depending on the vehicle type, the color of the vehicle body, the shooting angle and size of the subject, the illuminance, the presence or absence of a shielding area, and the like. Such differences in image features affect the recognition accuracy of the machine learning model and also differ in the difficulty of learning. Therefore, in order to construct a highly versatile model that is not affected by differences in image features, it is necessary to prepare optimal learning data. Optimal learning data generally refers to a data set that includes target images with various features and contains a sufficient amount for objects with a high difficulty of recognition. However, it is not easy to classify the features of target images and determine a quantitative balance.

[0003] To solve this problem, a method called active learning has been proposed. Active learning is a method in which the model itself selects learning data and requests human beings to perform labeling, and learning can be effectively advanced even with a small amount of data. However, existing active learning methods have problems such that the criteria are general, such as simply selecting image data with low recognition performance, and it cannot be said that the selection is optimal based on image features and performance.

[0004] Patent Document 1 describes a technique for determining the attributes of a target object using a multimodal base model of language and images. However, it is limited to determining general attributes such as type, shape, and color of fruits, fish, and meat, and does not describe classifying differences in visual features within the same attribute. Non-Patent Document 1 describes a technique for clustering images of target objects using a multimodal base model of language and images, associating them with linguistic meaning, and assigning correct labels and recognition performance for object recognition to the resulting feature quantities. This describes a method for extracting scenes prone to recognition errors and small amounts of data included in the training data. [Prior art documents] [Patent Documents]

[0005] [Patent Document 1] Japanese Patent Publication No. 2024-091177 [Non-patent literature]

[0006] [Non-Patent Document 1] Sabri Eyuboglu, 7 others, "DOMINO: DISCOVERING SYSTEMATIC ERRORS WITH CROSS-MODAL EMBEDDINGS", [online], March 24, 2022, Cornell University, [searched October 28, 2020], Internet<URL:https: / / arxiv.org / abs / 2203.14960> [Overview of the project] [Problems that the invention aims to solve]

[0007] The technologies described in Patent Document 1 and Non-Patent Document 1 classify images using their association with linguistic features, which is useful for understanding recognition performance and trends in the distribution of training data. However, since object recognition models do not recognize objects based on linguistic features, classification based solely on features that can be expressed in language cannot fully explain the distribution of training data or its correlation with recognition performance. Furthermore, it is difficult to perfectly express image features in language, making this classification method insufficient for images that may contain features that are difficult to verbalize or abstract concepts. Moreover, even if a necessary and sufficient classification method exists, in order to obtain higher learning effectiveness, it is necessary to meticulously adjust the distribution of images included in the dataset used for training, requiring more detailed classification and analysis than the generally assumed number of hierarchies or clusters. In other words, the inventions described in Patent Document 1 and Non-Patent Document 1 cannot provide a means for indexing the objects to be recognized in images using their linguistic meaning and non-linguistic features. [Means for solving the problem]

[0008] A first aspect of the present invention is an image evaluation method performed by a computer, comprising: a vectorization step of vectorizing a language query and an image to be recognized into language vectors and image vectors using a multimodal base model of language and images; a first evaluation index calculation step of calculating a first evaluation index based on the similarity between the language vectors and the image vectors; an orthogonal vector extraction step of extracting orthogonal vectors having an origin on the language vectors and pointing toward the image vectors; a second evaluation index calculation step of calculating a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vectors and is orthogonal to the language vectors and the orthogonal vectors; and an evaluation step of evaluating the image to be recognized using the first evaluation index and the second evaluation index. A second aspect of the present invention provides an image evaluation system comprising: a vectorization unit that vectorizes a language query and a target image into language vectors and image vectors using a multimodal base model of language and images; a first evaluation index calculation unit that calculates a first evaluation index based on the similarity between the language vectors and the image vectors; an orthogonal vector extraction unit that extracts orthogonal vectors having an origin on the language vectors and pointing toward the image vectors; a second evaluation index calculation unit that calculates a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vectors and is orthogonal to the language vectors and the orthogonal vectors; and an evaluation unit that evaluates the target image using the first evaluation index and the second evaluation index. [Effects of the Invention]

[0009] According to the present invention, a means can be provided to index objects to be recognized in an image based on their linguistic meaning and non-linguistic features. Details, configuration, and effects of the above means will be revealed by the following embodiments. [Brief explanation of the drawing]

[0010] [Figure 1] Configuration diagram of an image analysis device that performs image analysis. [Figure 2] Hardware configuration diagram of the image analysis device [Figure 3] Diagram of the image classification unit and image language model. [Figure 4] Flowchart showing the processing of the image classification unit [Figure 5] Configuration diagram of the nonverbal feature index extraction unit [Figure 6] Diagram showing the processing of the nonverbal feature index extraction unit. [Figure 7] Flowchart showing the processing of the nonverbal feature index extraction unit. [Figure 8] A diagram showing an example of a training image count map. [Figure 9] A diagram showing an example of a performance map. [Figure 10] A flowchart showing the process of creating a training image count map executed by the performance analysis unit. [Figure 11] Flowchart showing the performance map creation process executed by the Performance Analysis Department [Figure 12] Flowchart showing the number evaluation process by the Classification Evaluation Department [Figure 13] Flowchart showing the performance evaluation process by the Classification Evaluation Department [Figure 14] Flowchart showing the processing of the Image Selection Department [Figure 15] Diagram showing the configuration of the image analysis device from different viewpoints [Figure 16] Configuration diagram of the image analysis device in the second embodiment [Figure 17] Configuration diagram of the image classification unit and the image language model in the second embodiment [Figure 18] Diagram showing an example of arranging the configuration on a cloud environment [Figure 19] Diagram showing an example of arranging the configuration on a vehicle [Figure 20] Diagram showing a second example of arranging the configuration on a cloud environment

Mode for Carrying Out the Invention

[0011] —First Embodiment— Hereinafter, referring to FIGS. 1 to 15, the first embodiment of the image analysis method will be described. In this embodiment, a “processing image” is an image including one or more objects. A “recognition target image” is an image including one object included in the processing image. Since the processing image may contain only one object, the processing image and the recognition target image may be the same. An object is, for example, a car, a person, a traffic signal, a road sign, etc.

[0012] FIG. 1 is a configuration diagram of an image analysis device 1 that executes an image analysis method. The image analysis device 1 includes a storage device 4 and a processing unit 5. The storage device 4 includes an object recognition model 41, an image language model 42, a learned image recording unit 43, a verification image recording unit 44, and a learning image recording unit 45. The processing unit 5 includes an image extraction unit 51, an image classification unit 52, a performance analysis unit 53, a classification evaluation unit 54, and an image selection unit 55.

[0013] The object recognition model 41 is a pre-trained model that detects objects in an image. The input to the object recognition model 41 is the processed image, and the output of the object recognition model 41 is the predicted label. A predicted label is a combination of the object's name and data indicating the object's position in the image, such as a bounding box. As mentioned above, in this embodiment, the processed image contains one or more objects. Therefore, the predicted label may contain two or more combinations of object names and bounding boxes. The trained image recording unit 43, the verification image recording unit 44, and the training image recording unit 45 each store a combination of the processed image and the correct label. The correct label is data indicating the object's name and its position in the image, and is, so to speak, a pre-created predicted label. The correct label can also be called a "pre-created correct answer." The difference between the correct label and the predicted label is whether it was pre-created or created by the object recognition model 41. The processed image and correct label stored in the trained image recording unit 43 were used to train the object recognition model 41. The processed images and correct labels stored in the training image recording unit 45 will be used for future retraining of the object recognition model 41. The training image recording unit 45 stores the processed images and correct labels through a process described later.

[0014] The image extraction unit 51 extracts the image to be recognized from the processed image using the predicted label output by the object recognition model 41, or the correct label contained in the trained image recording unit 43 or the verification image recording unit 44. When using the object recognition model 41, the image extraction unit 51 first inputs the processed image to the object recognition model 41 to obtain the predicted label. Then, the image extraction unit 51 obtains the image to be recognized by cropping the image from the processed image input to the object recognition model 41 based on the bounding box contained in the predicted label. When using the trained image recording unit 43 or the verification image recording unit 44, the image extraction unit 51 targets the processed image combined with the correct label and obtains the image to be recognized by cropping the image based on the bounding box contained in the correct label. The processing of the image classification unit 52 and the performance analysis unit 53 will be described later.

[0015] The classification and evaluation unit 54 performs either a quantity evaluation process using the training image count map 161, which will be described later, or a performance evaluation process using the performance map 162, which will be described later. Whether the classification and evaluation unit 54 performs the quantity evaluation process or the performance evaluation process may be specified by, for example, operator 90, by someone other than operator 90, determined in advance by some criteria, or determined randomly. The image selection unit 55 selects an image dataset for retraining from the training image recording unit 45 that matches the selection criteria acquired by the classification and evaluation unit 54.

[0016] Figure 2 is a hardware diagram of the image analysis device 1. The image analysis device 1 comprises a CPU 501, which is a central processing unit; a ROM 502, which is a read-only storage device; a RAM 503, which is a read-write storage device; an input / output device 504, which is a user interface; and a communication device 505. The CPU 501 loads the program stored in the ROM 502 into the RAM 503 and executes it, thereby realizing various calculations in the processing unit 5. The image analysis device 1 may be implemented using an FPGA (Field Programmable Gate Array), which is a rewritable logic circuit, or an ASIC (Application Specific Integrated Circuit), which is an application-specific integrated circuit, instead of the combination of CPU 501, ROM 502, and RAM 503. Alternatively, the image analysis device 1 may be implemented using a different combination of configurations, such as a combination of CPU 501, ROM 502, RAM 503 and FPGA, instead of the combination of CPU 501, ROM 502, and RAM 503.

[0017] In Figure 2, for convenience, the image analysis device 1 is shown as being composed of a single hardware device, but the image analysis device 1 may be composed of multiple hardware devices. In this case, the hardware devices may be installed adjacent to each other, or they may be connected via a local area network or the internet. The input / output device 504 is, for example, a keyboard that receives input from the operator 90. The communication device 505 enables the image analysis device 1 to communicate with other components. The communication device 505 is, for example, a communication interface. Note that the input / output device 504 and the communication device 505 may be configured as a single unit, and input from the operator 90 may be received via communication.

[0018] Figure 3 is a diagram showing the configuration of the image classification unit 52 and the image language model 42. The image classification unit 52 comprises a query generation unit 521, a similarity calculation unit 522, and a non-verbal feature index extraction unit 523. The image language model 42 comprises a language encoder 421 and an image encoder 422. The image language model 42 is a multimodal base model of language and images. The language encoder 421 converts the input language query 83 into a language vector 84. The image encoder 422 converts the input recognition target image 81 into an image vector 82. The language encoder 421 and the image encoder 422 are pre-trained so that the language vector 84 and the image vector 82 match when the linguistic meaning of the language query 83 and the recognition target image 81 match. The language vector 84 and the image vector 82 are vector data represented in the same vector space.

[0019] Operator 90 specifies the object 80 to the image classification unit 52. The object 80 is data that identifies an object and is synonymous with the name of the object included in the correct label. The object 80 is, for example, the string "automobile" or a predetermined code that indicates "automobile". The image classification unit 52 outputs the specified object 80 to the image extraction unit 51 and obtains the recognition target image 81 from the image extraction unit 51. If the object 80 is "automobile", the recognition target image 81 is an image of various automobiles extracted from the processed image. In Figure 3, it is shown as if there is only one recognition target image 81, but in reality there are often multiple recognition target images 81.

[0020] The query generation unit 521 of the image classification unit 52 generates a language query 83 using the object 80 specified by the operator 90 and inputs it to the language encoder 421 of the image language model 42. The query generation unit 521 may use the object 80 itself as the language query 83, or it may generate the language query 83 using a template (not shown). The language query 83 may be, for example, "car" or "a photo of a car". It is preferable that the language query 83 be input in the language language encoder 421 supports, and the query generation unit 521 may have a dictionary, or the object 80 may be described in a specific language. One language query 83 is generated for each object 80.

[0021] The language encoder 421 generates a language vector 84 corresponding to the language query 83 and outputs it to the image classification unit 52. Strictly speaking, the language query 83 is not directly input to the language encoder 421; preprocessing as described below is required. However, the preprocessing described here is just one example, and other methods may be used. Here, we will describe an example of tokenization, ID conversion, padding, embedding, and positional encoding as preprocessing. First, in tokenization, the input text is divided into words and subwords. For example, the sentence "A cute dog playing on the grass" is divided into "A", "cute", "dog", "playing", "on", "the", "grass". Next, each token is converted into a number, i.e., an ID. This is necessary for the model to treat the text as a number. For example, "dog" is converted to ID 1234, "playing" to ID 5678, etc., based on a predefined dictionary.

[0022] Then, to maintain a consistent text length, specific IDs are added to shorter texts for padding. This enables batch processing. In the next embedding step, each token ID is converted into a high-dimensional vector. The embedding is a vector representation that contains the semantic information of the tokens. Next, a positional encoding is added to preserve the order information of the tokens. This allows the model to understand the order of the tokens. After these preprocessing steps, these vectors are input to the language encoder 421, and a vector representing the meaning of the text is output.

[0023] The image classification unit 52 outputs the image to be recognized 81 to the image encoder 422 of the image language model 42. The image encoder 422 generates an image vector 82 corresponding to the input image to be recognized 81 and outputs it to the image classification unit 52. The image classification unit 52 inputs the language vector 84 output by the language encoder 421 to the similarity calculation unit 522 and the non-verbal feature index extraction unit 523, and inputs the image vector 82 output by the image encoder 422 to the similarity calculation unit 522 and the non-verbal feature index extraction unit 523. In other words, the image vector 82 and the language vector 84 are input to the similarity calculation unit 522 and the non-verbal feature index extraction unit 523, respectively. The similarity calculation unit 522 calculates a language feature index 88 using the image vector 82 and the language vector 84. The non-verbal feature index extraction unit 523 calculates a non-verbal feature index 89 using the image vector 82 and the language vector 84. The processing of the nonverbal feature index extraction unit 523 will be explained in detail later with reference to Figures 5 to 7. In the following, the verbal feature index 88 may be referred to as the "first index," and the nonverbal feature index 89 may be referred to as the "second index."

[0024] The similarity calculation unit 522 calculates the similarity between two input vectors, namely the image vector 82 and the language vector 84. For example, the similarity calculation unit 522 calculates the cosine similarity between the two vectors, that is, the cosine value of the angle between the two vectors. Specifically, the similarity calculation unit 522 calculates the cosine similarity by dividing the dot product of the two vectors by the product of the magnitudes of the two vectors. The similarity calculation unit 522 outputs the calculated similarity as a linguistic feature index 88. As mentioned above, there may be multiple recognition target images 81 for a single object 80. In that case, the similarity calculation unit 522 receives two or more image vectors 82 and one language vector 84 as input. In this case, the similarity calculation unit 522 calculates the similarity between the language vector 84 and each image vector 82 and outputs them as a linguistic feature index 88.

[0025] Figure 4 is a flowchart of the processing of the image classification unit 52. First, in step S301, the image classification unit 52 receives an object 80, for example, the string "automobile," from the operator 90. In the following step S302, the image classification unit 52 has a query generation unit 521 that generates a language query 83 based on the object 80. In the following step S303, the image classification unit 52 outputs the language query 83 generated in step S302 to the language encoder 421 and obtains a language vector 84 from the language encoder 421. In the following step S304, the image classification unit 52 uses an image extraction unit 51 to obtain the recognition target image 81. In the following step S305, the image classification unit 52 inputs the recognition target image 81 obtained in step S304 to the image encoder 422 and obtains an image vector 82 from the image encoder 422. In the following step S306, the image classification unit 52 has a similarity calculation unit 522 that calculates the similarity between the image vector 82 and the language vector 84 as a linguistic feature index 88. In the following step S307, the image classification unit 52 has a non-linguistic feature index extraction unit 523 that uses the image vector 82 and the language vector 84 to calculate a non-linguistic feature index 89, and the process shown in Figure 4 is completed. The details of step S307 will be explained later with reference to Figures 5 and 7.

[0026] Figure 5 shows the correlation of data calculated by the non-verbal feature index extraction unit 523. The non-verbal feature index extraction unit 523 receives image vectors 82 and language vectors 84 as input. The non-verbal feature index extraction unit 523 generates a reference vector 153 orthogonal to the language vector 84. The non-verbal feature index extraction unit 523 also calculates an inter-vector angle 151, which is the angle between the language vector 84 and the image vector 82. Furthermore, the non-verbal feature index extraction unit 523 calculates a non-verbal feature vector 154 using the image vector 82, language vector 84, and inter-vector angle 151. Finally, the non-verbal feature index extraction unit 523 calculates the similarity between the reference vector 153 and the non-verbal feature vector 154 as a non-verbal feature index 89.

[0027] Figure 6 is a vector correlation diagram, and Figure 7 is a diagram showing the processing of the non-verbal feature index extraction unit 523. The non-verbal feature index extraction unit 523 receives a language vector 84 output by the language encoder 421 and an image vector 82 output by the image encoder 422 as inputs. In step S311, the non-verbal feature index extraction unit 523 first determines an arbitrary reference vector 153 orthogonal to the language vector 84. In the following step S312, the non-verbal feature index extraction unit 523 calculates the vector-to-vector angle 151 between the image vector 82 and the language vector 84.

[0028] In the following step S313, the non-verbal feature index extraction unit 523 uses the language vector 84, the image vector 82, and the inter-vector angle 151 to extract an orthogonal vector as the non-verbal feature vector 154 that has its origin on the language vector 84 and points toward the endpoint of the image vector 82. As shown in Figure 6, once the language vector 84, the image vector 82, and the inter-vector angle 151 are determined, the non-verbal feature vector 154 that satisfies the conditions is uniquely determined. Also, as shown in the lower part of Figure 6, the reference vector 153 and the non-verbal feature vector 154 lie on the same plane.

[0029] In the following step S314, the non-verbal feature index extraction unit 523 calculates the similarity between the reference vector 153 and the non-verbal feature vector 154. For example, the non-verbal feature index extraction unit 523 calculates cosine similarity by calculating the cosine value of the angle between the two vectors. The similarity obtained in this way is output as a non-verbal feature index 89. In the following step S315, it is determined whether all image vectors 82 have been processed. If it is determined that all image vectors 82 have been processed, the process shown in Figure 7 is terminated. If it is determined that there are unprocessed image vectors 82, the process returns to step S312.

[0030] Furthermore, as can be seen from the fact that the return destination is step S312, the reference vector 153 is not changed. In other words, even when one language vector 84 and multiple image vectors 82 are input to the non-verbal feature index extraction unit 523, the non-verbal feature index extraction unit 523 determines only one reference vector 153 and calculates all non-verbal feature indices 89 using that reference vector 153.

[0031] The linguistic feature index 88 is an index that expresses the similarity to a linguistic query 83 that has linguistic meaning, and can be said to be an index that reflects linguistic features. On the other hand, the same value of the linguistic feature index 88 can be extracted from image vectors 82 that have different vector components. This is because, when a linguistic vector 84 is defined, there can be multiple image vectors 82 that have the same inter-vector angle 151. In other words, even though the components of the image vector 82 that express the features of the image are different, the closeness to the linguistic meaning represented by the linguistic query 83 is treated as equivalent. This is equivalent to, for example, the existence of multiple images that are judged to be semantically equivalent to the word "car". Since there are various images of cars, it is natural to treat them all as images of cars linguistically, but this classification is insufficient for purposes such as analyzing the performance of the object recognition model 41 and linking it to the distribution of training data.

[0032] The non-verbal feature index 89 is calculated by separating the image vector 82 into components orthogonal to the language vector 84 and comparing them with a reference vector 153 that has equivalent features. This makes it possible to treat image vectors 82 having different vector components that show the same value of the language feature index 88 as different image vectors 82. Here, the components orthogonal to the language vector 84 can be considered as the components that least reflect the linguistic features represented by the language query 83, and are therefore called non-verbal features.

[0033] The performance analysis unit 53 uses the linguistic feature index 88 and non-linguistic feature index 89 calculated by the image classification unit 52 to analyze the correlation between the data distribution of the image to be recognized 81 and the recognition performance when using the object recognition model 41. Specifically, the performance analysis unit 53 defines an evaluation space 160 with the linguistic feature index 88 on one axis and the non-linguistic feature index 89 on the other axis, and maps the data distribution or index of multiple image to be recognized 81. In the evaluation space 160, image to be recognized 81 with similar linguistic and non-linguistic features can be mapped to similar spaces, so the similarity of multiple image to be recognized 81 can be quantitatively determined.

[0034] Figures 8 and 9 show examples of evaluations performed by the performance analysis unit 53, with Figure 8 being the training image map 161 and Figure 9 being the performance map 162. In this embodiment, the three-dimensional space where the evaluation values ​​created by the performance analysis unit 53 are shown, as in Figures 8 and 9, is called the evaluation space 160. In the evaluation space 160, the linguistic feature index 88 is set on the vertical axis and the non-linguistic feature index 89 is set on the horizontal axis, and the height axis can be one of two types as will be explained below. In this embodiment, the plane in the evaluation space 160 is divided into predetermined regions, and each region is called a "grid". For example, if the linguistic feature index 88 is divided into 10 and the non-linguistic feature index 89 is also divided into 10, 100 grids are defined.

[0035] Figure 8 shows the distribution of the number of training images, i.e., the number of recognition target images 81 extracted from the trained image recording unit 43, on the height axis. Figure 9 shows the performance index, i.e., the recognition rate when recognition is performed using the object recognition model 41, on the height axis. In Figure 8, areas with a large number of images are shown in white, and areas with a small number of images are shown in black. In Figure 8, the distribution of multiple recognition target images 81 is represented by a smooth 3D curved surface, and it can be seen that there are white areas where images are concentrated, i.e., peaks and surrounding areas with a relatively small number of images.

[0036] On the other hand, Figure 9 shows an example of the recognition rate of object recognition model 41, with higher recognition rates indicated in white and lower recognition rates in black. In Figure 8, the white areas where images are concentrated are represented as black valleys in Figure 9, indicating that the recognition rate in these areas is extremely low. In other words, despite the large number of images used for training, these are areas where recognition is difficult for object recognition model 41, and therefore areas with high data value for retraining.

[0037] On the other hand, it can be seen that there are regions where the recognition rate is high in areas with a relatively small number of trained image distributions. Such regions are easy for the object recognition model 41 to recognize, and can be judged as regions where high recognition performance can be expected with a small amount of training data. In other words, these are regions where the data value for retraining is not very high. Through this visualization and analysis, it becomes possible to evaluate each region of the evaluation space 160 and the images classified into those regions.

[0038] The training image map 161 maps each reference recognition target image 81A into the evaluation space 160 based on the linguistic and non-linguistic features of the images. Images mapped to the same region in the evaluation space 160 have similar linguistic and non-linguistic features. Therefore, the evaluation space 160 can also be viewed as a grid space divided by a certain range width using linguistic feature index 88 and non-linguistic feature index 89.

[0039] Figure 10 is a flowchart showing the process of creating a training image map executed by the performance analysis unit 53. First, in step S320, the performance analysis unit 53 initializes each grid of the evaluation space 160. For example, if 100 grids are defined, this process assigns zero to an integer array having 100 elements. In the following step S321, all processed images and correct labels are extracted from the trained image recording unit 43 and the process proceeds to step S322. The combinations of processed images and correct labels extracted in step S321 are used in order in step S322, as will be described later. In step S322, the performance analysis unit 53 outputs one set of unprocessed processed images and correct labels to the image extraction unit 51, and extracts all recognition target images 81 from those processed images based on the correct labels, proceeding to step S323. The recognition target images 81 extracted in step S322 are used in order in step S323, as will be described later.

[0040] In step S323, the performance analysis unit 53 uses the unprocessed recognition target image 81 extracted in step S322 to cause the image classification unit 52 to calculate two indices, namely a linguistic feature index 88 and a non-linguistic feature index 89. The processing of the image classification unit 52 at this time is generally as described with reference to Figure 4, but is modified as follows. That is, the image classification unit 52 uses the name of an object included in one of the correct labels output from the performance analysis unit 53 instead of the object 80 received from the operator 90 in step S301. Also, instead of executing step S304, it is treated as if the aforementioned unprocessed recognition target image 81 has been extracted.

[0041] In the following step S324, the performance analysis unit 53 maps the recognition target image 81 to the evaluation space 160 and proceeds to step S325. Specifically, the performance analysis unit 53 increases the grid value corresponding to the combination of the two indicators calculated in step S303 by 1. In the following step S325, the performance analysis unit 53 determines whether or not all of the recognition target images 81 extracted in step S322 have been processed. If the performance analysis unit 53 determines that all of the recognition target images 81 extracted in step S322 have been processed, it proceeds to step S326; if it determines that there are unprocessed recognition target images 81, it returns to step S323.

[0042] In step S326, the performance analysis unit 53 determines whether all the processing images extracted in step S321 have been processed. If the performance analysis unit 53 determines that all processing images have been processed, it terminates the process shown in Figure 10. If it determines that there are unprocessed processing images, it returns to step S322. The training image count map 161 can be created by counting the number of images corresponding to each grid, so data creation is completed by repeating the voting process in step S324 for all processing images for all recognition target images 81. In addition to the process shown in Figure 10, the performance analysis unit 53 may also create three-dimensional graph data as shown in Figure 8.

[0043] Figure 11 is a flowchart showing the performance map creation process executed by the performance analysis unit 53. First, in step S330, the performance analysis unit 53 initializes each grid in the evaluation space 160. For example, if 100 grids are defined, this process assigns zero to all elements of a multidimensional array having 100 elements, each element capable of storing any number of integers or decimals. In the following step S331, all processed images and correct labels are extracted from the verification image recording unit 44, and the process proceeds to step S332. The combinations of processed images and correct labels extracted in step S331 are used in order in step S332, as will be described later.

[0044] In step S332, the performance analysis unit 53 outputs an unprocessed set of processed images and a ground truth label to the image extraction unit 51, and uses the bounding box included in the ground truth label to extract all recognition target images 81 from the processed images, proceeding to step S333. However, in the performance image count map creation process, the recognition target images 81 extracted using the bounding box included in the ground truth label are called reference recognition target images 81A. If multiple reference recognition target images 81A are extracted in step S332, they are used in order in step S333 as described later. The processed images and ground truth labels output to the image extraction unit 51 in this step are called "specific processed images" and "specific ground truth labels." These specific processed images and specific ground truth labels remain unchanged until a positive determination is made in step S337, which will be described later.

[0045] In step S333, the performance analysis unit 53 uses the unprocessed reference recognition target image 81A extracted in step S332 to cause the image classification unit 52 to calculate two indices, namely a linguistic feature index 88 and a non-linguistic feature index 89. The processing of the image classification unit 52 at this time is generally as explained with reference to Figure 4, but is modified as follows. That is, the image classification unit 52 uses the name of the object included in the specific correct label instead of the object 80 received from the operator 90 in step S301. Also, instead of executing step S304, it is treated as if the aforementioned unprocessed reference recognition target image 81A has been extracted. In the following step S334, the performance analysis unit 53 uses the object recognition model 41 to extract the recognition target image 81. However, the recognition target image 81 extracted in this step is called the evaluation recognition target image 81B. Specifically, the performance analysis unit 53 inputs the names of the objects included in the specific correct labels and the specific processed images into the object recognition model 41 to obtain the recognition target image 81B for evaluation.

[0046] In the following step S335, the performance analysis unit 53 compares and evaluates the reference recognition target image 81A and the evaluation recognition target image 81B. This evaluation may be a recognition rate representing the success or failure of recognition, or an IoU (Intersection over Union) representing the degree of overlap between the two regions of the reference recognition target image 81A and the evaluation recognition target image 81B; it is not limited to a specific performance indicator. In the following step S336, the performance analysis unit 53 maps the evaluation results from step S335 to the evaluation space 160. Mapping in this step means storing the evaluation results from step S335 in the multidimensional array initialized in step S330. In step S324 of Figure 10, the value of the corresponding grid was increased by 1, but this step differs in that it saves the evaluation results.

[0047] In the following step S337, the performance analysis unit 53 determines whether all of the recognition target images 81 extracted in step S332 have been processed. If the performance analysis unit 53 determines that all of the recognition target images 81 extracted in step S332 have been processed, it proceeds to step S338; otherwise, it returns to step S333. In step S338, the performance analysis unit 53 determines whether all of the processing images extracted in step S331 have been processed. If the performance analysis unit 53 determines that all of the processing images have been processed, it proceeds to step S339; otherwise, it returns to step S332.

[0048] In step S339, the performance analysis unit 53 aggregates the evaluation results mapped to the evaluation space 160 and completes the process shown in Figure 11. Specifically, the performance analysis unit 53 calculates the average value of the evaluation results for each grid in the evaluation space 160. Since the individual evaluation results were saved in step S336, the average value can be calculated in this step. In addition to the process shown in Figure 11, the performance analysis unit 53 may also create three-dimensional graph data as shown in Figure 9.

[0049] Figure 12 is a flowchart showing the image count evaluation process by the classification evaluation unit 54. First, in step S341, the classification evaluation unit 54 calculates the total number of recognition target images 81 mapped to the training image count map 161. In the following step S342, the classification evaluation unit 54 extracts, in other words, counts, the number of recognition target images 81 in each grid of the training image count map 161. In the following step S343, the classification evaluation unit 54 calculates the ratio of the extracted number of images to the total number for each grid. In the following step S344, the classification evaluation unit 54 evaluates the data sufficiency of each grid in the evaluation space 160 according to the evaluation criteria. The evaluation criteria may be whether the number of images is 1 or more, or whether the calculated ratio is greater than or equal to a predetermined value.

[0050] In the following step S345, the classification and evaluation unit 54 labels each grid in the evaluation space 160 according to the evaluation results of step S344. Specifically, the classification and evaluation unit 54 assigns a label indicating that the fewer the number of recognition target images 81 included in the grid, the higher the value of the images in that grid. The label may be a binary value indicating high or low value, or it may be an identification value of 3 or more indicating high value, for example, a number from 1 to 10. In the following step S346, the classification and evaluation unit 54 sets the required amount of data for each grid in the evaluation space 160 based on the labeling results and terminates the process shown in Figure 12. For example, based on the labeling results, the classification and evaluation unit 54 can determine data selection criteria for each grid so that the number of recognition target images 81 is reduced in grids where the number of recognition target images 81 is already excessive, and so that the number of recognition target images 81 is increased in areas where there is a shortage of recognition target images 81. As a result, it is possible to provide guidance for a suitable data balance in the entire evaluation space 160.

[0051] Figure 13 is a flowchart of the performance evaluation process performed by the classification and evaluation unit 54. However, some of the processes shown in Figure 13 are similar to those in Figure 12. In step S351, the classification and evaluation unit 54 extracts the performance mapped to each grid in the performance map 162. This step S351 corresponds to steps S341 to S343 in Figure 12. In the following step S354, the classification and evaluation unit 54 evaluates the data sufficiency of each grid in the evaluation space 160 according to the evaluation criteria, similar to step S344 in Figure 12. The evaluation criteria used in this step are different from those used in step S344 in Figure 12. In the following step S355, the classification and evaluation unit 54 labels the data value of each grid in the evaluation space 160 according to the evaluation results, similar to step S345 in Figure 12. In the following step S356, the classification and evaluation unit 54 sets the required amount of data for each grid in the evaluation space 160 based on the labeling results, similar to step S346 in Figure 12, and terminates the process shown in Figure 13.

[0052] Figure 14 is a flowchart showing the processing of the image selection unit 55. First, in step S360, the image selection unit 55 extracts all processing images from the training image recording unit 45 and proceeds to step S361. The processing images extracted in this step are used sequentially in steps S361 to S364, as will be described later. In step S361, the image selection unit 55 uses the object recognition model 41 to extract all recognition target images 81 from the processing images and proceeds to step S362. However, if the processing images have correct labels, the recognition target images 81 may be extracted using the bounding boxes of the correct labels.

[0053] In step S362, the image selection unit 55 selects one of the unprocessed recognition target images 81 extracted in step S361 and has the image classification unit 52 calculate two indices, namely a linguistic feature index 88 and a non-linguistic feature index 89, before proceeding to step S363. If the recognition target image 81 was extracted in step S361 using the object recognition model 41, the name of the object included in the predicted label output by the object recognition model 41 is set as the target object 80. If the recognition target image 81 was extracted in step S361 using the correct label assigned to the processed image, the name of the object included in that correct label is set as the target object 80. In step S363, the image selection unit 55 determines whether all recognition target images 81 extracted in step S361 have been processed. If the image selection unit 55 determines that all recognition target images 81 have been processed, it proceeds to step S364; if it determines that there are unprocessed recognition target images 81, it returns to step S362.

[0054] In step S364, the image selection unit 55 aggregates the distribution of the number of recognition target images 81 included in the processing images and proceeds to step S365. Specifically, the image selection unit 55 aggregates the number of images included in each grid of the evaluation space 160. In the following step S365, the image selection unit 55 determines whether all the processing images extracted in step S360 have been processed. If the image selection unit 55 determines that all processing images have been processed, it proceeds to step S366; if it determines that there are unprocessed processing images, it returns to step S361.

[0055] In step S366, the image selection unit 55 selects a predetermined number of processed images from the processed images extracted in step S360. This number may be a predetermined number, or it may be a random value determined each time step S366 is executed. In the following step S367, the image selection unit 55 sums the distribution of the number of recognition target images 81 included in the selected processed images. Since the number of recognition target images 81 included in each grid was previously aggregated for each processed image in step S364, in this step it is only necessary to aggregate for the processed images selected in step S366.

[0056] In the following step S368, the image selection unit 55 determines whether the aggregation result in step S367 matches the selection criteria. If the image selection unit 55 determines that it matches the selection criteria, it proceeds to step S369; otherwise, it returns to step S366. In step S369, the image selection unit 55 adopts the processed images selected in step S366 as a dataset for retraining and terminates the process shown in Figure 14. This makes it possible to obtain a dataset suitable for retraining based on the linguistic and non-linguistic features of the images and the performance of the object recognition model 41.

[0057] Figure 15 is a diagram that organizes some of the components of the image analysis device 1 from a different perspective than Figures 1 and 3. The image analysis device 1 comprises a vectorization unit 71, a first evaluation index calculation unit 72, an orthogonal vector extraction unit 73, a second evaluation index calculation unit 74, and an evaluation unit 75. The vectorization unit 71 uses an image language model 42, which is a multimodal base model of language and images, to vectorize the language query 83 and the recognition target image 81 into a language vector 84 and an image vector 82. The vectorization unit 71 is a query generation unit 521 that outputs the language query 83 to a language encoder 421, and an image classification unit 52 that outputs the recognition target image 81 to an image encoder 422.

[0058] The first evaluation index calculation unit 72 calculates a first evaluation index, i.e., a linguistic feature index 88, based on the similarity between the language vector 84 and the image vector 82. The first evaluation index calculation unit 72 is the similarity calculation unit 522. The orthogonal vector extraction unit 73 extracts an orthogonal vector, i.e., a non-linguistic feature vector 154, which has its origin on the language vector 84 and points toward the image vector 82. The orthogonal vector extraction unit 73 corresponds to a part of the function of the non-linguistic feature index extraction unit 523, and specifically corresponds to the processing in step S313 shown in Figure 7. The second evaluation index calculation unit 74 calculates a second evaluation index, i.e., a non-linguistic feature index 89, based on the similarity between the non-linguistic feature vector 154 and a predetermined reference vector 153 on a plane that includes the non-linguistic feature vector 154 and is orthogonal to the language vector 84. The second evaluation index calculation unit 74 corresponds to a part of the function of the non-linguistic feature index extraction unit 523, and specifically corresponds to the processing in step S314 shown in Figure 7. The evaluation unit 75 evaluates the recognition target image 81 using linguistic feature index 88 and non-linguistic feature index 89. The evaluation unit 75 combines the functions of both the performance analysis unit 53 and the classification evaluation unit 54.

[0059] According to the first embodiment described above, the following effects and advantages can be obtained. (1) The image evaluation method performed by the computer-based image analysis device 1 includes a vectorization step (S303, S305 in Figure 4) in which a language query 83 and a target image 81 are vectorized into a language vector 84 and an image vector 82 using an image language model 42, which is a multimodal base model of language and images; a first evaluation index calculation step (S306 in Figure 4) in which a language query 83 and a target image 81 are vectorized into a language vector 84 and an image vector 82, respectively; and an orthogonal vector having its origin on the language vector 84 and pointing toward the image vector 82, i.e. The process includes an orthogonal vector extraction step (S313 in Figure 7) for extracting non-verbal feature vectors 154, a second evaluation index calculation step (S314 in Figure 7) for calculating a non-verbal feature index 89 based on the similarity between the non-verbal feature vector 154 and a predetermined reference vector 153 on a plane that includes the non-verbal feature vectors 154 and is orthogonal to the language vectors 84, and the non-verbal feature index 154, and an evaluation step (training image map creation process, performance map creation process, image count evaluation process, performance evaluation process) for evaluating the recognition target image 81 using the language feature index 88 and the non-verbal feature index 89. Therefore, recognition target objects contained in an image can be indexed by their language meaning and non-verbal features. Specifically, it is possible to calculate the numerical values ​​necessary for creating the training image map 161 shown in Figure 8 and the performance map 162 shown in Figure 9.

[0060] (2) The images to be recognized 81 are trained images used to train the object recognition model 41. In the training image count map creation process and the count evaluation process, the images to be recognized 81 are mapped onto an evaluation space 160 with linguistic feature index 88 and non-linguistic feature index 89 as axes (S324 in Figure 10), and the mapping results of multiple images to be recognized 81 are accumulated on the evaluation space 160 as a training data distribution (S324 in Figure 10) to evaluate the data distribution in each space. The training image count map 161 is a visualization of this data distribution.

[0061] (3) Based on the data distribution of the trained images, the data proportion included in each region within the defined evaluation space is extracted (S343 in Figure 12), the data sufficiency within the evaluation space is evaluated from the data proportion extraction results (S344), and each region is labeled according to the evaluation criteria (S345). Therefore, the value of each grid in the evaluation space 160, or more precisely, the value of the image corresponding to each grid, can be labeled based on the number of trained images.

[0062] (4) In the evaluation step, the verification images are input to the object recognition model 41, and the output of the object recognition model 41 is compared with the correct labels to calculate the performance of the object recognition model 41 as an evaluation value (S334, S335 in Figure 11). The evaluation value is mapped onto an evaluation space 160 with the first and second evaluation indicators as axes (S336), and the mapping results of multiple verification images are integrated as a performance distribution on the evaluation space 160 to evaluate the performance characteristics of each space (S339).

[0063] (5) Based on the performance distribution of the verification images, the performance included in each region within the defined evaluation space is extracted (S351 in Figure 13), the performance achievement status of each region within the evaluation space is evaluated from the performance extraction results (S354), and each region is labeled according to the evaluation criteria (S355).

[0064] (6) As shown in Figure 14, the image selection unit 55 extracts images for the object recognition model 41 to retrain based on the labeling results of each region within the defined evaluation space.

[0065] (7) The image language model 42 comprises a language encoder 421 and an image encoder 422. The image encoder 422 receives the entire processing image, or an extracted image cut out from the processing image using a pre-created bounding box, as the image to be recognized 81.

[0066] (Variation 1) In the first embodiment described above, either a count evaluation process to evaluate the training count map 161 or a performance evaluation process to evaluate the performance map 162 was performed. However, the training count map 161 and the performance map 162 may be analyzed in combination to determine the data selection criteria. The performance evaluation process using the performance map 162 reveals the performance achievement status of each grid in the evaluation space 160. In addition, the count evaluation process using the training count map 161 reveals the distribution of trained images in each grid. Therefore, data selection criteria can be determined for each grid based on the distribution of the training count map 161. For example, data selection criteria can be determined for each grid so that the proportion of areas that have achieved performance is less than the proportion of known trained images, and the proportion of areas that have not achieved performance is higher. This makes it possible to determine the sufficiency of each grid based on the features of the trained images and the performance of the object recognition model 41, and to provide guidance for a suitable data balance in the entire evaluation space 160.

[0067] (Modification 2) In the first embodiment described above, only one indicator was calculated in the performance evaluation process. However, multiple indicators may be calculated, and multiple performance maps 162 may be created. In the performance evaluation process described in the first embodiment, a suitable data balance within the evaluation space 160 can be shown for a given performance indicator. Therefore, a suitable data balance can also be calculated for different performance indicators and target values. For example, a suitable data balance can be calculated based on two performance indicators: recognition rate and IoU. When using multiple performance guidelines, a guideline for a suitable data balance that takes multiple performance indicators into account can be shown by methods such as summing, averaging, or weighted averaging the data balance of each guideline.

[0068] (Variation 3) In the first embodiment described above, the classification and evaluation unit 54 was capable of performing both a count evaluation process to evaluate the learning count map 161 and a performance evaluation process to evaluate the performance map 162. However, the classification and evaluation unit 54 may only be capable of performing either the count evaluation process or the performance evaluation process.

[0069] (Modification 4) In the first embodiment described above, the image selection unit 55 selected an image dataset for retraining from the training image recording unit 45, but the retraining of the object recognition model 41 was not performed. However, the processing unit 5 may further include a retraining unit that performs the retraining of the object recognition model 41. Once the retraining of the object recognition model 41 is completed by the retraining unit, the image dataset selected by the image selection unit 55 is stored in the trained image recording unit 43. This makes it possible to generate a training image count map 161 and a performance map 162 that are different from those before retraining. By reviewing the optimal data balance using the newly generated training image count map 161 and performance map 162, the data balance can be dynamically adjusted.

[0070] —Second Embodiment— A second embodiment of the image evaluation method will be described with reference to Figures 16 and 17. In the following description, the same reference numerals are used for components that are the same as in the first embodiment, and the differences will be mainly explained. Points that are not specifically explained are the same as in the first embodiment. This embodiment differs from the first embodiment mainly in that it generates various language queries 83 and identifies language queries 83 in which the training image count map 161 and performance map 162 show significant bias.

[0071] Figure 16 is a configuration diagram of the image analysis device 1A in the second embodiment. The difference from the first embodiment is that the processing unit 5 has a deviation language query identification unit 56. Also, the configuration and processing of the image classification unit 52A are slightly different from the first embodiment, as will be described later. The image classification unit 52A generates multiple language queries 83, as will be described later. In this embodiment, the processing of the non-language feature index extraction unit 523 shown in Figure 5, which includes the calculation of linguistic feature indexes 88 and the calculation of non-language feature vectors 154, also called orthogonal vectors, and the calculation of non-language feature indexes 89 are performed for each generated language query 83.

[0072] The deviation language query identification unit 56 compares multiple training image count maps 161 or multiple performance maps 162 corresponding to multiple language queries 83 to identify the language query 83 with the largest deviation in the data distribution in the training image count map 161 or the largest deviation in the performance features in the performance map 162. The deviation in the data distribution in the training image count map 161 is, for example, the difference between the minimum and maximum number of images in each grid. The deviation in the performance features in the performance map 162 is, for example, the difference between the minimum and maximum recognition rate in each grid.

[0073] Figure 17 is a diagram showing the configuration of the image classification unit 52 and the image language model 42 in the second embodiment. The difference from Figure 3 in the first embodiment is that the query generation unit 521 reads template 521A. Template 521A contains multiple templates. By using this template 521A, the query generation unit 521 generates language queries 83 such as "car", "A photo of a car", "An image of a car from any angle", "A photo of a car at night", and "An image of a car driving in the dark" when the object 80 is "car".

[0074] By changing the language query 83 in this way, the linguistic feature index 88 changes, and as a result, the non-linguistic feature index 89 extracted using the language vector 84 extracted from the language query 83 also changes. In other words, the distribution of the number of training images map 161 and the performance map 162 also changes. To put it another way, by preparing multiple language queries 83, it is possible to generate diverse numbers of training images map 161 and performance map 162. Therefore, by searching for the language query 83 in which the desired performance difference is most pronounced, it becomes possible to accurately analyze the conditions under which the object recognition model 41 can easily recognize objects or the conditions under which it has difficulty recognizing objects. It is also possible to analyze the conditions under which the bias of the trained images is most pronounced. The conditions obtained in this way provide useful information for constructing a dataset for retraining.

[0075] According to the second embodiment described above, the following effects and advantages can be obtained. (8) Includes a language query generation step that generates multiple language queries 83 by combining the input object 80 with multiple templates 521A. The processing of the non-language feature index extraction unit 523 shown in Figure 5, which includes the calculation of linguistic feature indexes 88 and non-language feature vectors 154, also called orthogonal vectors, and the calculation of non-language feature indexes 89 are performed for each generated language query 83. Includes a deviation language query identification process that identifies the language query 83 with the largest deviation in the data distribution in the training image count map 161 or the performance feature deviation in the performance map 162. This allows for accurate analysis of conditions under which the object recognition model 41 can easily recognize objects or conditions under which it has difficulty recognizing objects.

[0076] The technologies described in each of the embodiments and modifications above can be used for a variety of applications. For example, this technology can be used as an evaluation and retraining technology for an object recognition model 41 for autonomous driving of automobiles. In autonomous driving of automobiles, an object recognition model 41 for recognizing objects is deployed on the vehicle body. Generally, an object recognition model 41 is not a one-time deployment; it is often configured to be replaceable when defects, insufficient performance, or the need for functional updates arise. For example, it is conceivable to accumulate data collected from the automobile in a cloud environment, update the object recognition model 41 based on that data, and then redeploy the updated object recognition model 41 to the vehicle.

[0077] One concept known as a connected car involves a vehicle that is constantly connected to the internet, collecting data from sensors and cameras, as well as operational data, and distributing necessary information and software, including an object recognition model 41, to the vehicle. Using such a system, as shown in Figure 18, the collected images 116 from the vehicle 20 can be stored in the learning image recording unit 45, which stores candidate images for retraining on the cloud environment 9, and used for retraining using this technology. The retrained object recognition model 41 is then distributed back to the vehicle 20, and the updated new object recognition model 41A is used in the vehicle 20.

[0078] Alternatively, as shown in Figure 19, the functions of this technology can be deployed on the vehicle 20. In this case, an image acquisition unit 6, such as a camera, is mounted on the vehicle 20A, and the collected images 116 obtained by the image acquisition unit 6 are recorded in the learning image recording unit 45. With this configuration, the object recognition model 41 can be updated without connecting to the cloud environment 9. Furthermore, as shown in Figure 20, it is also possible to deploy all or part of the functions included in the processing unit 5 on the vehicle 20B and perform the analysis of the collected images 116 on the vehicle 20B. In this case, by sharing the data selection criteria 117 determined by the classification evaluation processing unit 124 from the cloud environment 9 with the vehicle 20B, only images with high data value can be selected from the collected images 116 and sent to the cloud environment 9. This reduces the amount of data of collected images 116 sent to the cloud environment 9, and allows for the efficient collection and storage of only data with high retraining value.

[0079] A similar mechanism can be applied to applications using the object recognition model 41 in factories, commercial facilities, buildings, train stations, etc. This technology can be applied to applications where the object recognition model 41 is retrained using images collected during the operation of the application, and the updated object recognition model 41A is reused.

[0080] In the embodiments and modifications described above, the configuration of the functional blocks is merely an example. Several functional configurations shown as separate functional blocks may be integrated, or a configuration represented in one functional block diagram may be divided into two or more functions. Furthermore, some of the functions of one functional block may be provided by other functional blocks.

[0081] In the embodiments and modifications described above, the program is stored in the ROM 502, but the program may also be stored in a non-volatile memory device. Furthermore, the image analysis device 1 may have an input / output interface (not shown), and the program may be read from another device via the input / output interface and a medium available to the image analysis device 1 when needed. Here, "medium" refers to, for example, a storage medium detachable from the input / output interface, or a communication medium, i.e., a wired, wireless, or optical network, or a carrier wave or digital signal propagating through such a network. Also, some or all of the functions realized by the program may be realized by hardware circuits or FPGAs.

[0082] The embodiments and modifications described above may be combined in any way. Although various embodiments and modifications have been described above, the present invention is not limited to these. Other embodiments that can be conceivable within the scope of the technical idea of ​​the present invention are also included within the scope of the present invention. [Explanation of symbols]

[0083] 1: Image analysis device 41: Object Recognition Models 42: Image Language Model 52: Image Classification Department 53:Performance analysis department 54: Classification and Evaluation Department 55: Image sorting section 56: Deviation Language Query Identification Unit 80: Object 81: Image to be recognized 82: Image Vector 83: Language queries 84: Language Vectors 88: Linguistic characteristic index 89: Nonverbal characteristic indicators 124: Classification and Evaluation Processing Unit 151: Angle between vectors 153: Reference vector 154: Non-verbal feature vectors 160: Evaluation Space 161: Map of learning sheets 162: Performance Map 421: Language Encoder 422: Image encoder 521: Query generation unit 521A: Template 522: Similarity calculation section 523: Nonverbal feature index extraction unit

Claims

1. A computer-based image evaluation method, Using a multimodal platform model of language and images, a vectorization step is performed to vectorize language queries and the image to be recognized into language vectors and image vectors, A first evaluation index calculation step, which calculates a first evaluation index based on the similarity between the language vector and the image vector, An orthogonal vector extraction step of extracting orthogonal vectors that have an origin on the language vector and point toward the image vector, A second evaluation index calculation step, which calculates a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vector and is orthogonal to the language vector and the orthogonal vector, An evaluation step of evaluating the image to be recognized using the first evaluation index and the second evaluation index, Image evaluation methods, including those mentioned above.

2. The image evaluation method according to claim 1, The aforementioned image to be recognized is a trained image used to train the object recognition model. In the aforementioned evaluation step, The image to be recognized is mapped onto an evaluation space with the first evaluation index and the second evaluation index as axes. An image evaluation method that evaluates the data distribution in each space by accumulating the mapping results of multiple images to be recognized as a training data distribution on the evaluation space.

3. The image evaluation method according to claim 2, The aforementioned evaluation step further: Based on the data distribution of the trained images, the proportion of data contained in each region within the defined evaluation space is calculated. The calculated proportion is used to evaluate the data sufficiency within the evaluation space. An image evaluation method that labels each of the aforementioned regions according to evaluation criteria.

4. The image evaluation method according to claim 1, The aforementioned image to be recognized is a verification image used to verify the performance of the object recognition model. In the aforementioned evaluation step, The verification image is input to the object recognition model, and the output of the object recognition model is compared with a pre-created correct answer to calculate the performance of the object recognition model as an evaluation value. The aforementioned evaluation values ​​are mapped onto an evaluation space with the first evaluation index and the second evaluation index as axes. An image evaluation method that explains the performance characteristics of each space by integrating the mapping results of multiple verification images as a performance distribution on the evaluation space.

5. The image evaluation method according to claim 4, The aforementioned evaluation step further: Based on the performance distribution of the verification images, the performance included in each region within the evaluation space is extracted. The performance achievement status of each domain is evaluated using the aforementioned evaluation values. An image evaluation method that labels each of the aforementioned regions according to evaluation criteria.

6. An image evaluation method according to claim 3 or claim 5, An image evaluation method further comprising a training data extraction step of extracting images for retraining the object recognition model based on the labeling results of each of the aforementioned regions.

7. The image evaluation method according to claim 1, The multimodal infrastructure model includes an image encoder and a language encoder. An image evaluation method in which, in the vectorization step, the entire processed image, or an extracted image cut out using a bounding box previously created from the processed image, is input to the image encoder as the image to be recognized.

8. The image evaluation method according to claim 2, The method further includes a language query generation step that generates multiple language queries by combining input objects with multiple templates, The first evaluation index calculation step, the orthogonal vector extraction step, the second evaluation index calculation step, and the evaluation step are performed for each language query. An image evaluation method further comprising a deviation language query identification step that identifies the language query with the largest deviation in the data distribution in the evaluation step.

9. The image evaluation method according to claim 4, The method further includes a language query generation step that generates multiple language queries by combining input objects with multiple templates, The first evaluation index calculation step, the orthogonal vector extraction step, the second evaluation index calculation step, and the evaluation step are performed for each language query. An image evaluation method further comprising a deviation language query identification step that identifies the language query with the largest deviation in the performance characteristics in the evaluation step.

10. A vectorization unit that uses a multimodal platform model of language and images to vectorize language queries and images to be recognized into language vectors and image vectors, A first evaluation index calculation unit calculates a first evaluation index based on the similarity between the language vector and the image vector, An orthogonal vector extraction unit that extracts orthogonal vectors having a starting point on the language vector and pointing toward the image vector, A second evaluation index calculation unit calculates a second evaluation index based on the similarity between a predetermined reference vector on a plane that includes the orthogonal vector and is orthogonal to the language vector and the orthogonal vector, An image evaluation system comprising an evaluation unit that evaluates the recognition target image using the first evaluation index and the second evaluation index.

11. The image evaluation system according to claim 10, An image evaluation system in which the vectorization unit, the first evaluation index calculation unit, the orthogonal vector extraction unit, the second evaluation index calculation unit, and the evaluation unit are all mounted on a vehicle or in a cloud environment.

12. The image evaluation system according to claim 10, The aforementioned image to be recognized is a verification image used to verify the performance of the object recognition model. The evaluation unit inputs the verification images into the object recognition model, compares the output of the object recognition model with a pre-created correct answer, calculates the performance of the object recognition model as an evaluation value, maps the evaluation value onto an evaluation space with the first evaluation index and the second evaluation index as axes, and integrates the mapping results of multiple verification images as a performance distribution on the evaluation space to show the performance characteristics of each space. The evaluation unit further extracts the performance included in each region within the evaluation space based on the performance distribution of the verification image, evaluates the performance achievement status of each region using the evaluation value, labels each region according to the evaluation criteria, and transmits the evaluation criteria to the vehicle. The vectorization unit, the first evaluation index calculation unit, the orthogonal vector extraction unit, the second evaluation index calculation unit, and the evaluation unit are installed in a cloud environment. The vehicle is an image evaluation system comprising an image collection unit that transmits images that meet the evaluation criteria received from the cloud environment to the cloud environment.