Information processing device, method, and program
The information processing apparatus addresses the challenge of providing detailed insights into image recognition results by generating an importance map and using a natural language model to explain the recognition basis, thereby enhancing the understanding of image recognition processes.
Patent Information
- Application Number
- JP2023200722
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2043-11-28
AI Technical Summary
Existing image recognition technologies, such as those using convolutional neural networks (CNNs), struggle to provide detailed insights into the basis of their recognition results, making it difficult to analyze incorrect results.
An information processing apparatus and method that generates a feature map from an input image, creates an importance map to highlight critical pixels, and uses a natural language model to explain the image recognition result by generating a prompt based on the image recognition result, coordinate values, and pixel values of important pixels.
Facilitates a deeper understanding of the image recognition process by providing a clear explanation of the recognition basis, enabling better analysis of correct and incorrect results.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, method, and program.
Background Art
[0002] In recent years, image recognition technologies using learning models such as convolutional neural networks (CNNs) have been widely used. Also, in image recognition, a technique for creating an importance map that indicates the positions of pixels serving as the basis for image recognition determination by a learning model within an input image, which is an image to be processed, according to the importance for the determination has also been used. For example, Patent Document 1 discloses a technique for determining defective portions by an image recognition technique. Further, Patent Document 1 discloses a technique for creating an importance map, which is a heat map, that indicates the positions of pixels serving as the basis for the determination of defective portions by the density of hatching according to the degree of contribution to the determination. Techniques such as class activation mapping (CAM) for generating an importance map using a heat map or the like are also called explainable AI (XAI) and have been attracting attention as techniques that enable understanding and trusting the prediction results and inference results by a machine learning model.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] According to the technology described in Patent Document 1, it is possible to visually grasp the basis for the image recognition result determined by image recognition. However, although the basis for the image recognition result can be visually captured by the importance map, in the technology described in Patent Document 1, it may be difficult to know the basis in detail. For example, when the image recognition result is incorrect, the importance map alone may not be sufficient to analyze the cause of the incorrect result.
[0005] Therefore, an object of the present invention is to provide a technology capable of facilitating the understanding of the basis for an image recognition result.
Means for Solving the Problems
[0006] One aspect of the present disclosure is an information processing apparatus that provides the basis for an image recognition result, which applies a learned image recognition model to a processing target image to generate a feature map including at least one attention area, and outputs the feature map and the image recognition result of the processing target image; a map generation unit that generates an importance map from the feature map based on the importance of each of a plurality of pixels constituting the processing target image; a template acquisition unit that acquires a prompt template including an instruction for a natural language model; a prompt generation unit that generates a prompt by fitting the output image recognition result, the coordinate values and pixel values of the pixels determined to have a relatively high importance in the importance map, and the number of pixels constituting the processing target image to the prompt template; and a prompt output unit that outputs the prompt to the natural language model.
[0007] Another aspect of the present disclosure is a method for providing the basis for recognition of an image recognition result by an information processing apparatus, including: an image recognition result output step of applying a learned image recognition model to a processing target image to generate a feature map including at least one region of interest, and outputting the feature map and the image recognition result of the processing target image; a map generation step of generating an importance map from the feature map based on the importance of each of a plurality of pixels constituting the processing target image; a template acquisition step of acquiring a prompt template including an instruction for a natural language model; a prompt generation step of generating a prompt by fitting the output image recognition result, the coordinate values and pixel values of pixels determined to have a relatively high importance in the importance map, and the number of pixels constituting the processing target image to the prompt template; and a prompt output step of outputting the prompt to the natural language model.
[0008] Yet another aspect of the present disclosure is a program for causing an information processing apparatus to execute: an image recognition result output step of applying a learned image recognition model to a processing target image to generate a feature map including at least one region of interest, and outputting the feature map and the image recognition result of the processing target image; a map generation step of generating an importance map from the feature map based on the importance of each of a plurality of pixels constituting the processing target image; a template acquisition step of acquiring a prompt template including an instruction for a natural language model; a prompt generation step of generating a prompt by fitting the output image recognition result, the coordinate values and pixel values of pixels determined to have a relatively high importance in the importance map, and the number of pixels constituting the processing target image to the prompt template; and a prompt output step of outputting the prompt to the natural language model.
Advantages of the Invention
[0009] According to the present invention, it is possible to provide a technique capable of facilitating the understanding of the basis for recognition of an image recognition result.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11A
Figure 11B
Figure 12
Figure 13
Figure 14
Figure 15
DETAILED DESCRIPTION OF THE INVENTION
[0011] Embodiments of the present invention will be described with reference to the accompanying drawings. In each figure, elements with the same reference numerals have the same or similar configurations.
[0012] <System Configuration> FIG. 1 shows a system configuration example of an image recognition result recognition basis providing system 1 including an information processing apparatus according to an embodiment of the present disclosure. The recognition basis providing system 1 according to the present embodiment includes an information processing apparatus 100 and a language model server 200. In the recognition basis providing system 1, the information processing apparatus 100 and the language model server 200 are connected to each other via, for example, the Internet or the like.
[0013] As will be described later, the information processing apparatus 100 is configured to execute an image recognition process on an input image, which is an image to be processed, and generate a prompt using a prompt template based on the image recognition result. Further, the information processing apparatus 100 is configured to output the generated prompt to the language model server 200. The information processing apparatus 100 may be a terminal device such as a personal computer (PC), a notebook PC, a tablet terminal, or a mobile phone such as a smartphone.
[0014] The language model server 200 is configured to generate a recognition basis for the image recognition result based on the prompt output from the information processing apparatus 100. The language model server 200 may include a generative AI having a natural language model that performs natural language processing such as a large language model (LLM), and may be a server capable of providing a question-and-answer service that generates a sentence as an answer to an input sentence. Hereinafter, the sentence input to the language model server 200 is called a prompt or an instruction sentence.
[0015] The recognition basis providing system 1 according to this embodiment may further include a database 800. As will be described later with reference to FIG. 2, the database 800 may have, for example, a template database 810 that stores a plurality of prompt templates. Further, the database 800 may have, for example, an image database 820 that stores image data used as teacher data in the learning phase of the image recognition model included in the image recognition result output unit 110 described later.
[0016] <Functional block configuration> With reference to FIG. 2, the information processing apparatus 100 according to this embodiment will be described. FIG. 2 is an example of a functional block diagram of the information processing apparatus 100 according to this embodiment. The information processing apparatus 100 includes an image recognition result output unit 110, a map generation unit 120, a template acquisition unit 130, a prompt generation unit 140, a prompt output unit 150, and a storage unit 160.
[0017] The image recognition result output unit 110 is configured to apply a learned image recognition model to a processing target image that is an input image to generate a feature map including at least one attention area, and output the feature map and the image recognition result of the processing target image. As the learned image recognition model, for example, a learned model machine-learned by a CNN (Convolutional Neural Network) can be used. As a CNN-based learned model, for example, any object detection algorithm such as YOLO (You Only Look Once) may be used.
[0018] The learned model machine-learned by the CNN used as the image recognition model has, for example, an input layer that receives an input image, an output layer that outputs a result of predicting the class of the attention area detected from the input image, and an intermediate layer that is a layer other than the input layer and the output layer. The intermediate layer is, for example, an intermediate layer that has been learned by the error backpropagation method (Back Propagation).
[0019] The input layer has a plurality of neurons that receive the input of the pixel values of each pixel included in the input image, and passes the input pixel values to the intermediate layer. The intermediate layer has a plurality of neurons that extract the image feature parameters of the input image, and passes the extracted image feature parameters to the output layer. For example, the intermediate layer may have a configuration in which a convolution layer that convolves the pixel values of each pixel input from the input layer and a pooling layer that maps the pixel values convolved by the convolution layer are alternately connected, and finally extracts the image feature parameters while compressing the pixel information of the input image. Thereafter, the intermediate layer predicts the class of the region of interest detected in the input image by a fully connected layer in which the parameters are learned by the error backpropagation method. The prediction result is output to an output layer having a plurality of neurons. The prediction result of image recognition is expressed probabilistically.
[0020] In a CNN-based image recognition model, for example, for an input image, by repeatedly performing scanning of a filter as a convolution process and scanning of a window as a pooling scan, a feature map with a gradually smaller size is calculated. For example, by scanning a filter for convolution processing or a window for pooling processing with a predetermined number of stride values, which is the amount by which the filter or window is slid, a feature map with a smaller size is calculated. The stride value is, for example, 2. When the stride value is 2, the feature map newly calculated each time the filter or window is scanned has a size of 1 / 2. Note that weight coefficients may be assigned to each of the plurality of filters. When the input image has a plurality of channels (for example, Red, Green, and Blue), filters are applied to each channel.
[0021] Among the intermediate layers of the CNN of the image recognition model, the final intermediate layer is connected to the output layer via a fully-connected connection. The feature map of the final intermediate layer can be obtained, for example, by scanning an average pooling window of the same size as the target feature map over the feature map of the previous intermediate layer (e.g., of size 7x7) to calculate the average value of all elements in each channel, i.e., by global average pooling. By performing global average pooling to calculate the overall average within each channel in this way, the feature map of the final intermediate layer can be treated as a feature vector of a dimension corresponding to the number of channels. For the feature map thus obtained, softmax processing or the like is performed as necessary using the weight coefficients assigned to each connection of the fully-connected connection to calculate the output layer. The output layer includes, for example, N elements (units) when performing N-class image recognition (N is a natural number), and based on the N elements, the inference value for each class is calculated as a probabilistic representation, for example, by normalization. Thereby, for example, the class corresponding to the maximum inference value is determined as the prediction result.
[0022] The CNN-based image recognition model may be prepared, for example, by learning an arbitrary image and label data as training data. The weight coefficients assigned to each connection of the fully-connected connection and the weight coefficients of each filter in the intermediate layer may be calculated in the learning phase, and the weight coefficients may be appropriately updated in the inference phase described above.
[0023] The map generation unit 120 is configured to generate an importance map from the feature map based on the importance of each of the plurality of pixels constituting the image to be processed. The map generation unit 120 generates, for example, a map that can visualize the importance of each pixel of the feature map based on each feature map of each intermediate layer calculated in the image recognition result output unit 110.
[0024] The map generation unit 120 generates an importance map, for example, using Class Activation Mapping (CAM). For example, for each feature map in the intermediate layer, an importance map may be generated by weighting and adding with the weight coefficients of the fully connected connections used in the calculation of the output layer.
[0025] Alternatively, the map generation unit 120 may generate an importance map using Gradient-weighted Class Activation Mapping (Grad-CAM). When Gradient-weighted Class Activation Mapping is used, for example, in each feature map, the gradient of the output value of the feature map with respect to the prediction result (value of the output layer) of the class is calculated, weighted with the weight coefficients of each feature map, and added to generate an importance map. In the calculation of the output value of each coordinate of the map, for example, an activation function (e.g., ReLU function) may be applied, and thereby features that do not contribute to the class determination may be deleted.
[0026] The importance map generated by Class Activation Mapping or Gradient-weighted Class Activation Mapping may be represented, for example, as a heatmap whose color changes according to the magnitude of the output value at each coordinate corresponding to the importance. As described above, when using a CNN-based pre-trained image recognition model, the size of the feature map in the intermediate layer becomes smaller each time a filter is applied, but the map generation unit 120 may generate an importance map scaled to the size of the original image. For example, when the map generation unit 120 uses Gradient-weighted Class Activation Mapping, the gradient may be backpropagated.
[0027] In the map generation unit 120, for example, Eigen-CAM may also be used to generate a heatmap by performing singular value decomposition on the feature map and applying it to the feature map of the intermediate layer for which visualization is desired among the obtained singular value vectors.
[0028] The template acquisition unit 130 is configured to acquire a prompt template including an instruction for the natural language model. The template acquisition unit 130 acquires, for example, the prompt template stored in the template database 810.
[0029] The recognition basis providing system 1 including the information processing apparatus 100 in the present embodiment provides the recognition result of the image recognition result by the image recognition result output unit 110 of the information processing apparatus 100 by expressing it in natural language. In the recognition basis providing system 1 according to the present embodiment, the prompt, which is an instruction sentence output from the information processing apparatus 100 to the natural language model of the language model server 200, is based on the image recognition result, the information of the pixels with high importance in the importance map, and the size of the processing target image. The prompt includes an instruction sentence for requesting the language model server 200 to explain in natural language the basis of the prediction result output by the image recognition result output unit 110. The prompt template is configured to include, for example, text information that can be output to the language model server by fitting information including the image recognition result, the information of the pixels with high importance in the importance map, and the size of the processing target image.
[0030] The template database 810 may store, for example, a plurality of prompt templates. The template database 810 may store a plurality of prompt templates prepared according to the image recognition result or a plurality of prompt templates prepared according to the use of the image recognition result, and the template acquisition unit 130 may select and acquire the prompt template used according to the image recognition result or the use of the image recognition result.
[0031] The prompt generation unit 140 is configured to generate a prompt by fitting the output image recognition result, the coordinate values and pixel values of the pixels determined to have a relatively high importance in the importance map, and the number of pixels constituting the image to be processed to the prompt template acquired by the prompt template acquisition unit 130. The prompt template, for example, fits information about the class of the attention area detected in the image to be processed as the image recognition result output from the image recognition result output unit 110 to the prompt template. The image recognition result is recognized in the attention area, such as "dog", "Dog", etc.
[0032] In this embodiment, the image to be processed, which is the input image, is, for example, a two-dimensional image. The coordinate values of the pixels determined to have a relatively high importance in the importance map are, for example, (x, y) coordinates. The pixel value of the pixel may be possessed by each pixel according to the color plane constituting the image to be processed. When the image to be processed includes, for example, three color planes of RGB (red, green, blue), each pixel has a pixel value (r, g, b). Also, as the number of pixels constituting the image to be processed, the number of pixels in the x direction and y direction, which are the image sizes of the image to be processed, may be used.
[0033] The prompt output unit 150 outputs the prompt generated by the prompt generation unit 140 to the language model server 200.
[0034] The storage unit 160 stores, for example, a learned model 162 that has been learned to output an image recognition result by inputting the image to be processed.
[0035] Note that in FIG. 2, it is exemplified that both the template database 810 and the image database 820 are provided in the same database 800, but the present invention is not limited to this. The template database 810 and the image database 820 may be physically configured as separate databases. Further, at least one and / or at least a part of the data stored in the template database 810 and the data stored in the image database 820 may be stored in the storage unit of the information processing apparatus 100.
[0036] Referring to FIG. 3, the language model server 200 will be described. FIG. 3 is a diagram showing a functional block configuration of the language model server 200.
[0037] As shown in FIG. 3, the language model server 200 of the recognition ground providing system 1 for image recognition results according to the present embodiment includes a natural language model 210.
[0038] The natural language model 210 is, for example, a generative AI having a large language model (Large Language Model) trained with a huge amount of text data by deep learning. As the large language model, for example, GPT (Generative Pre-Trained Transformer) of OpenAI may be used. For example, GPT-3 (the third generation of GPT) called davinci may be used. The natural language model 210 includes a prompt acquisition unit 212, an answer sentence generation unit 214, and an answer sentence output unit 216.
[0039] The prompt acquisition unit 212 is configured to acquire the prompt output by the prompt output unit 150 of the information processing apparatus 100. The answer sentence generation unit 214 is configured to estimate an answer to the prompt using the large language model based on the prompt acquired by the prompt acquisition unit 212 and generate an answer sentence. The answer sentence output unit 216 is configured to output the answer sentence to the prompt generated by the answer sentence generation unit 214.
[0040] <Hardware Configuration> FIG. 4 is a diagram showing a hardware configuration example of the information processing apparatus 100.
[0041] The information processing apparatus 100 includes a processor 11 such as a CPU (Central Processing Unit) and a GPU (Graphical Processing Unit), a memory, a storage device 12 such as an HDD (Hard Disk Drive) and / or an SSD (Solid State Drive), a communication IF (Interface) 13 that performs wired or wireless communication, an input device 14 that receives input operations, and an output device 15 that outputs information. The input device 14 is, for example, a keyboard, a touch panel, a mouse, and / or a microphone, etc. The output device 15 is, for example, a display, a touch panel, and / or a speaker, etc.
[0042] In the present embodiment, the language model server 200 may also have the same hardware configuration as the information processing apparatus 100 described with reference to FIG. 4. Further, the language model server 200 may be configured from one or more physical servers, etc., as necessary, or may be configured using a virtual server operating on a hypervisor.
[0043] The recognition basis providing program or a part thereof that provides the recognition basis of the image recognition result by the information processing apparatus 100 according to the present embodiment may be stored and provided in a computer-readable storage medium such as the storage device 12. Alternatively, it may be provided from outside the information processing apparatus 100 via a communication network to which the information processing apparatus 100 is connected. In the information processing apparatus 100 and / or the language model server 200, for example, the processor 11 executes the recognition basis providing program according to the present embodiment, thereby realizing various operations described later with reference to FIG. 5 and the like.
[0044] For example, the storage unit 160 of the information processing apparatus 100 described above can be realized by using the storage device 12 included in the information processing apparatus 100. Also, the image recognition result output unit 110, the map generation unit 120, the template acquisition unit 130, the prompt generation unit 140, the prompt output unit 150, and the storage unit 160 can be realized by the processor 11 of the information processing apparatus 100 executing a program stored in the storage device 12. Further, the program can be stored in a storage medium. The storage medium storing the program may be a computer-readable non-transitory storage medium (Non-transitory computer readable medium). The non-transitory storage medium is not particularly limited, and may be, for example, a storage medium such as a USB memory or a CD-ROM.
[0045] Note that these physical configurations are examples and do not necessarily have to be independent of each other. For example, the information processing apparatus 100 and / or the language model server 200 according to the present embodiment may include an LSI (Large-Scale Integration) in which the processor 11 and the storage device 12 are integrated. Also, as described above, the information processing apparatus 100 and / or the language model server 200 may include a GPU as the processor 11, and at this time, various operations described later with reference to FIG. 5 and the like may be realized by the GPU executing the recognition basis providing program.
[0046] Also, the information processing apparatus 100 and the language model server 200 are not limited to the above-described configurations. For example, some functions of the information processing apparatus 100 may be configured to be executed by other information processing apparatuses or servers. Also, both the information processing apparatus 100 and the language model server 200 may be configured using a cloud server.
[0047] Also, regarding the database 800, for example, some or all of the information and data stored in the database 800 may be stored in the storage unit provided in the information processing apparatus 100.
[0048] <Processing Procedure> Referring to FIG. 5, the outline of the process for providing the basis for the image recognition result according to this embodiment will be described. FIG. 5 is a diagram showing the outline of the process for providing the basis for recognition executed by the recognition basis providing system 1 according to this embodiment.
[0049] First, an input image, which is the image to be processed, is acquired (S502). The image to be processed may be input to the information processing apparatus 100 by the user, for example, or may be acquired when the information processing apparatus 100 receives it from another information processing apparatus such as a mobile terminal or a PC used by the user. In the example shown in FIG. 5, a photo of a dog is input as the image to be processed.
[0050] Next, the image to be processed is input to the learned image recognition model of the image recognition result output unit 110 (S504). In this embodiment, as the image recognition model, for example, a CNN-based image recognition model as described above is used.
[0051] The recognition result of the image to be processed is output by the image recognition model into which the image to be processed is input (S506). The image recognition model, for example, repeatedly executes convolution processing and the like on the image to be processed, calculates a feature vector, and thereby identifies at least one area to be noted. In the input image illustrated in FIG. 5, for example, the area where the dog appears is identified as the area to be noted, and a recognition result such as "dog", which is the class of the area to be noted, is output based on the result of weighted addition processing of the output values of the feature vector.
[0052] FIG. 6A shows another example of an input image. In the input image 300A shown in FIG. 6A, a person sitting at a table, a cup, a ball, a book, etc. are captured. FIG. 6B shows an example of the image recognition result of the input image 300A illustrated in FIG. 6A. As shown in FIG. 6B, in the image 300B in which the image recognition results are superimposed and shown, a plurality of regions of interest 310_1, 310_2, 310_3, 310_4, 310_5, 310_6, 310_7, 310_8, and 310_9 are detected. In the image 300B shown in FIG. 6B, the regions 310_1, 310_2, and 310_3 correspond to the class of person, as indicated by "person", the regions 310_4, 310_5, and 310_6 correspond to the class of cup, as indicated by "cup", the regions 310_7 and 310_8 correspond to the class of ball, as indicated by "bowl", and the region 310_9 corresponds to the class of book, as indicated by "book", and are respectively classified.
[0053] Returning to FIG. 5, next, an importance map is generated from the feature map based on the importance of each of the plurality of pixels constituting the image to be processed (S508). As the image recognition model, for example, a CNN-based image recognition model is used. When gradient class activation mapping is used by the map generation unit 120, as described above, among the output values of the output layer, based on the gradient of the output value of the feature map of the intermediate layer with respect to the output value of the class regarded as the recognition result, an importance map is generated.
[0054] For example, with the importance map, for the input image illustrated in FIG. 5, it can be visually grasped that for the pixels that are the basis for determining that the region of interest in the image recognition process is a dog, pixels with a relatively high contribution to the determination are distributed in the region near the dog's face and then in the region of the body including the tail, etc. Note that as the importance map, it may be output only as a heat map, as shown in the left part of S508 in FIG. 5, or as shown in the right part, by superimposing it on the input image, it may be output by a method that makes it easier to visually grasp which part of the input image contributed to the recognition of the class of dog.
[0055] Next, the template acquisition unit 130 acquires a prompt template from, for example, the template database 810 (S510). The prompt template has blanks where information such as, for example, the prediction result of image recognition, the image size, and the coordinates and pixel values of important pixels is to be input. In the present embodiment, the prompt template may be created in advance by, for example, a prompt engineer, or may be created by executing a program for creating the prompt template.
[0056] Subsequently, the prompt generation unit 140 generates a prompt (S512). The prompt generation unit 140 generates a prompt by fitting the image recognition result, the image size, and the information of important pixels, which are the information necessary for providing the basis for recognition of the image recognition result, to the prompt template acquired by the template acquisition unit 130. The image size may be, for example, the number of pixels constituting the image to be processed. For example, the number of pixels in the horizontal and vertical directions of the image to be processed may be fitted.
[0057] With reference to FIG. 7, an example of the generation of a prompt according to the present embodiment will be described. As shown in FIG. 7, as data to be fitted to the prompt template 400A, important pixel data 412, image size data 414, and prediction result data 416 are acquired. The important pixel data 412 includes the x-coordinates and y-coordinates of important pixels (pixels p00, p01,..., pkl) and, for example, RGB pixel values. The image size data 414 is, for example, the resolution of the input image and includes the number of pixels X in the x-direction and the number of pixels Y in the y-direction. The prediction result data 416 includes, for example, the predicted class and the position information of the region of interest including the x-coordinates and y-coordinates of the endpoints of the rectangular region, for example.
[0058] The prompt template 400A includes blanks into which, for example, important pixel data 412, image size data 414, and prediction result data 416 are fitted. The prompt generation unit 140 generates a prompt 400B by fitting, for example, the important pixel data 412, the image size data 414, and the prediction result data 416 into the blanks of the prompt template 400A. As the important pixels, for example, information of a predetermined number (e.g., 30 pixels) of pixels may be acquired in order from the pixels with high importance among the pixels determined to be important. Also, information of the important pixels acquired in order from the pixels with high importance may be fitted into the prompt template, or among the important pixels, they may be fitted into the prompt template in order from the pixels with smaller coordinate values.
[0059] In the recognition basis providing system 1 according to the present embodiment, the important pixel data 412, the image size data 414, and the prediction result data 416 are fitted into the prompt template, but the data fitted into the prompt template is not limited to these, and other information may be fitted. For example, when the predicted classes are finite, information regarding the fields to which each class belongs may be included in the prompt that is the instruction text. As the category of the predicted class, for example, when predicting an animal captured in the image to be processed, information that the category is an animal may be included. By adding this information, for example, it is possible to provide a recognition basis specialized for animal image recognition by the natural language model 210.
[0060] In addition, information on the image recognition model used and information on the model for generating the importance map (i.e., explainable AI (XAI)) may be included in the prompt. For example, as the image recognition model, information on models such as YOLOv5 and Swin Transformer, and as the importance map generation tool, information on models such as Eigen-CAM and Score-CAM may be included in the prompt. Alternatively, instead of the model name, the technical content underlying the model used, such as CNN, gradient class activation mapping, and error backpropagation method, may be included in the prompt. This information can also provide a more accurate basis for recognition.
[0061] Also, for example, information on pixels with relatively low importance may be included in the prompt. Information on pixels with relatively low importance can also be used in explaining the basis for the image recognition result. By analyzing the processing recognized as unimportant by the image recognition model based on the information of the pixels considered unimportant, the image recognition model can be verified.
[0062] Note that by increasing the information included in the prompt, the information that is a prerequisite for generating an answer by the natural language model 210 can be increased, so that a more accurate answer sentence can be obtained. However, if there is a character limit for the text that can be input as the prompt, the information that fits into the prompt template may be selected.
[0063] Returning to FIG. 5, next, the prompt generated by the prompt generation unit 140 is output to the natural language model 210 of the language model server 200 by the prompt output unit 150 (S514). The prompt acquisition unit 212 of the natural language model 210 acquires the prompt output by the prompt output unit 150 of the information processing apparatus 100.
[0064] Based on the acquired prompt, the natural language model 210 outputs a sentence in natural language expressing the basis for the image recognition result (S516).
[0065] Referring to FIGS. 8 to 10, an example of the basis for recognition generated by the natural language model 210 based on the prompt output from the prompt generation unit 140 will be described. FIG. 8 shows the basis for recognition 420 output by the natural language model 210 based on the prompt 400B when the class of the prediction result illustrated in FIG. 7 is "person". Further, FIG. 9 shows the basis for recognition 420 of the image recognition result when the class of the prediction result is "cup", and FIG. 10 shows the basis for recognition 420 of the image recognition result when the class of the prediction result is "book". As shown in FIGS. 8 to 10, in the basis for recognition of the image recognition result generated by the natural language model 210, for example, the influence of the coordinates of important pixels on the prediction result of the pixel value, the background area of the predicted object (such as a person, a cup, a book, etc.), and the relationship with areas other than the predicted object, etc. Based on this, the image recognition process that serves as the basis for the image recognition result from each pixel is analyzed and explained.
[0066] In the example described above with reference to FIG. 7, the prompt template and the prompt when the image recognition result is determined to be correct were exemplified. However, even when the image recognition result is determined to be incorrect, the process of providing the basis for recognition of the image recognition result according to the present embodiment may be executed. At this time, for example, a different prompt template may be prepared for the case where the image recognition result is determined to be incorrect.
[0067] That is, in the recognition basis providing system 1 according to the present embodiment, the prompt template may include a first prompt template used when the image recognition result is correct and a second prompt template used when the image recognition result is incorrect. Further, the second prompt template may include a prompt sentence that requests the natural language model 210 to output information regarding the method for obtaining a correct image recognition result. At this time, the prompt generation unit 140 may be configured to select either the first prompt template or the second prompt template in correspondence with the correctness information of the image recognition result.
[0068] FIG. 11A shows an example of the input image in this case. In the input image 320A shown in FIG. 11A, a person wearing a suit is captured. FIG. 11B shows an example of the image recognition result of the input image 320A illustrated in FIG. 11A. As shown in FIG. 11B, in the image 320B where the image recognition results are superimposed, a region of interest 330_1 is detected. In the image 320B shown in FIG. 11B, the region 330_1 is classified into a class corresponding to a suit, as indicated by "suit".
[0069] In the examples shown in FIGS. 11A and 11B, for example, when a prediction of "person" is expected, the prediction of "suit" is determined to be incorrect. By the process of providing the basis for the image recognition result according to this embodiment, it is possible to explain the reason for the prediction and propose improvement measures when an incorrect prediction is made, such as when a region that should be predicted as "person" is predicted as "suit".
[0070] FIG. 12 shows the prompt generation process in this case. As shown in FIG. 12, a prompt template 450A for the case where an incorrect prediction is made is acquired. Also, similar to the prompt generation process illustrated in FIG. 7, in addition to the important pixel data 412, the image size data 414, and the prediction result data 416, in the example shown in FIG. 12, for example, model information 418 may also be used. The model information 418 may include, for example, information on the image recognition model used in the image recognition process or the explainable AI model used in the importance map generation process. The model information 418 may be used, for example, for analyzing the suitability of the model in the executed image recognition process and map generation process. The prompt generation unit 140 generates a prompt 450B by fitting the important pixel data 412, the image size data 414, the prediction result data 416, and the model information 418 to the prompt template 450A. The prompt 450B includes, for example, questions regarding improvement measures in addition to the analysis of the reason for the prediction result being "suit".
[0071] FIG. 13 shows an example of the output of the recognition basis 460 generated based on the prompt 450B by the natural language model 210 when the image recognition result is incorrect. As shown in FIG. 13, the recognition basis 460 includes, for example, an explanation of the pixels serving as the basis for the erroneously predicted class "suit", an explanation regarding the appropriateness of the model selection, and also a proposal for improvement measures.
[0072] In the above-described embodiment, the case where a still image is used as the image to be processed has been described as an example. However, in the present embodiment, the image to be processed may be a moving image. That is, in the present embodiment, the image to be processed includes a moving image including a plurality of frame images, and the prompt generation unit 140 determines a frame of interest, which is a frame to be noted, based on the feature maps generated for the plurality of frame images, and further configures the frame identification information of the frame of interest, the frame rate of the moving image, and the number of pixels included in the frame to be noted to be fitted to the prompt template.
[0073] FIG. 14 shows the prompt generation process when a moving image is used as the image to be processed. As shown in FIG. 14, as data to be fitted to the prompt template 470A, important pixel data 472, moving image size data 474, and prediction result data 476 are acquired. The important pixel data 472 further includes the frame number as the frame identification information of the important pixel in addition to the x coordinate, y coordinate, and pixel value of the important pixel of the important pixel data 412 (FIG. 7). The moving image size data 474 further includes the moving image length T (s) and the frame rate F (fps) in addition to the number of pixels X in the x direction and the number of pixels Y in the y direction, which are the same as the image size data 414 (FIG. 7) as the resolution of the moving image. The prediction result data 476 includes, similar to the prediction result data 416 (FIG. 7), the predicted class and the information on the x coordinates and y coordinates of the endpoints of the rectangular region to be noted.
[0074] In the image recognition process of a moving image, for example, all frames included in the moving image may be regarded as target frames to be noted, or some frames may be regarded as target frames to be noted. The important pixel data 472 may include, for example, pixel information about frames in which the target of the predicted class is reflected among all frames.
[0075] As shown in FIG. 14, the prompt template 470A used for the process of providing the basis for the recognition result of the moving image recognition includes, for example, blanks into which the important pixel data 472, the video size data 474, and the prediction result data 476 are fitted. The prompt generation unit 140 generates a prompt 470B by fitting the important pixel data 472, the image size data 474, and the prediction result data 476 into the blanks of the prompt template 470A.
[0076] In the present embodiment, using the prompt 470B generated in this way, similar to the above-described process of providing the basis for the recognition result of the image recognition, a description text in natural language expression of the basis for the recognition is generated by the natural language model 210, thereby providing the basis for the prediction result (for example, the prediction result that the moving image is a video of "soccer") by the image recognition process for the moving image.
[0077] With reference to FIG. 15, a method for providing the basis for the recognition result of the image recognition according to an embodiment of the present disclosure will be described. FIG. 15 is a flowchart of the method for providing the basis for the recognition result of the image recognition according to an embodiment of the present disclosure.
[0078] First, an image recognition model is applied to the image to be processed, and an image recognition result is output (S1502). That is, a learned image recognition model is applied to the image to be processed to generate a feature map including at least one target region, and the generated feature map and the image recognition result of the image to be processed are output.
[0079] Next, an importance map is generated (S1504). That is, an importance map is generated from the above-described feature map based on the importance of each of the plurality of pixels constituting the image to be processed.
[0080] Subsequently, a prompt template is acquired (S1506). That is, a prompt template including an instruction for the natural language model is acquired.
[0081] Next, a prompt is generated (S1508). That is, a prompt is generated by fitting the output image recognition result, the coordinate values and pixel values of the pixels determined to have a relatively high importance in the importance map, and the number of pixels constituting the image to be processed to the acquired prompt template.
[0082] Subsequently, the prompt is output to the natural language model (S1510).
[0083] <Summary> According to the embodiment described above, in the recognition basis providing system 1 for the image recognition result, a prompt is generated by fitting the image recognition result by the image recognition model and the information of the importance map generated based on the feature map to the prompt template, and the prompt thus generated is output to the natural language model. As a result, the natural language model can generate, as a sentence explained in an easy-to-understand expression, the recognition basis of the prediction result of the image recognition. Therefore, according to the present embodiment, it is possible to provide a technology capable of facilitating the understanding of the recognition basis of the image recognition result.
[0084] Also, in the recognition basis providing system 1 for the image recognition result according to the present embodiment, for example, an interactive generation type AI may be used as the natural language model 210, and based on the provided recognition basis, further detailed analysis of the basis and the like can be made possible by continuing the dialogue. Further, the technology according to the present embodiment can contribute to the achievement of Goal 9, "Build the infrastructure for industry and technological innovation," of the Sustainable Development Goals (SDGs).
[0085] The process of providing the basis for the image recognition result according to this embodiment may be used, for example, in a manufacturing process. For example, this embodiment can be used for the analysis of defects of defective products, etc. in quality control and identification of defective products, and for the explanation of causes. Alternatively, this embodiment can be applied, for example, to the explanation of the basis for determining abnormal locations such as cancer in medical images. Further, this embodiment may be applied to the explanation of the basis for the analysis of the flow of customers in a retail store.
[0086] The embodiments described above are for facilitating the understanding of the present invention and are not for limiting and interpreting the present invention. The flowcharts, sequences, each element included in the embodiments, and their arrangements, materials, conditions, shapes, sizes, etc. described in the embodiments are not limited to those exemplified and can be appropriately changed. Also, the calculation methods described in the above-described embodiments are not limited to those exemplified. Further, it is possible to partially replace or combine the configurations shown in different embodiments.
Explanation of Reference Numerals
[0087] 1... Recognition basis providing system, 110... Image recognition result output unit, 120... Map generation unit, 130... Prompt template acquisition unit, 140... Prompt generation unit, 150... Prompt output unit, 160... Storage unit, 162... Learned model, 200... Language model server, 210... Natural language model, 212... Prompt acquisition unit, 214... Answer sentence generation unit, 216... Answer sentence output unit, 800... Database
Claims
1. An information processing apparatus that provides the basis for recognizing an image recognition result, comprising: an image recognition result output unit that applies a learned image recognition model to a processing target image to generate a feature map including at least one attention area, and outputs the feature map and the image recognition result of the processing target image; a map generation unit that generates an importance map from the feature map based on the importance of each of a plurality of pixels constituting the processing target image; a template acquisition unit that acquires a prompt template including an instruction for a natural language model; a prompt generation unit that generates a prompt by fitting the output image recognition result, the coordinate values and pixel values of pixels determined to have relatively high importance in the importance map, and the number of pixels constituting the processing target image to the prompt template; a prompt output unit that outputs the prompt to the natural language model; An information processing apparatus comprising:
2. The prompt template includes a first prompt template used when the image recognition result is correct and a second prompt template used when the image recognition result is incorrect, The second prompt template includes a prompt sentence that requests the natural language model to output information regarding a method for obtaining a correct image recognition result, The information processing apparatus according to claim 1, wherein the prompt generation unit selects either the first prompt template or the second prompt template according to the correctness information of the image recognition result.
3. The image recognition result output unit includes a CNN, The map generation unit includes Grad-cam, The natural language model includes a generative AI including a large language model, The information processing apparatus according to claim 1.
4. The processing target image includes a moving image including a plurality of frame images, The prompt generation unit determines an attention frame, which is a frame to be noted, based on the feature maps generated for the plurality of frame images, and further fits the frame identification information of the attention frame, the frame rate of the moving image, and the number of pixels included in the frame to be noted to the prompt template. The information processing apparatus according to claim 1.
5. A method for providing the basis for recognizing an image recognition result by an information processing apparatus, comprising: An image recognition result output step of applying a learned image recognition model to a target image to generate a feature map including at least one region of interest, and outputting the feature map and the image recognition result of the target image; A map generation step of generating an importance map from the feature map based on the importance of each of a plurality of pixels constituting the target image; A template acquisition step of acquiring a prompt template including an instruction for a natural language model; A prompt generation step of generating a prompt by fitting the output image recognition result, the coordinate values and pixel values of the pixels determined to have a relatively high importance in the importance map, and the number of pixels constituting the target image to the prompt template; A prompt output step of outputting the prompt to the natural language model; A method for providing an image recognition basis, including the above steps.
6. An image recognition result output step of applying a learned image recognition model to a target image to generate a feature map including at least one region of interest, and outputting the feature map and the image recognition result of the target image; A map generation step of generating an importance map from the feature map based on the importance of each of a plurality of pixels constituting the target image; A template acquisition step of acquiring a prompt template including an instruction for a natural language model; A prompt generation step of generating a prompt by fitting the output image recognition result, the coordinate values and pixel values of the pixels determined to have a relatively high importance in the importance map, and the number of pixels constituting the target image to the prompt template; A prompt output step of outputting the prompt to the natural language model; A program for causing an information processing apparatus to execute the above steps.
Citation Information
Patent Citations
Object recognition system, location information acquisition method, and program
JP2022079775A
Image processing device and image processing method
JP2023132481A
Information processing device, information processing method, and program
JP7370118B1
Mentoring System
JP7416390B1
Program, method, information processing device, and system
JP7488617B1