Program, information processing device, and information processing method

The described program and device evaluate the validity of gaze areas in image recognition by assessing the contribution of partial image regions, addressing the challenge of verifying prediction results in AI image recognition, thereby enhancing accuracy and reducing costs.

JP7798109B2Active Publication Date: 2026-01-14SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023545060
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-02
Filing Date
2022-03-22
Publication Date
2026-01-14
Estimated Expiration
2042-03-22

AI Technical Summary

Technical Problem

Existing technologies face challenges in verifying the validity of prediction results for image recognition using artificial intelligence due to the large number of images used for training, making it difficult to confirm the basis for these predictions efficiently.

Method used

A computing device executes a validity evaluation function to assess the validity of gaze areas in image recognition, determining whether the gaze area is valid or not, using a program that evaluates the contribution of partial image regions to the detection of recognition targets through image recognition.

Benefits of technology

This approach reduces the cost and complexity of verifying the validity of prediction results by identifying and visualizing the contribution of image areas to the recognition process, allowing for more accurate classification and prioritization of data based on the validity of the prediction basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798109000001
    Figure 0007798109000001
  • Figure 0007798109000002
    Figure 0007798109000002
  • Figure 0007798109000003
    Figure 0007798109000003
Patent Text Reader

Abstract

This program causes an operation processing device to execute: a validation evaluation function that evaluates validity for an interest area on the basis of the interest area, which is an image area that becomes the basis for prediction, and a prediction area that is an image area in which an object to be recognized is predicted to be present by recognizing an input image using artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to the technical fields of a program, an information processing device, and an information processing method that perform processing to determine whether the basis for a prediction result of image recognition using artificial intelligence is valid. [Background technology]

[0002] There is a technology that detects and classifies subjects by performing image recognition using artificial intelligence (AI). When evaluating the performance of AI models used in such artificial intelligence, it is important to consider not only the accuracy of the prediction results but also whether the grounds on which the prediction results were derived are valid.

[0003] There is a technology that visualizes the basis for prediction results of image recognition by artificial intelligence (for example, Patent Document 1). [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2021-093004 Summary of the Invention [Problem to be solved by the invention]

[0005] However, the number of images used to train AI models is enormous, and it is difficult to verify the validity of the basis for prediction results for all of those images.

[0006] This technology was developed in light of these problems, and aims to reduce the cost required to confirm the basis for the prediction results of image recognition using artificial intelligence. [Means for solving the problem]

[0007] The program related to the present technology causes a computing device to execute a validity evaluation function that evaluates the validity of a gaze area, which is an image area where a recognition target is predicted to exist through image recognition using artificial intelligence on an input image, based on the prediction area, which is an image area that serves as the basis for the prediction. For example, it may be determined whether or not the gaze area is valid, or it may be determined only that the gaze area is valid, or it may be determined only that the gaze area is invalid.

[0008] The information processing device according to the present technology includes a validity evaluation unit that evaluates the validity of a gaze area based on a prediction area, which is an image area where a recognition target is predicted to exist by image recognition using artificial intelligence on an input image, and a gaze area, which is an image area that serves as the basis for the prediction.

[0009] The information processing method according to the present technology involves a calculation processing device executing a validity evaluation process that evaluates the validity of a gaze area, based on a prediction area, which is an image area predicted to contain a recognition target through image recognition using artificial intelligence on an input image, and a gaze area, which is an image area that serves as the basis for the prediction. The above-mentioned effects can also be achieved by such an information processing device and information processing method. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a functional block diagram of an information processing device according to the present technology. [Figure 2] FIG. 10 is a diagram illustrating an example of an input image. [Figure 3] FIG. 10 is a diagram illustrating an example of a prediction region and a fixation region. [Figure 4] FIG. 4 is a functional block diagram of a gaze area specifying unit. [Figure 5] FIG. 10 is a diagram illustrating an example of a state in which an input image is divided into partial image regions. [Figure 6] FIG. 10 is a diagram illustrating an example of a mask image. [Figure 7] This is the first example of visualized contribution. [Figure 8] This is a second example of visualized contribution. [Figure 9] This is the third example of visualized contribution. [Figure 10] FIG. 10 is a diagram showing an example in which two fixation regions are identified for one prediction region. [Figure 11] FIG. 2 is a functional block diagram of a classification unit. [Figure 12] FIG. 10 is a diagram illustrating an example of a classification result. [Figure 13] FIG. 10 is a diagram showing a first example of a presentation screen. [Figure 14] FIG. 10 is a diagram showing a second example of a presentation screen. [Figure 15] FIG. 10 is a diagram showing a third example of a presentation screen. [Figure 16] FIG. 10 is a diagram showing a fourth example of a presentation screen. [Figure 17] FIG. 10 is a diagram showing a fifth example of a presentation screen. [Figure 18] FIG. 1 is a block diagram of a computer device. [Figure 19] 10 is a flowchart illustrating an example of processing executed by the information processing device until an evaluation result of the validity of the gaze area is presented to the user. [Figure 20] 10 is a flowchart illustrating an example of a contribution degree visualization process. [Figure 21] 10 is a flowchart illustrating an example of a classification process. [Figure 22] 10 is a flowchart illustrating another example of the classification process. [Figure 23] 10 is a flowchart showing the processes executed by each information processing device until an AI model is created for a user to achieve a goal. [Figure 24] 10 is a flowchart showing the processing executed by each information processing device when a recognition error occurs in the created AI model. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, with reference to the accompanying drawings, embodiments according to the present technology will be described in the following order. <1. Configuration of information processing device> <2. Computer equipment> <3. Processing flow> <4. Application Examples> <5. Summary> <6. This Technology>

[0012] <1. Configuration of information processing device> The functional configuration of an information processing device 1 according to this embodiment will be described with reference to FIG. The information processing device 1 is a device that performs various processes according to instructions from a user (operator) who confirms the validity of the processing results of image recognition using artificial intelligence (AI).

[0013] The information processing device 1 may be, for example, a computer device serving as a user terminal used by a user, or may be a computer device serving as a server device connected to a user terminal.

[0014] The information processing device 1 receives the image recognition processing results as input data and outputs information that the user is desired to check from the input data.

[0015] The information processing device 1 includes a contribution visualization processing unit 2, a gaze area identification processing unit 3, a classification unit 4, and a display control unit 5.

[0016] The contribution visualization processing unit 2 calculates and visualizes the contribution DoC for each predicted area FA where the recognition target RO is predicted to exist in the input image II. The calculated contribution DoC is used to identify the gaze area GA in a later stage.

[0017] Here, the input image II, the prediction area FA, and the gaze area GA will be described. Figure 2 is an example of an input image II in which American football players are captured. As shown in the figure, the input image II includes four players.

[0018] When the recognition target RO is set to "uniform number," predicted areas FA1, FA2, and FA3 are extracted by image recognition processing using an AI model. That is, predicted areas FA1, FA2, and FA3 are areas that the AI ​​has estimated to be highly likely to contain uniform numbers.

[0019] Each prediction area FA is determined to have a high probability of including a uniform number based on a different area. The area that is the basis for the prediction of the prediction area FA is defined as the gaze area GA.

[0020] 3 shows a gaze area GA1 for a prediction area FA1. Note that there may be multiple gaze areas GA for one prediction area FA. Also, the gaze area GA does not have to be rectangular.

[0021] For example, in the example shown in FIG. 3, two gaze areas GA1-1 and GA1-2 exist for the predicted area FA1.

[0022] In other words, the AI ​​model recognizes the uniform number as the target RO by taking into account not only the numerical part of the uniform number but also the part of the player's neck to identify the position on the back and infer that the numbers are a uniform number.

[0023] The contribution visualization processing unit 2 calculates the degree of contribution DoC for each prediction area FA on the input image II. In the following description, the process of calculating the degree of contribution DoC for the prediction area FA1 will be described.

[0024] FIG. 4 shows a detailed functional configuration of the contribution visualization processing unit 2. The contribution visualization processing unit 2 includes an area division processing unit 21 , a mask image generation unit 22 , an image recognition processing unit 23 , a contribution calculation unit 24 , and a visualization processing unit 25 .

[0025] The region division processing unit 21 divides the input image II into a plurality of partial image regions DI. As a division method, for example, superpixels or the like may be used. In this embodiment, the region division processing unit 21 divides the input image II into a grid so that rectangular partial image regions DI are arranged in a matrix.

[0026] An example of an input image II and partial image regions DI is shown in Fig. 5. As shown in the figure, the input image II is divided into a large number of partial image regions DI.

[0027] The mask image generating unit 22 generates a mask image MI by applying a mask pattern for masking a part of the partial image region DI to the input image II. The mask pattern is created by determining whether or not to mask each of the M partial image regions DI included in the input image II.

[0028] An example of a mask image MI obtained by applying one of the created mask patterns to an input image II is shown in Fig. 6. The masked partial image area DI is defined as a partial image area DIM. If the number of partial image regions DI is M, then there are 2^M types of mask images MI.

[0029] If the number of mask images MI used to calculate the degree of contribution DoC is too large, the amount of calculation increases excessively, so the mask image generating unit 22 generates, for example, several hundred to several tens of thousands of types of mask images MI.

[0030] The image recognition processing unit 23 performs image recognition processing using an AI model with the mask image MI generated by the mask image generating unit 22 as input data.

[0031] Specifically, the prediction result likelihood PLF (inference score) for the prediction region FA1 is calculated for each mask image MI.

[0032] The prediction result likelihood PLF is information indicating the likelihood of the estimation result that the recognition target RO (for example, a uniform number) is included in the prediction area FA1 that is the processing target, that is, the correctness of the inference.

[0033] For example, when a partial image region DI that is important in prediction is masked, the prediction result likelihood PLF becomes small. On the other hand, when a partial image region DI that is important in prediction is not masked, the prediction result likelihood PLF becomes large.

[0034] For partial image regions DI that are important in prediction, the difference between the prediction result likelihood PLF when masked and the prediction result likelihood PLF when not masked is large, whereas for partial image regions DI that are not important in prediction, the difference between the prediction result likelihood PLF when masked and the prediction result likelihood PLF when not masked is small. In other words, the difference between the prediction result likelihood PLF when each partial image area DI is masked and when it is not masked is calculated, and the larger the difference, the higher the importance of that partial image area DI in the prediction can be estimated to be.

[0035] The image recognition processing unit 23 calculates the prediction result likelihood PLF for each prepared mask image MI.

[0036] The contribution calculation unit 24 calculates the contribution DoC for each partial image region DI using the prediction result likelihood PLF for each mask image MI.

[0037] The degree of contribution DoC is information indicating the degree of contribution to the detection of the recognition target RO. That is, a partial image region DI with a high degree of contribution DoC is considered to be a region that makes a high contribution to the detection of the recognition target RO.

[0038] For example, if a partial image area DI is masked and the predicted result likelihood PLF is low, and if the predicted result likelihood PLF is high when the partial image area DI is not masked, the contribution DoC of the partial image area DI will be high.

[0039] An example of a method for calculating the degree of contribution (DoC) is given below.

[0040] An input image II is divided into M partial image regions DI, which are designated as partial image regions DI1, DI2, . . . DIM.

[0041] The predicted result likelihood PLF obtained as a result of performing image recognition processing on the mask image MI1 is set as a predicted result likelihood PLF1.

[0042] The contributions DoC of the partial image regions DI1, DI2, . . . DIM are defined as contributions DoC1, DoC2, . . . DoCM.

[0043] At this time, the following equation (1) is obtained for the prediction result likelihood PLF1.

[0044] PLF1=A1×DoC1+A2×DoC2+···+AM×DoCM···Formula (1)

[0045] Here, A1, A2, ... AM are coefficients for each partial image region DI, and are set to "0" if masked, and "1" if not masked.

[0046] For example, when 1000 types of mask images MI1 to MI1000 are used, 1000 equations can be obtained in which the patterns of the coefficients A1 to AM on the left and right sides of equation (1) are different.

[0047] By using a large number of formulas (1), it is possible to obtain optimal solutions for the degrees of contribution DoC1 to DoCM. That is, the more mask images MI are prepared, the higher the accuracy of the calculated degrees of contribution DoC.

[0048] Furthermore, by dividing the partial image area DI into a grid-like area, it is possible to divide what was previously a single area in the superpixel into multiple areas, making it possible to analyze areas with high contribution DoC in more detail.

[0049] The visualization processing unit 25 performs processing to visualize the degree of contribution DoC. There are several possible visualization methods.

[0050] A first example of visualized degree of contribution DoC is shown in Figure 7. In this example, the level of the degree of contribution DoC is shown by the shade of the fill color. That is, the higher the degree of contribution DoC of a partial image area DI, the darker the color it is filled in.

[0051] As shown in the figure, it can be seen that the degree of contribution DoC of the partial image area DI included in the prediction area FA1 and the partial image area DI around the player's neck is high.

[0052] A second example of visualized degree of contribution (DoC) is shown in Figure 8. In this example, only the partial image area (DI) where the degree of contribution (DoC) is equal to or greater than a certain value is filled in. The density of the fill color of the degree of contribution (DoC) is proportional to the magnitude of the degree of contribution (DoC).

[0053] A third example of visualized degree of contribution DoC is shown in Fig. 9. In this example, only partial image areas DI where the degree of contribution DoC is equal to or greater than a certain value are displayed with a frame, and the degree of contribution DoC is displayed as a numerical value (0 to 100) within the frame of the partial image area DI.

[0054] In either method, the partial image region DI with a high degree of contribution DoC is visualized in an easily understandable manner. An image in which the degree of contribution DoC is visualized, that is, an image of the degree of contribution DoC such as a heat map as shown in FIGS. 7 to 9, is presented to the user by the display control unit 5 at a subsequent stage.

[0055] The gaze area identification processing unit 3 performs processing to identify the gaze area GA by analyzing the contribution degree DoC as processing prior to classification processing by the classification unit 4 at the subsequent stage.

[0056] In the analysis process of the degree of contribution DoC, for example, a partial image area DI with a high degree of contribution DoC is identified as a gaze area GA. Note that if a group of partial image areas DI with high degrees of contribution DoC exists, the area is identified as one gaze area GA.

[0057] There are various methods for identifying the gaze area GA from the heat map of the degree of contribution DoC.

[0058] For example, when one partial image area DI is treated as one cell, the contribution degree DoC is smoothed using a maximum value filter consisting of three cells vertically and horizontally, and then the contribution degree DoC of the partial image area DI whose value did not change before and after smoothing is left unchanged, and the contribution degree DoC of the other partial image areas DI is set to 0.

[0059] Furthermore, small peaks are eliminated by changing the degree of contribution DoC that is less than a predetermined value to 0.

[0060] After the above processing, partial image regions DI remain that have a contribution DoC other than 0. Each cluster of the remaining partial image regions DI is treated as one gaze area GA, and the partial image region DI is treated as the representative region RP of the gaze area GA.

[0061] Furthermore, one fixation area GA can be configured to include the surrounding partial image areas DI centered around the representative area RP. For example, among the partial image areas DI surrounding the representative area RP, areas whose pre-processing contribution degree DoC is equal to or greater than a predetermined value are included in one fixation area GA centered around the representative area RP.

[0062] That is, one fixation area GA can include multiple partial image areas DI.

[0063] 10 shows a state in which gaze areas GA1-1 and GA1-2 are identified as gaze areas GA corresponding to a prediction area FA1 in an input image II. A representative area RP is set for each of the gaze areas GA1-1 and GA1-2.

[0064] Another method for identifying the gaze area GA is, for example, to first extract an area having a larger degree of contribution DoC than the adjacent partial image area DI.

[0065] Next, partial image regions DI whose contribution DoC is lower than a threshold value are excluded. Finally, of the remaining partial image areas DI, partial image areas DI that are close in distance are grouped together and treated as one fixation area GA.

[0066] The representative area (or representative point) of the gaze area GA is the center of gravity of the partial image area DI included in the gaze area GA. At this time, the center of gravity may be obtained using the respective contribution degrees DoC as weights.

[0067] The gaze area identification processing unit 3 analyzes the gaze area GA identified using various methods in this way.

[0068] Specifically, the number of gaze areas GA, the position of the gaze area GA relative to the prediction area FA, the difference between the contribution DoC of the prediction area FA and the contribution DoC outside the prediction area FA, and the like are taken into consideration.

[0069] 10, the number of gaze areas GA is set to "2," and the positions relative to the prediction area FA are such that gaze area GA1-1 is set to be inside the prediction area FA, and gaze area GA1-2 is set to be outside the prediction area FA. In addition, to calculate the difference between the contribution degree DoC of the prediction area FA and the contribution degree DoC outside the prediction area FA, the average value of the contribution degree DoC of the prediction area FA and the average value of the contribution degree DoC outside the prediction area FA are calculated.

[0070] These pieces of information are used in the classification process in the classification unit 4 at the subsequent stage.

[0071] The classification unit 4 uses each piece of information obtained by the gaze area identification processing unit 3 to evaluate the validity of the gaze area GA with respect to the predicted area FA and performs classification processing.

[0072] As shown in FIG. 11, the classification unit 4 includes a classification processing unit 41 and a priority determination unit .

[0073] The classification processing unit 41 evaluates the validity of the gaze area GA and classifies each data into a category based on the evaluation result. Specifically, the classification processing unit 41 assigns a "valid" category, a "needs confirmation" category, or a "use for analysis" category to the prediction result of the input image II.

[0074] The "valid" category is a category that is classified when the recognition target RO is detected based on correct evidence, and is a category that classifies data that does not require the user to confirm the validity of the gaze area GA. In other words, cases classified in the "valid" category are the cases with the lowest priority of presentation to the user.

[0075] The "needs confirmation" and "used for analysis" categories are used when it is not possible to determine whether the recognition target RO is valid. In other words, there is a possibility that the recognition target RO is detected based on correct grounds or incorrect grounds, and it is highly necessary for the user to confirm the detection.

[0076] The "needs confirmation" category is a category that is classified when it is not possible to determine whether the gaze area GA that serves as the basis for prediction is valid, and is a category that classifies data whose validity the user should confirm.

[0077] The "Used for analysis" category is used when the AI ​​model is unable to make predictions with a high degree of confidence, and classifies data for which it is advisable for the user to analyze the cause.

[0078] FIG. 12 shows an example of classification of an input image II based on the presence or absence of a prediction area FA, the position of the gaze area GA relative to the prediction area FA, and the accuracy of the prediction result.

[0079] When there is a prediction area FA, the accuracy of the prediction result and the positional relationship of the gaze area GA to the prediction area FA become important.

[0080] Specifically, the case where there is a prediction area FA will be described. If the prediction result is correct and the gaze area GA exists only within the prediction area FA, the gaze area GA is classified into the "valid" category as a validity evaluation.

[0081] In addition, if the prediction result is correct and the gaze area GA exists both inside and outside the prediction area FA (the case shown in Figure 10), the validity of the gaze area GA is classified into the "needs confirmation" category.

[0082] Furthermore, if the prediction result is correct and the gaze area GA exists only outside the prediction area FA, the gaze area GA is classified into the "needs confirmation" category as a validity assessment.

[0083] In addition, if the prediction result is incorrect and the gaze area GA exists only within the prediction area FA, the gaze area GA is classified into the "confirmation required" category as an evaluation of the validity of the gaze area GA.

[0084] In addition, if the prediction result is incorrect and the gaze area GA exists both inside and outside the prediction area FA, the gaze area GA is classified into the "confirmation required" category as an evaluation of the validity of the gaze area GA.

[0085] In addition, if the prediction result is incorrect and the gaze area GA exists only outside the prediction area FA, the gaze area GA is classified into the "needs confirmation" category as an evaluation of the validity of the gaze area GA.

[0086] Also, if there is no GA for the gaze area, it is classified into the "Used for analysis" category.

[0087] "No recognition target" means that the recognition target RO cannot be detected, which contradicts the existence of the predicted area FA. Since there is no such data, classification into categories is not performed.

[0088] Next, a case where there is no prediction area FA will be described. If there is no predicted area FA, the recognition target RO cannot be detected and there is no gaze area GA, so classification into categories based on the relationship between the gaze area GA and the predicted area FA is not performed.

[0089] If the prediction result is correct and the result is "no recognition target", this means that the recognition target RO cannot be detected for the input image II in which the recognition target RO does not exist, and the result is classified into the "valid" category.

[0090] If the prediction result is incorrect and "no recognition target" is given, it means that it is determined that the recognition target RO does not exist even though it does exist in the input image II, and the image is classified into the "used for analysis" category.

[0091] The priority determination unit 42 assigns a confirmation priority to each of the input image II and the data of the prediction result thereof. The priority determination unit 42 assigns a higher priority to data that requires more user confirmation.

[0092] Specifically, the priority determination unit 42 sets the lowest priority to data that has been assigned the "valid" category.

[0093] Furthermore, the priority determination unit 42 sets the highest priority (for example, the first priority) for data to which the "check required" category has been assigned.

[0094] Furthermore, the priority determination unit 42 sets the data assigned the "used for analysis" category as the next highest priority (for example, second priority) after the data assigned the "needs confirmation" category.

[0095] Here, there are various patterns of data assigned the "needs confirmation" category, as shown in Fig. 12. Therefore, it is possible to further differentiate the priority of data assigned the "needs confirmation" category.

[0096] The data that is given a high priority among the data that has been assigned the "needs confirmation" category varies depending on the situation.

[0097] For example, if the goal is to increase the accuracy rate of predictions made by an AI model, a high priority can be set for cases where the prediction results are incorrect, so that users can check them first.

[0098] On the other hand, if the AI ​​has a sufficient accuracy rate and you want to check the validity of the prediction basis, set a high priority to cases where the prediction result is correct and the gaze area GA exists outside the prediction area FA.

[0099] The priority setting by the priority determination unit 42 may be in the form of giving a score such as 0 to 100, or may be in the form of giving flag information indicating whether or not confirmation by the user is required. Also, when assigning flag information, it is sufficient to simply assign "1" to items that require confirmation. In other words, it is not necessary to assign "0" to items that do not require confirmation. Alternatively, a flag may be added only to those items for which confirmation is not required.

[0100] Also, here we have explained an example of classifying into three categories: "Valid", "Needs confirmation", and "Used for analysis", but it is also possible to classify into two categories: "Needs confirmation" and "Other", or "Valid" and "Other".

[0101] The display control unit 5 performs processing to display a heat map of the contribution degree DoC, the validity of the gaze area GA, and the like on the display unit so that the user can understand the priority of confirmation. The display unit may be provided in the information processing device 1, or may be provided in another information processing device configured to be able to communicate with the information processing device 1 (for example, a user terminal used by a user).

[0102] Some examples of presentation screens to be presented to the user are shown below. 13 shows a first example of the presentation screen. The presentation screen is provided with a data display section 51 that displays various information such as images and data to be presented to the user, and a change operation section 52 that changes the display mode of the data displayed on the data display section 51.

[0103] The data display section 51 displays the original image of the image recognition target on which one predicted area FA is superimposed, a heat map of the contribution degree DoC, and the gaze area.

[0104] In addition to these images, the data display unit 51 also displays the file name of the original image, the recognition target RO, the prediction result likelihood PLF, the average contribution DoC within the gaze area GA, the number of gaze areas GA, the average contribution DoC outside the gaze area GA, the category, a valid mark field and an invalid mark field for inputting the confirmation result, etc. In addition, the correctness or incorrectness of the prediction result may also be displayed.

[0105] In the state shown in FIG. 13, input image II and its data are displayed in descending order of presentation priority to the user.

[0106] The change operation section 52 includes a data number change section 61 for changing the number of data items displayed on one page, a search field 62 for searching for data, a data address display field 63 for displaying and changing the location of input data and output data, a display button 64 for displaying data with settings specified by the user, a reload button 65, and a filter condition change button 66 for changing the filter conditions.

[0107] By using the filter function, it is possible to display only data with a high presentation priority. The same applies to other presentation screens, which will be described later.

[0108] Furthermore, the data display unit 51 has a sorting function to change the display mode. For example, by selecting an item name in the table of the data display unit 51, the display order of the data display unit 51 is changed to match the selected item.

[0109] FIG. 14 shows a second example of the presentation screen. In the second example, the same information as in the first example is presented. In addition, the size of each image is changed and displayed according to the category into which it is classified. Specifically, images assigned the "needs confirmation" category are displayed larger. This makes it easier for the user to recognize data that requires confirmation.

[0110] In addition, in the state shown in FIG. 14, input image II and its data are displayed in descending order of presentation priority to the user.

[0111] In addition to changing the size of the image, data assigned the "check required" category may be emphasized by changing the color of the frame or by changing the color of the text.

[0112] Fig. 15 shows a third example of the presentation screen. In the third example, only one image is displayed for one piece of data. In the example shown, a heat map of the degree of contribution DoC is displayed.

[0113] In addition, when one image is selected from the multiple images shown in Figure 15, details of the data related to the selected image may be displayed, such as the file name of the original image, the recognition target RO, the predicted result likelihood PLF, the average contribution DoC within the gaze area GA, the number of gaze areas GA, the average contribution DoC outside the gaze area GA, and the category.

[0114] In the third example, the size of the images varies depending on the data, as shown in the figure. For example, the image of data with a high confirmation priority is displayed large, while the image of data with a low confirmation priority is displayed small. Note that the image of data assigned the "valid" category, which is deemed to have the lowest confirmation priority, may be omitted from display.

[0115] In the state shown in FIG. 15, images with a higher presentation priority to the user are displayed at the top.

[0116] 16 shows a fourth example of the presentation screen. In the fourth example, each piece of data is displayed for each classification result shown in FIG.

[0117] For example, this is suitable when you want to check the data for each classification result, not just the classified category. Also, as in the third example, when an image is selected, detailed data about that image may be displayed.

[0118] Fig. 17 shows a fifth example of the presentation screen. In the fifth example, only the images of each data are displayed in a matrix. Note that Fig. 17 shows only the outer frame of the image, and the contents of the image (input image II and the heat map of contribution DoC superimposed thereon) are omitted for ease of viewing. In this display mode, the user can check a large amount of data at once.

[0119] Note that images with a high confirmation priority may be displayed larger, or only images with a high presentation priority may be displayed.

[0120] Furthermore, similar to the third and fourth examples, when one image is selected, detailed data about that image may be displayed.

[0121] <2. Computer equipment> The information processing device 1 and the user terminal used by the user are configured as a computer device. A functional block diagram of the computer device is shown in FIG. It should be noted that each computer device does not need to have all of the components described below, and may have only some of them.

[0122] 18, a CPU (Central Processing Unit) 71 of each computer device executes various processes in accordance with programs stored in a ROM (Read Only Memory) 72 or a nonvolatile memory unit 74 such as an EEPROM (Electrically Erasable Programmable Read-Only Memory), or a program loaded from a storage unit 79 to a RAM (Random Access Memory) 73. The RAM 73 also stores data necessary for the CPU 71 to execute various processes as appropriate. The CPU 71, ROM 72, RAM 73, and nonvolatile memory unit 74 are interconnected via a bus 83. The bus 83 is also connected to an input / output interface 75.

[0123] The input / output interface 75 is connected to an input unit 76 that includes an operator and an operation device. For example, the input unit 76 may be various types of operators or operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, or a remote controller. An operation by the user is detected by the input unit 76, and a signal corresponding to the input operation is interpreted by the CPU 71.

[0124] Furthermore, the input / output interface 75 is connected integrally or separately to a display unit 77 made up of an LCD (Liquid Crystal Display) or an organic EL panel, and an audio output unit 78 made up of a speaker, etc. The display unit 77 is a display unit that displays various information, and is configured, for example, by a separate display device connected to a computer device. The display unit 77 displays images for various image processing, moving images to be processed, etc. on the display screen based on instructions from the CPU 71. Furthermore, the display unit 77 displays various operation menus, icons, messages, etc., i.e., GUI (Graphical User Interface), based on instructions from the CPU 71.

[0125] The input / output interface 75 may be connected to a storage unit 79 configured with a hard disk or solid-state memory, or a communication unit 80 configured with a modem or the like.

[0126] The communication unit 80 performs communication processing via a transmission path such as the Internet, and communication with various devices via wired / wireless communication, bus communication, and the like.

[0127] A drive 81 is also connected to the input / output interface 75 as required, and a removable storage medium 82 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory is appropriately mounted thereon. Drive 81 allows data files such as image files and various computer programs to be read from removable storage medium 82. The read data files are stored in storage unit 79, and images and sounds contained in the data files are output on display unit 77 and audio output unit 78. Computer programs and the like read from removable storage medium 82 are installed in storage unit 79 as needed.

[0128] In this computer device, for example, software for the processing of this embodiment can be installed via network communication by the communication unit 80 or via a removable storage medium 82. Alternatively, the software may be stored in advance in the ROM 72, the storage unit 79, etc.

[0129] The CPU 71 performs processing operations based on various programs, and thereby the information processing device 1 executes necessary communication processing. The computer device constituting the information processing device 1 is not limited to being configured as a single information processing device as shown in Fig. 18, but may be configured as a system of multiple information processing devices. The multiple information processing devices may be systemized using a LAN or the like, or may be located in a remote location via a VPN using the Internet or the like. The multiple information processing devices may include information processing devices as a server group (cloud) available through a cloud computing service.

[0130] <3. Processing flow> The process executed by the information processing device 1 to present to the user the evaluation result of the validity of the gaze area GA, which is the pixel area on which the AI ​​model bases its judgment in image recognition processing, will be described with reference to the attached drawings.

[0131] The CPU 71 of the information processing device 1 performs a visualization process of the degree of contribution DoC in step S101 of Fig. 19. The detailed process flow of this process will be described later.

[0132] The visualization process of the degree of contribution DoC outputs an image such as that shown in Fig. 7 or that shown in Fig. 8. Note that the frame representing the prediction area FA shown in each figure does not have to be superimposed on the image. That is, in the process of step S101, an image is generated that visualizes the degree of contribution DoC so that the level of the degree of contribution DoC for each partial image region DI can be seen, or so that partial image regions DI with high degrees of contribution DoC can be seen.

[0133] In step S102, the CPU 71 of the information processing device 1 executes a process of identifying a fixation area GA. In this process, an area with a high degree of contribution DoC is identified as the fixation area GA, as shown in FIG.

[0134] In step S103, the CPU 71 of the information processing device 1 executes classification processing. Through this processing, a label is assigned to the input image II according to the relationship between the prediction area FA and the gaze area GA, and the input image II is classified into a category. The label may be a "gaze area" label, a "gaze area outside" label, or a "gaze area not present" label shown in FIG. 12. A specific processing flow will be described later.

[0135] In step S104, the CPU 71 of the information processing device 1 performs a process of assigning a priority to each piece of data. The data classified into categories for each input image II is assigned a presentation priority according to the category.

[0136] Based on the presentation priority, the CPU 71 of the information processing device 1 executes a display control process in step S105. Through this process, presentation screens according to the various display modes shown in Figs. 13 to 17 are displayed on a display unit such as a monitor included in the information processing device 1 or on a display unit included in another information processing device.

[0137] Details of the contribution visualization process in step S101 are shown in FIG.

[0138] In the contribution visualization process, the CPU 71 of the information processing device 1 performs region division processing in step S201. By this processing, the input image II is divided into partial image regions DI (see FIG. 5). Note that the partial image regions DI may have a shape other than a rectangle by using superpixels or the like.

[0139] In step S202, the CPU 71 of the information processing device 1 executes a process of generating a mask image MI. Fig. 6 shows an example of the mask image MI.

[0140] In step S203, the CPU 71 of the information processing device 1 executes image recognition processing using an AI model. This processing executes image recognition processing for detecting the designated recognition target RO.

[0141] In step S204, the CPU 71 of the information processing device 1 executes a process of calculating the degree of contribution DoC. In this process, the degree of contribution DoC is calculated for each partial image region DI.

[0142] In step S205, the CPU 71 of the information processing device 1 performs processing to visualize the degree of contribution DoC. There are various possible visualization methods, and in the above description, examples thereof have been shown in FIGS.

[0143] FIG. 21 shows an example of the details of the classification process in step S103 in FIG. The classification process is performed as many times as the number of input images II. In step S301, the CPU 71 of the information processing device 1 determines whether or not a recognition target RO exists in the input image II.

[0144] If it is determined that the recognition target RO does not exist, that is, if the recognition target RO cannot be detected in the input image II, the CPU 71 of the information processing device 1 assigns a "no recognition target" label to the input image II in step S302.

[0145] Subsequently, in step S303, the CPU 71 of the information processing device 1 determines whether the prediction result that the recognition target RO could not be detected is correct or not. Whether the prediction result is correct or not may be determined and input by the user.

[0146] If the prediction result is correct, the CPU 71 of the information processing device 1 classifies the input image II into the "valid" category in step S304. This case corresponds to the case where the AI ​​model has drawn a correct conclusion based on correct evidence.

[0147] On the other hand, if it is determined in step S303 that the prediction result that the recognition target RO could not be detected is incorrect, i.e., if the recognition target RO was not detected despite being present in the input image II, the CPU 71 of the information processing device 1 classifies the input image II into the "Used for analysis" category in step S305.

[0148] After executing either step S304 or step S305, the CPU 71 of the information processing device 1 ends the classification process shown in FIG.

[0149] If it is determined in step S301 that the recognition target RO exists in the input image II, the CPU 71 of the information processing device 1 determines in step S306 whether or not the gaze area GA exists.

[0150] If it is determined that no gaze area GA exists, i.e., if no partial image area DI with a large contribution DoC exists, the CPU 71 of the information processing device 1 assigns a "no gaze area" label to the input image II in step S307 and classifies it into the "used for analysis" category.

[0151] After completing the process of step S307, the CPU 71 of the information processing device 1 ends the classification process shown in FIG.

[0152] On the other hand, if it is determined in step S306 that the gaze area GA exists, the CPU 71 of the information processing device 1 determines in step S308 whether the number of gaze areas GA is N or less. N is set to a value less than 10 at most, such as 4 or 5.

[0153] If the number of fixation areas GA is greater than N, the CPU 71 of the information processing device 1 proceeds to the process of step S307.

[0154] On the other hand, if the number of fixation areas GA is N or less, the CPU 71 of the information processing device 1 determines in step S309 whether or not the fixation areas GA exist only within the predicted area FA.

[0155] If the gaze area GA exists only within the prediction area FA, the CPU 71 of the information processing device 1 assigns the label "within the prediction area FA" to the input image II in step S310.

[0156] Next, in step S311, the CPU 71 of the information processing device 1 determines whether the prediction result is correct or not.

[0157] If the prediction result is correct, that is, if the recognition target RO has been properly detected, the CPU 71 of the information processing device 1 classifies the input image II into the "valid" category in step S312.

[0158] On the other hand, if the prediction result is incorrect, the CPU 71 of the information processing device 1 classifies the input image II into the "check required" category in step S313.

[0159] After completing the process of either step S312 or step S313, the CPU 71 of the information processing device 1 ends the classification process shown in FIG.

[0160] In step S309, if the gaze area GA does not exist only within the predicted area FA, that is, if the gaze area GA exists at least outside the predicted area FA, the CPU 71 of the information processing device 1 determines in step S314 whether the gaze area GA exists only outside the predicted area FA.

[0161] If it is determined that the fixation area GA exists only outside the prediction area FA, the CPU 71 of the information processing device 1 assigns a label "outside the prediction area" to the input image II and classifies it into the "check required" category in step S315.

[0162] On the other hand, if it is determined that the gaze area GA does not exist only outside the predicted area FA, that is, if it is determined that the gaze area GA exists both inside and outside the predicted area FA, the CPU 71 of the information processing device 1 assigns an "inside / outside predicted area" label to the input image II in step S316 and classifies it into the "needs confirmation" category.

[0163] After completing the process of either step S315 or step S316, the CPU 71 of the information processing device 1 ends the classification process shown in FIG.

[0164] Another example of the classification process in step S103 in Fig. 19 is shown in Fig. 22. Note that the same steps as in Fig. 21 are given the same step numbers and descriptions thereof will be omitted where appropriate.

[0165] In step S301, the CPU 71 of the information processing device 1 determines whether or not a recognition target RO exists in the input image II.

[0166] If it is determined that the recognition target RO does not exist, the CPU 71 of the information processing device 1 executes the processes from step S302 to step S305 as appropriate, and ends the series of processes shown in FIG.

[0167] On the other hand, if it is determined that a recognition target RO exists, the CPU 71 of the information processing device 1 determines in step S321 whether the average value of the contribution degree DoC within the prediction area FA is greater than the average value of the contribution degree DoC outside the prediction area FA, and whether the difference is greater than or equal to the first threshold value Th1.

[0168] If it is determined that the average value of the contribution DoC within the prediction area FA is less than or equal to the average value of the contribution DoC outside the prediction area FA, or the difference is less than the first threshold Th1, for example, if the average values ​​of the contribution DoC inside and outside the prediction area FA are approximately the same, or if the contribution DoC outside the prediction area FA is higher, or if the contribution DoC within the prediction area FA is slightly higher, the CPU 71 of the information processing device 1 determines in step S322 whether the gaze area GA is outside the prediction area FA.

[0169] If it is determined that the gaze area GA does not exist outside the predicted area FA, then the gaze area GA does not exist within the predicted area FA either, and therefore in step S307, the CPU 71 of the information processing device 1 assigns a "no gaze area" label to the input image II and classifies it into the "use for analysis" category.

[0170] On the other hand, if it is determined that the fixation area GA exists outside the prediction area FA, the CPU 71 of the information processing device 1 determines in step S323 whether the average value of the degree of contribution DoC within the prediction area FA is equal to or greater than the second threshold value Th2.

[0171] If it is determined that the contribution DoC within the prediction area FA is greater than or equal to the second threshold Th2, the average value of the contribution DoC outside the prediction area FA is also equal to or greater than the second threshold Th2, so in step S316, the CPU 71 of the information processing device 1 assigns the input image II a label of "inside / outside prediction area" and classifies it into the "needs confirmation" category.

[0172] If it is determined in step S323 that the contribution DoC within the predicted area FA is less than the second threshold Th2, it means that the gaze area GA does not exist within the predicted area FA, and therefore in step S315, the CPU 71 of the information processing device 1 assigns the input image II an "outside predicted area" label and classifies it into the "needs confirmation" category.

[0173] After executing the process of step S307, step S315, or step S316, the CPU 71 of the information processing device 1 ends the series of processes shown in FIG.

[0174] In step S321, if it is determined that the average value of the contribution DoC within the predicted area FA is greater than the average value of the contribution DoC outside the predicted area FA and that the difference is greater than or equal to the first threshold value Th1, the CPU 71 of the information processing device 1 determines in step S324 whether the gaze area GA is outside the predicted area FA.

[0175] If it is determined that the fixation area GA exists outside the prediction area FA, the CPU 71 of the information processing device 1 assigns a label "inside / outside prediction area" to the input image II and classifies it into the "check required" category in step S316.

[0176] On the other hand, if it is determined in step S324 that the fixation area GA does not exist outside the predicted area FA, the CPU 71 of the information processing device 1 executes the processes of steps S310 to S313 and ends the series of processes shown in FIG.

[0177] By executing either the classification process shown in FIG. 21 or FIG. 22, a label is assigned to each input image II input to the AI ​​model and the image is classified into a category. <4. Application Examples> A processing flow for a user to achieve a purpose by utilizing the above-described processes executed by the information processing device 1 will be described with reference to FIGS.

[0178] FIG. 23 shows an example of a processing flow when a user uses a user terminal to connect to the information processing device 1 as a server device and utilizes the AI ​​model generation function provided by the information processing device 1.

[0179] In the following description, the processes shown in FIGS. 23 and 24 are described as being executed in the information processing device 1, but some of the processes may be executed in the user terminal.

[0180] In step S401, the CPU 71 of the information processing device 1 sets and considers a problem. This process is, for example, a process for setting and considering a problem that the user wants to solve, such as analyzing the movement of customers visiting a store. Specifically, the CPU 71 of the information processing device 1 performs initial settings for generating an AI model according to the purpose specified by the user and the specification information of the device that operates the AI ​​model. In the initial settings, for example, the number of layers and the number of nodes of the AI ​​model are set.

[0181] In step S402, the CPU 71 of the information processing device 1 collects learning data. The learning data is a plurality of image data, and may be specified by the user or may be automatically acquired from an image DB (Database) by the CPU 71 of the information processing device 1 according to the purpose.

[0182] In step S403, the CPU 71 of the information processing device 1 performs learning using the learning data, thereby acquiring a trained AI model.

[0183] In step S404, the CPU 71 of the information processing device 1 evaluates the performance of the trained AI model. For example, the performance is evaluated using the accuracy rate of the recognition result of the image recognition process.

[0184] In step S405, the CPU 71 of the information processing device 1 evaluates the validity of the gaze area. This process executes at least the processes of steps S101, S102, and S103 in Fig. 19. In addition, the CPU 71 may execute the processes of steps S104 and S105 in Fig. 19 as processes for user confirmation.

[0185] In step S406, the CPU 71 of the information processing device 1 determines whether the target performance has been achieved. This determination process may be performed by the CPU 71 of the information processing device 1, or the CPU 71 of the information processing device 1 and the CPU 71 of the user terminal may perform a process of allowing the user to select whether the target performance has been achieved.

[0186] If it is determined that the target performance has been achieved, or if the user has selected that the target performance has been achieved, the CPU 71 of the information processing device 1 determines in step S407 whether or not the gaze area GA is appropriate.

[0187] If it is determined to be valid, operation of the AI ​​model is started. When starting operation of the AI ​​model, the CPU 71 of the information processing device 1 may perform processing to start operation of the AI ​​model. For example, it may perform processing to send the AI ​​model to a user terminal, or it may perform processing to store the completed AI model in a DB.

[0188] If it is determined in step S406 that the target performance has not been achieved, or if the user selects that the target performance has not been achieved, the CPU 71 of the information processing device 1 determines in step S408 whether performance improvement can be expected by randomly adding learning data.

[0189] For example, if a lack of learning iterations is suspected, it is determined in step S408 that performance can be expected to improve by randomly adding learning data.

[0190] In this case, the CPU 71 of the information processing device 1 returns to step S402 and collects learning data.

[0191] On the other hand, if it is determined that performance improvement cannot be expected by randomly adding learning data, the CPU 71 of the information processing device 1 performs an analysis based on the evaluation result of the validity of the attention area GA in step S409. That is, the CPU 71 performs an analysis process based on the evaluation result of the validity in step S405 described above.

[0192] Next, in step S410, the CPU 71 of the information processing device 1 determines whether or not additional data having characteristics to be collected has been identified, i.e., whether or not additional data to be collected has been identified. If additional data to be collected has been identified, the CPU 71 of the information processing device 1 returns to step S402 and collects learning data.

[0193] On the other hand, if it is determined that additional data to be collected has not been identified, the CPU 71 of the information processing device 1 returns to step S401 and starts over from setting and examining the problem.

[0194] The AI ​​model thus obtained is then operated by the user to achieve a desired goal.

[0195] The processing flow when a recognition error occurs during operation is shown in Fig. 24. Note that the same steps as in Fig. 23 are given the same step numbers and descriptions thereof will be omitted where appropriate.

[0196] In step S501, the CPU 71 of the information processing device 1 performs analysis processing of the gaze area GA. As described above, this processing is processing for labeling the recognition results focusing on the gaze area GA and classifying them into categories.

[0197] In step S502, the CPU 71 of the information processing device 1 analyzes the analysis result of the gaze area GA.

[0198] In step S408, the CPU 71 of the information processing device 1 determines whether performance can be improved by randomly adding learning data. If it is determined that performance can be improved by randomly adding learning data, the CPU 71 of the information processing device 1 proceeds to a learning data collection process in step S402.

[0199] Then, the CPU 71 of the information processing device 1 performs re-learning in step S503, and updates the AI ​​model in step S504. The updated AI model is deployed and used in the user environment.

[0200] On the other hand, if it is determined in step S408 that performance improvement cannot be expected by randomly adding learning data, the CPU 71 of the information processing device 1 determines in step S410 whether or not additional data having characteristics to be collected exists. If it is determined that additional data having characteristics to be collected exists, the CPU 71 of the information processing device 1 determines in step S505 whether or not data to be deleted exists.

[0201] If it is determined that there is data to be deleted, i.e., if there is an input image II that is not suitable for learning an AI model, the CPU 71 of the information processing device 1 deletes the corresponding input image II in step S506, and then proceeds to processing in step S503.

[0202] On the other hand, if it is determined in step S505 that there is no data to be deleted, the CPU 71 of the information processing device 1 reconsiders the AI ​​model in step S507. In this process, for example, the processes of steps S401, S402, and S403 in FIG. 23 are executed.

[0203] Next, in step S504, the CPU 71 of the information processing device 1 updates the AI ​​model. Through this process, for example, the AI ​​model newly acquired in step S507 is deployed to the user environment.

[0204] <5. Summary> As explained in each of the above examples, the program executed by the information processing device 1 as an arithmetic processing device has a validity evaluation function (function of the classification processing unit 41) that evaluates the validity of the gaze area GA based on the predicted area FA (FA1, FA2, FA3), which is an image area predicted to contain the recognition target RO by image recognition using artificial intelligence (AI) for the input image II, and the gaze area GA (GA1, GA1-1, GA1-2), which is an image area used as the basis for the prediction. For example, it may be determined whether or not the gaze area GA is appropriate, or it may be determined only that the gaze area GA is appropriate, or it may be determined only that the gaze area GA is inappropriate. Therefore, in order to improve the performance of the artificial intelligence, the input image that the worker should check and the predicted result can be identified, allowing the work to be done efficiently and reducing the human and time costs required to confirm the basis on which the predicted result was derived.

[0205] As described with reference to FIG. 21 and the like, in the evaluation of validity, it may be determined that the fixation area GA is valid. This makes it possible to identify cases where the recognition target RO is predicted based on an appropriate gaze area GA. In other words, it becomes possible to extract cases where the recognition target RO is predicted without being based on an appropriate gaze area GA, or cases where it is unclear whether the gaze area GA is appropriate in the first place. Therefore, the operator can specify the input image to be checked and the predicted result.

[0206] As described with reference to FIGS. 21 and 22, the validity evaluation function (the function of the classification processing unit 41) may perform evaluation based on a comparison between the predicted area FA and the gaze area GA. For example, the validity is evaluated based on the positional relationship and degree of overlap between the predicted area FA and the gaze area GA. This allows the validity of the gaze area GA to be properly evaluated, allowing the operator to properly identify the input image II and its prediction result that should be confirmed.

[0207] As described with reference to FIG. 21 and the like, the validity evaluation function (the function of the classification processing unit 41) may perform evaluation based on the positional relationship between the prediction area FA and the fixation area GA. As a result, for example, when the predicted area FA and the gaze area GA match, the gaze area GA is determined to be valid. Therefore, it is possible to identify the input image II and its prediction result, which do not require confirmation by the operator, thereby improving work efficiency.

[0208] As described with reference to FIG. 21 and the like, the validity evaluation function (the function of the classification processing unit 41) may perform evaluation based on whether or not the fixation area GA is located within the prediction area FA. Specifically, if the gaze area GA is included in the prediction area FA, it can be evaluated that the detection of the recognition target RO is performed based on an appropriate gaze area GA. In other words, it can be evaluated that the gaze area GA is valid.

[0209] As described with reference to FIG. 21 and the like, the validity evaluation function (the function of the classification processing unit 41) may perform evaluation based on the number of fixation areas GA. For example, if there is only one gaze area GA, the gaze area GA is likely to be valid. On the other hand, there may be cases where the degree of contribution DoC for the entire area of ​​the input image II is large and the number of gaze areas GA is large. In such cases, the gaze area GA is likely to be invalid. Therefore, by focusing on the number of gaze areas GA, it is possible to evaluate whether prediction (detection) of the recognition target RO is performed based on an appropriate gaze area GA.

[0210] As explained with reference to Figure 21 etc., the validity evaluation function (function of the classification processing unit 41) may determine that the gaze area GA is valid if the gaze area GA exists only within the prediction area FA and the prediction of the recognition target RO is correct. Such a prediction is highly likely to correctly predict the recognition target RO based on correct evidence. By evaluating such an input image II and the prediction result as valid, the efficiency of the confirmation work can be improved.

[0211] As explained with reference to Figure 21 etc., if the validity evaluation function (function of the classification processing unit 41) cannot determine that the gaze area GA is valid, the information processing device 1 as an arithmetic processing device may be caused to execute a classification function (function of the classification unit 4) that classifies the predicted results of image recognition depending on whether or not the gaze area GA exists. If the gaze area GA does not exist, it is impossible to determine whether the gaze area GA is valid. For such input images II, the predicted result likelihood PLF is also low, so it is desirable for the operator to analyze the cause. With this configuration, such input images II can be classified into the "Used for analysis" category, and the input images II to be used for analysis can be clarified.

[0212] As explained with reference to Figures 11 and 21, the information processing device 1 as an arithmetic processing device may be caused to execute a priority determination function (function of the priority determination unit 42) that determines the priority so that the confirmation priority is higher when the gaze area GA cannot be determined to be valid and the gaze area GA exists than when the gaze area GA cannot be determined to be valid and the gaze area GA does not exist. For example, the information processing device 1 as an arithmetic processing device may execute a priority determination function (function of the priority determination unit 42) that determines the priority of checking the predicted result of image recognition to be first priority when it cannot be determined that the gaze area GA is valid and the gaze area GA exists, and determines the priority of checking the predicted result of image recognition to be second priority when it cannot be determined that the gaze area GA is valid and the gaze area GA does not exist, and the first priority may be higher than the second priority. The input images II with the first priority include cases where the recognition target RO is incorrectly recognized based on the gaze area GA within the prediction area FA. In such cases, the AI ​​model confidently detects the wrong object as the recognition target RO. Such input image II can be used for machine learning retraining or additional training to reduce the possibility of false positives and improve the performance of the AI ​​model. Therefore, by setting the priority of such input image II as the first priority, which is higher than the second priority, efficient training of the AI ​​model can be achieved.

[0213] As explained with reference to Figures 19 and 20, the information processing device 1 as an arithmetic processing device may be made to execute a contribution calculation function (function of the contribution calculation unit 24) that calculates the contribution DoC to the prediction result by image recognition for each partial image area DI in the input image II, and a gaze area identification function (function of the gaze area identification processing unit 3) that identifies the gaze area GA based on the contribution DoC. By calculating the degree of contribution DoC for each predetermined image region, it becomes possible to identify the gaze region GA.

[0214] As explained with reference to Figure 22, etc., the validity evaluation function (a function of the classification processing unit 41) may perform evaluation based on the difference between the contribution degree DoC for the prediction area FA and the contribution degree DoC for areas other than the prediction area FA. For example, even if it is determined that the degree of contribution DoC for the predicted area FA is high and that the gaze area GA exists within the predicted area FA, the degrees of contribution DoC for areas other than the predicted area FA may also be generally high. In such a case, the recognition target RO is detected with a great deal of consideration given to areas other than the predicted area FA, and therefore, this is not necessarily an appropriate state. According to this configuration, validity is evaluated based on the difference between the degree of contribution DoC of the predicted area FA and the other areas, thereby making it possible to prevent the validity from being erroneously evaluated as high.

[0215] As explained with reference to Figures 4, 6, 20, etc., the contribution calculation function (function of the contribution calculation unit 24) may calculate the contribution DoC based on the prediction result likelihood PLF for the prediction area FA obtained as a result of making predictions on multiple mask images MI with different patterns of mask presence or absence for each partial image area DI in the input image II. That is, the degree of contribution DoC is an index showing how much the region contributes to the derivation of the prediction result, in other words, to the detection of the recognition target RO. By calculating the degree of contribution DoC for each partial image region DI, the gaze region GA can be appropriately identified.

[0216] As described with reference to FIG. 5 and the like, the partial image area DI may be a pixel area divided into a grid pattern. One possible method for dividing the input image II into partial image regions DI is to use superpixels, which group similar pixels together and treat them as a single region. However, with superpixels, the partial image regions DI may end up being large, and sufficient resolution may not be achieved. On the other hand, by dividing the image into a grid and determining the partial image regions DI without considering the similarity between pixels, it is possible to obtain sufficient resolution for calculating the degree of contribution DoC.

[0217] As described with reference to Figure 1 etc., the display control function (function of the display control unit 5) that executes display control for presenting the predicted results of image recognition may be executed by the information processing device 1 as an arithmetic processing device. By displaying the input image II that requires the worker's confirmation, the worker's work efficiency can be improved. In addition, by displaying information such as the predicted area FA, the gaze area GA, and the accuracy of the prediction result together with the input image II, an environment that makes it easy for the worker to perform confirmation work can be provided.

[0218] As explained with reference to each of Figures 13 to 17, the display control function (function of the display control unit 5) may perform display control so that an image in which the predicted area FA and the gaze area GA are superimposed on the input image II is displayed. This makes it easier to grasp the positions of the predicted area FA and the gaze area GA relative to the input image II, thereby improving the efficiency of the worker's work.

[0219] As explained with reference to each of Figures 13 to 17, the priority determination function (function of the priority determination unit 42) that determines the priority of confirmation for the predicted results of image recognition may be executed by the information processing device 1 as an arithmetic processing device, and the display control function (function of the display control unit 5) may be caused to perform display control so that the predicted results of image recognition are displayed based on the priority when presented. For example, display control may be performed so that input images II and prediction results are displayed in descending order of priority, or so that only input images II and prediction results with high priority are displayed, or so that input images II and prediction results with high priority are displayed in a conspicuous manner, thereby improving the efficiency of the confirmation work.

[0220] As described with reference to each of Figures 13 to 15, the display control function (the function of the display control unit 5) may execute display control so that the display is performed in a display order based on priority. This makes it easier for the operator to grasp the input images II and prediction results that have high priority.

[0221] As described with reference to each of Figures 13 to 17, the display control function (the function of the display control unit 5) may perform display control so that prediction results of image recognition with low priority are not displayed. This prevents the operator from being presented with input images II or prediction results that do not need to be confirmed, thereby improving work efficiency.

[0222] Such a program is executed by the information processing device 1 described above, and can be pre-recorded in a hard disk drive (HDD) as a storage medium built into a device such as a computer, or in a ROM in a microcomputer having a CPU. Alternatively, the program can be temporarily or permanently stored (recorded) on a removable storage medium such as a flexible disk, a CD-ROM (Compact Disk Read Only Memory), a Magneto Optical (MO) disk, a Digital Versatile Disc (DVD), a Blu-ray Disc (registered trademark), a magnetic disk, a semiconductor memory, or a memory card. Such removable storage media can be provided as so-called packaged software. Such programs can be installed onto a personal computer or the like from a removable storage medium, or can be downloaded from a download site via a network such as a LAN (Local Area Network) or the Internet.

[0223] The above-mentioned information processing device 1 is equipped with a validity evaluation unit (classification processing unit 41) that evaluates the validity of the gaze area GA based on a predicted area FA, which is an image area predicted to contain the recognition target RO through image recognition using artificial intelligence AI for the input image II, and the gaze area GA, which is an image area used as the basis for the prediction.

[0224] The information processing method executed by the information processing device 1 is a method in which an arithmetic processing device executes a validity evaluation process (processing by the classification processing unit 41) to evaluate the validity of the gaze area GA based on a predicted area FA, which is an image area predicted to contain the recognition target RO through image recognition using artificial intelligence AI for the input image II, and the gaze area GA, which is an image area used as the basis for the prediction.

[0225] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0226] Furthermore, the above-described examples may be combined in any manner, and even when various combinations are used, the various effects described above can be obtained.

[0227] <6. This Technology> This technology can also be configured as follows. (1) The calculation processing device is caused to execute a validity evaluation function for evaluating the validity of a gaze area based on a prediction area, which is an image area predicted to contain a recognition target by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area used as the basis for the prediction. program. (2) In the evaluation of the validity, it is determined whether the gaze area is valid. The program described in (1) above. (3) The validity evaluation function performs the evaluation based on a comparison between the predicted region and the gaze region. A program according to any one of (1) to (2) above. (4) The validity evaluation function performs the evaluation based on the positional relationship between the prediction area and the gaze area. The program described in (3) above. (5) The validity evaluation function performs the evaluation based on whether the gaze area is located within the predicted area. The program described in (4) above. (6) The validity evaluation function performs the evaluation based on the number of fixation areas. A program according to any one of (1) to (5) above. (7) The validity evaluation function determines that the gaze area is valid when the gaze area exists only within the prediction area and the prediction of the recognition target is correct. The program described in (2) above. (8) When the validity evaluation function fails to determine that the gaze area is valid, the calculation processing device executes a classification function of classifying the prediction result of the image recognition depending on whether the gaze area exists or not. The program described in (2) above. (9) A priority determination function is executed by the arithmetic processing device to determine a priority such that the priority of confirmation is higher in a case where the gaze area cannot be determined to be valid and the gaze area exists than in a case where the gaze area cannot be determined to be valid and the gaze area does not exist. The program described in (8) above. (10) a contribution calculation function for calculating a contribution of each partial image region in the input image to the prediction result by the image recognition; and a gaze area specifying function for specifying the gaze area based on the contribution degree. A program according to any one of (1) to (9) above. (11) The validity evaluation function performs the evaluation based on a difference between the contribution rate for the prediction region and the contribution rate for a region other than the prediction region. The program according to (10) above. (12) The contribution degree calculation function calculates the contribution degree based on a prediction result likelihood for the prediction region obtained as a result of performing the prediction on a plurality of mask images in which a pattern of mask presence or absence is different for each partial image region in the input image. The program according to any one of (10) to (11) above. (13) The partial image area is a pixel area divided into a grid pattern. The program according to any one of (10) to (12) above. (14) A display control function for causing a calculation processing device to execute display control for presenting the prediction result of the image recognition. A program according to any one of (1) to (13) above. (15) The display control function executes the display control so that an image in which the prediction region and the gaze region are superimposed on an input image is displayed. The program according to (14) above. (16) causing a processing unit to execute a priority determination function for determining a priority for confirmation of the predicted result of the image recognition; The display control function executes the display control so that the prediction result of the image recognition is displayed based on the priority. A program according to any one of (14) to (15) above. (17) The display control function executes the display control so that the display is performed in the display order based on the priority. The program according to (16) above. (18) The display control function executes the display control so that the prediction result of the image recognition with a low priority is not displayed. A program according to any one of (16) to (17) above. (19) A validity evaluation unit is provided that evaluates the validity of a gaze area based on a prediction area, which is an image area where a recognition target is predicted to exist by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area that serves as the basis for the prediction. Information processing device. (20) A calculation processing device executes a validity evaluation process for evaluating the validity of a gaze area based on a prediction area, which is an image area predicted to contain a recognition target by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area used as the basis for the prediction. Information processing methods. [Explanation of symbols]

[0228] 1. Information processing equipment 3. Gaze area identification processing unit (gaze area identification function) 4 Classification section (classification function) 5 Display control section (display control function) 24 Contribution calculation section (contribution calculation function) 41 Classification processing unit (validity evaluation function) 42 Priority determination unit (priority determination function) II Input image RO Recognition Target FA, FA1, FA2, FA3 predicted regions GA, GA1, GA1-1, GA1-2 Gaze area DI, DIM partial image area MI mask image PLF predicted outcome likelihood DoC Contribution

Claims

1. A validity evaluation function that evaluates the validity of a gaze area based on a prediction area, which is an image area where a recognition target is predicted to exist by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area that serves as the basis for the prediction; and a contribution calculation function for calculating a contribution of each partial image region in the input image to the prediction result by the image recognition; a gaze area specifying function for specifying the gaze area based on the contribution degree, The validity evaluation function performs the evaluation based on a difference between the contribution rate for the prediction region and the contribution rate for a region other than the prediction region. program.

2. In the evaluation of the validity, it is determined whether the gaze area is valid. The program according to claim 1.

3. The validity evaluation function performs the evaluation based on the positional relationship between the prediction area and the gaze area. The program according to claim 1.

4. The validity evaluation function performs the evaluation based on whether the gaze area is located within the predicted area. The program according to claim 3.

5. The validity evaluation function performs the evaluation based on the number of fixation areas. The program according to claim 1.

6. The validity evaluation function determines that the gaze area is valid when the gaze area exists only within the prediction area and the prediction of the recognition target is correct. The program according to claim 2.

7. When the validity evaluation function fails to determine that the gaze area is valid, the calculation processing device executes a classification function of classifying the prediction result of the image recognition depending on whether the gaze area exists or not. The program according to claim 2.

8. A priority determination function is executed by the arithmetic processing device to determine a priority such that the priority of confirmation is higher in a case where the gaze area cannot be determined to be valid and the gaze area exists than in a case where the gaze area cannot be determined to be valid and the gaze area does not exist. The program according to claim 7.

9. The partial image area is a pixel area divided into a grid pattern. The program according to claim 1.

10. A display control function for causing a calculation processing device to execute display control for presenting the prediction result of the image recognition. The program according to claim 1.

11. The display control function executes the display control so that an image in which the prediction region and the gaze region are superimposed on an input image is displayed. The program according to claim 10.

12. causing a processing unit to execute a priority determination function for determining a priority for confirmation of the predicted result of the image recognition; The display control function executes the display control so that the prediction result of the image recognition is displayed based on the priority. The program according to claim 10.

13. The display control function executes the display control so that the display is performed in the display order based on the priority. The program according to claim 12.

14. The display control function executes the display control so that the prediction result of the image recognition with a low priority is not displayed. The program according to claim 12.

15. a validity evaluation unit that evaluates the validity of a gaze area based on a prediction area, which is an image area predicted to contain a recognition target by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area that serves as the basis for the prediction; a contribution calculation unit that calculates a contribution of each partial image region in the input image to a prediction result obtained by the image recognition; a gaze area identification processing unit that identifies the gaze area based on the contribution degree, The validity evaluation unit performs the evaluation based on a difference between the degree of contribution for the prediction region and the degree of contribution for a region other than the prediction region. Information processing device.

16. a validity evaluation process for evaluating the validity of a gaze area based on a prediction area, which is an image area predicted to contain a recognition target by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area used as the basis for the prediction; a contribution calculation process for calculating a contribution to the prediction result by the image recognition for each partial image region in the input image; a gaze area specifying process for specifying the gaze area based on the contribution degree, In the validity evaluation process, the evaluation is performed based on a difference between the contribution rate for the prediction region and the contribution rate for a region other than the prediction region. Information processing methods.

17. A validity evaluation function that evaluates the validity of a gaze area based on a prediction area, which is an image area where a recognition target is predicted to exist by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area that serves as the basis for said prediction; a contribution calculation function for calculating a contribution of each partial image region in the input image to the prediction result by the image recognition; a gaze area specifying function for specifying the gaze area based on the contribution degree, The contribution degree calculation function calculates the contribution degree based on a prediction result likelihood for the prediction region obtained as a result of performing the prediction on a plurality of mask images in which a pattern of mask presence or absence is different for each partial image region in the input image. program.

18. A calculation processing device is caused to execute a validity evaluation function for evaluating the validity of a gaze area based on a prediction area, which is an image area where a recognition target is predicted to exist by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area that is used as the basis for said prediction; The validity evaluation function performs the evaluation based on the number of fixation areas. program.

19. A validity evaluation function that determines the validity of a gaze area based on a prediction area, which is an image area where a recognition target is predicted to exist by image recognition using artificial intelligence for an input image, and a gaze area, which is an image area that serves as the basis for said prediction; a classification function of classifying the prediction result of the image recognition depending on whether or not the gaze area exists when the validity evaluation function fails to determine that the gaze area is valid. program.

Citation Information

Patent Citations

  • Information processing device, processor for endoscope, information processing method, and program

    JP2020089712A

  • Electronic apparatus, control unit, control method, and control program

    JP2020102657A

  • Explanation support device, and explanation support method

    JP2021022159A

  • Image generation device, image generation method, and program

    JP2021093004A

  • Information processing device, information processing method, and computer program

    US20210019656A1