Information processing system, control method therefor, and program

The proposed mechanism addresses the reliability issue of AI systems by comparing inference results from two evaluation datasets, enabling effective visualization and analysis of AI performance and reliability.

JP2025095151APending Publication Date: 2025-06-26CANON MARKETING JAPAN INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023210965
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing AI systems face reliability issues as high accuracy in evaluation datasets does not guarantee performance in real-world field operations, leading to decreased AI reliability.

Method used

A mechanism is provided to extract and compare data with different inference results from two evaluation datasets, allowing for the visualization and analysis of AI reliability by displaying the distribution of values in these datasets.

Benefits of technology

This mechanism enables the assessment of AI reliability by comparing inference results across different datasets, providing insights into the AI's performance and helping to identify areas for improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025095151000001_ABST
    Figure 2025095151000001_ABST
Patent Text Reader

Abstract

To provide a mechanism for grasping reliability of AI.SOLUTION: An information processing system provided herein is configured to provide control to extract data with different inference results by a trained model from a first evaluation dataset and second evaluation dataset used to evaluate the trained model, and compare and display a distribution of values of a first variable in the first evaluation dataset and a distribution of values of the first variable in the data with different inference results.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system, a control method thereof, and a program.

Background Art

[0002] In recent years, AI generated by deep learning or machine learning using learning data has been used in various fields. When determining whether AI is practical, it is common to prepare an evaluation dataset and measure the accuracy with respect to it. However, the quality of AI may not be reliable based solely on the performance with respect to the evaluation dataset.

[0003] Patent Document 1 discloses a technique for visualizing a prediction situation in order to grasp the factors that cause an error between a prediction and an actual result.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Disclosure of the Invention

Problems to be Solved by the Invention

[0005] Even if the accuracy with respect to the evaluation dataset is high, high accuracy may not be obtained when actually operating in the field. In such a case, the reliability of AI decreases, so a mechanism for developing reliable AI is necessary.

[0006] Therefore, an object of the present invention is to provide a mechanism capable of grasping the reliability of AI.

Means for Solving the Problems

[0007] In order to solve the above problems, the present invention Extraction means for extracting data with different inference results by the learned model in the first evaluation dataset and the second evaluation dataset used for evaluating the learned model, Display control means for controlling to compare and display the distribution of the values of the first variable in the first evaluation dataset and the distribution of the values of the first variable in the data with different inference results, Characterized by comprising.

Effect of the Invention

[0008] According to the present invention, it becomes possible to provide a mechanism capable of grasping the reliability of AI.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0011] FIG. 1 is a diagram showing an example of the system configuration of a visualization system for the nature of inferences of an AI (trained model generated by deep learning / machine learning) in an embodiment of the present invention.

[0012] The client terminal 101 is configured to be connected via the network 100. The client terminal is, for example, a personal computer (hereinafter referred to as a PC). The client terminal 101 performs the main processing for visualizing the nature of inferences by the AI. The network 100 can take forms such as a wired LAN, a wireless LAN, or a USB according to the physical interface of the client terminal 101. A server 102 may be placed on the network 100. The client terminal 101 may read data from the server 102.

[0013] FIG. 2 is a block diagram showing an example of the hardware configuration of the client terminal 101 in an embodiment of the present invention.

[0014] As shown in FIG. 2, in the client terminal, a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an input controller 205, a video controller 206, a memory controller 207, and a communication I / F controller 208 are connected via a system bus 204.

[0015] The CPU 201 comprehensively controls each device and controller connected to the system bus 204.

[0016] The ROM 202 or the external memory 211 holds the BIOS (Basic Input / Output System), the OS (Operating System), which are control programs executed by the CPU 201, a computer-readable and executable program for implementing this information processing method, and various necessary data (including data tables).

[0017] The RAM 203 functions as the main memory, work area, etc. of the CPU 201. When executing a process, the CPU 201 loads a program, etc. necessary for the execution from the ROM 202 or the external memory 211 into the RAM 203, and realizes various operations by executing the loaded program.

[0018] The input controller 205 controls the input from input devices such as a keyboard 209 and a pointing device such as a mouse (not shown). When the input device is a touch panel, it is assumed that the user can give various instructions by pressing (touching with a finger, etc.) in accordance with icons, cursors, buttons, etc. displayed on the touch panel.

[0019] Also, the touch panel may be a touch panel capable of detecting the positions touched by a plurality of fingers, such as a multi-touch screen.

[0020] The video controller 206 controls the display to an external output device such as a display 210. The display includes the display of a notebook personal computer integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. Also, for a device capable of receiving the above-described touch operation, an input device is also provided.

[0021] The video controller 206 can control a video memory (VRAM) for performing display control, and can use a part of the RAM 203 as a video memory area, or can also provide a dedicated video memory separately.

[0022] The memory controller 207 controls access to the external memory 211. As the external memory, an external storage device (hard disk) that stores a boot program, various applications, font data, user files, edited files, and various data, a flexible disk (FD), or a compact flash (registered trademark) memory connected via an adapter to a PCMCIA card slot can be used.

[0023] The communication I / F controller 208 is for connecting and communicating with an external device via a network, and executes communication control processing on the network. For example, communication using TCP / IP, a telephone line such as ISDN, and communication using a 3G line of a mobile phone are possible.

[0024] Further, the CPU 201 enables display on the display 210 by executing an outline font expansion (rasterization) process on, for example, a display information area in the RAM 203. In addition, the CPU 201 enables a user instruction using a mouse cursor (not shown) on the display 210.

[0025] Next, an example of the functional configuration of the client terminal 101 of the present invention will be described with reference to FIG. 3.

[0026] The input reception unit receives an input instruction from the user through the input controller 205.

[0027] The data analysis unit acquires data stored in the database 340 and the file storage 350 according to the received instruction, and processes it for inspection by the AI quality inspection unit.

[0028] The AI quality inspection department performs AI quality inspection using the analyzed data. It executes the process in Figure 4 to inspect whether the AI to be inspected meets each inspection item.

[0029] The result output unit displays the inspected results on the display of the client terminal 101. For example, it displays inspection result screens such as those from Figure 11 to Figure 21 on the display. Hereinafter, the display destination in this embodiment shall be the display of the client terminal 101.

[0030] In the database 340 and the file storage 350, an evaluation data table 1500, a learning data table 1550, etc. as shown in Figure 23 are stored. In this embodiment, the CPU 201 acquires the data table 1550 etc. from the database 340 and the file storage 350 and uses them for the quality evaluation of the AI.

[0031] Next, with reference to the flowcharts in Figures 4 to 11, the process executed by the client terminal 101 in the embodiment of the present invention will be described.

[0032] When evaluating an AI, multiple viewpoints can be considered. For example, there are "explanability", "bias", "sensitivity", "error analysis", and "backward compatibility".

[0033] Explanability is a viewpoint for evaluating whether the inferences made by the AI are reliable. Since an AI with unreliable inference results is difficult to utilize, it is evaluated from this viewpoint.

[0034] Bias is a viewpoint for evaluating whether there is any bias in the dataset used for inferences, learning, and evaluation by the AI. By evaluating whether there is bias in the dataset, it is evaluated whether an AI that is poor at inferring a specific label has been constructed.

[0035] Robustness is a perspective for evaluating the stability of AI learning. For example, robustness is evaluated by verifying whether an AI with good accuracy has not been built only for a specific evaluation dataset.

[0036] Error analysis is a perspective for evaluating whether it is easy to analyze errors. When the performance of an AI is not good, it is evaluated whether the points for improvement can be immediately understood.

[0037] Backward compatibility is a perspective for evaluating changes in the nature of inference between the pre-improvement model and the post-improvement model. When the performance of an AI is improved, there may be cases where data that could be correctly inferred before the improvement can no longer be inferred, so it is evaluated from this perspective.

[0038] The flowchart of FIG. 4 is a process in which the CPU 201 of the client terminal 101 reads and executes a predetermined control program, and is a flowchart showing a process for visualizing the nature of inference by AI. Note that the flowchart of FIG. 4 is executed as an internal process of the operation screen 1200 of FIG. 18.

[0039] In step S401, the CPU 201 of the client terminal 101 receives an input of an AI job number from the user.

[0040] The AI job number is information (identifier) that uniquely identifies the generated AI. Also, using the AI job number as a key, data related to the AI identified by the AI job number can be obtained from the database 340 or the file storage 350. Data related to the AI is, for example, the evaluation data table 1500 of the inference results for the AI evaluation dataset of FIG. 22. The evaluation data table 1500 consists of items such as the row number 1510 of the data, the explanatory variable column 1511, the correct label column 1515, and the inference label (inference result by AI) column 1516.

[0041] The line number 1510 of the data is an item for uniquely identifying the data lines, and numbers are registered.

[0042] The explanatory variable column 1511 is an item of explanatory variables used in inference by AI, and floating-point (float) type numerical values, integer (int type) numerical values, and strings representing categories are registered.

[0043] In the correct label column 1515, there is an item of the correct target variable, and strings representing class labels or integer values representing class indices are registered.

[0044] In the inference label column 1516 by AI, there is an item of the target variable inferred by AI, and strings representing class labels or integer values representing class indices are registered.

[0045] In addition to the above, there is also data related to AI, such as the learned AI, accuracy information for the evaluation data set, and the learning data table 1550 which is the data set used for AI learning. The learning data set is in the same format as the evaluation data set, but since it is the data set used for learning, it does not have the inference label column 1516 by AI.

[0046] Also, an example of receiving the input of the AI job number from the user is shown in FIG. 18. FIG. 18 is an AI inspection selection screen for receiving the input of the AI job number from the user and performing an inspection of the AI. The AI inspection selection screen consists of a pull-down list 1201, the selected AI job number 1202, and an inspection result display area 1210.

[0047] The pull-down list 1201 is for receiving the selection of the AI job number from the user. In the pull-down list 1201, multiple AI job numbers can be selected as shown in FIG. 19. In that case, the inspection result display area 1210 in FIG. 18 is divided, and the inspection results can be displayed in parallel like the inspection result display area A 1210 and the inspection result display area B 1211 in FIG. 19.

[0048] The selected AI job number 1202 displays the AI job number input by the user.

[0049] The inspection result display area 1210 is an area for displaying the inspection results of the AI, and it displays the inspection results of the inference characteristics by the AI after step S402.

[0050] In FIG. 18, the AI job number is displayed in a pull-down list for selection, but any method may be used as long as the user can specify the AI to be inspected, such as a form that accepts input of the AI job number from the user.

[0051] In step S402, the CPU 201 of the client terminal 101 executes a subprogram that visualizes the explanatory variables with low-accuracy intervals for the AI specified by the AI job number accepted in step S401. This subprogram is one of the inspection items for understanding the nature of the inference by the AI. By calculating the accuracy for each interval of the explanatory variables, it visualizes the low-accuracy intervals and aims to provide information for analyzing the nature of the AI to the AI developer.

[0052] FIG. 6 is a flowchart showing the flow of the process in S402 (the process of visualizing the explanatory variables with low-accuracy intervals).

[0053] In step S601, the CPU 201 of the client terminal 101 acquires the evaluation data table 1500 associated with the AI job number input in step S401 and the accuracy information of the entire dataset. The acquisition sources are the database 340 and the file storage 350.

[0054] In step S602, the CPU 201 of the client terminal 101 acquires the accuracy reduction threshold value used for determining OK (qualified) / NG (unqualified) in the inspection. The acquisition sources are the database 340 and the file storage 350.

[0055] In step S603, the CPU 201 of the client terminal 101 acquires the number of intervals when dividing the interval in the subsequent step S606. The acquisition source is the database 340 or the file storage 350.

[0056] In step S604, the CPU 201 of the client terminal 101 acquires each column of the explanatory variable column 1511 in the evaluation data table 1500 one by one, and starts a loop to combine the correct label column 1515 and the inference label column 1516 by AI. An example after combination is shown in Fig. 23(a). The combined table 1600 is a table obtained by combining the explanatory variable var1 column 1512, the correct label column 1515, and the inference label column 1516 by AI. This is repeated for all columns of explanatory variables, and when the processing for all explanatory variables is completed, the loop is terminated.

[0057] In step S605, the CPU 201 of the client terminal 101 determines whether the explanatory variable for one column acquired in step S604 is a float-type explanatory variable. If YES, it proceeds to step S606; if NO, it proceeds to S607.

[0058] In step S606, the CPU 201 of the client terminal 101 divides the float-type explanatory variable for one column acquired in step S604 into intervals by the number of intervals acquired in step S603 and assigns them to the intervals.

[0059] The method of dividing the interval is to acquire the minimum value and the maximum value in one column of the explanatory variable, and divide them into intervals with equal intervals. For example, when the minimum value of one column of the explanatory variable is 0, the maximum value is 10, and the number of intervals is 5, the following five intervals are generated. That is, the interval from 0 to 2, the interval from 2 to 4, the interval from 4 to 6, the interval from 6 to 8, and the interval from 8 to 10. Then, each piece of the float-type explanatory variable for one column is assigned to the divided intervals.

[0060] An example of interval division is shown in Fig. 23(a). By the method described above, the interval is divided into intervals 1610 from 4 to 6, interval 1620 from 6 to 8, interval 1630 from 8 to 10, etc. Among the records 1605 in the join table 1600, since the value in the explanatory variable var1 column 1512 is 4.8, it is assigned to the interval 1610 from 4 to 6.

[0061] In step S607, the CPU 201 of the client terminal 101 calculates the accuracy based on the records for each interval. The accuracy is indicated, for example, by accuracy which is a ratio when the number of data is the denominator and the number of correct answers that match the inference by AI is the numerator.

[0062] Note that explanatory variables other than the float type are regarded as categorical variables, and the accuracy is calculated for each category that the categorical variable takes. An example in the case of an explanatory variable other than the float type is shown in Fig. 23(b). When the value taken by the explanatory variable var3 column 1517 is a string such as red or blue, the records 1640 with the value red and the records 1650 with the value blue are grouped by string, and the accuracy of each is calculated. Also, in the case where the value taken by the explanatory variable var4 column 1518 is an integer such as 1, 2.., similar to this, the records 1660 with the value 1 and the records 1670 with the value 2 are grouped by value, and the accuracy of each is calculated.

[0063] In step S608, the CPU 201 of the client terminal 101 sets a threshold based on the accuracy information of the entire dataset acquired in step S601 and the decrease threshold acquired in step S602. Based on the set threshold, the accuracy for each interval or each category obtained in step S607 is evaluated.

[0064] For example, if the accuracy information of the entire dataset obtained in step S601 is 60% and the accuracy degradation threshold is 5 percentage points, the threshold will be 55%. It is determined whether it is below this threshold. In the example of FIG. 23(c), the inspection determination table 1680 for each section is shown. For example, in the section from 0 to 2, the accuracy is 70%, which is above the threshold, so the determination is OK. However, in the section from 6 to 8, it is 20%, which is below the threshold, so the determination is NG. Based on the inspection determination table 1680, the bar graph 713 in FIG. 11 is created.

[0065] In step S609, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, information regarding the accuracy of the entire dataset and the degradation threshold is displayed in the result explanation column 702, and a message is created for the explanation of the section of the explanatory variable that did not meet the threshold.

[0066] In step S610, the CPU 201 of the client terminal 101 reads out an advice message on the actions to be taken by future AI developers according to the inspection result from the database 340 or the file storage 350. Note that a pre-created message may be obtained from the database 340 or the file storage 350 for display, or an appropriate message may be displayed based on past data by utilizing AI.

[0067] In step S611, the CPU 201 of the client terminal 101 displays the accuracy inspection result 700 for each section in FIG. 11 in the inspection result display area 1210 in FIG. 18. Note that the inspection result for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.

[0068] Fig. 11 shows an example of a result screen visualizing explanatory variables with low-accuracy intervals. The result screen consists of inspection result information 701, result explanation column 702, pull-down list 703 for selecting explanatory variables, accuracy graph 710 for each interval, vertical axis 711 of the graph, horizontal axis 712 of the graph, bar graph 713, breakdown of accuracy 714, bar graph 715 that did not meet the standard, and future action information 730.

[0069] The inspection result information 701 displays the result of the inspection (OK / NG). If there is no interval lower than the threshold value in step S608, it is displayed as OK; otherwise, it is displayed as NG.

[0070] The result explanation column 702 is the explanation of the result. The message generated in step S609 is displayed here.

[0071] The pull-down list 703 for selecting explanatory variables is a pull-down list for selecting explanatory variables.

[0072] The accuracy graph 710 for each interval is a graph of the accuracy for each interval in the explanatory variable.

[0073] The vertical axis 711 of the graph shows the numerical value of accuracy on the vertical axis.

[0074] The horizontal axis 712 of the graph shows intervals or categorical variable names on the horizontal axis.

[0075] The bar graph 713 shows the accuracy in the interval.

[0076] The breakdown of accuracy 714 shows the accuracy, the number of learning cases in the interval, and the number of correct data. For example, it shows that out of 10 learning cases, 6 were correct. In addition, the breakdown of 1 correct out of 5 is also displayed, allowing one to understand which intervals have been learned more and which are less.

[0077] The bar graph 715 of the section that did not meet the standard is highlighted conspicuously. Also, an exclamation mark 716 appears next to the accuracy information and is highlighted. As a result, it becomes easier to recognize which section has low accuracy, enabling verification of whether there are areas where AI struggles or whether the number of datasets is sufficient. For example, it can be determined that it is necessary to increase the amount of learning data for sections with low accuracy. Conversely, if the learning data is more abundant in a particular section than in others, it can be recognized that overfitting has occurred in that section, allowing confirmation of whether the learning amount is uneven among any of the explanatory variables.

[0078] The future action information 730 displays the future actions based on the inspection results obtained in step S610.

[0079] Returning to the description of FIG. 4. In step S403, the CPU 201 of the client terminal 101 executes a subprogram that performs a process of calculating and visualizing the degree of coincidence between the correct distribution and the inference distribution. This subprogram is one of the inspection items for grasping the nature of the inference by AI, and its purpose is to inspect and visualize whether there are any abnormalities in the tendency of the inference by AI by checking whether there is a deviation between the histogram of the labels inferred by AI and the histogram of the correct labels.

[0080] FIG. 7 is a flowchart showing the flow of the process in S403 (the process of calculating and visualizing the degree of coincidence between the correct distribution and the inference distribution).

[0081] In step S701, the CPU 201 of the client terminal 101 acquires the evaluation data table 1500 associated with the AI job number input in step S401 from the database 340 or the file storage 350.

[0082] In step S702, the CPU 201 of the client terminal 101 acquires the similarity threshold used for determining OK (pass) / NG (fail) of the inspection. The acquisition source is the database 340 or the file storage 350.

[0083] In step S703, the CPU 201 of the client terminal 101 processes the evaluation data table 1500 acquired in step S701 according to the current processing. Specifically, a table is created by extracting the correct label column 1515 and the inference label column 1516 by AI from the evaluation data table 1500.

[0084] In step S704, the CPU 201 of the client terminal 101 aggregates the histogram of the correct label column 1515. That is, for the correct label column 1515, the number of each class is aggregated.

[0085] In step S705, the CPU 201 of the client terminal 101 aggregates the histogram of the inference label column 1516. That is, for the inference label column 1516, the number of each class is aggregated.

[0086] In step S706, the CPU 201 of the client terminal 101 regards the histogram of the correct label obtained in step S704 and the histogram of the inference label obtained in step S705 as probability distributions, and calculates the degree of coincidence between the probability distributions. For example, JS divergence can be used to calculate the degree of coincidence. Note that the smaller the value of the JS divergence, the more similar the probability distributions are.

[0087] In step S707, the CPU 201 of the client terminal 101 compares the degree of coincidence obtained in step S706 with the threshold value acquired in S702, and outputs the inspection result. When the calculation method is JS divergence, if the degree of coincidence is less than or equal to the threshold value, it is OK; otherwise, it is NG.

[0088] By calculating the degree of agreement between the correct distribution and the inference distribution, it is possible to check whether there are any abnormalities in the inference tendency. That is, if there is a problem with the learning method, there is a possibility that the inference tendency will be biased. Explaining using the distribution shown in Figure 12, the correct distribution is a distribution in the order of Class 1, 3, 2, 4 from the one with the largest number of data, while the inference distribution is in the order of Class 2, Class 3, Class 1, Class 4 from the one with the largest number of data. Since there is a difference in the distribution of data like this, when calculating the distance between the correct distribution and the inference distribution using JS divergence, it becomes "0.334". This value exceeds the threshold (0.3). This indicates that the AI does not have the correct inference tendency with respect to the correct distribution, and it is necessary to adjust the learning data.

[0089] In step S708, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, a message to be displayed in the result explanation column 802 is created regarding the degree of agreement obtained in step S706, the information of the threshold value acquired in S702, and the result of comparing those pieces of information.

[0090] In step S709, the CPU 201 of the client terminal 101 reads out an advice message on the actions to be taken by future AI developers from the database 340 or the file storage 350. Note that it may be possible to obtain and display a pre-created message from the database 340 or the file storage 350, or it may be possible to display an appropriate message based on past data by utilizing AI.

[0091] In step S710, the CPU 201 of the client terminal 101 displays the inspection result 800 of the correct and inference distributions in Figure 12 in the inspection result display area 1210 in Figure 18. Note that the inspection result for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.

[0092] Fig. 12 shows an example of a result screen visualizing the degree of coincidence between the correct distribution and the inference distribution. The result screen consists of inspection result information 801, result explanation column 802, histogram plot 810, vertical axis 811 of the histogram, horizontal axis 812 of the histogram, bar graph 813 of the correct label, bar graph 814 of the inference label, legend 820 of the histogram, and future action information 830.

[0093] The inspection result information 801 displays the result (OK / NG) of the inspection. The result determined in step S707 is displayed.

[0094] The result explanation column 802 displays the message generated in step S708 here.

[0095] The histogram plot 810 is a plot that arranges the histograms of the correct label and the inference label side by side.

[0096] The vertical axis 811 of the histogram represents the number of labels in the evaluation dataset.

[0097] The horizontal axis 812 of the histogram represents the types of labels of the target variable.

[0098] The bar graph 813 of the correct label is a bar graph of the number of correct labels. That is, it is a bar graph showing the number aggregated for each class in step S704.

[0099] The bar graph 814 of the inference label is a bar graph of the number of inference labels. That is, it is a bar graph showing the number aggregated for each class in step S705.

[0100] The legend 820 is the legend of the histogram.

[0101] The future action information 830 displays the future actions based on the inspection results obtained in step S709.

[0102] Return to the description of FIG. 4. In step S404, the CPU 201 of the client terminal 101 executes a subprogram that visualizes the outliers of the explanatory variables in the dataset. This subprogram is one of the inspection items for grasping the nature of the dataset used for AI learning and evaluation. By examining the distribution of possible values of each explanatory variable in the dataset, it is possible to check whether there are extremely large (or small) values in the explanatory variables and inspect the quality of the dataset. If there are outliers, it may lead to a decrease in the inference accuracy or learning speed of the AI. Therefore, it is necessary to detect the presence or absence of outliers and exclude them if necessary.

[0103] FIG. 8 is a flowchart showing the flow of the process in S404 (the process of visualizing the outliers of the explanatory variables in the dataset).

[0104] In step S801, the CPU 201 of the client terminal 101 acquires the evaluation data table 1500 associated with the AI job number input in step S401 from the database 340 or the file storage 350.

[0105] In step S802, the CPU 201 of the client terminal 101 acquires the learning data table 1550 associated with the AI job number input in step S401 from the database 340 or the file storage 350.

[0106] In step S803, for each of the evaluation data table 1500 and the learning data table 1550, the CPU 201 of the client terminal 101 acquires the explanatory variable columns of the float type among the explanatory variable columns 1511.

[0107] In step S804, the CPU 201 of the client terminal 101 acquires the float-type explanatory variable column acquired in step S803 one column at a time from the explanatory variable column 1511, and starts the loop of the subsequent processing. The subsequent processing repeats for the number of float-type explanatory variable columns acquired in step S803, and when the processing for all explanatory variables is completed, the loop is terminated. Also, the subsequent processing will be described taking the evaluation data table 1500 as an example, but the same processing is executed for the learning data table 1550 as well.

[0108] In step S805, the CPU 201 of the client terminal 101 calculates the quartiles for the float-type explanatory variable column acquired in step S804.

[0109] In step S806, the CPU 201 of the client terminal 101 creates a box plot based on the quartiles of the explanatory variables obtained in step S805.

[0110] In step S807, the CPU 201 of the client terminal 101 determines whether there are outliers in the explanatory variables. For example, it determines whether there is a value that is more than 1.5 times the vertical width of the box away from the end of the box, and determines a value that is more than 1.5 times away as an outlier. Although the value "more than 1.5 times the vertical width of the box away from the end of the box" is used as the outlier, it is assumed that the criteria used to determine whether it is an outlier can be set arbitrarily.

[0111] In step S808, the CPU 201 of the client terminal 101 determines the inspection result. Here, if an outlier is found in step S807, it is NG; otherwise, it is OK.

[0112] In step S809, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, for the explanation of whether an explanatory variable containing an outlier is found in the dataset, a message to be displayed in the result explanation column 902 is created.

[0113] In step S810, the CPU 201 of the client terminal 101 reads an advice message on the actions that future AI developers should take from the database 340 or the file storage 350. Note that it is also possible to obtain and display a pre-created message from the database 340 or the file storage 350, or to display an appropriate message based on past data by utilizing AI.

[0114] In step S811, the CPU 201 of the client terminal 101 displays the outlier inspection result 900 in Fig. 13 in the inspection result display area 1210 of Fig. 18. Note that the inspection result for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.

[0115] Fig. 13 shows an example of a result screen visualizing the outliers of the explanatory variables of the dataset. The result screen consists of inspection result information 901, result description column 902, box plot area 910, vertical axis 911 of the box plot, horizontal axis 912 of the box plot, outlier 914, and future action information 930.

[0116] The inspection result information 901 displays the result (OK / NG) of the inspection. The result determined in step S808 is displayed.

[0117] The result description column 902 displays the message generated in step S809 here.

[0118] The box plot area 910 displays the box plot of the explanatory variables created in step S906. In Fig. 13, only the box plots of var1 and var2 are displayed, but it is also possible to display all the evaluated explanatory variable columns, or to display the explanatory variable columns selected by a pull-down list or the like.

[0119] The vertical axis 911 of the box plot represents the value range of one explanatory variable.

[0120] The horizontal axis 912 of the box-and-whisker plot indicates the labels of the training dataset and the evaluation dataset. For each explanatory variable column, as shown in Figure 13, the training dataset and the evaluation dataset are displayed side by side.

[0121] If there are outliers in the training dataset, it may lead to a decrease in the inference accuracy of the AI, etc., so it is necessary to display the presence or absence of outliers. Similarly, if there are outliers in the evaluation dataset, it may also lead to a decrease in the inference accuracy of the AI, etc. In this case, if the accuracy of the AI for the evaluation dataset is low, it is possible that the evaluation dataset contains many outliers rather than the training dataset, which is the cause of the accuracy decrease. Therefore, it is necessary to investigate not only the training dataset but also how many outliers exist in the evaluation dataset.

[0122] Also, by displaying the training dataset and the evaluation dataset side by side, for example, there is no need to switch the screen for comparison, etc., and it becomes easy to check whether the explanatory variables to be evaluated contain outliers.

[0123] The outlier 914 represents the outlier found in step S808.

[0124] The future action information 930 displays the future actions based on the inspection results obtained in step S810.

[0125] Although an example where outliers can be visualized by a box-and-whisker plot has been described, any method that can visualize outliers, such as standard deviation or cluster analysis, is acceptable.

[0126] Return to the description of FIG. 4. In step S405, the CPU 201 of the client terminal 101 executes a subprogram for performing a noise tolerance test process. This subprogram is one of the inspection items for grasping the nature of AI inference, and aims to inspect how robust the AI is when a minute value is given to each explanatory variable. That is, when there are some changes in the explanatory variables, if the inference result is greatly affected, the AI becomes difficult to use, so the robustness of the AI is evaluated.

[0127] FIG. 9 is a flowchart showing the flow of the process of S405 (the process of performing a noise tolerance test).

[0128] In step S901, the CPU 201 of the client terminal 101 acquires the AI, the evaluation data table 1500, and the accuracy in the evaluation data set associated with the AI job number input in step S401 from the database 340 or the file storage 350. In step S902, the CPU 201 of the client terminal 101 acquires the parameter of the noise to be applied to the evaluation data set (for example, the magnification factor multiplied by the variance of the normal distribution) from the database 340 or the file storage 350.

[0129] In step S903, the CPU 201 of the client terminal 101 acquires the threshold value that is the allowable amount of accuracy degradation used for the determination of OK (qualified) / NG (unqualified) of the inspection. The acquisition source is the database 340 or the file storage 350.

[0130] In step S904, the CPU 201 of the client terminal 101 processes the evaluation data table 1500 acquired in step S901 according to the current process. Specifically, the row number 1510 of the data, the explanatory variable column 1511, and the correct label column 1515 are extracted from the evaluation data table 1500, and an extraction table 1800 as shown in FIG. 24 is created.

[0131] In step S905, the CPU 201 of the client terminal 101 acquires the explanatory variable columns of the float type from among the explanatory variable columns 1511.

[0132] In step S906, the CPU 201 of the client terminal 101 acquires the columns of the float-type explanatory variables acquired in step S905 from among the explanatory variable columns 1511 one by one, and starts the loop of the subsequent processing. The subsequent processing is repeated for the number of columns of the float-type explanatory variables acquired in step S905, and when the processing for all the explanatory variables is completed, the loop is terminated.

[0133] In step S907, the CPU 201 of the client terminal 101 calculates the variance of the data in the explanatory variable var1 column 1512.

[0134] In step S908, the CPU 201 of the client terminal 101 generates noise 1810 for the number of records in the explanatory variable var1 column 1512.

[0135] The noise is a value sampled from a normal distribution. The mean, which is a parameter of the normal distribution, is 0, and the variance is the value calculated in step S907 multiplied by the magnification factor acquired in step S902. The reason for multiplying by the magnification factor is that it is necessary to change the amount of noise according to the problem handled by the AI, and a parameter for adjusting the magnitude of the noise is required. In this way, noise 1810 is generated. Note that sampling from the normal distribution may be done randomly or in a non-random manner.

[0136] In step S909, the CPU 201 of the client terminal 101 copies the extraction table 1800 created in step S904, and adds the noise 1810 to each record in the explanatory variable var1 column 1512. The table after adding noise to the extraction table 1800 is the table 1820 after noise addition, and the explanatory variable var1 column 1512 becomes the explanatory variable var1 column 1821 after noise addition. For example, the explanatory variable with the data row number 1510 being 0 is "4.8", and when noise is added to this, it is changed to "4.9".

[0137] In step S910, the CPU 201 of the client terminal 101 performs inference by the AI acquired in step S901 on the dataset of the post-noise addition table 1820 generated in step S909, and calculates the accuracy. That is, using the same AI as before the noise addition, the inference accuracy for the post-noise addition table 1820 is calculated.

[0138] In step S911, the CPU 201 of the client terminal 101 compares the accuracy of the pre-noise addition evaluation dataset acquired in step S901 with the accuracy of the post-noise addition evaluation dataset acquired in S910, and calculates the amount of accuracy degradation.

[0139] In step S912, the CPU 201 of the client terminal 101 determines whether the amount of accuracy degradation obtained in step S911 is within the threshold value of the allowable amount of accuracy degradation acquired in step S903. If it is within the threshold value, it is considered OK; otherwise, it is considered NG.

[0140] In step S913, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, for the amount of accuracy degradation before and after the noise addition and the explanation of whether it is within the threshold value, a message to be displayed in the result explanation column 1002 is created.

[0141] In step S914, the CPU 201 of the client terminal 101 reads out an advice message for the actions that future AI developers should take from the database 340 or the file storage 350. Note that a pre-created message may be acquired and displayed from the database 340 or the file storage 350, or an appropriate message may be displayed based on past data using AI.

[0142] In step S915, the CPU 201 of the client terminal 101 displays the inspection result 1000 after noise addition in FIG. 14 in the inspection result display area 1210 in FIG. 18. Note that the inspection result for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.

[0143] FIG. 14 shows an example of the result screen of the noise tolerance test. The result screen consists of inspection result information 1001, result description column 1002, comparison table 1010 of accuracies before and after noise addition, and future action information 1030.

[0144] The inspection result information 1001 displays the result (OK / NG) of the inspection. The determination result of step S912 is displayed.

[0145] The result description column 1002 displays the message generated in step S913 here.

[0146] The comparison table 1010 of accuracies before and after noise addition consists of an accuracy explanation column 1015 and an actual accuracy column 1020. The accuracy explanation column 1015 is a column that explains the result of the noise added to each explanatory variable column or the result without adding noise. The actual accuracy column 1020 is a column that displays the difference from the case without accuracy and noise. In the first row 1021, the accuracy of the evaluation data set before noise addition obtained in step S901 is displayed. Since no noise is added, no value indicating the difference is displayed in the parentheses. From the next row, the difference between the accuracy when noise is added for each explanatory variable and the accuracy before noise addition is displayed respectively. For example, the third row 1022 displays the accuracy obtained in step S910, and the difference output in step S911 is displayed in the parentheses. Note that when it is determined in step S912 that it does not fall within the threshold value, it may be identified and displayed, for example, by displaying the displayed value in red.

[0147] In this way, by adding noise and comparing the accuracies before and after it, the robustness of the AI to be evaluated can be measured. That is, it is possible to grasp in advance how much the accuracy deteriorates when noise enters the input data in the production environment. If it exceeds the threshold value, since the inference result is likely to be affected when noise enters, for example, it can be recognized that it is necessary to take measures such as increasing the learning data so as not to obtain an unintended inference result.

[0148] The future action information 1030 displays the future actions based on the inspection results obtained in step S914.

[0149] Return to the explanation of FIG. 4. In step S406, the CPU 201 of the client terminal 101 determines whether multiple jobs are selected in the pull-down list 1201. If YES, it proceeds to step S407, and if NO, it proceeds to step S408.

[0150] In step S407, the CPU 201 of the client terminal 101 executes a subprogram that performs a process of visualizing the change in the inference performance between models. This subprogram is one of the inspection items for grasping the change in the nature of the inference by the AI, and aims to investigate the effect of the AI improvement activity by comparing the inference results of two AIs and visualizing the changed records.

[0151] FIG. 10 is a flowchart showing the flow of the process in S407 (the process of visualizing the change in the inference performance between models).

[0152] In step S1001, the CPU 201 of the client terminal 101 acquires the evaluation data table A1900 associated with the first AI job number among the plurality of AI job numbers input in step S401.

[0153] In step S1002, the CPU 201 of the client terminal 101 acquires the evaluation data table B1930 associated with the second AI job number among the plurality of AI job numbers input in step S401.

[0154] In this embodiment, the AI related to the first AI job number is described as the pre-improvement AI, and the AI related to the second AI job number is described as the post-improvement AI.

[0155] In step S1003, the CPU 201 of the client terminal 101 divides the records (evaluation data) in the evaluation data table A and the evaluation data table B into records where the correct answer and the inference match (correct answer) and records where they do not match (incorrect answer). This will be specifically described using the example of FIG. 25. In FIG. 25(a), the evaluation data table A1900 is divided into a correct answer table A1910 obtained by extracting records where the correct label column 1515 and the inference label column 1516 match, and an incorrect answer table A1920 obtained by extracting records where they do not match.

[0156] Also, as shown in FIG. 25(b), the same process is performed on the evaluation data table B1930 to create a correct answer table B1940 obtained by extracting records where the correct label column 1515 and the inference label column 1516 of the AI match, and an incorrect answer table B1950 obtained by extracting records where they do not match.

[0157] In step S1004, the CPU 201 of the client terminal 101 acquires the data index of the records that were incorrect in the evaluation data table A but are now correct in the evaluation data table B. That is, the indexes that exist in the index 1921 of the incorrect answer table A1920 and also exist in the index 1941 of the correct answer table B1940 are extracted, and a table 1960 is created as shown in FIG. 25(c).

[0158] In step S1005, the CPU 201 of the client terminal 101 obtains the data index of the record that is correct in the data table A but incorrect in the data table B. That is, an index that exists in the index 1911 of the correct table A1910 and also exists in the index 1951 of the incorrect table B1950 is extracted, and a table 1970 is created as shown in Fig. 25(c).

[0159] In step S1006, the CPU 201 of the client terminal 101 creates a distribution diagram of the whole and data for each explanatory variable. This is a function for investigating the cause of the transition of the inference results between the data table A and the data table B from the perspective of the explanatory variables. For example, it is possible to investigate how the inference results have changed between the AI before improvement and the AI after improvement.

[0160] Fig. 15 is an example of a screen visualizing the explanatory variables of the records whose inferences have changed between the evaluation data table A and the evaluation data table B.

[0161] The visualization screen includes a pull-down list 1105, a graph display area 1110 showing the distribution of the explanatory variables (the explanatory variables selected in 1105. Hereinafter referred to as explanatory variable A.) of the records that were incorrect in the AI related to the first AI job number (the AI before improvement) but correct in the AI related to the second AI job number (the AI after improvement), the vertical axis 1111 of the histogram, the horizontal axis 1112 of the histogram, a histogram 1113 showing the distribution of the explanatory variable A in the evaluation data table A1900, a histogram 1114 showing the distribution of the explanatory variable A of the records that were incorrect in the AI before improvement but correct in the AI after improvement, a legend 1119, and a graph display area 1120 showing the distribution of the explanatory variable A of the records that were correct in the AI before improvement but incorrect in the AI after improvement.

[0162] The pull-down list 1105 is for selecting the explanatory variable to be displayed. The displayed histograms 1113 to 1114 change according to the selected explanatory variable.

[0163] The vertical axis 1111 of the histogram represents the ratio of the number of records with the numerical value (interval) specified on the horizontal axis to the total number of records for the explanatory variable A. For example, if the total number of evaluation data is 100 records and the number of records where the explanatory variable A is between 4.0 and 4.9 is 3, then with the value of 4.0 - 4.9 on the horizontal axis, the vertical axis is graphically displayed as 0.03.

[0164] The horizontal axis 1112 of the histogram represents the numerical value (interval) of the explanatory variable A.

[0165] By superimposing and displaying the distribution of the explanatory variable A in the entire evaluation data and the distribution of the explanatory variable A related to the records that were incorrect with the previous AI but correct with the improved AI, it is possible to recognize in which numerical value / interval of the explanatory variable A the inference accuracy has improved. For example, in a histogram like 1115, if the portion where the correct answers have newly increased is large compared to the overall histogram, it can be considered that the AI has become better at (the inference accuracy has increased).

[0166] Similar to the graph displayed in the area 1120, by superimposing and displaying the distribution of the explanatory variable A in the entire evaluation data and the distribution of the explanatory variable A related to the records that were correct with the previous AI but incorrect with the improved AI, it is possible to recognize in which numerical value / interval of the explanatory variable A the inference accuracy has deteriorated. For example, as in histogram 1125, if the portion where there are newly inconsistent parts has become large compared to the overall histogram, it can be considered that the AI has become less proficient at (the inference accuracy has decreased).

[0167] In this way, by comparing the trends of the entire explanatory variables with the trends of the cases that have changed from incorrect to correct (or from correct to incorrect), it supports the investigation of the causes of the changes in the inference results. For example, if there are many cases that have changed from incorrect to correct, it can be judged that the improvement has been successful. Conversely, if there are cases that have changed from correct to incorrect, it becomes clear which parts have become less proficient, thus providing a trigger for considering the investigation of the causes.

[0168] In step S1007, the CPU 201 of the client terminal 101 creates a confusion matrix of records. This is a function for investigating the tendency of changes in the inference results between the evaluation data table A and the evaluation data table B from the perspective of the target variable.

[0169] FIG. 16 is an example of a screen showing a confusion matrix of the tendency of changes in the inference results. The confusion matrix screen consists of a confusion matrix 1131 that has changed from incorrect to correct and a confusion matrix 1141 that has changed from correct to incorrect.

[0170] The confusion matrix 1131 that has changed from incorrect to correct is a confusion matrix that displays, for each class of the target variable, the number of inferences that were incorrect in the evaluation data table A but correct in the evaluation data table B. It consists of a confusion matrix 1134 that shows how many cases there are where the inference result of class 1132 in the evaluation data table A has changed to the inference result of class 1133 in the evaluation data table B.

[0171] For example, the "25" in cell 1135 of class 2 in the evaluation data table A and class 1 in the evaluation data table B means that there were a total of 25 records that were misclassified as class 2 in the evaluation data table A (the correct label and the inference label did not match), but were correctly classified as class 1 in the evaluation data table B (the correct label and the inference label matched). From this, it can be understood that the measures in the evaluation data table B have improved the errors in class 1.

[0172] The confusion matrix 1141 that has changed from correct to incorrect is a confusion matrix that displays, for each class of the target variable, the number of inferences that were correct in the evaluation data table A but incorrect in the evaluation data table B. It consists of a confusion matrix 1144 that shows how many cases there are where the inference result of class 1142 in the evaluation data table A has changed to the inference result of class 1143 in the evaluation data table B.

[0173] For example, the fact that cell 1145 in class 4 of evaluation data table A and class 1 of evaluation data table B is "50" means that there were a total of 50 records that were correctly classified as class 4 in evaluation data table A (the correct label and the inferred label matched), but were misclassified as class 1 in evaluation data table B (the correct label and the inferred label did not match). From this, it can be grasped that the mistakes in class 4 have worsened with the measures in data table B.

[0174] By visualizing the changes in the inferences by AI in this way, it supports the qualitative evaluation of the measures. Specifically, before and after the improvement of the AI, it becomes clear which classes have become better or worse at making inferences. Therefore, it becomes easier to make judgments regarding the improvement activities of the AI. Also, it becomes easier to notice the changes in nature due to improving the AI. Even if an area that was good at making inferences before the improvement becomes difficult after the improvement, it is not necessarily the case that the improvement activities have failed. If the improvement has made the AI better at making inferences in another area, the nature of the AI has changed, and the developer can consider the usage scenarios according to the changes in nature.

[0175] In step S1008, the CPU 201 of the client terminal 101 creates a data table for checking the details of the records that have changed. This is information for checking the details of the changes visualized in steps S1006 and S1007.

[0176] FIG. 17 is a result screen of a data table for checking the details of the records that have changed. The result screen consists of the number-of-records information 1151 of records that have changed from incorrect to correct, the records 1152 that have changed from incorrect to correct, the number-of-records information 1155 of records that have changed from correct to incorrect, and the records 1156 that have changed from correct to incorrect.

[0177] The number-of-records information 1151 of records that have changed from incorrect to correct is a message representing the number of records in which inferences that did not match the correct answer have become matching.

[0178] The record 1152 that changed from incorrect to correct is table data composed of records having the index extracted in step S1004.

[0179] The record count information 1155 that changed from correct to incorrect is a message representing the number of records for which the inference that was correct has become incorrect.

[0180] The record 1156 that changed from correct to incorrect is table data composed of records having the index extracted in S1005.

[0181] In step S1009, the CPU 201 of the client terminal 101 displays the results of FIGS. 15, 16, and 17 in the inspection result display area 1210 of FIG. 18. Note that the comparison evaluation results inspected this time may be written to the database 340 or the file storage 350 so that they can be quickly displayed next time.

[0182] Returning to the description of FIG. 4, in step S408, the CPU 201 of the client terminal 101 creates a summary table of the inspection results and displays it in the inspection result display area 1210 of FIG. 18.

[0183] FIG. 20 is a screen of a summary table of inspection results. The screen of the summary table consists of a summary table 1301.

[0184] The summary table 1301 displays a list of OK / NG determinations of inspection results performed in the previous processing. It consists of an inspection number column 1310, an inspection item name column 1311, an inspection result column 1312, and a simple result explanation column 1314. In the inspection result column 1312, for example, in the case of NG, it may be highlighted, and it may be highlighted according to the OK / NG result. Note that the summary table for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.

[0185] Next, a function for searching for jobs that have recorded inspection results will be described. The flowchart in FIG. 5 is a flowchart of the function for searching for jobs based on the OK / NG results of inspection items. For example, when receiving the pressing of a button for using the search function from the menu screen, the screen in FIG. 21 is opened, and the processing of this flowchart is executed.

[0186] In step S501, the CPU 201 of the client terminal 101 receives the input of the check box 1410 for the inspection number in FIG. 21. For each inspection item described in this embodiment, it is possible to search for AI jobs whose inspection results are OK. Although only numbers 1 to 3 are shown for the check box 1410 for the inspection number, it is displayed as many as the number of inspection items.

[0187] In step S502, the CPU 201 of the client terminal 101 determines whether the "Search Jobs" button 1415 has been pressed. If YES, it proceeds to step S503. If NO, it returns to step S502.

[0188] In step S503, the CPU 201 of the client terminal 101 searches for jobs whose inspection results of the inspection items checked in the check box in step S501 are OK. When multiple inspection items are checked, an AND search is performed. When nothing is checked, all jobs are output as search results.

[0189] In step S504, the CPU 201 of the client terminal 101 displays the job list 1420 of the search results on the screen. The job list 1420 is a data table of the search results of AI jobs. It consists of a check box 1421 for detailed viewing of inspection results, an AI job number column 1422, and an inspection number column 1423.

[0190] In step S505, the CPU 201 of the client terminal 101 receives the input of the check box 1421 for detailed viewing of AI jobs.

[0191] In step S506, the CPU 201 of the client terminal 101 determines whether the "View Details of Inspection Results of Selected Job" button 1430 has been pressed. If YES, it proceeds to step S507; if NO, it returns to step S506.

[0192] In step S507, the CPU 201 of the client terminal 101 transitions to the screen of FIG. 18 and displays the details of the inspection results of the AI jobs with checkboxes filled in.

[0193] By searching for the inspection result summary, the characteristics of each AI can be grasped. In the example of FIG. 21, it can be seen that the AI of AI job number 3 has less bias in inference to a specific label than inspection number 2 and is a relatively robust AI compared to the result of inspection number 3. On the other hand, there is a section within a certain explanatory variable that the AI is not good at according to the result of inspection number 1. Since it is rare for all items to be OK, the areas where each AI is good and the areas where it is not good can be confirmed. Also, when wanting to select which AI to use depending on the application scenario, by searching for the inspection items that are OK, the user can easily find the AI they need.

[0194] As described above, according to this embodiment, a mechanism capable of grasping the reliability of AI can be provided.

[0195] Note that in this embodiment, an AI related to data analysis is described as an example, but it is not limited to this. For example, it can be applied to the evaluation of various AIs such as those used in image recognition, appearance inspection, demand prediction, etc. Also, not limited to the evaluation of AI, as long as it is created based on a dataset, it may be used for the evaluation of that dataset.

[0196] The present invention can be implemented, for example, as an embodiment such as a system, device, method, program, or recording medium. Specifically, it may be applied to a system composed of a plurality of devices, or also to a device consisting of a single device.

[0197] Note that each of the various controls described as being performed by the CPU 201 may be performed by one piece of hardware, or the entire apparatus may be controlled by a plurality of pieces of hardware (for example, a plurality of processors or circuits) sharing the processing.

[0198] Also, although the present invention has been described in detail based on its preferred embodiments, the present invention is not limited to these specific embodiments, and various forms within the scope not departing from the gist of this invention are also included in the present invention. Furthermore, each of the above-described embodiments merely shows one embodiment of the present invention, and it is also possible to appropriately combine the embodiments.

[0199] Also, in the above-described embodiments, the case where the present invention is applied to a PC has been described as an example, but this is not limited to this example, and it is applicable to any apparatus capable of displaying the evaluation result of an inference model. For example, it is applicable to a PDA, a mobile phone terminal (smartphone), a tablet terminal, and the like.

[0200] (Other Embodiments) The present invention is also realized by executing the following processing. That is, software (program) that realizes the functions of the above-described embodiments is supplied to a system or apparatus via a network or various storage media, and a computer (or CPU, MPU, etc.) of the system or apparatus reads and executes the program code. In this case, the program and the storage medium storing the program constitute the present invention.

Explanation of Reference Numerals

[0201] 100 Network 101 Client Terminal 102 Server

Claims

1. In a first evaluation dataset and a second evaluation dataset used for evaluating a learned model, an extraction means for extracting data with different inference results by the learned model; A display control means for controlling to comparatively display the distribution of the values of the first variable in the first evaluation dataset and the distribution of the values of the first variable in the data with different inference results; An information processing system characterized by comprising the above.

2. The information processing system according to claim 1, wherein the extraction means extracts data in which the inference result by the learned model for the first evaluation dataset is incorrect and the inference result by the learned model for the second evaluation dataset is correct.

3. The information processing system according to claim 1, wherein the extraction means extracts data in which the inference result by the learned model for the first evaluation dataset is correct and the inference result by the learned model for the second evaluation dataset is incorrect.

4. In a first evaluation dataset and a second evaluation dataset used for evaluating a learned model, an extraction means for extracting data with different inference results by the learned model; A display control means for controlling to comparatively display the inference result by the learned model for the first evaluation dataset and the inference result by the learned model for the second evaluation dataset; An information processing system characterized by comprising the above.

5. The information processing system according to claim 4, wherein the display control means controls to comparatively display the number of cases where the inference result by the learned model for the first evaluation dataset is incorrect and the number of cases where the inference result by the learned model for the second evaluation dataset is correct.

6. The information processing system according to claim 4, wherein the display control means controls to comparatively display the number of cases where the inference result by the learned model for the first evaluation dataset is correct and the number of cases where the inference result by the learned model for the second evaluation dataset is incorrect.

7. In a first evaluation dataset and a second evaluation dataset used for evaluating a learned model, an extraction step of extracting data for which inference results by the learned model are different; A display control step of controlling to comparatively display the distribution of values of a first variable in the first evaluation dataset and the distribution of values of the first variable in the data for which the inference results are different; A control method for an information processing system, comprising the above.

8. In a first evaluation dataset and a second evaluation dataset used for evaluating a learned model, an extraction step of extracting data for which inference results by the learned model are different; A display control step of controlling to comparatively display the inference result by the learned model for the first evaluation dataset and the inference result by the learned model for the second evaluation dataset; A control method for an information processing system, comprising the above.

9. A program for causing at least one computer to function as the information processing system according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Predicted situation visualization device, predicted situation visualization method, and predicted situation visualization program

    JP7095744B2