Information processing system, control method therefor, and program
By dividing datasets into groups based on variable values and assessing inference results, the proposed mechanism addresses the reliability issues of AI systems, ensuring effective evaluation of AI performance in real-world scenarios.
Patent Information
- Application Number
- JP2023210957
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-26
AI Technical Summary
Existing AI systems face reliability issues as high accuracy on evaluation datasets does not guarantee performance in real-world field operations, leading to decreased AI reliability.
A mechanism that divides datasets into groups based on variable values, identifies and displays groups where inference results fall below a threshold, enabling the assessment of AI reliability.
This approach allows for the effective evaluation of AI reliability, providing insights into performance gaps between evaluation datasets and real-world operations.
Smart Images

Figure 2025095145000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system, a control method thereof, and a program.
Background Art
[0002] In recent years, AI generated by deep learning or machine learning using learning data has been used in various fields. When determining whether AI is practical, it is common to prepare an evaluation dataset and measure the accuracy thereof. However, the quality of AI may not be reliable only based on the performance on the evaluation dataset.
[0003] Patent Document 1 discloses a technique for visualizing a prediction situation in order to grasp the factors that cause an error between a prediction and an actual result.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Disclosure of the Invention
Problems to be Solved by the Invention
[0005] Even if the accuracy with respect to the evaluation dataset is high, high accuracy may not be obtained when actually operating in the field. In such a case, the reliability of AI decreases, and thus a mechanism for developing reliable AI is necessary.
[0006] Therefore, an object of the present invention is to provide a mechanism capable of grasping the reliability of AI.
Means for Solving the Problems
[0007] In order to solve the above problems, the present invention Control means for controlling to divide into a plurality of groups based on the value of the first variable for the dataset used for evaluating the learned model, display control means for controlling to identify and display a group among the plurality of groups in which the inference result by the learned model is equal to or less than a threshold value; characterized by comprising.
Effect of the Invention
[0008] According to the present invention, it becomes possible to provide a mechanism capable of grasping the reliability of AI.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Embodiments for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0011] FIG. 1 is a diagram showing an example of the system configuration of a visualization system for the nature of inferences of an AI (a learned model generated by deep learning / machine learning) in an embodiment of the present invention.
[0012] The client terminal 101 is configured to be connected via the network 100. The client terminal is, for example, a personal computer (hereinafter referred to as a PC). The client terminal 101 performs the main processing for visualizing the nature of inferences by the AI. The network 100 can take forms such as a wired LAN, a wireless LAN, or a USB according to the physical interface of the client terminal 101. A server 102 may be placed on the network 100. The client terminal 101 may read data from the server 102.
[0013] FIG. 2 is a block diagram showing an example of the hardware configuration of the client terminal 101 in an embodiment of the present invention.
[0014] As shown in FIG. 2, in the client terminal, a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an input controller 205, a video controller 206, a memory controller 207, and a communication I / F controller 208 are connected via a system bus 204.
[0015] The CPU 201 comprehensively controls each device and controller connected to the system bus 204.
[0016] The ROM 202 or the external memory 211 holds the BIOS (Basic Input / Output System), the OS (Operating System), which are control programs executed by the CPU 201, a computer-readable and executable program for realizing this information processing method, and various necessary data (including data tables).
[0017] The RAM 203 functions as the main memory, work area, etc. of the CPU 201. When executing a process, the CPU 201 loads a program, etc. necessary for the execution from the ROM 202 or the external memory 211 into the RAM 203, and realizes various operations by executing the loaded program.
[0018] The input controller 205 controls the input from input devices such as a keyboard 209 and a pointing device such as a mouse (not shown). When the input device is a touch panel, it is assumed that the user can give various instructions by pressing (touching with a finger, etc.) in accordance with the icons, cursors, buttons, etc. displayed on the touch panel.
[0019] Also, the touch panel may be a touch panel capable of detecting the positions touched by a plurality of fingers, such as a multi-touch screen.
[0020] The video controller 206 controls the display to an external output device such as a display 210. The display includes the display of a notebook personal computer integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. Also, for a device capable of receiving the above-described touch operation, an input device is also provided.
[0021] The video controller 206 can control a video memory (VRAM) for display control, and can use a part of the RAM 203 as a video memory area, or can separately provide a dedicated video memory.
[0022] The memory controller 207 controls access to the external memory 211. As the external memory, an external storage device (hard disk) that stores a boot program, various applications, font data, user files, edited files, and various data, a flexible disk (FD), or a compact flash (registered trademark) memory connected via an adapter to a PCMCIA card slot can be used.
[0023] The communication I / F controller 208 is connected to and communicates with an external device via a network, and executes communication control processing on the network. For example, communication using TCP / IP, a telephone line such as ISDN, and communication using a 3G line of a mobile phone are possible.
[0024] In addition, the CPU 201 enables display on the display 210 by executing an outline font expansion (rasterization) process on, for example, a display information area in the RAM 203. Further, the CPU 201 enables user instructions using a mouse cursor (not shown) on the display 210.
[0025] Next, with reference to FIG. 3, an example of the functional configuration of the client terminal 101 of the present invention will be described.
[0026] The input reception unit receives an input instruction from the user through the input controller 205.
[0027] The data analysis unit acquires data stored in the database 340 and the file storage 350 according to the received instruction, and processes it for inspection by the AI quality inspection unit.
[0028] The AI quality inspection department performs AI quality inspection using the analyzed data. It executes the process of FIG. 4 to inspect whether the AI to be inspected meets each inspection item.
[0029] The result output unit displays the inspected results on the display of the client terminal 101. For example, it displays inspection result screens such as those from FIG. 11 to FIG. 21 on the display. Hereinafter, the display destination in this embodiment is the display of the client terminal 101.
[0030] In the database 340 and the file storage 350, an evaluation data table 1500, a learning data table 1550, etc. as shown in FIG. 23 are stored. In this embodiment, the CPU 201 acquires data tables 1550, etc. from the database 340 and the file storage 350 and uses them for AI quality evaluation.
[0031] Next, with reference to the flowcharts of FIGS. 4 to 11, the processing executed by the client terminal 101 in the embodiment of the present invention will be described.
[0032] When evaluating an AI, multiple viewpoints are considered. For example, there are "explanability", "bias", "sensitivity", "error analysis", and "backward compatibility".
[0033] Explanability is a viewpoint for evaluating whether the inferences made by the AI are reliable. If the inference results by the AI are not reliable, it becomes difficult to utilize that AI, so it is evaluated from this viewpoint.
[0034] Bias is a viewpoint for evaluating whether there is no bias in the dataset used for inferences, learning, and evaluation by the AI. By evaluating whether there is bias in the dataset, it is evaluated whether an AI that is poor at inferring a specific label has not been constructed.
[0035] Robustness is a perspective for evaluating the stability of AI learning. For example, robustness is evaluated by verifying whether an AI with good accuracy has not been built only for a specific evaluation dataset.
[0036] Error analysis is a perspective for evaluating whether it is easy to analyze errors. When the performance of an AI is not good, it is evaluated whether the points for improvement can be immediately understood.
[0037] Backward compatibility is a perspective for evaluating changes in the nature of inference between the pre-improvement model and the post-improvement model. When the performance of an AI is improved, there may be a case where data that could be correctly inferred before the improvement can no longer be inferred, so it is evaluated from this perspective.
[0038] The flowchart of FIG. 4 is a process in which the CPU 201 of the client terminal 101 reads and executes a predetermined control program, and is a flowchart showing a process for visualizing the nature of inference by an AI. Note that the flowchart of FIG. 4 is executed as an internal process of the operation screen 1200 of FIG. 18.
[0039] In step S401, the CPU 201 of the client terminal 101 receives an input of an AI job number from the user.
[0040] The AI job number is information (identifier) that uniquely identifies the generated AI. Also, using the AI job number as a key, data related to the AI identified by the AI job number can be obtained from the database 340 or the file storage 350. Data related to the AI is, for example, the evaluation data table 1500 of the inference results for the evaluation dataset of the AI in FIG. 22. The evaluation data table 1500 consists of items such as a data row number 1510, an explanatory variable column 1511, a correct label column 1515, and an AI inference label (AI inference result) column 1516.
[0041] The line number 1510 of the data is an item for uniquely identifying the data row, and the number is registered.
[0042] The explanatory variable column 1511 is an item of the explanatory variable used in the inference by AI, and a floating-point (float) type numerical value, an integer (int type) numerical value, or a character string representing a category is registered.
[0043] In the correct label column 1515, there is an item of the correct target variable, and a character string representing a class label or an integer value representing a class index is registered.
[0044] In the inference label column 1516 by AI, there is an item of the target variable inferred by AI, and a character string representing a class label or an integer value representing a class index is registered.
[0045] In addition to the above, there is also data related to AI, such as the learned AI, accuracy information for the evaluation dataset, and the learning data table 1550 which is the dataset used for AI learning. The learning dataset is in the same format as the evaluation dataset, but since it is the dataset used for learning, it does not have the inference label column 1516 by AI.
[0046] Also, an example of receiving the input of the AI job number from the user is shown in FIG. 18. FIG. 18 is an AI inspection selection screen for receiving the input of the AI job number from the user and performing the inspection of the AI. The AI inspection selection screen consists of a pull-down list 1201, the selected AI job number 1202, and an inspection result display area 1210.
[0047] The pull-down list 1201 is for receiving the selection of the AI job number from the user. In the pull-down list 1201, multiple AI job numbers can be selected as shown in FIG. 19. In that case, the inspection result display area 1210 in FIG. 18 is divided, and the inspection results can be displayed in parallel like the inspection result display area A 1210 and the inspection result display area B 1211 in FIG. 19.
[0048] The selected AI job number 1202 displays the AI job number input by the user.
[0049] The inspection result display area 1210 is an area for displaying the inspection results of the AI, and it displays the inspection results of the inference characteristics by the AI after step S402.
[0050] In addition, in FIG. 18, the AI job number is displayed in a pull-down list for accepting selection, but any method may be used as long as it is a form in which the user can specify the AI to be inspected, such as a form for accepting input of the AI job number from the user.
[0051] In step S402, the CPU 201 of the client terminal 101 executes a subprogram that visualizes the explanatory variables having a low-accuracy interval for the AI specified by the AI job number accepted in step S401. This subprogram is one of the inspection items for grasping the nature of the inference by the AI, and by calculating the accuracy for each interval of the explanatory variables, it visualizes the low-accuracy intervals and aims to provide information for analyzing the nature of the AI to the AI developer.
[0052] FIG. 6 is a flowchart showing the flow of the process in S402 (the process of visualizing the explanatory variables having a low-accuracy interval).
[0053] In step S601, the CPU 201 of the client terminal 101 acquires the evaluation data table 1500 associated with the AI job number input in step S401 and the accuracy information of the entire dataset. The acquisition sources are the database 340 and the file storage 350.
[0054] In step S602, the CPU 201 of the client terminal 101 acquires the accuracy reduction threshold value used for determining OK (qualified) / NG (unqualified) of the inspection. The acquisition sources are the database 340 and the file storage 350.
[0055] In step S603, the CPU 201 of the client terminal 101 acquires the number of intervals when dividing the interval in the subsequent step S606. The acquisition source is the database 340 or the file storage 350.
[0056] In step S604, the CPU 201 of the client terminal 101 acquires each column of the explanatory variable column 1511 in the evaluation data table 1500 one by one, and starts a loop to combine the correct label column 1515 and the inference label column 1516 by AI. An example after combination is shown in Fig. 23(a). The combined table 1600 is a table obtained by combining the explanatory variable var1 column 1512, the correct label column 1515, and the inference label column 1516 by AI. This is repeated for all columns of explanatory variables, and when the processing for all explanatory variables is completed, the loop is terminated.
[0057] In step S605, the CPU 201 of the client terminal 101 determines whether the explanatory variable for one column acquired in step S604 is a floating-point type explanatory variable. If YES, it proceeds to step S606; if NO, it proceeds to S607.
[0058] In step S606, the CPU 201 of the client terminal 101 divides the floating-point type explanatory variable for one column acquired in step S604 into intervals by the number of intervals acquired in step S603 and assigns them to the intervals.
[0059] The method of dividing the interval is to acquire the minimum value and the maximum value in one column of the explanatory variable, and divide it into intervals with equal intervals. For example, when the minimum value of one column of the explanatory variable is 0, the maximum value is 10, and the number of intervals is 5, the following five intervals are generated. That is, the interval from 0 to 2, the interval from 2 to 4, the interval from 4 to 6, the interval from 6 to 8, and the interval from 8 to 10. Then, each piece of the floating-point type explanatory variable for one column is assigned to the divided interval.
[0060] An example of interval division is shown in Fig. 23(a). By the method as described above, the interval is divided into intervals 1610 from 4 to 6, interval 1620 from 6 to 8, interval 1630 from 8 to 10, etc. Among the combined tables 1600, since the value of the explanatory variable var1 column 1512 of the record 1605 is 4.8, it is assigned to the interval 1610 from 4 to 6.
[0061] In step S607, the CPU 201 of the client terminal 101 calculates the accuracy based on the records for each interval. The accuracy is indicated, for example, by accuracy which is a ratio when the number of data is the denominator and the number of times the correct answer and the inference by AI match is the numerator.
[0062] Note that explanatory variables other than the float type are regarded as categorical variables, and the accuracy is calculated for each category that the categorical variable takes. An example in the case of explanatory variables other than the float type is shown in Fig. 23(b). When the value taken by the explanatory variable var3 column 1517 is a string such as red or blue, the records are grouped by string like record 1640 with value red and record 1650 with value blue, and the accuracy of each is calculated. Also, in the case where the value taken by the explanatory variable var4 column 1518 is an integer such as 1, 2.., it is the same. The records are grouped by value like record 1660 with value 1 and record 1670 with value 2, and the accuracy of each is calculated.
[0063] In step S608, the CPU 201 of the client terminal 101 sets a threshold based on the accuracy information of the entire dataset acquired in step S601 and the decrease threshold acquired in step S602. Based on the set threshold, the accuracy for each interval or each category obtained in step S607 is evaluated.
[0064] For example, if the accuracy information of the entire dataset obtained in step S601 is 60% and the accuracy degradation threshold is 5 percentage points, the threshold will be 55%. It is determined whether it is below this threshold. In the example of FIG. 23(c), the inspection determination table 1680 for each section is shown. For example, in the section from 0 to 2, since the accuracy is 70%, it exceeds the threshold and the determination is OK, while in the section from 6 to 8, it is 20% and is below the threshold and is determined as NG. Based on the inspection determination table 1680, the bar graph 713 in FIG. 11 is created.
[0065] In step S609, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, information regarding the accuracy of the entire dataset and the degradation threshold is displayed in the result explanation column 702, and a message is created for the explanation of the section of the explanatory variable that did not meet the threshold.
[0066] In step S610, the CPU 201 of the client terminal 101 reads out an advice message on the actions to be taken by future AI developers from the database 340 or the file storage 350 according to the inspection result. Note that it may be possible to obtain and display a pre-created message from the database 340 or the file storage 350, or it may be possible to display an appropriate message based on past data by utilizing AI.
[0067] In step S611, the CPU 201 of the client terminal 101 displays the accuracy inspection result 700 for each section in FIG. 11 in the inspection result display area 1210 in FIG. 18. Note that the inspection result for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.
[0068] Fig. 11 shows an example of a result screen visualizing explanatory variables with low-precision intervals. The result screen consists of inspection result information 701, result description column 702, pull-down list 703 for selecting explanatory variables, graph 710 of precision for each interval, vertical axis 711 of the graph, horizontal axis 712 of the graph, bar graph 713, breakdown of precision 714, bar graph 715 that did not meet the standard, and future action information 730.
[0069] The inspection result information 701 displays the result of the inspection (OK / NG). If there is no interval lower than the threshold value in step S608, it is displayed as OK; otherwise, it is displayed as NG.
[0070] The result description column 702 provides an explanation of the result. The message generated in step S609 is displayed here.
[0071] The pull-down list 703 for selecting explanatory variables is a pull-down list for selecting explanatory variables.
[0072] The graph 710 of precision for each interval is a graph of precision for each interval in the explanatory variable.
[0073] The vertical axis 711 of the graph shows the numerical value of precision on the vertical axis.
[0074] The horizontal axis 712 of the graph shows intervals or category variable names on the horizontal axis.
[0075] The bar graph 713 shows the precision in the interval.
[0076] The breakdown of precision 714 shows the precision, the number of learning cases in the interval, and the number of correctly answered data. For example, it shows that out of 10 learning cases, 6 were answered correctly. In addition, the breakdown of 1 correct answer out of 5 is also displayed, allowing one to understand which intervals have been learned more and which have been learned less.
[0077] The bar graph 715 of the section that did not meet the standard is highlighted conspicuously. Also, an exclamation mark 716 appears next to the accuracy information and is highlighted. As a result, it becomes easier to recognize which section has low accuracy, enabling verification of whether there are areas where AI struggles or whether the number of datasets is sufficient. For example, it can be determined that it is necessary to increase the amount of learning data for sections with low accuracy. Conversely, if the learning data is more abundant in a particular section compared to others, it can be recognized that overfitting has occurred in that section, allowing confirmation of whether the learning amount is evenly distributed among the explanatory variables.
[0078] The future action information 730 displays the future actions based on the inspection results obtained in step S610.
[0079] Returning to the description of FIG. 4. In step S403, the CPU 201 of the client terminal 101 executes a subprogram that performs a process of calculating and visualizing the degree of coincidence between the correct distribution and the inference distribution. This subprogram is one of the inspection items for understanding the nature of the inference by AI, and its purpose is to inspect and visualize whether there are any abnormalities in the tendency of the inference by AI by checking whether there is a deviation between the histogram of the labels inferred by AI and the histogram of the correct labels.
[0080] FIG. 7 is a flowchart showing the flow of the process in S403 (the process of calculating and visualizing the degree of coincidence between the correct distribution and the inference distribution).
[0081] In step S701, the CPU 201 of the client terminal 101 acquires the evaluation data table 1500 associated with the AI job number input in step S401 from the database 340 or the file storage 350.
[0082] In step S702, the CPU 201 of the client terminal 101 acquires the similarity threshold used for determining OK (qualified) / NG (unqualified) in the inspection. The acquisition source is the database 340 or the file storage 350.
[0083] In step S703, the CPU 201 of the client terminal 101 processes the evaluation data table 1500 obtained in step S701 according to the current process. Specifically, a table is created by extracting the correct label column 1515 and the inference label column 1516 by AI from the evaluation data table 1500.
[0084] In step S704, the CPU 201 of the client terminal 101 aggregates the histogram of the correct label column 1515. That is, for the correct label column 1515, the number of each class is aggregated.
[0085] In step S705, the CPU 201 of the client terminal 101 aggregates the histogram of the inference label column 1516. That is, for the inference label column 1516, the number of each class is aggregated.
[0086] In step S706, the CPU 201 of the client terminal 101 regards the histogram of the correct label obtained in step S704 and the histogram of the inference label obtained in step S705 as probability distributions, and calculates the degree of coincidence between the probability distributions. For example, JS divergence can be used to calculate the degree of coincidence. Note that the smaller the value of the JS divergence, the more similar the probability distributions are.
[0087] In step S707, the CPU 201 of the client terminal 101 compares the degree of coincidence obtained in step S706 with the threshold value obtained in S702, and outputs the inspection result. When the calculation method is JS divergence, if the degree of coincidence is less than or equal to the threshold value, it is OK; otherwise, it is NG.
[0088] By calculating the degree of agreement between the correct distribution and the inference distribution, it is possible to check whether there are any abnormalities in the inference tendency. That is, if there is a problem with the learning method, there is a possibility that the inference tendency will be biased. Using the distribution shown in FIG. 12 for explanation, the correct distribution is a distribution in the order of class 1, 3, 2, 4 from the one with the largest number of data, while the inference distribution is in the order of class 2, class 3, class 1, class 4 from the one with the largest number of data. Since there is a difference in the distribution of data in this way, when calculating the distance between the correct distribution and the inference distribution using JS divergence, it becomes "0.334". This value exceeds the threshold (0.3). This indicates that the AI does not have the correct inference tendency with respect to the correct distribution, and it is necessary to adjust the learning data.
[0089] In step S708, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, regarding the degree of agreement obtained in step S706, the information of the threshold value obtained in S702, and the explanation of the result of comparing these information, a message to be displayed in the result explanation column 802 is created.
[0090] In step S709, the CPU 201 of the client terminal 101 reads out an advice message on the actions that future AI developers should take from the database 340 or the file storage 350. Note that a pre-created message may be obtained from the database 340 or the file storage 350 and displayed, or an appropriate message may be displayed based on past data by utilizing the AI.
[0091] In step S710, the CPU 201 of the client terminal 101 displays the inspection result 800 of the correct and inference distributions in FIG. 12 in the inspection result display area 1210 in FIG. 18. Note that the inspection result for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.
[0092] Fig. 12 shows an example of a result screen visualizing the degree of coincidence between the correct distribution and the inference distribution. The result screen consists of inspection result information 801, result explanation column 802, histogram plot 810, vertical axis 811 of the histogram, horizontal axis 812 of the histogram, bar graph 813 of the correct label, bar graph 814 of the inference label, legend 820 of the histogram, and future action information 830.
[0093] The inspection result information 801 displays the result (OK / NG) of the inspection. The result determined in step S707 is displayed.
[0094] The result explanation column 802 displays the message generated in step S708 here.
[0095] The histogram plot 810 is a plot with the histogram of the correct label and the histogram of the inference label arranged side by side.
[0096] The vertical axis 811 of the histogram represents the number of labels in the evaluation dataset.
[0097] The horizontal axis 812 of the histogram represents the types of labels of the target variable.
[0098] The bar graph 813 of the correct label is a bar graph of the number of correct labels. That is, it is a bar graph showing the number aggregated for each class in step S704.
[0099] The bar graph 814 of the inference label is a bar graph of the number of inference labels. That is, it is a bar graph showing the number aggregated for each class in step S705.
[0100] The legend 820 is the legend of the histogram.
[0101] The future action information 830 displays the future actions based on the inspection results obtained in step S709.
[0102] Return to the description of FIG. 4. In step S404, the CPU 201 of the client terminal 101 executes a subprogram for performing a process of visualizing outliers of the explanatory variables of the data set. This subprogram is one of the inspection items for grasping the nature of the data set used for AI learning and evaluation. By examining the distribution of possible values of each explanatory variable of the data set, it is possible to check whether there are extremely large (or small) values in the explanatory variables and inspect the quality of the data set. If there are outliers, it may lead to a decrease in the inference accuracy and learning speed of the AI. Therefore, it is necessary to detect the presence or absence of outliers and exclude them if necessary.
[0103] FIG. 8 is a flowchart showing the flow of the process in S404 (the process of visualizing outliers of the explanatory variables of the data set).
[0104] In step S801, the CPU 201 of the client terminal 101 acquires the evaluation data table 1500 associated with the AI job number input in step S401 from the database 340 or the file storage 350.
[0105] In step S802, the CPU 201 of the client terminal 101 acquires the learning data table 1550 associated with the AI job number input in step S401 from the database 340 or the file storage 350.
[0106] In step S803, for each of the evaluation data table 1500 and the learning data table 1550, the CPU 201 of the client terminal 101 acquires the explanatory variable columns of the float type among the explanatory variable columns 1511.
[0107] In step S804, the CPU 201 of the client terminal 101 acquires the float-type explanatory variable columns obtained in step S803 from the explanatory variable column group 1511 one by one, and starts the loop of the subsequent processing. The subsequent processing repeats for the number of float-type explanatory variable columns obtained in step S803, and when the processing for all explanatory variables is completed, the loop is terminated. Also, the subsequent processing will be described by taking the evaluation data table 1500 as an example, but the same processing is executed for the learning data table 1550 as well.
[0108] In step S805, the CPU 201 of the client terminal 101 calculates the quartiles for the float-type explanatory variable columns obtained in step S804.
[0109] In step S806, the CPU 201 of the client terminal 101 creates a box plot based on the quartiles of the explanatory variables obtained in step S805.
[0110] In step S807, the CPU 201 of the client terminal 101 determines whether there are outliers in the explanatory variables. For example, it determines whether there is a value that is more than 1.5 times the vertical width of the box away from the end of the box, and determines a value that is more than 1.5 times away as an outlier. Although the value "more than 1.5 times the vertical width of the box away from the end of the box" is used as the outlier, it is assumed that the criterion used to determine whether it is an outlier can be set arbitrarily.
[0111] In step S808, the CPU 201 of the client terminal 101 determines the inspection result. Here, if an outlier is found in step S807, it is NG; otherwise, it is OK.
[0112] In step S809, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, it creates a message to be displayed in the result explanation column 902 regarding whether an explanatory variable containing an outlier is found in the dataset.
[0113] In step S810, the CPU 201 of the client terminal 101 reads advice messages on actions that future AI developers should take from the database 340 and the file storage 350. Note that it is also possible to obtain and display pre-created messages from the database 340 and the file storage 350, or to display appropriate messages based on past data by utilizing AI.
[0114] In step S811, the CPU 201 of the client terminal 101 displays the outlier inspection result 900 in Fig. 13 in the inspection result display area 1210 in Fig. 18. Note that the inspection result for the job number inspected this time may be written to the database 340 and the file storage 350 so that it can be quickly displayed next time.
[0115] Fig. 13 shows an example of a result screen visualizing the outliers of the explanatory variables of the dataset. The result screen consists of inspection result information 901, result explanation column 902, box-and-whisker plot area 910, vertical axis 911 of the box-and-whisker plot, horizontal axis 912 of the box-and-whisker plot, outlier 914, and future action information 930.
[0116] The inspection result information 901 displays the result (OK / NG) of the inspection. The result determined in step S808 is displayed.
[0117] The result explanation column 902 displays the message generated in step S809 here.
[0118] The box-and-whisker plot area 910 displays the box-and-whisker plot of the explanatory variables created in step S906. In Fig. 13, only the box-and-whisker plots of var1 and var2 are displayed, but it is also possible to display all the evaluated explanatory variable columns, or to display the explanatory variable columns selected by a pull-down list or the like.
[0119] The vertical axis 911 of the box-and-whisker plot represents the value range of one explanatory variable.
[0120] The horizontal axis 912 of the box-and-whisker plot indicates the labels of the training dataset and the evaluation dataset. For each explanatory variable column as shown in FIG. 13, the training dataset and the evaluation dataset are displayed side by side.
[0121] If there are outliers in the training dataset, it may lead to a decrease in the inference accuracy of the AI, etc., so it is necessary to display the presence or absence of outliers. Similarly, if there are outliers in the evaluation dataset, it may also lead to a decrease in the inference accuracy of the AI, etc. In this case, if the accuracy of the AI for the evaluation dataset is low, it is possible that the reason for the accuracy decrease is that the evaluation dataset contains many outliers rather than the training dataset. Therefore, it is necessary to investigate not only the training dataset but also how many outliers exist in the evaluation dataset.
[0122] Also, by displaying the training dataset and the evaluation dataset side by side, for example, there is no need to switch the screen etc. to compare, and it becomes easy to check whether the explanatory variable to be evaluated contains outliers.
[0123] The outlier 914 represents the outlier found in step S808.
[0124] The future action information 930 displays the future actions based on the inspection results obtained in step S810.
[0125] Although an example in which outliers can be visualized by a box-and-whisker plot has been described, any method that can visualize outliers such as standard deviation or cluster analysis is acceptable.
[0126] Return to the description of FIG. 4. In step S405, the CPU 201 of the client terminal 101 executes a subprogram for performing a noise tolerance test process. This subprogram is one of the inspection items for grasping the nature of AI inference, and is intended to inspect how robust the AI is when a minute value is given to each explanatory variable. That is, when there is a slight change in the explanatory variable, if the inference result is greatly affected, the AI becomes difficult to use, so the robustness of the AI is evaluated.
[0127] FIG. 9 is a flowchart showing the flow of the process in S405 (the process of performing a noise tolerance test).
[0128] In step S901, the CPU 201 of the client terminal 101 acquires the AI, the evaluation data table 1500, and the accuracy in the evaluation data set associated with the AI job number input in step S401 from the database 340 or the file storage 350. In step S902, the CPU 201 of the client terminal 101 acquires the parameter of the noise to be applied to the evaluation data set (for example, the magnification factor multiplied by the variance of the normal distribution) from the database 340 or the file storage 350.
[0129] In step S903, the CPU 201 of the client terminal 101 acquires the threshold value that is the allowable amount of accuracy degradation used for determining OK (qualified) / NG (unqualified) of the inspection. The acquisition source is the database 340 or the file storage 350.
[0130] In step S904, the CPU 201 of the client terminal 101 processes the evaluation data table 1500 acquired in step S901 according to the current process. Specifically, the row number 1510 of the data, the explanatory variable column 1511, and the correct label column 1515 are extracted from the evaluation data table 1500, and an extraction table 1800 as shown in FIG. 24 is created.
[0131] In step S905, the CPU 201 of the client terminal 101 acquires the explanatory variable columns of the float type among the explanatory variable columns 1511.
[0132] In step S906, the CPU 201 of the client terminal 101 acquires the columns of the float-type explanatory variables acquired in step S905 one by one from the explanatory variable columns 1511, and starts the loop of the subsequent processing. The subsequent processing repeats for the number of columns of the float-type explanatory variables acquired in step S905, and when the processing for all the explanatory variables is completed, the loop is terminated.
[0133] In step S907, the CPU 201 of the client terminal 101 calculates the variance of the data in the explanatory variable var1 column 1512.
[0134] In step S908, the CPU 201 of the client terminal 101 generates noise 1810 for the number of records in the explanatory variable var1 column 1512.
[0135] The noise is a value sampled from a normal distribution. The mean, which is a parameter of the normal distribution, is 0, and the variance is the value calculated in step S907 multiplied by the magnification factor acquired in step S902. The reason for multiplying by the magnification factor is that it is necessary to change the amount of noise according to the problem handled by the AI, and a parameter for adjusting the magnitude of the noise is required. In this way, noise 1810 is generated.
[0136] In step S909, the CPU 201 of the client terminal 101 copies the extraction table 1800 created in step S904, and adds the noise 1810 to each record in the explanatory variable var1 column 1512. The table after adding noise to the extraction table 1800 is the table 1820 after noise addition, and the explanatory variable var1 column 1512 becomes the explanatory variable var1 column 1821 after noise addition. For example, the explanatory variable with the row number 1510 of the data being 0 is "4.8", and when noise is added to this, it is changed to "4.9".
[0137] In step S910, the CPU 201 of the client terminal 101 performs inference by the AI acquired in step S901 on the dataset of the table 1820 after noise addition generated in step S909, and calculates the accuracy. That is, using the same AI as before noise addition, the inference accuracy for the table 1820 after noise addition is calculated.
[0138] In step S911, the CPU 201 of the client terminal 101 compares the accuracy of the evaluation dataset before noise addition acquired in step S901 with the accuracy of the evaluation dataset after noise addition acquired in S910, and calculates the amount of accuracy degradation.
[0139] In step S912, the CPU 201 of the client terminal 101 determines whether the amount of accuracy degradation obtained in step S911 is within the threshold value of the allowable amount of accuracy degradation acquired in step S903. If it is within the threshold value, it is OK; otherwise, it is NG.
[0140] In step S913, the CPU 201 of the client terminal 101 creates a message to be displayed on the screen according to the inspection result. Specifically, for the description of the amount of accuracy degradation before and after noise addition and whether it is within the threshold value, a message to be displayed in the result description column 1002 is created.
[0141] In step S914, the CPU 201 of the client terminal 101 reads an advice message on the actions to be taken by future AI developers from the database 340 or the file storage 350. Note that a pre-created message may be acquired and displayed from the database 340 or the file storage 350, or an appropriate message may be displayed based on past data using AI.
[0142] In step S915, the CPU 201 of the client terminal 101 displays the inspection result 1000 after noise addition in FIG. 14 in the inspection result display area 1210 in FIG. 18. Note that the inspection result for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.
[0143] FIG. 14 shows an example of the result screen of the noise tolerance test. The result screen consists of inspection result information 1001, result description column 1002, comparison table 1010 of accuracy before and after noise addition, and future action information 1030.
[0144] The inspection result information 1001 displays the result (OK / NG) of the inspection. The determination result of step S912 is displayed.
[0145] The result description column 1002 displays the message generated in step S913 here.
[0146] The comparison table 1010 of accuracy before and after noise addition consists of an accuracy description column 1015 and an actual accuracy column 1020. The accuracy description column 1015 is a column that explains the result of the noise added to each explanatory variable column or the result without noise addition. The actual accuracy column 1020 is a column that displays the difference from the case without accuracy and noise. In the first row 1021, the accuracy of the evaluation data set before noise addition obtained in step S901 is displayed. Since no noise is added, no value indicating the difference is displayed in the parentheses. From the next row, the difference between the accuracy when noise is added for each explanatory variable and the accuracy before noise addition is displayed respectively. For example, the third row 1022 displays the accuracy obtained in step S910, and the difference output in step S911 is displayed in the parentheses. Note that when it is determined in step S912 that it does not fall within the threshold value, it may be identified and displayed, for example, by displaying the displayed value in red.
[0147] In this way, by adding noise and comparing the accuracies before and after it, the robustness of the AI to be evaluated can be measured. That is, it is possible to grasp in advance a measure of how much the accuracy deteriorates when noise enters the input data in the production environment. If it exceeds the threshold value, since the inference result is likely to be affected when noise enters, for example, it can be recognized that it is necessary to take measures such as increasing the learning data so as not to obtain an unintended inference result.
[0148] The future action information 1030 displays the future action based on the inspection result obtained in step S914.
[0149] Return to the description of FIG. 4. In step S406, the CPU 201 of the client terminal 101 determines whether a plurality of jobs are selected in the pull-down list 1201. If YES, it proceeds to step S407, and if NO, it proceeds to step S408.
[0150] In step S407, the CPU 201 of the client terminal 101 executes a subprogram that performs a process of visualizing the change in the inference performance between models. This subprogram is one of the inspection items for grasping the change in the nature of the inference by the AI, and aims to investigate the effect of the AI improvement activity by comparing the inference results of two AIs and visualizing the changed records.
[0151] FIG. 10 is a flowchart showing the flow of the process in S407 (the process of visualizing the change in the inference performance between models).
[0152] In step S1001, the CPU 201 of the client terminal 101 acquires the evaluation data table A1900 associated with the first AI job number among the plurality of AI job numbers input in step S401.
[0153] In step S1002, the CPU 201 of the client terminal 101 acquires the evaluation data table B1930 associated with the second AI job number among the plurality of AI job numbers input in step S401.
[0154] In this embodiment, the AI related to the first AI job number will be described as the pre-improvement AI, and the AI related to the second AI job number will be described as the post-improvement AI.
[0155] In step S1003, the CPU 201 of the client terminal 101 divides the records (evaluation data) in the evaluation data table A and the evaluation data table B into records where the correct answer and the inference match (correct answer) and records where they do not match (incorrect answer). This will be specifically described using the example in FIG. 25. In FIG. 25(a), the evaluation data table A1900 is divided into a correct answer table A1910 obtained by extracting the records where the correct label column 1515 and the inference label column 1516 match, and an incorrect answer table A1920 obtained by extracting the records where they do not match.
[0156] Also, as shown in FIG. 25(b), the same process is performed on the evaluation data table B1930 to create a correct answer table B1940 obtained by extracting the records where the correct label column 1515 and the inference label column 1516 of the AI match, and an incorrect answer table B1950 obtained by extracting the records where they do not match.
[0157] In step S1004, the CPU 201 of the client terminal 101 acquires the data index of the records that were incorrect in the evaluation data table A but are now correct in the evaluation data table B. That is, the indexes that exist in the index 1921 of the incorrect answer table A1920 and also exist in the index 1941 of the correct answer table B1940 are extracted, and a table 1960 is created as shown in FIG. 25(c).
[0158] In step S1005, the CPU 201 of the client terminal 101 obtains the data indexes of the records that are correct in the data table A but incorrect in the data table B. That is, it extracts the indexes that exist in the index 1911 of the correct table A1910 and also exist in the index 1951 of the incorrect table B1950, and creates a table 1970 as shown in Fig. 25(c).
[0159] In step S1006, the CPU 201 of the client terminal 101 creates a distribution diagram of the whole and data for each explanatory variable. This is a function for investigating the cause of the transition of the inference results between the data table A and the data table B from the perspective of the explanatory variables. For example, it is possible to investigate how the inference results have changed between the AI before improvement and the AI after improvement.
[0160] Fig. 15 is an example of a screen visualizing the explanatory variables of the records whose inferences have changed between the evaluation data table A and the evaluation data table B.
[0161] The visualization screen includes a pull-down list 1105, a graph display area 1110 showing the distribution of the explanatory variables (the explanatory variables selected in 1105. Hereinafter, referred to as explanatory variable A.) of the records that were incorrect in the AI related to the first AI job number (the AI before improvement) but correct in the AI related to the second AI job number (the AI after improvement), the vertical axis 1111 of the histogram, the horizontal axis 1112 of the histogram, a histogram 1113 showing the distribution of the explanatory variable A in the evaluation data table A1900, a histogram 1114 showing the distribution of the explanatory variable A of the records that were incorrect in the AI before improvement but correct in the AI after improvement, a legend 1119, and a graph display area 1120 showing the distribution of the explanatory variable A of the records that were correct in the AI before improvement but incorrect in the AI after improvement.
[0162] The pull-down list 1105 is for selecting the explanatory variable to be displayed. The displayed histograms 1113 to 1114 change according to the selected explanatory variable.
[0163] The vertical axis 1111 of the histogram represents the ratio of the number of records with the numerical value (interval) specified on the horizontal axis to the total number of records for the explanatory variable A. For example, if the total number of evaluation data is 100 records and the number of records where the explanatory variable A is between 4.0 and 4.9 is 3, then with the value on the horizontal axis being 4.0 - 4.9, the vertical axis is graphically displayed as 0.03.
[0164] The horizontal axis 1112 of the histogram represents the numerical value (interval) of the explanatory variable A.
[0165] By superimposing and displaying the distribution of the explanatory variable A in the entire evaluation data and the distribution of the explanatory variable A related to the records that were incorrect with the previous AI but correct with the improved AI, it is possible to recognize in which numerical value / interval of the explanatory variable A the inference accuracy has improved. For example, in a histogram like 1115, if the part that has newly become correct is larger for the overall histogram, it is considered that the AI has become better at (the inference accuracy has increased for) that part.
[0166] By superimposing and displaying the distribution of the explanatory variable A in the entire evaluation data and the distribution of the explanatory variable A related to the records that were correct with the previous AI but incorrect with the improved AI, as in the graph displayed in the area 1120, it is possible to recognize in which numerical value / interval of the explanatory variable A the inference accuracy has deteriorated. For example, in a histogram like 1125, if the part that has newly become inconsistent is larger for the overall histogram, it is considered that the AI has become less good at (the inference accuracy has decreased for) that part.
[0167] In this way, by comparing the tendency of the entire explanatory variable and the tendency of the cases that have changed from incorrect to correct (or from correct to incorrect), it supports the investigation of the cause of the change in the inference result. For example, if there are many cases that have changed from incorrect to correct, it can be judged that the improvement has been successful. Conversely, if there are cases that have changed from correct to incorrect, it becomes clear which part has become less good at, thus providing a trigger for considering the investigation of the cause.
[0168] In step S1007, the CPU 201 of the client terminal 101 creates a confusion matrix for records. This is a function for investigating the tendency of changes in the inference results of the evaluation data table A and the evaluation data table B from the perspective of the target variable.
[0169] FIG. 16 is an example of a screen showing a confusion matrix of the tendency of changes in the inference results. The confusion matrix screen consists of a confusion matrix 1131 that has changed from incorrect to correct and a confusion matrix 1141 that has changed from correct to incorrect.
[0170] The confusion matrix 1131 that has changed from incorrect to correct is a confusion matrix that displays, for each class of the target variable, the number of inferences that were incorrect in the evaluation data table A but became correct in the evaluation data table B. It consists of a confusion matrix 1134 that shows how many cases there are where the inference result of class 1132 in the evaluation data table A has changed to the inference result of class 1133 in the evaluation data table B.
[0171] For example, "25" in cell 1135 of class 2 in the evaluation data table A and class 1 in the evaluation data table B means that there were a total of 25 records that were misclassified as class 2 in the evaluation data table A (the correct label and the inference label did not match) but were correctly classified as class 1 in the evaluation data table B (the correct label and the inference label matched). From this, it can be understood that the measures in the evaluation data table B have improved the errors in class 1.
[0172] The confusion matrix 1141 that has changed from correct to incorrect is a confusion matrix that displays, for each class of the target variable, the number of inferences that were correct in the evaluation data table A but became incorrect in the evaluation data table B. It consists of a confusion matrix 1144 that shows how many cases there are where the inference result of class 1142 in the evaluation data table A has changed to the inference result of class 1143 in the evaluation data table B.
[0173] For example, the fact that cell 1145 in class 4 of evaluation data table A and class 1 of evaluation data table B is "50" means that there were a total of 50 records that were correctly classified as class 4 in evaluation data table A (the correct label and the inference label matched), but were misclassified as class 1 in evaluation data table B (the correct label and the inference label did not match). From this, it can be understood that the mistakes in class 4 have worsened with the measures in data table B.
[0174] By visualizing the changes in inferences by AI in this way, it supports the qualitative evaluation of measures. Specifically, before and after the improvement of AI, it becomes clear which classes have become better or worse at making inferences. Therefore, it becomes easier to make judgments regarding the improvement activities of AI. Also, it becomes easier to notice the changes in nature due to improving AI. Even if an area that was good at making inferences before improvement becomes difficult after improvement, it is not necessarily the case that the improvement activities have failed. If inferences in another area have become better due to the improvement, the nature of the AI has changed, and developers can consider the usage scenarios according to the changes in nature.
[0175] In step S1008, the CPU 201 of the client terminal 101 creates a data table for checking the details of the records that have changed. This is information for checking the details of the changes visualized in steps S1006 and S1007.
[0176] Figure 17 is the result screen of a data table for checking the details of the records that have changed. The result screen consists of the number of records information 1151 that have changed from incorrect to correct, the records 1152 that have changed from incorrect to correct, the number of records information 1155 that have changed from correct to incorrect, and the records 1156 that have changed from correct to incorrect.
[0177] The number of records information 1151 that have changed from incorrect to correct is a message representing the number of records in which inferences that did not match the correct answer have become matching.
[0178] The 1152 records that changed from incorrect to correct are table data composed of records having the indexes extracted in step S1004.
[0179] The record count information 1155 of records that changed from correct to incorrect is a message indicating the number of records for which the inference that was correct has become incorrect.
[0180] The 1156 records that changed from correct to incorrect are table data composed of records having the indexes extracted in S1005.
[0181] In step S1009, the CPU 201 of the client terminal 101 displays the results of FIGS. 15, 16, and 17 in the inspection result display area 1210 of FIG. 18. Note that the comparison and evaluation results inspected this time may be written to the database 340 or the file storage 350 so that they can be quickly displayed next time.
[0182] Returning to the description of FIG. 4. In step S408, the CPU 201 of the client terminal 101 creates a summary table of the inspection results and displays it in the inspection result display area 1210 of FIG. 18.
[0183] FIG. 20 is a screen of the summary table of the inspection results. The screen of the summary table consists of a summary table 1301.
[0184] The summary table 1301 displays a list of OK / NG determinations of the inspection results performed in the previous processing. It consists of an inspection number column 1310, an inspection item name column 1311, an inspection result column 1312, and a simple result explanation column 1314. In the inspection result column 1312, for example, in the case of NG, it may be highlighted according to the OK / NG result. Note that the summary table for the job number inspected this time may be written to the database 340 or the file storage 350 so that it can be quickly displayed next time.
[0185] Next, a function for searching for jobs that have recorded inspection results will be described. The flowchart in FIG. 5 is a flowchart of the function for searching for jobs from the OK / NG results of inspection items. For example, when receiving a press of a button for using the search function from the menu screen, the screen in FIG. 21 is opened and the processing of this flowchart is executed.
[0186] In step S501, the CPU 201 of the client terminal 101 receives the input of the check box 1410 for the inspection number in FIG. 21. For each inspection item described in this embodiment, it is possible to search for AI jobs whose inspection results are OK. Although only the check boxes 1410 for inspection numbers from 1 to 3 are shown in the figure, they are displayed as many as the number of inspection items.
[0187] In step S502, the CPU 201 of the client terminal 101 determines whether the "Search for Jobs" button 1415 has been pressed. If YES, proceed to step S503. If NO, return to step S502.
[0188] In step S503, the CPU 201 of the client terminal 101 searches for jobs whose inspection results of the inspection items checked in the check box in step S501 are OK. When multiple inspection items are checked, an AND search is performed. When nothing is checked, all jobs are output as search results.
[0189] In step S504, the CPU 201 of the client terminal 101 displays the job list 1420 of the search results on the screen. The job list 1420 is a data table of the search results of AI jobs. It consists of a check box 1421 for detailed viewing of inspection results, an AI job number column 1422, and an inspection number column 1423.
[0190] In step S505, the CPU 201 of the client terminal 101 receives the input of the check box 1421 for detailed viewing of the AI job.
[0191] In step S506, the CPU 201 of the client terminal 101 determines whether the "View Details of Inspection Results of Selected Job" button 1430 has been pressed. If YES, it proceeds to step S507; if NO, it returns to step S506.
[0192] In step S507, the CPU 201 of the client terminal 101 transitions to the screen of FIG. 18 and displays the details of the inspection results of the AI jobs with checkboxes filled in.
[0193] By searching for the inspection result summary, the characteristics of each AI can be grasped. In the example of FIG. 21, it can be seen that the AI of AI job number 3 has less bias in inferences to specific labels than inspection number 2 and is a relatively robust AI compared to the results of inspection number 3. On the other hand, there is a section within a certain explanatory variable that the AI is not good at according to the results of inspection number 1. Since it is rare for all items to be OK, the areas where each AI is good and the areas where it is not good can be confirmed. Also, when wanting to select which AI to use depending on the application scenario, by searching for the inspection items that are OK, the user can easily find the AI they need.
[0194] As described above, according to this embodiment, a mechanism capable of grasping the reliability of the AI can be provided.
[0195] Note that in this embodiment, an AI related to data analysis is used as an example for explanation, but it is not limited to this. For example, it can be applied to the evaluation of various AIs such as those used in image recognition, appearance inspection, demand prediction, etc. Also, not limited to the evaluation of AIs, if it is created based on a dataset, it may be used for the evaluation of that dataset.
[0196] The present invention can be implemented, for example, as an embodiment such as a system, device, method, program, or recording medium. Specifically, it may be applied to a system composed of a plurality of devices, or it may also be applied to a device consisting of a single device.
[0197] Note that each of the various controls described as being performed by the CPU 201 may be performed by one piece of hardware, or the entire apparatus may be controlled by a plurality of pieces of hardware (e.g., a plurality of processors or circuits) sharing the processing.
[0198] Moreover, although the present invention has been described in detail based on its preferred embodiments, the present invention is not limited to these specific embodiments, and various forms within the scope not departing from the gist of this invention are also included in the present invention. Furthermore, each of the above-described embodiments merely shows one embodiment of the present invention, and it is also possible to appropriately combine the embodiments.
[0199] Also, in the above-described embodiments, the case where the present invention is applied to a PC has been described as an example, but this is not limited to this example, and it is applicable to any apparatus capable of displaying the evaluation result of the inference model. For example, it is applicable to a PDA, a mobile phone terminal (smartphone), a tablet terminal, and the like.
[0200] (Other Embodiments) The present invention is also realized by executing the following processing. That is, software (program) that realizes the functions of the above-described embodiments is supplied to a system or apparatus via a network or various storage media, and a computer (or CPU, MPU, etc.) of the system or apparatus reads out and executes the program code. In this case, the program and the storage medium storing the program constitute the present invention.
Description of Reference Numerals
[0201] 100 Network 101 Client Terminal 102 Server
Claims
1. Control means for controlling to divide a dataset used for evaluating a learned model into a plurality of groups based on the value of a first variable; Display control means for controlling to identify and display a group among the plurality of groups in which the inference result by the learned model is equal to or less than a threshold value; An information processing system comprising the same.
2. The information processing system according to claim 1, wherein the display control means controls to display the number of times the learned model has been learned and the ratio of the number of times the inference result was correct.
3. The learned model includes determination means for determining whether or not it includes a group equal to or less than the threshold value; The information processing system according to claim 1, wherein the learned model determined by the determination means to include a group equal to or less than the threshold value is evaluated as not satisfying the inspection result.
4. The information processing system according to claim 1, wherein the threshold value is determined based on the inference result of the learned model and the allowable reduction amount of the inference result.
5. The information processing system according to claim 1, wherein the display control means controls to display text generated according to the inspection result.
6. A control step for controlling to divide a dataset used for evaluating a learned model into a plurality of groups based on the value of a first variable; A display control step for controlling to identify and display a group among the plurality of groups in which the inference result by the learned model is equal to or less than a threshold value; A control method for an information processing system comprising the same.
7. A program for causing at least one computer to function as the information processing system according to any one of claims 1 to 5.
Citation Information
Patent Citations
Detection system, information processing apparatus, evaluation method, and program
JP2019106119A
Machine learning model development and optimization process that ensures performance validation and data sufficiency for regulatory approval
US20210201190A1
Model analysis device, model analysis method, and recording medium
WO2023181245A1
Predicted situation visualization device, predicted situation visualization method, and predicted situation visualization program
JP7095744B2