Computer system and data analysis method

The computer system addresses the challenge of evaluating teacher data influence on decision tree-based models by calculating similarity and influence scores, achieving efficient and accurate assessments without excessive processing time.

JP7699529B2Active Publication Date: 2025-06-27HITACHI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021191403
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-06-27
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently evaluating the degree of influence of teacher data on the prediction accuracy of decision tree-based machine learning models without significantly increasing processing time, and are limited in applicability to specific types of machine learning models.

Method used

A computer system that evaluates each piece of teacher data by calculating a similarity score based on its similarity to other teacher data in a learned decision tree model, and uses this score to select target data for calculating an influence score that assesses its impact on the model's accuracy.

Benefits of technology

Enables efficient evaluation of the influence of each teacher data point on the accuracy of a decision tree-based machine learning model, while keeping processing time manageable, and allows for broader applicability beyond deep learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699529000001
    Figure 0007699529000001
  • Figure 0007699529000002
    Figure 0007699529000002
  • Figure 0007699529000003
    Figure 0007699529000003
Patent Text Reader

Abstract

To provide a computer system capable of evaluating an influence degree of teacher data while suppressing an increase in a processing time to a machine learning model of a decision tree system.SOLUTION: A similar score calculation part 21 calculates a similar score obtained by evaluating similarity between each teacher data in a learned model and another teacher data about the teacher data used for learning of the learned model by using a tree structure of the learned model of an object predictor 13. An evaluation part selects object data being teacher data of an evaluation object from a teacher data group on the basis of the similar score, and calculates an influence score obtained by evaluating an influence degree of the object data to the accuracy of the learned model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a computer system and a data analysis method.

Background Art

[0002] Generally, in order to improve the prediction accuracy of a machine learning model, it is considered effective to increase the number of pieces of teacher data used for learning the machine learning model. However, there may be harmful data mixed in the teacher data that rather reduces the prediction accuracy of the machine learning model learned from it. Examples of harmful data include mislabeled data in which an incorrect value is set for the target variable, and outlier data indicating a special situation with a low reproducibility rate.

[0003] Non-Patent Document 1 discloses a technique for evaluating the degree of influence of target data on the prediction accuracy of a reference model by comparing the prediction errors for specific evaluation data in each of a reference model learned from all n pieces of teacher data and a reference model learned from n - 1 pieces of teacher data obtained by removing one piece of target data from the n pieces of teacher data. In this technique, by using all the teacher data as target data, learning the reference model respectively, and comparing the prediction errors, it is possible to evaluate the degree of influence of all the teacher data on the prediction accuracy of the reference model.

[0004] Non-Patent Document 2 discloses a technique for approximately evaluating the degree of influence of each piece of teacher data on the prediction accuracy of a deep learning model for specific evaluation data based on the characteristics of the deep learning model, which is a type of machine learning model.

[0005] Patent Document 1 discloses a technique for analyzing the degree of influence of each piece of teacher data on the prediction accuracy of a deep learning model calculated for a plurality of evaluation data using the technique described in Non-Patent Document 2, and identifying harmful data that reduces the prediction accuracy of the deep learning model.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Non-Patent Document

[0007]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0008] In the technology described in Non-Patent Document 1, although it can be applied to any machine learning model, since it is necessary to perform machine learning processing for generating a reference model for each teacher data, there is a problem that the processing time becomes enormous in proportion to the number of teacher data.

[0009] In the technology described in Non-Patent Document 2, since the degree of influence on the prediction accuracy of teacher data is evaluated using the characteristics of the deep learning model, there is a problem that the applicable machine learning model is limited to the deep learning model. In particular, there is a problem that it cannot be applied to a decision tree-based machine learning model, which is an effective machine learning model for inference problems dealing with structured data.

[0010] In the technology described in Patent Document 1, since the degree of influence evaluated by the technology described in Non-Patent Document 2 is used, the applicable machine learning models are limited, similar to the technology described in Non-Patent Document 2. Note that by using the degree of influence evaluated by the technology described in Non-Patent Document 1 instead of the degree of influence evaluated by the technology described in Non-Patent Document 2, generality can be improved. However, in this case, similar to the technology described in Non-Patent Document 1, there is a problem that when the number of pieces of teacher data increases, the processing time becomes enormous.

[0011] An object of the present disclosure is to provide a computer system and a data analysis method capable of evaluating the degree of influence of teacher data on the prediction accuracy of a decision tree-based machine learning model while suppressing an increase in processing time.

Means for Solving the Problems

[0012] A computer system according to an aspect of the present disclosure is a computer system that evaluates each piece of teacher data included in a teacher data group used for learning a learned model having a tree structure by a decision tree, and uses the tree structure to calculate, for each piece of teacher data, a similarity score for evaluating the similarity between the teacher data and other teacher data in the learned model; and an evaluation unit that selects target data, which is the teacher data to be evaluated, from the teacher data group based on the similarity score and calculates an influence score for evaluating the degree of influence of the target data on the accuracy of the learned model.

Effects of the Invention

[0013] According to the present invention, it becomes possible to evaluate the degree of influence of each piece of teacher data on the accuracy of a learned model while suppressing an increase in processing time for a decision tree-based machine learning model.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Best Mode for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0016] FIG. 1 is a configuration diagram showing a computer system according to an embodiment of the present disclosure. The computer system 100 shown in FIG. 1 includes computers 1 to 3, and each of the computers 1 to 3 is communicably connected to each other via a network 10. Further, each of the computers 1 to 3 is connected to a terminal 4 via the network 10. The terminal 4 is a terminal device operated by a user who uses the computer system 100. Note that the computer system 100 shown in FIG. 1 is merely an example, and a configuration having one, two, or four or more computers may be used.

[0017] The computer 1 is a computer that predicts a value related to a desired event using a learned model that is a learned machine learning model, and includes a teacher data storage unit 11, an evaluation data storage unit 12, and a target predictor 13.

[0018] The teacher data storage unit 11 stores a group of teacher data, which is a plurality of teacher data used for learning the learned model. The evaluation data storage unit 12 stores a group of evaluation data, which is a plurality of evaluation data for evaluating the prediction accuracy of the learned model.

[0019] The target predictor 13 is a predictor that predicts a value related to a desired event based on the input data, and is realized by a learned model obtained by machine learning using the teacher data stored in the teacher data storage unit 11. The learned model of the present embodiment is a decision tree-based machine learning model (a machine learning model including a tree structure formed by decision trees).

[0020] The computer 2 is a computer that evaluates the influence degree of each teacher data stored in the teacher data storage unit 11 of the computer 1 on the prediction accuracy of the target predictor 13 of the computer 1, and includes a similarity score calculation unit 21, a data removal unit 22, a predictor generation unit 23, an accuracy evaluation unit 24, an influence score calculation unit 25, and a result output unit 26.

[0021] The similarity score calculation unit 21 calculates a similarity score, which is a value obtained by evaluating the similarity between each teacher data included in the teacher data group stored in the teacher data storage unit 11 of the computer 1 and other teacher data, and outputs the similarity score for each teacher data as similarity score data. Note that since the lower the similarity, the higher the rarity of the teacher data compared to other teacher data, the similarity score can also be said to be a value that evaluates the rarity of the teacher data in the teacher data group.

[0022] The data removal unit 22, the predictor generation unit 23, the accuracy evaluation unit 24, and the influence score calculation unit 25 constitute an evaluation unit that selects target data from the teacher data group stored in the teacher data storage unit 11 based on the similarity score calculated by the similarity score calculation unit 21 and calculates an influence score that evaluates the influence degree of the target data on the accuracy of the target predictor 13.

[0023] The data removal unit 22 selects target data from the teacher data group based on the similarity score, and generates a temporary teacher data group obtained by removing the target data from the teacher data group for each target data. The target data is, for example, the teacher data to be evaluated for calculating the influence score, and is, for example, the teacher data whose similarity score is less than or equal to the threshold value.

[0024] The predictor generation unit 23 is a generation unit that generates a temporary predictor using a temporary learned model obtained by learning the temporary teacher data group for each temporary teacher data group generated by the data removal unit 22 using the learning algorithm for generating the target predictor 13.

[0025] The accuracy evaluation unit 24 generates and outputs an evaluation result that evaluates the prediction accuracy of the target predictor 13 and each temporary predictor based on each evaluation data included in the evaluation data group stored in the evaluation data storage unit 12. Specifically, for each evaluation data, the accuracy evaluation unit 24 compares the prediction result of the target predictor 13 with respect to the explanatory variable of the evaluation data with the objective variable of the evaluation data, evaluates the prediction accuracy of the target predictor 13, and outputs the target predictor accuracy evaluation result that is the evaluation result. Similarly, for each evaluation data, the accuracy evaluation unit 24 compares the prediction result of each temporary predictor with respect to the explanatory variable of the evaluation data with the objective variable of the evaluation data, evaluates the prediction accuracy of each temporary predictor, and outputs the temporary predictor accuracy evaluation result that is the evaluation result.

[0026] Based on the evaluation result output from the accuracy evaluation unit 24, the influence score calculation unit 25 calculates an influence score that evaluates the influence degree of the target data on the accuracy of the target predictor 13 for each target data. Specifically, for each temporary predictor, the influence score calculation unit 25 calculates the comparison result of comparing the target predictor accuracy evaluation result, which is the evaluation result, with the temporary predictor accuracy evaluation result as the influence score of the target data excluded in the temporary teacher data group used for generating the temporary predictor. Then, the influence score calculation unit 25 outputs the influence score for each target data as influence score data.

[0027] The result output unit 26 outputs the data based on the influence score data to the terminal 4 as analysis result data indicating the analysis result by the computer system 100.

[0028] The computer 3 is a third computer that stores the data calculated by the computer 1, and has a similarity score storage unit 31 and an influence score storage unit 32.

[0029] The similarity score storage unit 31 stores the similarity score data output from the similarity score calculation unit 21 of the computer 2. The influence score storage unit 32 stores the influence score data output from the influence score calculation unit 25 of the computer 2.

[0030] Figure 2 is a diagram showing the hardware configuration of each of the computers 1 to 3. As shown in Figure 2, each of the computers 1 to 3 has a secondary storage device 101, a main memory device 102, a processor 103, an input device 104, an output device 105, and a network interface 106.

[0031] The secondary storage device 101 is a device that stores various data. For example, it stores a program (computer program) that defines the operation of the processor 103, and data used or generated by the processor 103 or other computers. The teacher data storage unit 11, the evaluation data storage unit 12, the similarity score storage unit 31, and the influence score storage unit 32 in Figure 1 are realized, for example, by the secondary storage device 101. The main memory device 102 is a memory that functions as a work area for processing by a program.

[0032] The processor 103 reads the program stored in the secondary storage device 101 into the main memory device 102, and uses the main memory device 102 to execute processing according to the program. Each part 13, 21 to 26 of the computer 1 shown in Figure 1 is realized by the processor 103.

[0033] The input device 104 is a device into which various information is input from an operator of the computer system or the like, and the input information is used for processing by the processor 103. The output device 105 is a device that outputs (for example, displays) various information. The network interface 106 is a communication device that is communicably connected to external devices such as other computers and terminals 4, and transmits and receives data to and from the external devices.

[0034] Figure 3 is a diagram showing an example of a teacher data group stored in the teacher data storage unit 11. In the example of Figure 3, the teacher data group is stored in the teacher data storage unit 11 as a teacher data table 300 having a table structure, and each record of the teacher data table 300 corresponds to individual teacher data.

[0035] The teacher data table 300 includes fields 301 to 303. Field 301 stores a teacher ID which is identification information for identifying teacher data. Field 302 stores the explanatory variables of the teacher data. When there are multiple explanatory variables, field 302 is provided for each explanatory variable, and each field 302 stores a different explanatory variable. Field 303 stores the target variable of the teacher data.

[0036] In this embodiment, the teacher data is data related to concrete. The explanatory variables of each teacher data are variables that affect the strength of the concrete (for example, the amount of water, the amount of cement, the number of days elapsed since the concrete was produced, etc.), and the target variable is the strength of the concrete.

[0037] FIG. 4 is a diagram showing an example of an evaluation data group stored in the evaluation data storage unit 12. In the example of FIG. 4, the evaluation data group is stored in the evaluation data storage unit 12 as an evaluation data table 400 having a table structure, and each record of the evaluation data table 400 corresponds to individual evaluation data.

[0038] The evaluation data table 400 includes fields 401 to 403. Field 401 stores an evaluation ID which is identification information for identifying evaluation data. Field 402 stores the explanatory variables of the evaluation data. When there are multiple explanatory variables, field 402 is provided for each explanatory variable, and each field 402 stores a different explanatory variable. Field 403 stores the target variable of the evaluation data. Note that the evaluation data is the same type of data as the teacher data, and in this embodiment, it is data related to the strength of concrete.

[0039] FIG. 5 is a diagram showing an example of similarity score data stored in the similarity score storage unit 31 of FIG. 1. The similarity score data 500 shown in FIG. 5 includes fields 501 and 502. Field 501 stores a teacher ID. Field 502 stores the similarity score of the teacher data identified by the teacher ID. The detailed calculation method of the similarity score will be described later.

[0040] FIG. 6 is a diagram showing an example of influence score data stored in the influence score storage unit 32 of FIG. 1. The influence score data 600 shown in FIG. 6 includes fields 601 and 602. Field 601 stores the teacher ID of the target data. Field 602 stores the influence score of the target data identified by the teacher ID. The detailed calculation method of the influence score will be described later.

[0041] FIG. 7 is a diagram showing the internal configuration of the target predictor 13. The target predictor 13 shown in FIG. 7 is a predictor realized by an ensemble tree model which is a kind of decision tree-based machine learning model.

[0042] The target predictor 13 shown in FIG. 7 has a plurality of decision trees 131 that predict values related to a desired event based on the input data, and calculates a prediction value as the target predictor 13 based on the prediction values predicted by each decision tree 131. Hereinafter, the prediction value predicted by the decision tree 131 is referred to as an individual prediction value, and the prediction value predicted by the target predictor 13 is simply referred to as a prediction value. The prediction value is, for example, a statistical value (e.g., the most frequent value or the average value, etc.) of the individual prediction values of each decision tree 131. The number of decision trees 131 is not particularly limited.

[0043] The decision tree 131 includes a plurality of nodes 131a, and each node 131a is linked by a determination condition for an explanatory variable. Among the nodes 131a of the decision tree 131, the node having no link destination is called a leaf node 131b, and a value related to the desired event is associated therewith. Therefore, the value corresponding to the leaf node 131b reached by the determination condition of each node 131a of the decision tree 131 becomes the individual prediction value.

[0044] Note that the node configurations of the respective decision trees 131 are different from each other. In addition, each decision tree 131 is assigned a decision tree ID for identifying the decision tree, and each leaf node 131b is assigned a leaf node ID for identifying the leaf node. The leaf node ID is uniquely set for each decision tree 131. That is, even if the values of the leaf node IDs are the same, different leaf nodes 131b are indicated if the decision trees 131 are different.

[0045] FIG. 8 is a diagram for explaining an example of a target predictor accuracy evaluation process for evaluating the accuracy of the target predictor 13, and FIG. 9 is a flowchart for explaining an example of the target predictor accuracy evaluation process.

[0046] In the target predictor accuracy evaluation process, first, the target predictor 13 acquires the explanatory variables of the evaluation data for each evaluation data from the evaluation data storage unit 12 (step S101).

[0047] The target predictor 13 calculates a predicted value obtained by predicting the value of the target variable from the explanatory variables of the evaluation data for each evaluation data (step S102). The target predictor 13 outputs the predicted value for each evaluation data as predicted value data 700 (step S103).

[0048] Thereafter, the accuracy evaluation unit 24 acquires the predicted value data 700 output from the target predictor 13 and also acquires an evaluation data group from the evaluation data storage unit 12 (step S104).

[0049] The accuracy evaluation unit 24 evaluates the prediction accuracy of the target predictor 13 based on the acquired predicted value data 700 and each evaluation data of the evaluation data group, and outputs the evaluation result as the target predictor accuracy evaluation result 710 (step S105), and ends the target predictor accuracy evaluation process. The prediction accuracy is, for example, a statistical value of the difference between the actual value and the predicted value of the target variable in each evaluation data. Examples of the statistical value include, for example, the mean error or the root mean square error.

[0050] FIG. 10 is a diagram showing an example of the predicted value data 700. The predicted value data 700 shown in FIG. 7 includes fields 701 and 702. Field 701 stores the evaluation ID. Field 702 stores the predicted value of the target variable of the evaluation data identified by the evaluation ID.

[0051] FIG. 11 is a diagram showing an example of the target predictor accuracy evaluation result 710. The target predictor accuracy evaluation result 710 shown in FIG. 11 has a field 711. The field 711 stores the accuracy of the prediction accuracy of the target predictor 13.

[0052] FIG. 12 is a diagram for explaining an example of the similarity score process for generating similarity score data, and FIG. 13 is a flowchart for explaining an example of the similarity score process.

[0053] In the similarity score process, first, the similarity score calculation unit 21 acquires the learned model that realizes the target predictor 13 and the teacher data group stored in the teacher data storage unit 11 (step S201).

[0054] The similarity score calculation unit 21 executes a similarity score calculation process (see FIGS. 14 and 15) for calculating the similarity score of each teacher data included in the teacher data group based on the acquired learned model and teacher data group (step S202).

[0055] The similarity score calculation unit 21 stores the similarity score of each teacher data as similarity score data in the similarity score storage unit 31 (step S203), and ends the similarity score process.

[0056] FIG. 14 is a diagram for explaining an example of the similarity score calculation process in step S202 of FIG. 13, and FIG. 15 is a flowchart for explaining an example of the similarity score calculation process. As shown in FIG. 14, the similarity score calculation unit 21 includes a tree structure extraction processing unit 211, a data application processing unit 212, a reached leaf node aggregation processing unit 213, and a similarity score calculation processing unit 214.

[0057] In the similarity score calculation process, first, the tree structure extraction processing unit 211 of the similarity score calculation unit 21 extracts the tree structure of the learned model from the learned model of the target predictor 13 (step S301). Specifically, the tree structure shows the links between nodes in each decision tree 131 included in the learned model.

[0058] Based on the tree structure extracted by the tree structure extraction unit 211, for each teacher data included in the teacher data group, the data application processing unit 212 identifies, for each decision tree 131 included in the learned model, the reached leaf node 131b that the teacher data reaches when the teacher data is input to the decision tree 131. The data application processing unit 212 outputs, as the reached leaf node data 800, the leaf node ID that identifies the reached leaf node for each decision tree 131 of each teacher data (step S302).

[0059] Based on the reached leaf node data 800, for each teacher data, the reached leaf node aggregation processing unit 213 aggregates, for each reached leaf node that the teacher data of each decision tree 131 reaches, the reach rate, which is the ratio of the teacher data that reaches the reached leaf node to the teacher data included in the teacher data group. The reached leaf node aggregation processing unit 213 outputs the aggregated data as the reached leaf node aggregated data 810 (step S303).

[0060] Based on the reached leaf node aggregated data 810, for each teacher data, the similarity score calculation processing unit 214 calculates and outputs a similarity score that evaluates the similarity of the teacher data to other teacher data (step S304), and ends the similarity score calculation process. The similarity score is, for example, a statistical value of the reach rate of each reached leaf node. Examples of the statistical value include an average value and a median value. Note that the reached leaf node aggregation processing unit 213 and the similarity score calculation processing unit 214 constitute a calculation processing unit that calculates the similarity score of each teacher data based on the reached leaf node data 800.

[0061] FIG. 16 is a diagram showing an example of the reached leaf node data 800. The reached leaf node data 800 shown in FIG. 16 includes fields 801 and 802. Field 801 stores the teacher ID. Field 802 is provided for each decision tree 131 and stores the leaf node ID of the reached leaf node that the teacher data identified by the teacher ID reaches in the corresponding decision tree 131.

[0062] FIG. 17 is a diagram showing an example of the reach leaf node aggregation data 810. FIG. 17 shows that the reach leaf node aggregation data 810 includes fields 811 and 812. Field 811 stores the teacher ID. Field 812 is provided for each decision tree 131 and stores the reach rate at which teacher data identified by the teacher ID reaches the reach leaf node in the corresponding decision tree 131.

[0063] For example, in the example of FIG. 17, in the decision tree 131 with the decision tree ID "Tree 1", it is shown that 0.5% of the total teacher data has reached the reach leaf node (leaf node ID "Leaf 3": see FIG. 16) where the teacher data with the teacher ID "1" reaches. Note that in each decision tree 131, the reach rates of the same reach leaf node are all the same value.

[0064] FIG. 18 is a diagram for explaining an example of the influence score calculation process for calculating the influence score, and FIG. 19 is a flowchart for explaining an example of the influence score calculation process.

[0065] In the influence score calculation process, first, the data removal unit 22 acquires the teacher data group stored in the teacher data storage unit 11 and the similarity score data stored in the similarity score storage unit 31. The data removal unit 22 uses the teacher data with the i-th lowest similarity score as the target data, and generates and outputs a temporary teacher data group 900 obtained by removing the target data from the teacher data group (step S401). Here, i is a counter value for counting the target data, and its initial value is 1.

[0066] The predictor generation unit 23 generates a temporary predictor 910, which is a temporary learned model obtained by learning the temporary teacher data group 900 generated in step S401, using the learning algorithm that generated the learned model of the target predictor 13 (step S402).

[0067] The temporary predictor 910 acquires an evaluation data group from the evaluation data storage unit 12, calculates a predicted value using the explanatory variables of each evaluation data included in the evaluation data group as inputs, and outputs the predicted value for each evaluation data as temporary predicted value data 920 (step S403).

[0068] The accuracy evaluation unit 24 acquires the temporary predicted value data 920 and the evaluation data group from the temporary predictor 910, evaluates the prediction accuracy of the temporary predictor 910 based on the acquired temporary predicted value data 920 and each evaluation data of the evaluation data group, and outputs the evaluation result as the temporary predictor accuracy evaluation result 930 (step S404). Similar to the target predictor accuracy evaluation result 710, the temporary predictor accuracy evaluation result 930 indicates, for example, the statistical value of the difference between the actual value and the predicted value of the target variable in each evaluation data as the prediction accuracy.

[0069] The influence score calculation unit 25 acquires the target predictor accuracy evaluation result 710 and the temporary predictor accuracy evaluation result 930 output in the target predictor accuracy evaluation process (see FIGS. 8 and 9), calculates the comparison result of comparing them as the influence score of the target data excluded in the temporary teacher data used for generating the temporarily learned model, and stores it in the influence score storage unit 32 (step S405). The influence score is, for example, the difference between the target predictor accuracy evaluation result 710 and the temporary predictor accuracy evaluation result 930.

[0070] The influence score calculation unit 25 determines whether or not an end condition for ending the influence score calculation process is satisfied (step S406). The end condition is that i, which is the number of temporary teacher data created, is equal to or greater than a threshold value, etc. The threshold value may be set by a user or an operator, for example, or may be predetermined.

[0071] If the end condition is not satisfied (step S406: No), the influence score calculation unit 25 increments i (step S407) and returns to the process of step S401. On the other hand, if the end condition is satisfied (step S407: Yes), the influence score calculation process ends.

[0072] FIG. 20 is a diagram for explaining an example of result output processing that outputs an influence score, and FIG. 21 is a flowchart for explaining an example of result output processing.

[0073] In the result output processing, first, the result output unit 26 acquires the teacher data stored in the teacher data storage unit 11, the similarity score data stored in the similarity score storage unit 31, and the influence score stored in the influence score storage unit 32 (step S501).

[0074] The result output unit 26 generates analysis result data obtained by combining the respective data acquired in step S501 using the teacher ID as a key, displays an analysis screen showing the branch result data on the terminal 4 (step S502), and ends the result output processing.

[0075] The result output unit 26 may extract target data for which the influence score indicates a decrease in the accuracy of the learned model as harmful data and include it in the analysis result data. For example, assume that the prediction accuracy of the learned model is the root mean square error of the actual value and the predicted value of the target variable in each evaluation data, and the influence score of each target data is a value obtained by subtracting the target predictor accuracy evaluation result 710 from the temporary predictor accuracy evaluation result 930. In this case, if the influence score is negative, it means that the prediction accuracy has been improved by excluding the target data. Therefore, the result output unit 26 extracts the target data as harmful data that degrades the accuracy of the learned model of the target predictor 13.

[0076] FIG. 22 is a diagram showing an example of an analysis screen. The analysis screen 1000 shown in FIG. 22 is a screen displayed on the terminal 4 and includes input boxes 1001 to 1004, an execution button 1005, and a display area 1006.

[0077] The input box 1001 is a box for specifying a target model which is a learned model for constructing the target predictor 13. The input box 1002 is a box for specifying teacher data. The input box 1003 is a box for specifying evaluation data. The input box 1004 is a box for specifying a search range. The search range is a range of similarity scores for specifying teacher data to be selected as target data, and a ratio starting from the lower similarity score of the teacher data or the number starting from the lower similarity score, etc. are specified.

[0078] The execution button 1005 is a button for executing the evaluation of teacher data, and when pressed, the processing by the computer system 100 is started. The display area 1006 is an area for displaying analysis result data, and in the example of FIG. 22, a list of harmful data is displayed.

[0079] As described above, according to the present embodiment, the similarity score calculation unit 21 uses the tree structure of the learned model of the target predictor 13 to calculate a similarity score for evaluating the similarity between each teacher data used for learning the learned model and other teacher data in the learned model. The evaluation units (22 to 25) select target data, which is the teacher data to be evaluated, from the teacher data group based on the similarity score, and calculate an influence score for evaluating the influence degree of the target data on the accuracy of the learned model. Therefore, it is possible to exclude teacher data that is considered to have little influence on the accuracy because it is not rare for a learned model with a high similarity score, and to evaluate only the influence degree on the accuracy of the learned model for teacher data that is likely to reduce the accuracy of the learned model. Thus, it is possible to evaluate the influence degree of teacher data while suppressing an increase in processing time.

[0080] In addition, in the present embodiment, for each decision tree included in the learned model, a similarity score is calculated based on the reached leaf node, which is the leaf node reached by the teacher data when the teacher data is input to the decision tree. Therefore, it is possible to more appropriately calculate the similarity score according to the learning content of the learned model, and thus it is possible to more accurately evaluate the similarity for the learned model.

[0081] In addition, in the present embodiment, for each teacher data, based on the aggregated data obtained by aggregating the reach rates, which are the ratios of the teacher data that reached the reached leaf node to the teacher data included in the teacher data group, for each reached leaf node of each decision tree, a similarity score is calculated. In particular, for each teacher data, the statistical value of the reach rate of each reached leaf node is calculated as the similarity score. Therefore, it is possible to more appropriately calculate the similarity score according to the learning content of the learned model, and thus it is possible to more accurately evaluate the similarity for the learned model.

[0082] In addition, in the present embodiment, an influence score is calculated based on the evaluation result of evaluating the accuracy of the learned model of the target predictor 13 and the evaluation result of evaluating the accuracy of the temporarily learned model obtained by learning the temporary teacher data group with the target data removed. Therefore, it is possible to more accurately evaluate the influence score of the target data.

[0083] In addition, in the present embodiment, the comparison result of comparing the evaluation result data of the learned model with the evaluation result of the temporarily learned model is calculated as the influence score of the target data excluded in the temporary teacher data group used for generating the temporarily learned model. Therefore, it is possible to more accurately evaluate the influence degree of the teacher data.

[0084] In addition, in the present embodiment, since the target data for which the influence score indicates an improvement in the accuracy of the learned model is extracted, it is possible to easily identify the teacher data that is harmful to the learned model.

[0085] Also, in the present embodiment, teacher data with a similarity score equal to or less than a threshold value is selected as target data. Therefore, it becomes possible to appropriately select the target data.

[0086] The above-described embodiments of the present disclosure are examples for explaining the present disclosure, and are not intended to limit the scope of the present disclosure only to those embodiments. A person skilled in the art can implement the present disclosure in various other modes without departing from the scope of the present disclosure.

Description of Reference Numerals

[0087] 1 to 3: Computers 4: Terminal 11: Teacher Data Storage Unit 12: Evaluation Data Storage Unit 13: Target Predictor 21: Similarity Score Calculation Unit 22: Data Removal Unit 23: Predictor Generation Unit 24: Accuracy Evaluation Unit 25: Influence Score Calculation Unit 26: Result Output Unit 31: Similarity Score Storage Unit 32: Influence Score Storage Unit 100: Computer System 211: Tree Structure Extraction Processing Unit 212: Data Application Processing Unit 213: Reached Leaf Node Aggregation Processing Unit 213: Reached Leaf Node Aggregation Processing 214: Similarity Score Calculation Processing Unit

Claims

1. A computer system for evaluating each piece of teacher data included in a group of teacher data used for training a trained model having a tree structure by a decision tree, comprising: a similarity score calculation unit that uses the tree structure to calculate a similarity score for evaluating the similarity between the piece of teacher data and other pieces of teacher data in the trained model for each piece of teacher data; an evaluation unit that selects target data, which is the teacher data to be evaluated, from the group of teacher data based on the similarity score, and calculates an impact score for evaluating the degree of influence of the target data on the accuracy of the trained model. A computer system having the above.

2. The similarity score calculation unit includes: a data application processing unit that, for each piece of teacher data, identifies a reached leaf node, which is the leaf node reached by the piece of teacher data when the piece of teacher data is input to the decision tree, for each decision tree included in the trained model; a calculation processing unit that calculates the similarity score based on the reached leaf node. The computer system according to claim 1.

3. The calculation processing unit includes: a aggregation processing unit that, for each piece of teacher data, generates aggregation data obtained by aggregating the reach rate, which is the ratio of the teacher data that has reached the reached leaf node to the teacher data included in the group of teacher data, for each reached leaf node of each decision tree; a similarity score calculation processing unit that calculates the similarity score based on the aggregation data. The computer system according to claim 2.

4. The similarity score calculation processing unit calculates a statistical value of the reach rate of each reached leaf node as the similarity score for each piece of teacher data. The computer system according to claim 3.

5. The evaluation unit includes: a data removal unit that selects the target data based on the similarity score and generates a temporary group of teacher data obtained by removing the target data from the group of teacher data for each piece of target data; a generation unit that generates a temporarily trained model obtained by training the temporary group of teacher data using the learning algorithm that generated the trained model for each temporary group of teacher data; an accuracy evaluation unit that generates an evaluation result for evaluating the accuracy of the trained model and each temporarily trained model based on evaluation data; an impact score calculation unit that calculates the impact score based on the evaluation result. The computer system according to claim 1.

6. The influence score calculation unit calculates, for each of the temporarily trained models, a comparison result obtained by comparing the evaluation result data of the trained model with the evaluation result of the temporarily trained model as the influence score of the target data excluded from the temporary teacher data group used for generating the temporarily trained model. The computer system according to claim 5.

7. The computer system according to claim 1, further comprising a result output unit that extracts and outputs target data among the target data whose influence score indicates a decrease in the accuracy of the trained model.

8. The evaluation unit uses, as the target data, the teacher data whose similarity score is equal to or lower than a threshold value. The computer system according to claim 1.

9. A data analysis method by a computer system for evaluating each teacher data included in a teacher data group used for training a trained model having a tree structure by a decision tree, The computer system includes a processor and a storage device that stores the teacher data group, The processor acquires the teacher data group from the storage device, The processor calculates, for each teacher data, a similarity score for evaluating the similarity between the teacher data and other teacher data in the trained model using the tree structure, The processor selects, based on the similarity score, target data that is the teacher data to be evaluated from the teacher data group, and calculates an influence score for evaluating the degree of influence of the target data on the accuracy of the trained model. A data analysis method.

Citation Information

Patent Citations

  • Method for analyzing learning data and computing system

    JP2020030738A

  • Learning data refining method and computer system

    JP2021033544A

  • Inference method, inference program and information processing apparatus

    JP2021099640A

  • Chained influence scores for improving synthetic data generation

    US20200334557A1