Teacher data editing support system, method, and program
The system addresses biased machine learning models by calculating and visualizing the contribution of discriminatory factors to correct answers, enabling targeted data editing to reduce bias and improve fairness.
Patent Information
- Application Number
- JP2022209963
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2042-12-27
AI Technical Summary
Existing machine learning models trained on data containing discriminatory factors can make biased decisions, and existing methods to address this, such as increasing data volume or binary classification adjustments, are insufficient or limited in applicability.
A system and method that calculates the contribution of discriminatory factors to correct answers, visually presents the relationship between answer changes and discrimination levels, and edits the data based on user specifications to reduce bias.
Reduces discriminatory decisions in machine learning models by allowing for targeted data editing that minimizes the impact of discriminatory factors on predictions.
Smart Images

Figure 0007798761000001 
Figure 0007798761000002 
Figure 0007798761000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a machine learning teacher data editing support system, a teacher data editing support method, and a teacher data editing support program. [Background technology]
[0002] In machine learning, various past human activity histories are sometimes used as training data. In the past, people may have been treated discriminatory based on various differences in their attributes. Therefore, past activity histories may contain information that includes such discriminatory treatment. For example, past credit histories at financial institutions may contain traces of discrimination based on race or gender. AI (Artificial Intelligence) models generated through machine learning using training data that includes such discrimination may make discriminatory decisions. Therefore, it is desirable to reduce discriminatory decisions made by AI models and improve fairness.
[0003] Patent Document 1 discloses a technology that, under the assumption that increasing the number of training data items improves the model's predictive accuracy and fairness, improves fairness in the field of images by generating perturbed images of images with attribute information that are relatively rare in the training data and adding them to the training data.
[0004] Non-patent document 1 discloses a method in which variables such as attributes that may lead to discrimination are used as discriminatory factors, and the case is considered in which the discriminatory factors and the correct answer are both binary (two-valued), in which the proportion of correct answers for each discriminatory factor that are in a desirable state is calculated as an index of fairness, and the correct answer is rewritten to improve that index. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. WO2022 / 123907A1 [Non-patent literature]
[0006] [Non-Patent Document 1] Kamiran, Faisal, and Toon Calders. “Data preprocessing techniques for classification without discrimination.” Knowledge and information systems 33.1 (2012):1-33 Summary of the Invention [Problem to be solved by the invention]
[0007] The technology disclosed in Patent Document 1 assumes that increasing the number of training data items will improve the model's prediction accuracy and fairness, but this is not necessarily the case. For example, if the original image used to generate a perturbed image is affected by a discriminant factor, adding the perturbed image to increase the training data may not reduce the influence of the discriminant factor from the model. The method disclosed in Non-Patent Document 1 targets binary classification problems where the correct answer is expressed as two values, and cannot be applied to other problems such as regression problems.
[0008] One objective of the present disclosure is to provide a technique that helps reduce discriminatory decisions made by machine learning models. [Means for solving the problem]
[0009] A teacher data editing support system according to one aspect of the present disclosure has an input of teacher data including a discrimination factor, which is a variable that may cause discrimination, a feature, which is a variable used for prediction, and a correct answer, and a determination unit that calculates a contribution level, which is an index showing the degree to which the discrimination factor contributed to the correct answer; a display unit that visibly presents evaluation information that represents the relationship between the degree to which the correct answer in the teacher data is changed and the degree of deviation from the initial value of the correct answer, or the degree of discrimination based on the contribution level; and an editing unit that accepts a specification of how much to change the correct answer, changes the correct answer in the teacher data based on the specification, and outputs the changed teacher data.
[0010] One aspect of the present disclosure provides a method for supporting editing of teacher data, which is a method for supporting editing of teacher data using a device having a processing device, in which the processing device calculates a contribution, which is an index showing the degree to which the discrimination factor contributed to the correct answer, in response to input of teacher data including a discrimination factor, which is a variable that may cause discrimination, a feature, which is a variable used for prediction, and a correct answer, and visually presents evaluation information that represents the relationship between the degree to which the correct answer in the teacher data is changed and the degree of deviation from the initial value of the correct answer, or the degree of discrimination based on the contribution, accepts a specification of how much to change the correct answer, changes the correct answer in the teacher data based on the specification, and outputs the changed teacher data.
[0011] A teacher data editing support program according to one aspect of the present disclosure inputs teacher data including a discrimination factor, which is a variable that may cause discrimination, a feature, which is a variable used for prediction, and a correct answer, into a device having a processing device; calculates a contribution, which is an index showing the degree to which the discrimination factor contributed to the correct answer; visibly presents evaluation information that represents the relationship between the degree to which the correct answer in the teacher data is changed and the degree of deviation from the initial value of the correct answer, or the degree of discrimination based on the contribution; accepts a specification of how much to change the correct answer; changes the correct answer in the teacher data based on the specification; and outputs the changed teacher data. [Effects of the Invention]
[0012] One aspect of the present disclosure allows for the reduction of discriminatory decisions made by machine learning models. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a functional block diagram showing an example configuration of a teacher data editing support system. [Figure 2] FIG. 10 is a conceptual diagram illustrating an example of a format of training data. [Figure 3] FIG. 10 is a conceptual diagram illustrating an example of a format of a determination result. [Figure 4] A conceptual diagram illustrating the format of edited teacher data. [Figure 5] 10 is a flowchart illustrating information processing performed by a determination unit. [Figure 6] 10 is a flowchart illustrating information processing performed by an editorial department. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example of a format of determination history data. [Figure 8] FIG. 4 is a conceptual diagram showing a first display example by the display unit. [Figure 9] FIG. 10 is a conceptual diagram showing a second display example by the display unit. [Figure 10] FIG. 10 is a conceptual diagram showing a third example of display by the display unit. [Figure 11] FIG. 10 is a conceptual diagram illustrating an example of a format of training data. [Figure 12] FIG. 10 is a conceptual diagram illustrating an example of a format of a determination result. [Figure 13] A conceptual diagram illustrating the format of edited teacher data. [Figure 14] FIG. 1 is a functional block diagram showing an example configuration of a teacher data editing support system. [Figure 15] FIG. 10 is a conceptual diagram illustrating an example of a format of requirement information. [Figure 16] 10 is a flowchart illustrating an example of information processing performed by a suggestion generating unit. [Figure 17] FIG. 1 is a functional block diagram showing an example configuration of a teacher data editing support system. [Figure 18] 10 is a flowchart illustrating information processing performed by an affiliated aggregation unit. [Figure 19] FIG. 10 is a conceptual diagram illustrating the format of a combination mask. [Figure 20] FIG. 10 is a conceptual diagram illustrating an example of a format for the partnership summary results. [Figure 21] 10 is a flowchart illustrating information processing performed by a contribution calculation unit. [Figure 22] FIG. 10 is a conceptual diagram illustrating an example of a format for the partnership summary results. [Figure 23] FIG. 10 is a conceptual diagram illustrating an example format of a provisional contribution result. [Figure 24] FIG. 1 is a block diagram illustrating an example of the configuration of a teacher data editing support system. [Figure 25] FIG. 2 is a conceptual diagram illustrating an example of a hardware configuration of a computer. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [Example]
[0015] FIG. 1 is a functional block diagram showing an example of the configuration of a training data editing support system.
[0016] The teacher data editing support system 1 includes at least a processing device and a storage device (not shown). The teacher data editing support system 1 may further include a communication device, an input device, an output device, etc.
[0017] The processing device is configured, for example, with a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), FPGA (Field-Programmable Gate Array), etc. The various functions of the teacher data editing support system 1 are realized by the processing device reading and executing various programs and data stored in the storage device.
[0018] More specifically, the processing device implements a determination unit 102, a display unit 104, and an editing unit 105 by reading and executing various programs and data stored in the storage device.
[0019] A storage device is a device that stores programs and data, and is, for example, a random access memory (RAM), a read only memory (ROM), or a non-volatile semiconductor memory (Non-Volatile RAM (NVRAM)).
[0020] The storage device may be, for example, a hard disc drive (HDD), a solid state drive (SSD), a storage system, a device for reading and writing recording media such as an integrated circuit (IC) card, a secure digital (SD) memory card, or an optical recording medium (e.g., a compact disc (CD), a digital versatile disc (DVD)), or the storage area of a cloud server.
[0021] The storage device may be a combination of multiple of the various storage devices described above.
[0022] Various programs and data are stored in the storage device. Specifically, teacher data 101, judgment results 103, and edited teacher data 106 are stored in the storage device. Note that these data may be stored in separate storage devices or may be stored in a single storage device.
[0023] A communication device is a wired or wireless communication interface that enables communication with other devices via communication means such as a Local Area Network (LAN) or the Internet, and is, for example, a Network Interface Card (NIC), a wireless communication module, a Universal Serial Interface (USB) module, or a serial communication module.
[0024] An input device is a device that accepts input from a user, and examples of the input device include a keyboard, a mouse, a touch panel, a card reader, and a voice input device.
[0025] The output device is a device that provides various information such as the progress and results of processing to the user. Examples of the output device include a screen display device (such as a Liquid Crystal Display (LCD) or a Head Mounted Display (HMD)), an audio output device, or a printer. Note that the teacher data editing support system 1 may be configured to input and output information to and from other devices via a communication device.
[0026] The judgment unit 102 receives as input training data including discriminatory factors, which are variables that may cause discrimination, features, which are variables used for prediction, and the correct answer, and calculates the contribution, which is an index showing the degree to which the discriminatory factors contributed to the correct answer.
[0027] The display unit 104 visually displays evaluation information indicating the relationship between the degree of change in the correct answer in the training data, the degree of deviation from the initial value of the correct answer, and the degree of discrimination based on the contribution. The system user 107 confirms the displayed content.
[0028] The editing unit 105 receives a specification as to how much to change the correct answer, changes the correct answer in the teacher data based on the specification, and outputs the changed teacher data as edited teacher data 106.
[0029] FIG. 2 is a conceptual diagram illustrating an example format of training data. The training data includes a data ID 200, discriminatory factor information 201, an input feature 202, and a correct answer 203. The discriminatory factor information 201 includes information such as gender and age as discriminatory factors, which are variables that may cause discrimination. The input feature 202 is a variable used for prediction, and includes, for example, annual income (unit: 10,000 yen) and address. In this embodiment, the correct answer is the credit amount (unit: 10,000 yen), which is a one-dimensional value. Note that Example 1 is an example regarding regression.
[0030] 3 is a conceptual diagram illustrating an example of the format of the determination result. The determination result has a data ID, a discriminatory factor contribution 302, and an input feature contribution 303. The above-mentioned contribution degree corresponds to the discriminatory factor contribution 302 and the input feature contribution 303. The contribution degree is an index indicating the degree to which the discriminatory factor contributed to the correct answer.
[0031] Figure 4 is a conceptual diagram illustrating the format of edited teacher data. The format of the edited teacher data is basically the same as the format of the teacher data, but the correct answer values have been edited. The edited correct answer column is represented as edited correct answer 403.
[0032] 5 is a flowchart illustrating information processing performed by the determination unit 102. The determination unit 102 performs processing from step S102 to step S105 for each edit count, which means the number of times the correct answer value has been edited (a loop of steps S101 and S106).
[0033] The determination unit 102 performs the process of step S103 for each training data (a loop of steps S102 and S104). In step S103, the determination unit 102 calculates the contribution of the discriminant factors and feature amounts of the training data to the correct answer value.
[0034] An example of an algorithm for calculating the degree of contribution is the Shapley method. When the Shapley method is used, the determination unit 102 creates a prediction model from training data and calculates a Shapley value for the predicted value. In this case, the degree of contribution is the Shapley value for the discriminant factor in the correct answer. Alternatively, the algorithm for calculating the degree of contribution may be the CohortShapley method, which calculates the degree of contribution directly from training data. However, the algorithm for calculating the degree of contribution is not limited to this.
[0035] In step S105, the determination unit 102 edits the value of the correct answer based on the contribution of the discriminatory factor of each training data.
[0036] In addition, the determination unit 102 may consider subtracting the contribution value from the value indicating the correct answer as one edit, and may repeat this edit to calculate the degree of deviation from the initial value of the correct answer and the degree of discrimination based on the contribution value.
[0037] 6 is a flowchart illustrating information processing performed by the editing unit 105. In step S201, the editing unit 105 extracts correct answer information for a specified number of edits from the judgment history data, and overwrites the correct answer value.
[0038] Fig. 7 is a conceptual diagram illustrating an example of the format of the judgment history data. The judgment history data includes the discriminant factor contribution, the input feature contribution, and the edited answer for each edit. The edited answer value changes as the data is edited, but the data before and after the change can be saved.
[0039] 8 is a conceptual diagram showing a first display example by the display unit. Display unit 104 displays edit count 501, edit target 502, judgment start button 503, and data display area 504 that displays data. The user selects the edit count. The user also inputs items such as gender and age as the edit target. A table based on the original data of the training data is displayed in data display area 504. When the user presses judgment start button 503, the judgment process begins.
[0040] 9 is a conceptual diagram showing a second display example on the display unit. Display unit 104 of second screen 600 displays a pull-down selection box 601 for discrimination risk index, a pull-down selection box 602 for proofreading tendency information, and a pull-down selection box 603 for optimal number of edits. Display unit 104 also displays a graph 604 showing discrimination risk and proofreading tendency for each number of edits, an edited data output button 605, and a detailed report display button 606.
[0041] The user operates the discrimination risk index pull-down selection box 601 to select the discrimination risk index, such as "gender" or "age," that the user wants to display in the graph 604. The user operates the proofreading tendency information pull-down selection box 602 to select the proofreading tendency information, such as "gender" or "age," that the user wants to display in the graph 604.
[0042] Graph 604 displays the content selected in the pull-down selection box as a line. The horizontal axis of the graph is the number of edits. The solid broken line indicates the value of the discrimination risk index, and the dashed curved line indicates the value of the proofreading tendency information. Note that as the number of edits increases, the discrimination risk tends to decrease while the number of edits is small, and eventually the amount of reduction in discrimination risk decreases. As the number of edits increases, the value of the proofreading tendency information, i.e., the degree of deviation from the initial value of the correct answer, tends to increase.
[0043] Display unit 104 displays the degree of deviation from the initial value of the correct answer relative to the number of edits and the degree of discrimination based on the degree of contribution. The dashed curve in graph 604 indicates the degree of deviation from the initial value of the correct answer relative to the number of edits. The solid broken line in graph 604 indicates the degree of discrimination based on the degree of contribution relative to the number of edits.
[0044] When the user presses the edited data output button 605, the editing unit 105 executes processing for the number of edits selected in the optimal number of edits pull-down selection box 603. When the user presses the detailed report display button 606, details of the judgment result for the number of edits selected in the optimal number of edits pull-down selection box 603 are displayed.
[0045] Fig. 10 is a conceptual diagram showing a third display example by the display unit. When the user presses the detailed report display button 606 on the second screen 600 shown in Fig. 6, a third screen 700 is displayed. The third screen 700 displays a distribution display pull-down selection box 701, a pull-down selection box 702 for the optimal number of edits, a graph 703 showing the distribution of counts by contribution level, a table 704 of training data, a table 705 of contribution level information, and a table 706 of edited training data.
[0046] The user operates the distribution display pull-down selection box 701 to select the subject to be displayed in the graph 703, such as "gender" or "age." The horizontal axis of the graph 703 is the item selected in the distribution display pull-down selection box 701, and in this example, the contribution of "gender" is the horizontal axis. The vertical axis of the graph 703 is the number of cases.
[0047] The table for teacher data 704 displays the teacher data before editing. For example, the credit amount of teacher data with data ID = 1 is the original value of 500. The table for edited teacher data 706 displays edited teacher data corresponding to the number of edits selected in the optimal number of edits 702. In this example, edited control data for a single edit is displayed. After one edit, the credit amount of the teacher data with data ID = 1 is 470. This is because, when looking at the row with data ID 1 in the contribution information, the contributions calculated using the Shapley algorithm were +20 for gender and +10 for age, respectively, so +20 and +10 were subtracted from the original credit amount of 500. In other words, 500 - 20 - 10 = 470 is the credit amount for the row of edited teacher data with data ID 1. Furthermore, if we look at each row with data ID = 2, the original credit amount of the training data is 331, and since gender and age in the discriminant factor contribution are -10 and -20, respectively, the credit amount of the edited training data is 331-(-10)-(-20) = 361. [Example]
[0048] In Example 2, we will explain a case where the training data is training data for a classification problem. In this case, the training data is training data used in machine learning for a classification problem that classifies data into multiple categories, and the initial value of the correct answer is 1 for one of the categories and 0 for all other categories.
[0049] FIG. 11 is a conceptual diagram illustrating an example format of training data. The training data includes a data ID 200, discriminatory factor information 201, input features 202, and a correct answer 203. The discriminatory factor information 201 includes information such as gender and age as discriminatory factors, which are variables that may cause discrimination. The input features 202 are variables used for prediction, including annual income (unit: 10,000 yen) and address. The correct answer 203 is a one-hot vector based on one-hot encoding consisting of multiple categories.
[0050] FIG. 12 is a conceptual diagram illustrating an example of the format of the determination result. The determination result includes a data ID, a discriminatory factor contribution for each category, and an input feature contribution. The above-mentioned contribution degree corresponds to the discriminatory factor contribution and the input feature contribution. The contribution degree is an index indicating the degree to which the discriminatory factor contributed to the correct answer. For example, in one edit, the determination unit 102 calculates the contribution degree of the discriminatory factor for each category and subtracts the contribution degree of the discriminatory factor in that category from the value of the category in the correct answer.
[0051] FIG. 13 is a conceptual diagram illustrating the format of edited teacher data. The format of the edited teacher data is basically the same as the format of the teacher data, but the correct answer values have been edited. The edited correct answer column is shown as edited correct answers 403. Edited correct answers 403 also includes the correct answer values for each category.
[0052] (Generating Suggestions) Fig. 14 is a functional block diagram showing an example of the configuration of a teacher data editing support system. The configuration of the teacher data editing support system 1A shown in Fig. 14 is almost the same as the configuration of the teacher data editing support system 1 shown in Fig. 1, so only the differences will be explained.
[0053] The teacher data editing support system 1A includes a processing device. The processing device further realizes a suggestion generation unit 109 by reading and executing various programs and data stored in a storage device. The storage device further stores requirement information 108.
[0054] The suggestion generating unit 109 receives specification of requirement information that the contribution degree of the discriminatory factor should satisfy, and calculates the degree to which the contribution degree satisfies the requirement information each time editing is performed.
[0055] FIG. 15 is a conceptual diagram illustrating an example format of requirement information. Requirement information 108 has a requirement ID, a discriminatory factor, an input feature, and a correct answer. For each of the case information of the discriminatory factor, the input feature, and the correct answer, a condition for the contribution to satisfy the requirement condition is defined. For example, for requirement information with requirement ID 1, a requirement is defined that the factor contributions of "male" and "female" are less than 20. For requirement information with requirement ID 2, a requirement is defined that the age is over 60 and the factor contribution is less than 20. In the table shown in FIG. 15, Null indicates that no condition is set for that column.
[0056] The display unit 104 displays the degree of deviation from the initial value of the correct answer and the degree of discrimination based on the contribution level for the number of edits in which the degree of fulfillment of the requirement information exceeds a predetermined threshold. The degree of fulfillment of the requirement information means, for example, the degree to which multiple requirements are specified and how many of the multiple requirements are fulfilled. The degree may be the number of times or the rate at which the requirement is fulfilled.
[0057] FIG. 16 is a flowchart illustrating information processing performed by the suggestion generation unit. The suggestion generation unit 109 performs the process of step S302 for each number of edits (a loop of steps S301 and S303). In step S302, the suggestion generation unit 109 evaluates the degree to which the edited correct value satisfies the requirement information. Here, evaluation may mean calculation or computation. The suggestion generation unit 109 displays information about the number of edits that have a high degree of requirement satisfaction on the display unit (step S304).
[0058] Fig. 17 is a functional block diagram showing an example of the configuration of a teacher data editing support system. The configuration of teacher data editing support system 1B shown in Fig. 17 is almost the same as the configuration of teacher data editing support system 1 shown in Fig. 1, so only the differences will be explained.
[0059] The teacher data editing support system 1B includes a processing device. As described above, the processing device implements a determination unit 102, a display unit 104, and an editing unit 105 by reading and executing various programs and data stored in the storage device. Here, the determination unit 102 includes an alliance aggregation unit 110 and a contribution calculation unit 111.
[0060] The alliance counting unit 110 sets all subsets of a set whose elements are discriminatory factors and features as alliances, and for each of all alliances, identifies other teacher data whose elements included in the alliance in the teacher data are similar to that teacher data as similar data of the teacher data. The contribution calculation unit 111 calculates the average value of correct answers for similar data of the teacher data for each of all alliances, and calculates the difference between the average values of correct answers for each combination of two alliances whose only difference is the presence or absence of the discriminatory factor as a provisional contribution, and calculates the average value of the provisional contribution as the contribution of the discriminatory factor.
[0061] The criteria for determining similarity for the alliance counting unit 110 may be based on a threshold value, a match, or the like. For example, if an element included in an alliance in certain teacher data is a continuous value A, a similarity range can be set as a threshold. For example, values between continuous value A-100 and continuous value A+100 may be determined to be similar, and other values may be determined to be dissimilar. If an element included in an alliance in certain teacher data is a categorical value, it may be determined to be similar if the categories match. The criteria for determining similarity are not limited to those described above.
[0062] 18 is a flowchart illustrating information processing performed by the alliance counting unit 110. The alliance counting unit 110 generates all conceivable combinations of the discriminatory factors and the sum of the dimensions of the feature amounts as combination masks (S401). The combination masks will be described later with reference to FIG. 19.
[0063] The alliance tabulation unit 110 performs the processes from step S403 to step S406 for each training data (a loop of steps S402 and S407). The alliance tabulation unit 110 performs the processes of step S404 and step S405 for each alliance ID (a loop of steps S403 and S406).
[0064] In step S404, the alliance aggregation unit 110 extracts similar data based on the values of the discriminant factors and features to be included in the alliance (S404). Note that for discriminant factors and features that take continuous values, a threshold for determining a similar state may be determined in advance based on the distribution of values in the entire training data.
[0065] In step S405, the alliance tabulation unit 110 stores the ID information of the teacher data that is determined to be similar data as the alliance tabulation result. The alliance tabulation result will be described later with reference to FIG.
[0066] FIG. 19 is a conceptual diagram illustrating an example of the format of a combination mask. The combination mask has information items (columns) including an alliance ID 1600, a discriminatory factor mask 1601, and an input feature mask 1602. The alliance ID is identification information that uniquely identifies an alliance. The discriminatory factor mask 1601 includes items indicating discriminatory factors such as gender and age. The input feature mask includes items indicating input features such as annual income and address. A value of 0 or 1 is set in the combination mask. A value of 0 means that the item is not included in the alliance. A value of 1 means that the item is included in the alliance. For example, an alliance with an alliance ID of 2 includes the address item in the input feature mask 1602. In step S401, the alliance aggregation unit 110 generates all possible combinations in which the value of each column is 0 or 1.
[0067] 20 is a conceptual diagram illustrating an example of the format of the alliance tabulation result. The alliance tabulation result is data that stores as a history which other data has been extracted as similar data by each teacher data in each alliance pattern.
[0068] The alliance aggregation result has an alliance ID 1600 and a similar dataset 1700 as information items (columns). The alliance ID 1600 is the same as that described in FIG. 19, so a detailed description will be omitted. The similar dataset 1700 includes multiple types of data indicating which data each training data extracted as similar data in each alliance pattern. For example, data with alliance ID=1 and data ID=1 is data indicating that #5, #6, and so on have been extracted as similar data. Data with alliance ID=1 and data ID=2 is data indicating that #3, #8, and so on have been extracted as similar data.
[0069] FIG. 21 is a flowchart illustrating information processing performed by the contribution calculation unit.
[0070] Contribution calculation unit 111 performs the processes from step S502 to step S505 for each training data (loop of steps S501 and S506). Contribution calculation unit 111 performs the processes of step S503 and step S504 for each affiliation ID (loop of steps S502 and S505).
[0071] In step S503, the contribution calculation unit 111 calculates the average value of the correct answers of similar data for each data and each alliance. In step S504, the contribution calculation unit 111 calculates the difference in the average correct answer value from the difference between alliances as the discriminant factor and the provisional contribution of the input feature.
[0072] In step S507, the contribution calculation unit 111 calculates the contribution from the history of the provisional contribution of each discriminatory factor and input feature amount. For example, the contribution of the discriminatory factor and input feature amount is calculated by calculating the average value for all combination patterns of the affiliation ID.
[0073] FIG. 22 is a conceptual diagram illustrating an example format of the alliance aggregation result. The format of the alliance aggregation result is the same as that of the alliance aggregation result described with reference to FIG. 20. In the case of FIG. 20, similar data was extracted for each alliance ID and data ID. For example, the similar data for alliance ID=1 and data ID=1 are #5, #6, and so on (see FIG. 20). In step S503, the contribution calculation unit 111 calculates the average value between the correct value of data #5, the correct value of #6, and so on. For example, the average correct value for similar data for alliance ID=1 and data ID=1 is 231. By performing this average value calculation for each alliance ID and data ID, the correct average value result shown in FIG. 22 is calculated.
[0074] FIG. 23 is a conceptual diagram illustrating the format of the provisional contribution result. As described above, in step S504, the contribution calculation unit 111 calculates the difference between the average correct answer value and the difference between the alliances as the provisional contribution of the discriminant factor / input feature. FIG. 23 shows the calculated provisional contribution 2000.
[0075] In step S504, the contribution calculation unit 111 calculates the difference between, for example, the first data having alliance ID=1 and data ID=1 and the second data having alliance ID=2 and data ID=1. In the example shown, there is no difference between the first data and the second data in terms of gender, age, annual income, etc., so the difference value is 0. There is a difference between the first data and the second data in terms of address, so the difference value is -10. In other words, the difference between the correct average value 231 of the first data and the correct average value 221 of the second data is calculated. The calculated correct average value 221 of the second data - correct average value 231 of the first data = -10 indicates the provisional contribution degree due to including "address" in the alliance.
[0076] Similarly, the contribution calculation unit 111 calculates the difference between the second data with alliance ID=2 and data ID=1 and the third data with alliance ID=3 and data ID=1. In this case, since there is no difference between the second data and the third data in terms of gender, age, and address, the difference value is 0. Since there is a difference between the second data and the third data in terms of annual income, the difference value is +20.
[0077] Fig. 24 is a block diagram showing an example of the configuration of a teacher data editing support system. The functional units and data that make up the teacher data editing support system 1 may be integrated into one device, or may be distributed across multiple devices. Fig. 24 shows an example of a distributed arrangement.
[0078] The teacher data editing support system 1C shown in Fig. 24 includes a computer 100-1, a computer 100-2, and a computer 100-3. These computers are connected to each other so as to be able to communicate with each other via a communication line NW such as the Internet.
[0079] In the teacher data editing support system 1C, the computer 100-1 corresponds to a server. The computer 100-2 corresponds to a user terminal. The computer 100-3 corresponds to a data server. Each of the computers 100-1, 100-2, and 100-3 has a processing device and a storage device.
[0080] The processing device of computer 100-1 reads and executes various programs and data stored in the storage device, thereby realizing a judgment unit 102 and an editing unit 105. The storage device of computer 100-1 stores judgment results 103 and edited teacher data 106. The processing device of computer 100-2 reads and executes various programs and data stored in the storage device, thereby realizing a display unit 104. Teacher data 101 is stored in the storage device of computer 100-3.
[0081] FIG. 25 is a conceptual diagram showing an example of the hardware configuration of a computer. Computer 2500 corresponds to each of computers 100-1, 100-2, and 100-3 shown in FIG. 24. Computer 2500 has a processor 2501, a main memory device 2502, a secondary memory device 2503, and a network interface 2504. Processor 2501 corresponds to the above-mentioned processing device. Main memory device 2502 and secondary memory device 2503 correspond to the above-mentioned storage device. Network interface 2504 is a device for communicating with external devices, etc. via network NW shown in FIG. 24.
[0082] The above-described embodiments of the present invention are merely illustrative examples of the present invention, and are not intended to limit the scope of the present invention to these embodiments alone. Those skilled in the art can implement the present invention in various other forms without departing from the scope of the present invention.
[0083] As described above, the teacher data editing support system has a determination unit that inputs teacher data including discrimination factors, which are variables that may cause discrimination, features, which are variables used for prediction, and a correct answer, and calculates the contribution, which is an index showing the degree to which the discrimination factors contributed to the correct answer; a display unit that visibly presents evaluation information that represents the relationship between the degree to which the correct answer in the teacher data is changed, the degree of deviation from the initial value of the correct answer, and the degree of discrimination based on the contribution; and an editing unit that accepts a specification of how much to change the correct answer, changes the correct answer in the teacher data based on the specification, and outputs the changed teacher data.
[0084] A method for supporting editing of teacher data using an apparatus having a processing device includes a determination step of calculating a contribution level, which is an index showing the degree to which a discrimination factor contributed to a correct answer, in response to input of teacher data including a discrimination factor, which is a variable that may cause discrimination, a feature, which is a variable used for prediction, and a correct answer; a display step of visibly presenting evaluation information that represents the relationship between the degree to which the correct answer in the teacher data is changed, the degree of deviation from the initial value of the correct answer, and the degree of discrimination based on the contribution level; and an editing step of accepting a specification of how much to change the correct answer, changing the correct answer in the teacher data based on the specification, and outputting the changed teacher data.
[0085] The teacher data editing support program inputs teacher data including discriminatory factors, which are variables that may cause discrimination, features, which are variables used for prediction, and a correct answer, into a device having a processing device, and provides the following: a judgment function that calculates the contribution, which is an index showing the degree to which the discriminatory factors contributed to the correct answer; a display function that visibly presents evaluation information that shows the relationship between the degree to which the correct answer in the teacher data is changed, the degree of deviation from the initial value of the correct answer, and the degree of discrimination based on the contribution; and an editing function that accepts a specification of how much to change the correct answer, changes the correct answer in the teacher data based on the specification, and outputs the changed teacher data.
[0086] Based on the above, it is possible to help reduce discriminatory judgments made by machine learning models.
[0087] The contribution degree is a numerical value indicating the portion of the numerical value indicating the correct answer that is caused by a discriminatory factor, and the determination unit considers subtracting the contribution degree numerical value from the numerical value indicating the correct answer to be one edit, and repeats the process of performing edits and calculating the degree of deviation from the initial value of the correct answer and the degree of discrimination based on the contribution degree, and the display unit displays the degree of deviation from the initial value of the correct answer and the degree of discrimination based on the contribution degree for each edit. This makes it possible to visualize the degree of deviation from the initial value of the correct answer and the degree of discrimination based on the contribution degree according to the number of edits and provide them to the user.
[0088] The display unit displays a graph showing the degree of deviation from the initial value of the correct answer versus the number of edits, and the degree of discrimination based on the contribution versus the number of edits. This makes it possible to visualize the degree of deviation from the initial value of the correct answer and the degree of discrimination based on the contribution as a graph according to the number of edits, and provide it to the user.
[0089] The training data is training data used in machine learning for classification problems that classify data into multiple categories, and the initial value of the correct answer is a value of 1 for one of the categories and a value of 0 for all other categories. In one edit, the judgment unit calculates the contribution of the discriminatory factor for each category and subtracts the contribution of the discriminatory factor for that category from the value of the category in the correct answer. This makes it possible to help reduce discriminatory judgments made by machine learning models, even when using training data for classification problems.
[0090] The system further includes a suggestion generation unit that receives specification of requirement information that the contribution of the discrimination factor must satisfy, and calculates the degree to which the contribution satisfies the requirement information for each edit, and the display unit displays, for the number of edits in which the degree to which the requirement information is satisfied exceeds a predetermined threshold, the degree of deviation from the initial value of the correct answer for the number of edits, and the degree of discrimination based on the contribution. This makes it possible to visualize and present to the user the number of edits that have a high degree of satisfaction of the requirement information based on the specification of the requirement information.
[0091] The contribution is the Shapley value for the discriminatory factor in the correct answer, which can help reduce discriminatory decisions made by machine learning models based on the Shapley value.
[0092] The determination unit has an alliance aggregation unit that treats all subsets of a set whose elements are discriminatory factors and features as alliances, and for all alliances, identifies, for each of all teacher data, other teacher data that have similar elements included in the alliance in the teacher data as similar data of the teacher data, and a contribution calculation unit that calculates, for each alliance, the average value of correct answers for similar data of the teacher data for all teacher data, and for each discriminatory factor, calculates the difference between the average values of correct answers for each combination of two alliances in which the only difference is the presence or absence of the discriminatory factor as a provisional contribution, and calculates the average value of the provisional contributions as the contribution of the discriminatory factor, thereby making it possible to calculate the contribution taking into account the alliances. [Explanation of symbols]
[0093] 1...teacher data editing support system, 100...computer, 101...teacher data, 102...judgment unit, 103...judgment result, 104...display unit, 105...editing unit, 106...teacher data, 107...system user, 108...requirements information, 109...suggestion generation unit, 110...affiliation aggregation unit, 111...contribution calculation unit, 503...judgment start button, 504...data display area, 2500...computer, 2501...processor, 2502...main memory unit, 2503...secondary memory unit, 2504...network interface
Claims
1. a determination unit that receives training data including a discriminatory factor, which is a variable that may cause discrimination, a feature, which is a variable used for prediction, and a correct answer, and calculates a contribution, which is an index showing the degree to which the discriminatory factor contributed to the correct answer; a display unit that visibly presents evaluation information that indicates a relationship between a degree of change in the correct answer in the training data and a degree of deviation from an initial value of the correct answer, or a degree of discrimination based on the degree of contribution; an editing unit that receives a specification of how much to change the correct answer, changes the correct answer in the training data based on the specification, and outputs the training data after the change; A teacher data editing support system with
2. The contribution is a numerical value indicating a portion of the numerical value indicating the correct answer that is caused by the discriminatory factor, the determination unit performs the editing by subtracting the numerical value of the degree of contribution from the numerical value indicating the correct answer, and repeatedly calculates the degree of deviation from the initial value of the correct answer and the degree of discrimination based on the degree of contribution; the display unit displays a degree of deviation from an initial value of the correct answer with respect to the number of edits, and a degree of discrimination based on the degree of contribution. The teacher data editing support system according to claim 1 .
3. the display unit displays a graph showing a degree of deviation from the initial value of the correct answer relative to the number of edits, and a degree of discrimination based on the degree of contribution relative to the number of edits. The teacher data editing support system according to claim 2 .
4. The training data is training data used in machine learning for a classification problem of classifying data into a plurality of categories, and the initial value of the correct answer is a value of 1 for any one category and a value of 0 for all other categories; the determination unit calculates a contribution of a discriminatory factor for each category in the single edit, and subtracts the contribution of the discriminatory factor for the category from the value of the category in the correct answer; The teacher data editing support system according to claim 2 .
5. The suggestion generation unit further includes: a suggestion generation unit that receives specification of requirement information that the contribution of the discriminatory factor should satisfy, and calculates a degree to which the contribution satisfies the requirement information each time editing is performed; the display unit displays, for the number of edits in which the degree of fulfillment of the requirement information exceeds a predetermined threshold, a degree of deviation from the initial value of the correct answer for the number of edits and a degree of discrimination based on the degree of contribution. The teacher data editing support system according to claim 2 .
6. The contribution is a Shapley value for the discrimination factor in the correct answer. The teacher data editing support system according to claim 2 .
7. The determination unit an alliance counting unit that sets all subsets of a set whose elements are the discrimination factors and the features as alliances, and for all alliances, identifies, for each of all teacher data, other teacher data whose elements included in the alliance in the teacher data are similar to the teacher data, as similar data of the teacher data; a contribution calculation unit that calculates the average value of correct answers for similar data of all training data for each of all alliances, calculates the difference between the average values of correct answers for each of two combinations of alliances in which the only difference is the presence or absence of the training data, as a provisional contribution, and calculates the average value of the provisional contribution as the contribution of the training data; having The teacher data editing support system according to claim 6.
8. A method for supporting teacher data editing by a device having a processing device, comprising: The processing device Calculating a contribution level, which is an index showing the degree to which the discriminatory factor contributed to the correct answer, in response to input of training data including a discriminatory factor, which is a variable that may cause discrimination, a feature, which is a variable used for prediction, and a correct answer; visually presenting evaluation information that indicates the relationship between the degree of change in the correct answer in the training data and the degree of deviation from the initial value of the correct answer, or the degree of discrimination based on the degree of contribution; The processing device executes the steps of: accepting a specification of how much to change the correct answer; changing the correct answer in the training data based on the specification; and outputting the training data after the change. A method for supporting teacher data editing.
9. In an apparatus having a processing device, A training data set including a discriminatory factor, which is a variable that may cause discrimination, a feature value, which is a variable used for prediction, and a correct answer is input, and a contribution level, which is an index showing the degree to which the discriminatory factor contributed to the correct answer, is calculated; visually presenting evaluation information that indicates the relationship between the degree of change in the correct answer in the training data and the degree of deviation from the initial value of the correct answer, or the degree of discrimination based on the degree of contribution; accepting a specification of how much to change the correct answer, changing the correct answer in the training data based on the specification, and outputting the training data after the change; A teacher data editing support program to make this happen.
Citation Information
Patent Citations
Deep learning model depolarization method and device based on data enhancement
CN114638374A
Deep learning model depolarization method and device based on mask shielding
CN115131816A
Training data generation program, device, and method
WO2021260945A1
Information processing device, information processing method, computer program, imaging device, vehicle device, and medical robot device
WO2022123907A1