Data Conversion Program, Apparatus, and Method

By specifying differences and determining application probabilities for conversion rules, the technology minimizes distribution changes in data conversion, ensuring fair and accurate machine learning model predictions.

JP7705622B2Active Publication Date: 2025-07-10FUJITSU LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022038624
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-07-10
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

When data distribution changes significantly before and after data conversion for bias removal, the prediction accuracy of a machine learning model deteriorates.

Method used

The technology specifies the difference between pre-conversion and post-conversion data for each conversion rule, determines the application probability based on data bias and distance, and applies these rules to generate post-conversion data, minimizing distribution changes and ensuring fairness.

Benefits of technology

This approach effectively suppresses changes in data distribution during bias removal, maintaining prediction accuracy and enhancing interpretability of the data conversion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705622000003
    Figure 0007705622000003
  • Figure 0007705622000004
    Figure 0007705622000004
  • Figure 0007705622000005
    Figure 0007705622000005
Patent Text Reader

Abstract

To suppress changes in data distribution due to data conversion to remove bias.SOLUTION: A data conversion device identifies, for each of plural conversion rules, the distance between pre-conversion data and post-conversion data generated by applying the plural conversion rules respectively to the pre-conversion data, determines application probabilities of the plural conversion rules respectively, in order to perform data conversion that minimizes the distance between the pre-conversion data and the post-conversion data, which is data conversion for making the distribution fair in the number of data when sensitive attributes are used as a basis, and generates the post-conversion data by applying the plural conversion rules to the pre-conversion data based on the determined application probabilities.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed technology relates to a data conversion program, a data conversion device, and a data conversion method.

Background Art

[0002] The value of a specific attribute included in training data used for training a machine learning model may be biased, and the determination result by that machine learning model may be discriminatory. For example, assume a case where a machine learning model for predicting the pass / fail result from a person's attributes is trained using training data with the values of attributes such as a person's gender, age, place of origin, etc. as explanatory variables and the pass / fail result of that person for employment, testing, etc. as the objective variable. In this case, if past history where being female is disadvantaged in the pass / fail result is used as training data, the machine learning model trained using that training data will make discriminatory predictions that disadvantage women.

[0003] Techniques have been proposed to remove the above bias by converting data. For example, techniques have been proposed to convert data so that the data distribution is the same whether there is an attribute that may cause discriminatory behavior or not. Also, techniques have been proposed to convert data that conforms to a predetermined conversion rule according to that conversion rule. Further, techniques have been proposed to convert from an arbitrary data X1 to an arbitrary data X2 with a probability P(X1,X2) while imposing constraints to suppress the degree of change in the distribution.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

[0005] When the data distribution changes significantly before and after data conversion for bias removal, there is a problem that the prediction accuracy of a machine learning model trained using the converted data as training data deteriorates.

[0006] As one aspect, the disclosed technology aims to suppress changes in the distribution of data by data conversion for removing bias.

Means for Solving the Problem

[0007] As one aspect, for each of a plurality of conversion rules, the disclosed technology specifies the difference between the data before conversion and the data after conversion generated by applying each of the plurality of conversion rules to the data before conversion. Further, the disclosed technology determines the application probability of each of the plurality of conversion rules based on the bias of the first plurality of data when based on the first attribute of the first plurality of data and the difference of each of the plurality of conversion rules. Then, the disclosed technology applies the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data.

Advantages of the Invention

[0008] As one aspect, it has the effect of being able to suppress changes in the distribution of data by data conversion for removing bias.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0010] Hereinafter, an example of an embodiment according to the disclosed technology will be described with reference to the drawings.

[0011] First, before describing the details of the embodiment, the removal of bias by data conversion will be described.

[0012] The pre-conversion data 100 shown in FIG. 1 has "gender" and "employment" as attributes. The value of the attribute "gender" is 1 if the gender of the person corresponding to each data is male, and 0 if female. Also, the attribute "employment" is an attribute indicating the employability of the person corresponding to each data, and the value of the attribute "employment" is 1 if employment is possible and 0 if employment is not possible. The same applies to the post-conversion data 102. As shown in the upper part of FIG. 1, in the pre-conversion data 100, when gender = male, the probability of employment = possible is 2 / 3, and when gender = female, the probability of employment = possible is 1 / 3. Thus, in the pre-conversion data 100, the probability of employment = possible varies greatly depending on gender, that is, there is a bias. In this example, 2 / 3 - 1 / 3 = 1 / 3 corresponds to the amount of bias. A machine learning model trained with biased data may exhibit discriminatory behavior such as a large change in prediction depending on a sensitive attribute (here, gender). Therefore, by converting the data as shown in the lower part of FIG. 1 (the dashed line part in FIG. 1), in the post-conversion data 102, whether gender = male or gender = female, the probability of employment = possible is 2 / 3, the amount of bias is 2 / 3 - 2 / 3 = 0, and the bias due to gender is removed.

[0013] Here, as for the above-described data conversion, it is desirable that the distribution of the data does not change significantly before and after the conversion. This is because if the distribution changes significantly, the prediction accuracy of a machine learning model trained using the converted data as training data may deteriorate. Also, it is desirable that the data conversion is interpretable by humans, that is, it has interpretability. This is because if it is not interpretable, it becomes difficult to manually check the validity of the conversion for the converted data. As an interpretable data conversion, a method of converting data based on predetermined conversion rules can be considered. Therefore, in the present embodiment, data bias is removed from the data by data conversion based on conversion rules that suppress changes in the distribution of the data after conversion. Hereinafter, the data conversion device according to the present embodiment will be described in detail.

[0014] As shown in FIG. 2, a plurality of pre-conversion data and a plurality of conversion rules are input to the data conversion device 10. Then, the data conversion device 10 performs data conversion on the pre-conversion data and outputs the converted data. The data included in each of the pre-conversion data and the converted data includes values for each of a plurality of attributes, similar to the example of FIG. 1. In the present embodiment, the types of attributes include general attributes, target attributes, and sensitive attributes. The target attribute is an attribute that is the judgment result in a task using the data, such as "employment" in the above example. The sensitive attribute is an attribute that can be a bias, such as "gender" in the above example. The general attribute is an attribute other than the target attribute and the sensitive attribute, and is, for example, educational background, age, etc. There may be a plurality of general attributes in the data, but hereinafter, for simplicity of explanation, the case where there is one general attribute will be described.

[0015] Functionally, as shown in FIG. 2, the data conversion device 10 includes a specifying unit 12, a determining unit 14, a generating unit 16, and an output unit 18.

[0016] The specification unit 12 specifies, for each of the plurality of conversion rules, a distance (difference) between the pre-conversion data and the post-conversion data generated by applying each of the plurality of conversion rules to the pre-conversion data. Here, data X k Set the value of the general attribute to x k , the value of the target attribute is y k , the value of the sensitive attribute is s k Let data X k (x k ,y k ,s k The determination unit 12 expresses the arbitrary data X k =(x k ,y k ,s k ), and data X m =(x m ,y m ,s m ) about X k and X m Distance c(X k ,X m For example, the definition of distance c(X k ,X m ) is X k and X m It can be considered as the Euclidean distance with

[0017] X1=(20,1,1), X2=(50,1,1), c(X1,X2)=30 X1=(20,1,1), X3=(25,1,1), c(X1,X3)=5 In this case, the larger the distance, the more different the data is. For example, in the above example, data X2 is more different from data X1 than data X3. That is, this distance c(X k ,X m ) is the data X k Data X m The identification unit 12 calculates the distance c(X k ,X m ) to identify the

[0018] The determination unit 14 determines the application probability of each of a plurality of conversion rules based on the bias of the data when the sensitive attribute is used as a reference and the difference between the data before and after the conversion. Specifically, the determination unit 14 determines the application probability of each of the plurality of conversion rules such that the bias of the data before and after the conversion and the difference between the data before and after the conversion are minimized when the sensitive attribute is used as a reference. This will be described in more detail below.

[0019] A conversion rule is a rule for converting data that matches a condition into new data, and is expressed, for example, as follows. Conversion rule r = ((x’, y’, s’), (x”, y”, s”)) if (x, y, s) = (x’, y’, s’) return (x”, y”, s”) That is, (x’, y’, s’) is the condition and (x”, y”, s”) is the result of the conversion. However, x’, y’, and s’ may be specific values or wildcards “*” that match all values. x”, y”, and s” are only specific values and do not include wildcards.

[0020] The determination unit 14 obtains a set R of conversion rules r that match the data X = (x, y, s), and determines an application probability p(r) that represents the ratio of the data to which the conversion rule r ∈ R is applied out of the total number of data X. Here, in order to remove the bias from the data before conversion, it is necessary to perform data conversion so that the number of data in which the target attribute indicates a predetermined value is fair regardless of the value of the sensitive attribute in the data after conversion. For example, the number of data corresponding to the sensitive attribute and the target attribute is expressed as follows.

[0021]

Number

[0022] However, each (x, y, s) is a discrete value. Also, 1(y n = j) is y nA function that returns 1 when it is equal to j and 0 otherwise. That is, equation (1) represents the number of data in the dataset for which the target attribute has a predetermined value. Equation (2) represents the number of data in the dataset for which the sensitive attribute has a predetermined value. Equation (3) represents the number of data in the dataset for which the target attribute has a predetermined value and the sensitive attribute has a predetermined value.

[0023] Also, in order to perform fair data conversion, the probability that the value of the target attribute becomes a predetermined value does not change depending on the sensitive attribute. Therefore, in the converted data, data conversion may be performed so that the following equation (4) and the following equation (5) are equal, that is, so that the following equation (6) is satisfied.

[0024]

Equation

[0025] The determination unit 14 determines the application probability p(r) for each conversion rule so as to suppress the change in the distribution of the data before and after the conversion while performing the fair data conversion as described above. In the present embodiment, the problem of determining the application probability p(r) for each conversion rule is formulated as a minimum cost flow problem. Specifically, as shown in FIG. 3, the determination unit 14 creates a network including a source node, a plurality of first nodes, a plurality of second nodes, a plurality of third nodes, and a sink node. In FIG. 3, the source node is represented by a white circle, the sink node is represented by a shaded circle, the first node is represented by a solid white square with rounded corners, the second node is represented by a double-line white square with rounded corners, and the third node is represented by a solid shaded square with rounded corners. To the branches connecting the nodes (arrows in FIG. 3), the cost per data required to flow data through the branch and the capacity indicating the maximum number of data that can flow through the branch are set. In FIG. 3, the cost and capacity set for each branch are represented in the notation of (cost, capacity).

[0026] The source node corresponds to the supply point of the flow in the minimum cost flow problem, and the sink node corresponds to the demand point. The determination unit 14 flows the number of data included in the data set D (pre-conversion data) from the source node toward the sink node. Each of the first nodes is a node corresponding to each combination (x’, y’, s’) of the values of each attribute of the pre-conversion data. The determination unit 14 connects each of the source node and the first nodes with a branch, and sets (0, N x’y’s’ ) to that branch. N x’y’s’ is the number of data where x = x’, y = y’, and s = s’ among the data X = (x, y, s) included in the data set D.

[0027] Each of the second nodes is a node corresponding to each of the conversion rules r. The determination unit 14 connects with a branch the first node and the second node corresponding to the conversion rule r to which the data corresponding to the first node matches, and sets (c((x’, y’, s’), (x”, y”, s”)), ∞) to that branch. c((x’, y’, s’), (x”, y”, s”)) is the distance between the pre-conversion and post-conversion data by the conversion rule r corresponding to the second node connected by the branch.

[0028] The third node is a node corresponding to the group representing the pair of the value y of the target attribute and the value s of the sensitive attribute. The determination unit 14 connects with a branch the second node and the third node corresponding to the group to which the post-conversion data by the conversion rule r corresponding to the second node belongs, and sets (0, ∞) to that branch. Also, the determination unit 14 connects each of the third nodes and the sink node with a branch, and sets (0, N s” N y” / N) to that branch. The determination unit 14 sets the value of N s” N y” / N so that the post-conversion data is fair, specifically, so as to satisfy the above formula (6).

[0029] As described above, by setting nodes, branches, and the cost and capacity for each branch, the solution to the minimum-cost flow problem of this network represents the conversion process in which the dataset D becomes fair using the conversion rules, and represents the conversion that minimizes the change in the distribution before and after the conversion. The determination unit 14 extracts a flow for flowing the data included in the dataset D from the source node to the sink node at the minimum cost by solving the minimum-cost flow problem of the network as shown in FIG. 3. The flow is the number of data flowing through each branch. For example, assume that the conversion rules matching the data X = (a, 0, 0) are ri (i = 1, 2, 3, 4), and the flow is extracted as shown in FIG. 4. In FIG. 4, the flow flowing to the second node corresponding to the conversion rule ri is represented by fi. The determination unit 14 determines the application probability p(ri) of each conversion rule ri according to p(ri) = fi / Σfi so that Σ r∈R p(r) = 1.

[0030] The generation unit 16 applies a plurality of conversion rules to the pre-conversion data based on the application probabilities determined by the determination unit 14 to generate post-conversion data. In the case of the example in FIG. 4, the generation unit 16 applies the conversion rule r1 to the data X = (a, 0, 0) with an application probability of 0.1, applies the conversion rule r3 with an application probability of 0.75, and applies the conversion rule r4 with an application probability of 0.15 to generate post-conversion data. For example, when there are 10 pieces of data X = (a, 0, 0), the generation unit 16 applies the conversion rule r1 to 1 piece of the data X, applies the conversion rule r3 to 7 or 8 pieces, and applies the conversion rule r4 to 1 or 2 pieces to generate post-conversion data.

[0031] The output unit 18 outputs the plurality of post-conversion data generated by the generation unit 16. Further, the output unit 18 may also output the application probability for each conversion rule applied by the generation unit 16. Thereby, the interpretability of data conversion is further improved.

[0032] The data conversion device 10 may be implemented by, for example, a computer 40 shown in FIG. 5. The computer 40 includes a CPU (Central Processing Unit) 41, a memory 42 as a temporary storage area, and a non-volatile storage unit 43. The computer 40 also includes input / output devices 44 such as an input unit and a display unit, and an R / W (Read / Write) unit 45 that controls reading and writing of data to and from a storage medium 49. The computer 40 further includes a communication I / F (Interface) 46 connected to a network such as the Internet. The CPU 41, the memory 42, the storage unit 43, the input / output devices 44, the R / W unit 45, and the communication I / F 46 are connected to each other via a bus 47.

[0033] The storage unit 43 may be implemented by an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, or the like. A data conversion program 50 for causing the computer 40 to function as the data conversion device 10 is stored in the storage unit 43 as a storage medium. The data conversion program 50 has a specific process 52, a determination process 54, a generation process 56, and an output process 58.

[0034] The CPU 41 reads the data conversion program 50 from the storage unit 43 and expands it in the memory 42, and sequentially executes the processes included in the data conversion program 50. By executing the specific process 52, the CPU 41 operates as the specific unit 12 shown in FIG. 2. Also, by executing the determination process 54, the CPU 41 operates as the determination unit 14 shown in FIG. 2. Further, by executing the generation process 56, the CPU 41 operates as the generation unit 16 shown in FIG. 2. Additionally, by executing the output process 58, the CPU 41 operates as the output unit 18 shown in FIG. 2. As a result, the computer 40 that executes the data conversion program 50 functions as the data conversion device 10. Note that the CPU 41 that executes the program is hardware.

[0035] Incidentally, the functions realized by the data conversion program 50 can also be realized by, for example, a semiconductor integrated circuit, more specifically, an ASIC (Application Specific Integrated Circuit), etc.

[0036] Next, the operation of the data conversion device 10 according to this embodiment will be described. When a plurality of pre-conversion data and a plurality of conversion rules are input to the data conversion device 10, the data conversion process shown in FIG. 6 is executed in the data conversion device 10. Note that the data conversion process is an example of the data conversion method of the disclosed technology.

[0037] In step S10, the specifying unit 12 acquires a plurality of pre-conversion data and a plurality of conversion rules input to the data conversion device 10. Next, in step S12, the specifying unit 12 specifies the distance between the pre-conversion data and the post-conversion data generated by applying each of the plurality of conversion rules to the pre-conversion data for each of the plurality of conversion rules.

[0038] Next, in step S14, the determination unit 14 determines the application probability of each of the plurality of conversion rules such that the bias between the pre-conversion and post-conversion data and the distance between the pre-conversion and post-conversion data are minimized based on the sensitive attribute. Next, in step S16, the generation unit 16 applies the plurality of conversion rules to the pre-conversion data based on the application probability determined in step S14 above to generate post-conversion data. Next, in step S18, the output unit 18 outputs the plurality of post-conversion data generated in step S16 above, and the data conversion process ends.

[0039] As described above, for each of the plurality of conversion rules, the data conversion device according to the present embodiment specifies the distance between the pre-conversion data and the post-conversion data generated by applying each of the plurality of conversion rules to the pre-conversion data. Further, the data conversion device determines the application probability of each of the plurality of conversion rules based on the bias of the data when the sensitive attribute is used as a reference and the distance between the pre-conversion and post-conversion data. Then, the data conversion device applies the plurality of conversion rules to the pre-conversion data based on the determined application probability to generate post-conversion data. Thereby, the data conversion device can suppress the change in the distribution of data due to data conversion for removing bias.

[0040] Note that, in the above embodiment, the case where the minimum cost flow problem is applied to determine the application probability has been described, but it is not limited thereto. For example, the data conversion device may comprehensively specify the distance between the pre-conversion and post-conversion data for the distribution pattern of the number of data that results in fair data conversion, that is, satisfies the above equation (6), and determine the application probability for each conversion rule based on the pattern with the minimum distance. However, by applying the minimum cost flow problem as in the above embodiment, the application probability can be determined efficiently.

[0041] Also, in the above embodiment, the mode in which the data conversion program is pre-stored (installed) in the storage unit has been described, but it is not limited thereto. The program according to the disclosed technology can also be provided in a form stored in a storage medium such as a CD-ROM, DVD-ROM, or USB memory.

[0042] Regarding the above embodiments, the following supplementary notes are further disclosed.

[0043] (Supplementary Note 1) For each of the plurality of conversion rules, specify the difference between the pre-conversion data and the post-conversion data generated by applying each of the plurality of conversion rules to the pre-conversion data. Based on the bias of the first plurality of data when based on the first attribute of the first plurality of data and the difference of each of the plurality of conversion rules, determine the application probability of each of the plurality of conversion rules, Apply the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data. A data conversion program characterized by causing a computer to execute processing.

[0044] (Appendix 2) The bias is the bias of the number of the first plurality of data for each combination of the first attribute and the second attribute. The data conversion program according to Appendix 1.

[0045] (Appendix 3) Each of the plurality of conversion rules is represented by a combination of data before conversion and data after conversion. The process of determining the application probability includes determining the application probability based on the number of data when distributing the first plurality of data to each of the conversion rules to which the first plurality of data corresponds so that the bias and the difference are minimized. The data conversion program according to Appendix 2.

[0046] (Appendix 4) The process of determining the application probability includes determining so that the sum of the application probabilities of each of the plurality of conversion rules to which the first plurality of data corresponds becomes 1. The data conversion program according to Appendix 3.

[0047] (Appendix 5) The process of determining the application probability includes applying the minimum cost flow problem to a network including a source node, a first node corresponding to the first plurality of data, a second node corresponding to the plurality of conversion rules, a third node corresponding to a combination of the first attribute and the second attribute, and a sink node, a first branch connecting the source node and the first node and having the number of data corresponding to the first node as its capacity, a second branch connecting the first node and the second node and having, as its cost, the difference when the data corresponding to the first node is converted by the conversion rule corresponding to the second node, a third branch connecting the second node and the third node corresponding to the combination for the data after conversion indicated by the conversion rule corresponding to the second node, and a fourth branch connecting the third node and the sink node and having, as its capacity, the number of data set so that the bias becomes fair, and determining the application probability so that the bias and the difference are minimized. The data conversion program according to Appendix 3 or Appendix 4.

[0048] (Appendix 6) The process of generating the second plurality of data includes applying the conversion rule to a number of data among the first plurality of data according to the determined application probability for the conversion rule for each conversion rule to which the first plurality of data apply. The data conversion program according to any one of Appendices 1 to 5.

[0049] (Appendix 7) Outputting the generated second plurality of data and the application probability for each conversion rule. The data conversion program according to any one of Appendices 1 to 6, characterized in that the computer is caused to execute the process.

[0050] (Appendix 8) For each of the plurality of conversion rules, specifying the difference between the data before conversion and the data after conversion generated by applying each of the plurality of conversion rules to the data before conversion. Based on the bias of the first plurality of data when based on the first attribute, and the difference of each of the plurality of conversion rules, determine the application probability of each of the plurality of conversion rules, Apply the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data, A data conversion device characterized by including a control unit that executes processing.

[0051] (Appendix 9) The bias is the bias of the number of the first plurality of data for each combination of the first attribute and the second attribute. The data conversion device according to Appendix 8.

[0052] (Appendix 10) Each of the plurality of conversion rules is represented by a combination of data before conversion and data after conversion. The process of determining the application probability includes determining the application probability based on the number of data when distributing the first plurality of data to each of the conversion rules to which the first plurality of data corresponds so that the bias and the difference are minimized. The data conversion device according to Appendix 9.

[0053] (Appendix 11) The process of determining the application probability includes determining so that the sum of the application probabilities of each of the plurality of conversion rules to which the first plurality of data corresponds is 1. The data conversion device according to Appendix 10.

[0054] (Appendix 12) The process of determining the application probability includes applying a minimum cost flow problem to a network including a source node, a first node corresponding to the first plurality of data, a second node corresponding to the plurality of conversion rules, a third node corresponding to a combination of the first attribute and the second attribute, and a sink node, a first branch connecting the source node and the first node and having the number of data corresponding to the first node as a capacity, a second branch connecting the first node and the second node and having the difference when the data corresponding to the first node is converted by the conversion rule corresponding to the second node as a cost, a third branch connecting the second node and the third node corresponding to the combination for the converted data indicated by the conversion rule corresponding to the second node, and a fourth branch connecting the third node and the sink node and having the number of data set so that the bias becomes fair as a capacity, and determining the application probability so that the bias and the difference are minimized. The data conversion device according to Appendix 10 or Appendix 11.

[0055] (Appendix 13) The process of generating the second plurality of data includes applying the conversion rule to a number of data among the first plurality of data according to the determined application probability for each conversion rule to which the first plurality of data applies. The data conversion device according to any one of Appendices 8 to 12.

[0056] (Appendix 14) Outputting the generated second plurality of data and the application probability for each conversion rule. The data conversion device according to any one of Appendices 8 to 13, wherein the control unit executes the process.

[0057] (Appendix 15) For each of the plurality of conversion rules, identify the difference between the data before conversion and the data after conversion generated by applying each of the plurality of conversion rules to the data before conversion. Based on the bias of the first plurality of data when based on the first attribute, and the difference of each of the plurality of conversion rules, determine the application probability of each of the plurality of conversion rules, Apply the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data. A data conversion method characterized in that a computer executes the process.

[0058] (Appendix 16) The bias is the bias of the number of the first plurality of data for each combination of the first attribute and the second attribute. The data conversion method according to Appendix 15.

[0059] (Appendix 17) Each of the plurality of conversion rules is represented by a combination of data before conversion and data after conversion. The process of determining the application probability includes determining the application probability based on the number of data when distributing the first plurality of data to each of the conversion rules to which the first plurality of data belongs so that the bias and the difference are minimized. The data conversion method according to Appendix 16.

[0060] (Appendix 18) The process of determining the application probability includes determining so that the sum of the application probabilities of each of the plurality of conversion rules to which the first plurality of data belongs is 1. The data conversion method according to Appendix 17.

[0061] (Appendix 19) The process of determining the application probability includes applying a minimum cost flow problem to a network including a source node, a first node corresponding to the first plurality of data, a second node corresponding to the plurality of conversion rules, a third node corresponding to a combination of the first attribute and the second attribute, and a sink node, a first branch connecting the source node and the first node and having the number of data corresponding to the first node as a capacity, a second branch connecting the first node and the second node and having the difference when the data corresponding to the first node is converted by the conversion rule corresponding to the second node as a cost, a third branch connecting the second node and the third node corresponding to the combination for the converted data indicated by the conversion rule corresponding to the second node, and a fourth branch connecting the third node and the sink node and having the number of data set so that the bias becomes fair as a capacity, and determining the application probability so that the bias and the difference are minimized. The data conversion method according to Appendix 17 or Appendix 18.

[0062] (Appendix 20) For each of the plurality of conversion rules, identify the difference between the data before conversion and the data after conversion generated by applying each of the plurality of conversion rules to the data before conversion. Based on the bias of the first plurality of data when based on the first attribute of the first plurality of data and the difference of each of the plurality of conversion rules, determine the application probability of each of the plurality of conversion rules. Apply the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data. A non-transitory storage medium storing a data conversion program characterized by causing a computer to execute the process.

Explanation of Signs

[0063] 10 Data conversion device 12 Specifying unit 14 Determining unit 16 Generating unit 18 Output unit 40 Computer 41 CPU 42 Memory 43 Storage unit 44 Input / output device 45 R / W unit 46 Communication I / F 47 Bus 49 Storage medium 50 Data conversion program 52 Specific process 54 Decision process 56 Generation process 58 Output process

Claims

1. For each of a plurality of conversion rules, identify the difference between the data before conversion and the data after conversion generated by applying each of the plurality of conversion rules to the data before conversion, Based on the bias of the first plurality of data when the first attribute of the first plurality of data is used as a reference and the difference for each of the plurality of conversion rules, determine the application probability for each of the plurality of conversion rules, Apply the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data, A data conversion program characterized by causing a computer to execute the processing.

2. The bias is a bias in the number of the first plurality of data for each combination of the first attribute and the second attribute, The data conversion program according to Claim 1.

3. Each of the plurality of conversion rules is represented by a combination of data before conversion and data after conversion, The process of determining the application probability includes determining the application probability based on the number of data when distributing the first plurality of data to each of the conversion rules corresponding to the first plurality of data such that the bias and the difference are minimized. The data conversion program according to Claim 2.

4. The process of determining the application probability includes determining such that the sum of the application probabilities of each of the plurality of conversion rules corresponding to the first plurality of data is 1. The data conversion program according to Claim 3.

5. The process of determining the application probability includes applying a minimum cost flow problem to a network including a source node, a first node corresponding to the first plurality of data, a second node corresponding to the plurality of conversion rules, a third node corresponding to a combination of the first attribute and the second attribute, and a sink node, and a first branch connecting the source node and the first node and having the number of data corresponding to the first node as a capacity, a second branch connecting the first node and the second node and having, as a cost, the difference when the data corresponding to the first node is converted by the conversion rule corresponding to the second node, a third branch connecting the second node and the third node corresponding to the combination for the converted data indicated by the conversion rule corresponding to the second node, and a fourth branch connecting the third node and the sink node and having the number of data set so that the bias becomes fair as a capacity, and determining the application probability so that the bias and the difference are minimized. The data conversion program according to claim 3 or claim 4.

6. The process of generating the second plurality of data includes applying the conversion rule to a number of data among the first plurality of data according to the application probability determined for the conversion rule for each conversion rule to which the first plurality of data applies. The data conversion program according to any one of claims 1 to 5.

7. Outputting the generated second plurality of data and the application probability for each conversion rule. The data conversion program according to any one of claims 1 to 6, characterized in that the computer is caused to execute the process.

8. For each of the plurality of conversion rules, specifying the difference between the data before conversion and the data after conversion generated by applying each of the plurality of conversion rules to the data before conversion. Based on the bias of the first plurality of data with respect to the first attribute of the first plurality of data and the difference of each of the plurality of conversion rules, determining the application probability of each of the plurality of conversion rules. Applying the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data. A data conversion device characterized by including a control unit that executes the process.

9. For each of a plurality of conversion rules, identify a difference between data before conversion and data after conversion generated by applying each of the plurality of conversion rules to the data before conversion. Based on the bias of the first plurality of data when the first attribute of the first plurality of data is used as a reference and the difference of each of the plurality of conversion rules, determine the application probability of each of the plurality of conversion rules. Apply the plurality of conversion rules to the first plurality of data based on the application probability to generate a second plurality of data. A data conversion method, characterized in that a computer executes the processing.

Citation Information

Patent Citations

  • Data generation method, data generation program, and information processing device

    JP2021193532A

  • Machine learning model surety

    US20200387836A1

  • Method for detecting and mitigating bias and weakness in artificial intelligence training data and models

    US20220012591A1

  • Search control program, search control method, and search control device

    WO2020144842A1

  • Trajectory linking device, trajectory linking method, and non-temporary computer-readable medium with program stored therein

    WO2020261378A1