Annotation support program, method, and device

The annotation support device addresses the integration of diverse stakeholder preferences by calculating and proposing label changes based on preference rankings, ensuring a machine learning model reflects multiple perspectives effectively.

WO2025182042A1PCT designated stage Publication Date: 2025-09-04FUJITSU LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/007634
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing machine learning systems fail to integrate diverse stakeholder preferences, leading to low satisfaction and unclear reflection of individual preferences in annotation results, as they either ignore stakeholder differences or result in non-convergent annotations.

Method used

An annotation support device that acquires labels from multiple stakeholder groups, calculates preference rankings, and proposes label changes based on these rankings to align with stakeholder preferences, using a training unit, extraction unit, acquisition unit, calculation unit, and proposal unit to optimize annotation results.

Benefits of technology

Enables the creation of a machine learning model that reflects diverse stakeholder preferences, improving satisfaction by clearly explaining how annotations contribute to indicator changes, thus enhancing the model's effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024007634_04092025_PF_FP_ABST
    Figure JP2024007634_04092025_PF_FP_ABST
Patent Text Reader

Abstract

This annotation support involves acquiring annotation results obtained by aggregating, for each of a plurality of stakeholder groups, labels annotated by annotators included in the stakeholder group for a plurality of instances to be used for training a machine learning model and ranks of metrics preferred by each stakeholder group during annotation, and when stakeholder groups have different annotation results, proposing a label change to a stakeholder group satisfying a predetermined condition on the basis of the ranks of metrics preferred by each stakeholder group.
Need to check novelty before this filing date? Find Prior Art

Description

Annotation support program, method, and device

[0001] The disclosed technology relates to an annotation support program, an annotation support method, and an annotation support device.

[0002] As machine learning technology spreads throughout society, a variety of applications using machine learning models are becoming increasingly common. It is believed that having a human being act as an oracle is effective in training a machine learning model that is convenient for individuals. To achieve human-in-the-loop machine learning, where a human is the oracle, fields such as active learning and interactive machine learning exist. Active learning is a method in which a learning algorithm interactively queries the user and other information sources to prioritize, select, generate, and label data that is useful for learning. Interactive machine learning is a method that aims to provide a better user experience and more effective learning mechanisms by directly interacting with the user during the model building process.

[0003] Furthermore, as a technology related to annotation by an annotator, for example, a data processing device has been proposed that selects data awaiting audit from an annotated data set based on data annotation information and annotator information, and performs quality inspection on the data awaiting audit.

[0004] Also proposed is a system that collects attention information associated with user responses to consuming media content, where the attention information is used to create attention-labeled behavioral data about the user responses, and an attention model is generated by applying machine learning techniques to a set of attention-labeled behavioral data from multiple users. The system includes an annotation tool that uses the attention data to facilitate human labeling of the user responses.

[0005] Also proposed is a machine learning based visual device selection apparatus that considers a user's face in the context of a database of labeled faces and visual devices, each of the labeled images reflecting the aesthetic value of a proposed visual device for the patient's or user's face.

[0006] For example, a system has been proposed for scoring autonomous vehicle trajectories using rational crowd data. The system generates a set of trajectories for a vehicle operating in an environment. Each trajectory in the set of trajectories is associated with a traffic scenario. The system predicts a rationality score for each trajectory in the set of trajectories. The rationality score is obtained from a machine learning model trained using inputs obtained from multiple human annotators and a loss function that penalizes predictions of rationality scores that violate a rulebook structure.

[0007] Japanese Patent Application Publication No. 2022-77969 Japanese Patent Publication No. 2021-525424 Japanese Patent Application Publication No. 2022-531413 US Patent Application Publication No. 2022 / 0063666

[0008] Radwan, "Human Active Learning," Active Learning-Beyond the Future, IntechOpen, 26 August 2019.Fails, JA, and Olsen Jr, D. "Interactive Machine Learning," Proceedings of the 8th International Conference on Intelligent User Interfaces, Association for Computing Machinery, Pages 39-45, 12 January 2003.

[0009] In machine learning systems used for social decision-making, decisions made by a single machine learning model affect many people, including various stakeholders. Because the interests of these various stakeholders do not coincide, a framework in which each person trains a different machine learning model cannot reflect the different preferences of these various stakeholders in a single machine learning model.

[0010] For example, in active learning, annotators are considered as oracles with accurate information, and machine learning models are trained by treating all annotations as providing the same reliable correct answer, regardless of who the annotator is. In interactive machine learning, each user is treated as an annotator, and the machine learning model is adjusted to suit each individual user. Therefore, a separate machine learning model is created for each user, making it impossible to create a single machine learning model that integrates them.

[0011] One aspect of the disclosed technology is to propose appropriate and convincing annotations in accordance with the preferences of various stakeholders.

[0012] In one aspect, the disclosed technology acquires labels annotated by annotators from multiple stakeholder groups for multiple instances used to train a machine learning model. The disclosed technology also acquires annotation results by aggregating the labels for each stakeholder group and a ranking of indicators preferred by each stakeholder group when annotating. If the annotation results differ among the stakeholder groups, the technology proposes a change to the labels for the stakeholder groups that satisfy a predetermined condition based on the ranking of the indicators preferred by each stakeholder group.

[0013] One aspect is that it has the effect of being able to propose appropriate and convincing annotations in accordance with the preferences of various stakeholders.

[0014] FIG. 1 is a functional block diagram of an annotation support device. FIG. 2 is a diagram for explaining extraction of uncertain instances. FIG. 3 is a diagram for explaining label change proposal. FIG. 4 is a block diagram showing a schematic configuration of a computer functioning as an annotation support device. FIG. 5 is a flowchart showing an example of annotation support processing. FIG. 6 is a diagram showing an outline of an example of annotation results. FIG. 7 is a diagram for explaining calculation of preference ranking. FIG. 8 is a diagram for explaining setting of a best pattern. FIG. 9 is a flowchart showing an example of proposal processing. FIG. 10 is a diagram for explaining determination of a decided label. FIG. 11 is a diagram for explaining determination of a proposed label. FIG. 12 is a diagram showing an example of proposal content.

[0015] Hereinafter, an example of an embodiment of the disclosed technology will be described with reference to the drawings.

[0016] In this embodiment, it is assumed that an annotator views the labels output by the machine learning model, corrects the labels, and retrains the machine learning model based on the corrected labels.

[0017] In such cases, one easy way to build a machine learning model that reflects the preferences of various stakeholders is to have multiple annotators decide on labels by majority vote. For example, in this method, multiple annotators annotate the same instance, and then the majority vote determines whether the instance is a positive or negative example. Then, retraining data is created using the results of the majority vote, and the machine learning model is retrained.

[0018] However, this method does not take into account the preferences of each stakeholder group. Furthermore, neither the annotators nor the machine learning model can determine how the preferences of each stakeholder were taken into account and reflected in the final machine learning model. This results in low satisfaction from the annotators, and there is also no explanation for how diverse preferences were taken into account.

[0019] Another easily conceivable method is to have annotators repeatedly revise labels until all annotators' results are available. For example, with this method, multiple annotators annotate the same instance. If different annotations are made for the same instance by different annotators, the instance is requested to be re-annotated, and this process is repeated until all annotation results for all instances are available. However, with this method, the annotations made by the annotators may end up being the same no matter how many times they are made, and the annotation results may not converge.

[0020] An instance is an individual decision contained in data. For example, an instance for a task of determining whether to approve a loan application using a machine learning model is the result of a loan application approval / disapproval decision for each customer based on the attribute information of each customer. For example, an instance for a task of predicting recidivism is the prediction result of whether each suspect will reoffend based on the attribute information of each suspect. For example, an instance for a task of diagnosing a disease is the diagnosis result for each patient based on the attribute information of each patient.

[0021] An annotator is a person who looks at the labels output by a machine learning model and re-labels instances. Annotation is the act of labeling instances. A stakeholder is a person who has a specific interest in decisions made in a system that includes a machine learning model. In the example of using a machine learning model to determine whether or not to approve a loan application, these would include bank staff, customers, and auditing institutions. A stakeholder group is a group of people who share the same interests. In the following, a stakeholder group may simply be referred to as a "group."

[0022] In consideration of the problems with the above-mentioned easily conceivable methods, this embodiment makes appropriate proposals that are most convincing to all stakeholders and encourage the stakeholders to agree on the annotated labels. The annotation support device according to this embodiment will be described in detail below.

[0023] 1 , the annotation support device 10 functionally includes a training unit 12, an extraction unit 14, an acquisition unit 16, a calculation unit 18, and a proposal unit 20. The acquisition unit 16 and the calculation unit 18 are examples of the “acquisition unit” of the disclosed technology.

[0024] The training unit 12 acquires a set of instances input to the annotation support device 10 and trains a machine learning model using the acquired set of instances. For example, when training a machine learning model using a support vector machine (SVM), the training unit 12 determines a decision boundary that best classifies positive examples and negative examples for each instance distributed in a feature space in which each attribute included in the instance is set on each axis.

[0025] The extraction unit 14 extracts instances within a predetermined range from the decision boundary of the machine learning model as instances to be annotated by an annotator. Hereinafter, the instances extracted by the extraction unit 14 are referred to as "uncertain instances." Instances that exist near the decision boundary of the machine learning model are instances that are evaluated with low certainty by the machine learning model, and are instances for which annotation by an annotator is effective.

[0026] Figure 2 shows a schematic diagram of the extraction of uncertain instances in a machine learning model using an SVM. In Figure 2, white circles represent positive example instances, and diagonally shaded circles represent negative example instances. Note that for simplicity of explanation, Figure 2 shows an example of a two-dimensional feature space, but the feature space may be three or more dimensions depending on the number of attributes included in the instances. In this example, 10 instances present within a certain distance from the decision boundary (the range indicated by the dashed line in Figure 2) are extracted as uncertain instances (instances within the dashed-dotted line in Figure 2).

[0027] The acquiring unit 16 acquires labels annotated by annotators included in each of the multiple stakeholder groups for the uncertain instances extracted by the extracting unit 14. The acquiring unit 16 aggregates the acquired labels for each stakeholder group and acquires them as the annotation result for each stakeholder group. Specifically, the acquiring unit 16 acquires, as the annotation result for each stakeholder group, a label obtained by majority vote of the labels annotated by each annotator for each instance within each stakeholder group.

[0028] The calculation unit 18 calculates the ranking of the indicators preferred by each of the stakeholder groups when annotating.

[0029] In this embodiment, it is assumed that the preferences of annotators within each stakeholder group are similar, while annotators belonging to different stakeholder groups have different preferences. In an example where a machine learning model is used to determine whether or not to grant a loan application, stakeholder groups include loan appraisers, audit organizations, and loan applicants. It is assumed that loan appraisers have a preference for making accurate decisions, audit organizations have a preference for fair loan decisions, and loan applicants have a preference for not being mistakenly judged as denied.

[0030] Therefore, in this embodiment, the values ​​of each stakeholder group are expressed as preferences for machine learning indices. For example, in the example of using the above machine learning model to determine whether or not to approve a loan application, it is assumed that the loan examiner wants to improve accuracy, the audit organization wants to improve fairness, and the loan applicant wants to reduce the false negative rate.

[0031] Based on the above, the calculation unit 18 calculates the ranking of preferred indicators for each stakeholder group (hereinafter referred to as "preference ranking") based on the annotation results for each stakeholder group. Specifically, the calculation unit 18 introduces the concept of a utility value that indicates the degree to which each annotator prefers a certain annotation set. An annotation set is a set of labels annotated for each uncertain instance. The utility value is expressed using preference values ​​that indicate the degree to which each annotator prefers each individual machine learning indicator (such as accuracy, fairness, and false negative rate) and the values ​​of each indicator.

[0032] For example, the calculation unit 18 determines the utility value as the inner product of the vector of preference values ​​for each index and the vector of the values ​​of each index, as shown in the following equation (1): V i =Σβ j x j +α (1) V i is the utility value for the annotation set i annotated by the annotator, β j is the preference value for the jth index, x j is the value of the jth index, and α is a constant term. j is adjusted so that the better the index, the larger the value. j may be normalized, for example, adjusted to a value between 0 and 1.

[0033] The calculation unit 18 calculates the probability P(V) that the annotation set i is selected from all the annotation sets assumed for the uncertain instance, as shown in the following formula (2): i ) is maximized by maximum likelihood estimation. j Estimate argmax βj P (V i ) (2)

[0034] The calculation unit 18 determines the preference ranking for each stakeholder group based on the estimated preference value for each index. Specifically, the calculation unit 18 calculates the β j The above are summed up, and the preference ranking is calculated as 1st, 2nd, etc., starting with the index with the largest total.

[0035] When the annotation results differ between stakeholder groups, the proposing unit 20 proposes a label change to stakeholder groups that satisfy a predetermined condition based on the preference ranking of each stakeholder group. A stakeholder group that satisfies the predetermined condition is a stakeholder group whose annotation result label differs from a proposed label (described in detail below, hereinafter referred to as a "proposed label").

[0036] Specifically, the proposing unit 20 sets, for each stakeholder group, a label pattern of each uncertain instance that optimizes each of the indicators as the best pattern. For example, the proposing unit 20 sets, as the best pattern, a label set that minimizes the difference from the annotation result and optimizes the indicators.

[0037] The proposing unit 20 determines instances for which annotation results differ between stakeholder groups as instances for which a label change is proposed. The proposing unit 20 also determines proposed labels based on the preference rankings and best patterns of each stakeholder group.

[0038] In this embodiment, the system encourages each stakeholder group to make a satisfactory annotation change based on the preference ranking of each stakeholder group. Specifically, if the annotation result satisfies the preference ranking of each stakeholder group but differs between stakeholder groups, it is desirable to select and propose a label that can be changed without changing the satisfying preference ranking. Furthermore, if the original annotation result does not satisfy the indicator of the stakeholder group's top preference ranking, it is desirable to propose how to change the annotation to optimize the indicator of the top preference ranking.

[0039] Therefore, for example, the proposing unit 20 determines, as the proposed label, a candidate that minimizes the sum of the preference rankings of the stakeholder group for the indicator when the proposed label candidate matches the label of the best pattern. Furthermore, when the sum of the preference rankings is the same, the proposing unit 20 determines, as the proposed label, a candidate that minimizes the change from the annotation result.

[0040] Furthermore, the proposing unit 20 identifies a reason for encouraging a change to the proposed label based on the proposed label, the preference rankings of each stakeholder, and the best pattern. Specifically, the proposing unit 20 identifies a reason for encouraging a change based on whether the proposed label is a label that minimizes the sum of the preference rankings or a label that minimizes the change, and whether the change is to a label that optimizes the index ranked first in the preference index of the stakeholder group.

[0041] A more specific description will be given with reference to Figure 3. The example in Figure 3 shows the annotation results for each of groups 1 to 3, the preference ranking for each group, and the best pattern for instances with instance IDs (identifiers of the instances) ranging from 1 to 5. The label "1" indicates a positive example, and the label "0" indicates a negative example. Below, an instance with an instance ID of a will be referred to as "instance a." The instance ID may also be simply referred to as "ID."

[0042] In the example of Fig. 3, for instance, if the label "1" that minimizes the sum of preference rankings is determined as the proposed label for instance 1, group 2 will be prompted to change to label "1." In this case, since the change is to a label that optimizes index 2, which has the highest preference ranking in group 2, the suggestion unit 20 identifies the reason as being a change to a label that optimizes the index with the highest preference ranking.

[0043] Furthermore, for example, if the label "1" that minimizes the sum of preference rankings is determined as the proposed label for instance 4, group 1 will be prompted to change to label "1." In this case, the change is to a label that optimizes index 2, which is ranked second in preference ranking for group 1. Therefore, taking into account the annotations of other stakeholder groups, the proposing unit 20 determines that the reason for the change is that the index ranked first in preference ranking cannot be optimized, but the label should be changed to optimize an index ranked in another order (second in this example).

[0044] Furthermore, although not shown in the figure, if the proposed label is the label that will cause the least change, the proposed label is often selected in the annotations of other stakeholder groups, which is why this is identified.

[0045] The proposing unit 20 generates a proposed message based on the proposed label and the identified reason, for example, by using a template. The proposing unit 20 transmits the generated message to the target stakeholder group. Note that the target to which the message is sent may be an information processing device used by a representative of the target stakeholder group, or may be an information processing device used by each annotator belonging to the target stakeholder group. When transmitting the message to the annotator's information processing device, the label annotated by each annotator acquired by the acquiring unit 16 may be referenced, and an annotator who has annotated a label different from the proposed label may be selected and transmitted.

[0046] Furthermore, when the annotation results repeatedly obtained have become stable from the previous annotation result, the proposing unit 20 determines whether any annotations of the uncertain instances are consistent among the stakeholder groups. If there is a mismatch, the proposing unit 20 automatically changes the annotation. The target to be changed and the label to be changed are determined in the same manner as in the proposed method described above.

[0047] The annotation support device 10 may be realized by, for example, a computer 40 shown in Fig. 4. The computer 40 includes a CPU (Central Processing Unit) 41, a GPU (Graphics Processing Unit) 42, a memory 43 as a temporary storage area, and a non-volatile storage device 44. The computer 40 also includes an input / output device 45 such as an input device and a display device, and an R / W (Read / Write) device 46 that controls reading and writing of data from and to a storage medium 49. The computer 40 also includes a communication I / F (Interface) 47 that is connected to a network such as the Internet. The CPU 41, GPU 42, memory 43, storage device 44, input / output device 45, R / W device 46, and communication I / F 47 are connected to one another via a bus 48.

[0048] The storage device 44 is, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage device 44 serving as a storage medium stores an annotation support program 50 for causing the computer 40 to function as the annotation support device 10. The annotation support program 50 includes a training process control command 52, an extraction process control command 54, an acquisition process control command 56, a calculation process control command 58, and a proposal process control command 60.

[0049] The CPU 41 reads the annotation support program 50 from the storage device 44, loads it into the memory 43, and sequentially executes the control instructions of the annotation support program 50. The CPU 41 operates as the training unit 12 shown in FIG. 1 by executing the training process control instruction 52. The CPU 41 operates as the extraction unit 14 shown in FIG. 1 by executing the extraction process control instruction 54. The CPU 41 operates as the acquisition unit 16 shown in FIG. 1 by executing the acquisition process control instruction 56. The CPU 41 operates as the calculation unit 18 shown in FIG. 1 by executing the calculation process control instruction 58. The CPU 41 operates as the proposal unit 20 shown in FIG. 1 by executing the proposal process control instruction 60. As a result, the computer 40 that executes the annotation support program 50 functions as the annotation support device 10. The CPU 41 that executes the program is hardware. A part of the program may be executed by the GPU 42.

[0050] The functions realized by the annotation support program 50 may be realized by, for example, a semiconductor integrated circuit, more specifically, an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or the like.

[0051] Next, the operation of the annotation support device 10 according to this embodiment will be described. Here, in an example where a loan application is approved or rejected using a machine learning model, the stakeholder groups are loan appraisers, audit organizations, and loan applicants, and accuracy, fairness, and false negative rate are used as machine learning indices.

[0052] When an instance set is input to the annotation support device 10, the annotation support process shown in Fig. 5 is executed in the annotation support device 10. Note that the annotation support process is an example of an annotation support method of the disclosed technology.

[0053] In step S10, the training unit 12 acquires the set of instances input to the annotation support device 10 and trains the machine learning model using the acquired set of instances. For example, in the case of a machine learning model using an SVM as shown in FIG. 2, attributes such as the applicant's annual income and the amount of default payment are set on each axis of the feature space, and instances in which the loan application is approved are positive examples, and instances in which the loan application is rejected are negative examples. Next, in step S12, the extraction unit 14 extracts uncertain instances that exist within a predetermined range from the decision boundary of the machine learning model as instances to be annotated by an annotator.

[0054] Next, in step S14, the acquisition unit 16 presents the uncertain instances to annotators included in each of the multiple stakeholder groups, and acquires labels annotated for the uncertain instances, for example, as shown in FIG. 6. The example in FIG. 6 shows an example in which labels annotated by five annotators included in each stakeholder group are acquired for ten uncertain instances. Note that at this stage, the "Result" column in FIG. 6 is blank. Furthermore, each "Amn" column indicates a label by the nth annotator of group m. Hereinafter, the nth annotator will be referred to as "annotator n."

[0055] Next, in step S16, the acquisition unit 16 acquires the label obtained by majority vote among the labels annotated by each annotator for each instance within each stakeholder group as the annotation result for that stakeholder group. The "Result" column in Fig. 6 corresponds to the annotation result.

[0056] Next, in step S18, the calculation unit 18 calculates the preference ranking of each stakeholder group. As described above, the calculation unit 18 estimates preference values ​​by maximum likelihood estimation. Maximum likelihood estimation is a method in which a person (an annotator in this embodiment) selects one item from a plurality of items and determines the person's preferences based on that selection. Here, it is assumed that there are 10 uncertain instances, and the annotator selects annotation set i from all possible annotation sets (annotation sets). Since two labels, 0 or 1, are assigned to the 10 uncertain instances, there are potentially two possible values. 10 = 1024 possible annotation sets. Let us consider that annotator n selects one annotation set i from 1024 items. For each annotation set, an index value can be defined, and this value is considered an attribute of each annotation set (item).

[0057] When the above formula (1) is applied to this example, the utility value formula shown in the following formula (3) can be formulated. V n,i = β n,acc ・x i,acc +β n,nfr ・x i,nfr +β n,fair ・x i,fair +α (3) V n,i is the utility value for annotation set i annotated by annotator n. n,j is the preference value of annotator n for index j, x i,j is the value of index j for annotation set i, and α is a constant term. In equation (3), j is acc, nfr, and fair, which represent the indices of accuracy, false negative rate, and fairness, respectively.

[0058] The calculation unit 18 calculates V n,iUsing the above, the probability that annotator n selects annotation set i is P n,i is calculated by the following formula (4): Note that the reason why V is placed on the shoulder of the base of the natural logarithm is because of the probability distribution (Gumbel distribution) used.

[0059]

[0060] The calculation unit 18 calculates the probability P n,i are summed by annotators belonging to the same stakeholder group, and the probability P n,i The set of β that maximizes the sum of argmax is found by maximum likelihood estimation. Note that the annotation set i is different for each annotator n because each annotator uses a different annotation method. The set of β obtained by this maximum likelihood estimation is the preference value that best matches the selection of the stakeholder group. argmax βj Σ n P n (5)

[0061] As shown in FIG. 7, the calculation unit 18 calculates the preference ranking of each index calculated for each stakeholder group in descending order of preference value β, such as 1st, 2nd, . . .

[0062] Next, in step S20, the proposing unit 20 sets, for each stakeholder group, a set of labels for each uncertain instance that optimizes each of the indicators as the best pattern. The best pattern is a set of labels set so that the value of an indicator that should be increased, such as accuracy and fairness, is maximized, and the best pattern is a set of labels set so that the value of an indicator that should be decreased, such as the false negative rate, is minimized.

[0063] As shown in FIG. 8 , for accuracy, a unique best pattern is determined for each uncertain instance, so the proposing unit 20 sets the same best pattern for each stakeholder group. On the other hand, for false negative rate, the false negative rate may decrease for each label, so a unique best pattern cannot be determined. Furthermore, for fairness, a unique best pattern cannot be determined because there are multiple patterns of the best label set. In such cases, the proposing unit 20 sets the label set that is the smallest difference from the annotation results of each stakeholder group and that produces the best result for each indicator as the best pattern for that indicator for that stakeholder group. In FIG. 8 , the shaded labels are labels that cannot be uniquely determined, so the labels of the annotation results of each stakeholder group are set.

[0064] Next, in step S22, the proposing unit 20 determines whether the annotation result acquired in step S16 has changed from the previous annotation result. If there has been a change, the process proceeds to step S30, and if there has not been a change, the process proceeds to step S50.

[0065] In step S30, a proposal process is executed, which will now be described with reference to FIG.

[0066] In step S32, the proposing unit 20 selects one uncertain instance. Next, in step S34, the proposing unit 20 determines whether the annotation results of all stakeholder groups for the selected uncertain instance are the same. If they are the same, the proposing unit 20 determines the labels annotated by all stakeholder groups for that uncertain instance as determined labels, as shown in the upper diagram of Figure 10, and proceeds to step S36. On the other hand, if the annotation results of each stakeholder group include different labels, the proposing unit 20 proceeds to step S38.

[0067] In step S36, the proposing unit 20 determines whether or not the determined labels have been determined for all uncertain instances. If there are uncertain instances for which the determined labels have not been determined, the process returns to step S32. On the other hand, as shown in the lower diagram of FIG. 10, if the determined labels have been determined for all uncertain instances, a consensus has been reached among the stakeholder groups for all uncertain instances, and the annotation support process ends.

[0068] In step S38, the sum of the preference rankings of each stakeholder group for the indicators that match the labels of the best pattern when the labels of all stakeholder groups are set to 0 or 1 is calculated, i.e., the sum of the preference rankings that are satisfied for the labels of 0 or 1. Next, in step S40, the proposing unit 20 determines whether the sum of the preference rankings when the labels of all stakeholder groups are set to 0 is the same as the sum of the preference rankings when the labels of all stakeholder groups are set to 1. If there is a tie, the process proceeds to step S42, and if there is no tie, the process proceeds to step S44.

[0069] For example, if all the labels of instance 2 shown in FIG. 11 are set to 0, no label exists that matches the best pattern in any of the groups, and therefore the sum of the preference rankings is calculated to be an extremely large value (e.g., infinity). On the other hand, if all the labels are set to 1, the preference ranking matches the label that provides the best result for the index with the highest preference ranking in all of groups 1, 2, and 3, and therefore the suggestion unit 20 calculates the sum of the preference rankings as 1 + 1 + 1 = 3. Note that if there are multiple indexes in each group that match 0 or 1 with the label of the best pattern, the preference ranking with the highest ranking is adopted. In this case, in step S44, the suggestion unit 20 determines the label "1" with the smaller sum of preference rankings as the proposed label.

[0070] Also, for example, if all the labels of instance 3 shown in FIG. 11 are set to 0, in group 1, it matches the label that provides the best result for accuracy with a preference ranking of third place. In group 2, it matches the label that provides the best result for fairness with a preference ranking of first place. In group 3, it matches the label that provides the best result for accuracy with a preference ranking of first place. Therefore, the proposer 20 calculates the sum of preference rankings as 3 + 1 + 1 = 5. On the other hand, if all the labels are set to 1, in group 1, it matches the label that provides the best result for false negative rate with a preference ranking of first place. In group 2, there is no label that matches the best pattern. In group 3, it matches the label that provides the best result for false negative rate with a preference ranking of second place. Therefore, the proposer 20 calculates the sum of preference rankings as 1 + ∞ + 2 = ∞. In this case, in step S44, the proposer 20 determines the label "0" with the smaller sum of preference rankings as the proposed label.

[0071] Furthermore, for example, if all the labels of instance 7 shown in FIG. 11 are set to 0, in group 1, it matches the label that provides the best result for fairness with a preference ranking of second place. In group 2, it matches the label that provides the best result for accuracy with a preference ranking of second place. In group 3, it matches the label that provides the best result for accuracy with a preference ranking of first place. Therefore, the suggestion unit 20 calculates the sum of the preference rankings as 2 + 2 + 1 = 5. On the other hand, if all the labels are set to 1, in group 1, it matches the label that provides the best result for false negative rate with a preference ranking of first place. In group 2, it matches the label that provides the best result for fairness with a preference ranking of first place. In group 3, it matches the label that provides the best result for fairness with a preference ranking of third place. Therefore, the suggestion unit 20 calculates the sum of the preference rankings as 1 + 1 + 3 = 5.

[0072] In this case, since the sums of the preference rankings are the same, the suggestion unit 20 determines the label that requires the least change from the current annotation result as the proposed label in step S42. For the annotation result of instance 7, if all values ​​are set to 0, two annotations will be changed, but if all values ​​are set to 1, only one annotation will be changed, so the suggestion unit 20 determines the proposed label to be "1".

[0073] In this example, when the number of groups is odd, the label with the least change from the current annotation result is uniquely determined, but when the number of groups is even, it may not be uniquely determined. In such cases, the proposed label may be determined to be either 0 or 1, or 0 or 1 may be selected randomly and determined as the proposed label.

[0074] In FIG. 11, among the labels of the best pattern that match 0, the label with the highest preference ranking in each group is shown shaded, and among the labels of the best pattern that match 1, the label with the highest preference ranking in each group is shown framed in bold.

[0075] Next, in step S46, the proposing unit 20 identifies the reason for prompting a change to the proposed label based on the proposed label, the preference ranking of each stakeholder, and the best pattern, and temporarily stores the reason together with the proposed label in association with the instance ID. Next, in step S48, it is determined whether all uncertain instances have been selected in step S32. If there are any unselected uncertain instances, the process returns to step S32; if all have been selected, the process proceeds to step S50.

[0076] In step S50, the proposing unit 20 generates a proposed message based on the proposed label stored in step S46 and the reason identified, for example, by using a template, and then sends the generated message to the target stakeholder group, returns to the annotation support process ( FIG. 5 ), and returns to step S14.

[0077] FIG. 12 shows an example of the annotation results and proposed labels at a certain stage, along with the proposed content. In FIG. 12, the shaded labels are the labels for which a change is proposed. When proposing instance 2 to group 1, the proposed label is the one with the smallest sum of preference rankings, and is a proposed change that will optimize the false negative rate for the first preference ranking in group 1. Therefore, the suggestion unit 20 generates and sends a message such as, for example, "For instance 2, people in your group place the most importance on the false negative rate. If you want to maximize the false negative rate, you will need to change the label. Do you want to change it?"

[0078] Furthermore, when proposing instance 3 to group 1, the proposed label is selected with the smaller sum of preference rankings, and is a proposal for a change that will optimize the index of a ranking other than group 1's first preference ranking. Therefore, the proposing unit 20 generates and sends a message such as, for example, "For instance 3, people in your group place the most importance on the false negative rate, but considering the opinions of other groups, it seems difficult to maximize the false negative rate. If you change the label, the accuracy of third place will be maximized. Do you want to change the label?"

[0079] Furthermore, when proposing instance 7 to group 2, the proposed label with the least change from the current annotation result is selected because the sum of the preference rankings was the same. Therefore, the proposing unit 20 generates and transmits a message such as, for example, "The current labeling for instance 7 is different, but when opinions from other groups are taken into account, it seems that many opinions support a different label. Would you like to change the label?"

[0080] If it is determined in step S22 that there is no change from the previous annotation result, it is determined that all stakeholder groups have no intention of changing the label any further, and the process proceeds to step S60. In step S60, the proposing unit 20 automatically annotates instances for which the final label has not been determined, and notifies the relevant stakeholder groups of the label changes and the reasons for the label changes. The label changes and the reasons for the label changes may be determined by the same process as in steps S14 to S30 (including steps S32 to S48). The annotation support process then ends.

[0081] As described above, the annotation support device according to this embodiment acquires labels annotated by annotators belonging to multiple stakeholder groups for multiple instances used in training a machine learning model. The annotation support device also acquires annotation results by aggregating the acquired labels for each stakeholder group. The annotation support device also acquires the ranking of indicators preferred by each stakeholder group when annotating. If the annotation results differ among stakeholder groups, the annotation support device proposes label changes to stakeholder groups that satisfy certain conditions based on the ranking of indicators preferred by each stakeholder group. This allows for the proposal of appropriate and highly convincing annotations in accordance with the preferences of various stakeholders.

[0082] For example, if a label is changed by simple majority vote, it is unclear how the annotations of the stakeholder group have been reflected in the label change, and it is unclear how the indicators will change as a result of the annotation, which reduces the sense of satisfaction. On the other hand, according to this embodiment, a change to a label determined based on the utility of each stakeholder group is proposed. Therefore, it is possible to understand what meaning one's own annotation has for the same stakeholder group, and how the indicators will change as a result of the annotation, thereby improving the sense of satisfaction.

[0083] By retraining a machine learning model using instances in which the label changes proposed by this embodiment have been made, it is possible to build a machine learning model that reflects the preferences of various stakeholders.

[0084] In the above embodiment, the preference ranking of each stakeholder is calculated from the annotation results of each stakeholder, but the present invention is not limited to this. For example, the preference ranking specified by a stakeholder group may be acquired.

[0085] In the above embodiment, the uncertain instances near the decision boundary of the machine learning model are used as the annotation target instances, but this is not limiting. For example, randomly selected instances may be used as the annotation target instances.

[0086] Furthermore, the instances to be annotated are not limited to training data used for training machine learning data, but may also be test data. In this case, the parameters of the machine learning model can be adjusted so that the labels output by the machine learning model match the labels of the instances after annotation is complete.

[0087] Furthermore, in the above embodiment, a machine learning model with a binary output has been described, but this embodiment can also be applied to a machine learning model that outputs labels with three or more values.

[0088] Furthermore, the determination of proposed labels, identification of the reason for change, and generation of messages in the above embodiment are merely examples, and may be based on the preference ranking of each stakeholder group.

[0089] In the above embodiment, the annotation support program is stored (installed) in advance in a storage device, but this is not limiting. The program according to the disclosed technology may be provided in a form stored in a storage medium such as a CD-ROM, a DVD-ROM, or a USB memory.

[0090] REFERENCE SIGNS LIST 10 Annotation support device 12 Training unit 14 Extraction unit 16 Acquisition unit 18 Calculation unit 20 Proposal unit 40 Computer 41 CPU 42 GPU 43 Memory 44 Storage device 45 Input / output device 46 R / W device 47 Communication I / F 48 Bus 49 Storage medium 50 Annotation support program 52 Training process control command 54 Extraction process control command 56 Acquisition process control command 58 Calculation process control command 60 Proposal process control command

Claims

1. An annotation support program that causes a computer to execute a process including: obtaining annotation results for each stakeholder group, in which labels annotated by annotators belonging to each of the multiple stakeholder groups for multiple instances used in training a machine learning model are aggregated, and obtaining the rankings of indicators preferred by each of the stakeholder groups when annotating; and, if the annotation results differ among the stakeholder groups, proposing changes to the labels to the stakeholder groups that satisfy predetermined conditions based on the rankings of the preferred indicators for each stakeholder group.

2. The annotation support program according to claim 1, wherein the ranking of the preferred indicators for each stakeholder group is calculated based on the annotation results for each stakeholder group.

3. The annotation support program described in claim 2, wherein the program estimates the preference value of each of the indicators for each stakeholder group so that a utility value for the multiple instances, represented by each value of the indicator calculated from a set of labels annotated for the multiple instances by each of the annotators included in the stakeholder group and a preference value indicating the degree of preference for each of the indicators, is maximized with respect to the sum of the utility values ​​of all sets of annotation labels assumed as annotations for the multiple instances, and determines a ranking of the preferred indicators for each stakeholder group based on the estimated preference value for each indicator.

4. The annotation support program according to any one of claims 1 to 3, wherein the process of proposing a label change includes setting a label pattern for each of the multiple instances to optimize each of the indicators for each stakeholder group, comparing the annotation results with the pattern, and determining the instances for which a label change is proposed and the proposed labels.

5. The annotation support program according to claim 4, wherein the pattern is a pattern that has the smallest difference from the annotation result and maximizes the index.

6. The annotation support program of claim 4, wherein the instance for which the label change is proposed is an instance for which the annotation results differ between the stakeholder groups, and the proposed label is the label of the instance that, for the indicators that match the label of the pattern, results in the smallest sum of the rankings of the preferred indicators of each of the stakeholder groups, or the label that results in the smallest change from the annotation result when the sum of the rankings for each of the instance labels is the same.

7. The annotation support program according to claim 6, wherein the predetermined condition is that the annotation result is different from the proposed label, and the process of suggesting a change to the label includes suggesting a reason for encouraging a change to the proposed label based on the proposed label, the ranking of the preference indicators, and the pattern.

8. The annotation support program of claim 7, wherein the process of proposing the reason includes identifying the reason based on whether the proposed label is the label that minimizes the sum of the rankings or the label that minimizes the change, and whether the change is to a label that optimizes the index that is ranked first in the ranking of the preferred index of the stakeholder group.

9. The annotation support program according to claim 8, wherein the process of identifying the reason includes: if the proposed label is the label with the smallest sum of the ranks and the change is to a label that optimizes the index that is ranked first among the preferred indicators of the stakeholder group, identifying as the reason that the preferred indicator is being changed to a label that optimizes the index that is ranked first; if the proposed label is the label with the smallest sum of the ranks and the change is not to a label that optimizes the index that is ranked first among the preferred indicators of the stakeholder group, identifying as the reason that the preferred indicator is being changed to a label that optimizes an indicator that is ranked other than the index that is ranked first, although the preferred indicator cannot be optimized in terms of the index that is ranked first, taking into account the annotations of other stakeholder groups; and if the proposed label is the label with the smallest change, identifying as the reason that the proposed label is often selected in the annotations of other stakeholder groups.

10. An annotation support program according to any one of claims 1 to 3, wherein the multiple instances are instances within a predetermined range from the decision boundary of the machine learning model.

11. An annotation support method in which a computer executes a process including: obtaining annotation results for each stakeholder group, in which labels annotated by annotators belonging to each stakeholder group for multiple instances used in training a machine learning model are aggregated, and obtaining a ranking of indicators preferred by each stakeholder group when annotating; and, if the annotation results differ among the stakeholder groups, proposing changes to the labels to the stakeholder groups that satisfy predetermined conditions based on the ranking of the preferred indicators for each stakeholder group.

12. The annotation support method according to claim 11, wherein the ranking of the preferred indicators for each stakeholder group is calculated based on the annotation results for each stakeholder group.

13. The annotation support method described in claim 12, further comprising estimating the preference value of each of the indicators for each stakeholder group so that a utility value for the multiple instances, represented by each value of the indicator calculated from a set of labels annotated for the multiple instances by each of the annotators included in the stakeholder group and a preference value indicating the degree of preference for each of the indicators, is maximized with respect to the sum of utility values ​​of all sets of annotation labels assumed as annotations for the multiple instances, and determining a ranking of the preferred indicators for each stakeholder group based on the estimated preference value for each indicator.

14. An annotation support method according to any one of claims 11 to 13, wherein the process of proposing a label change includes setting, for each stakeholder group, a label pattern for each of the multiple instances that will optimize each of the indicators, comparing the annotation results with the pattern, and determining the instances for which a label change is proposed and the proposed labels.

15. The annotation support method according to claim 14, wherein the pattern is a pattern that has the smallest difference from the annotation result and that maximizes the index.

16. The annotation support method according to claim 14, wherein the instance for which the label change is proposed is an instance for which the annotation results differ between the stakeholder groups, and the proposed label is the label of the instance for which the sum of the rankings of the preferred indicators of each of the stakeholder groups is the smallest for the indicators when the indicators match the label of the pattern, or the label for which the change from the annotation result is the smallest when the sum of the rankings for each of the labels of the instance is the same.

17. The annotation support method according to claim 16, wherein the predetermined condition is that the annotation result is different from the proposed label, and the process of suggesting a change to the label includes suggesting a reason for encouraging a change to the proposed label based on the proposed label, the ranking of the preference indicators, and the pattern.

18. The annotation support method described in claim 17, wherein the process of proposing the reason includes identifying the reason based on whether the proposed label is the label that minimizes the sum of the rankings or the label that minimizes the change, and whether the change is to a label that optimizes the index that is ranked first in the ranking of the preferred index of the stakeholder group.

19. The annotation support method according to claim 18, wherein the process of identifying the reason includes: if the proposed label is the label with the smallest sum of the ranks and the change is to a label that best ranks the preferred indicator of the stakeholder group with the number one preferred indicator, identifying as the reason that the change is to a label that best ranks the preferred indicator with the number one preferred indicator; if the proposed label is the label with the smallest sum of the ranks and the change is not to a label that best ranks the preferred indicator of the stakeholder group with the number one preferred indicator, identifying as the reason that the change is to a label that best ranks an indicator with another rank, even though the preferred indicator with the number one preferred indicator cannot be best ranked, taking into account annotations of other stakeholder groups; and if the proposed label is the label with the smallest change, identifying as the reason that the proposed label is often selected in annotations of other stakeholder groups.

20. An annotation support device comprising: an acquisition unit that acquires annotation results for each stakeholder group, in which labels annotated by annotators included in each of a plurality of stakeholder groups for a plurality of instances used to train a machine learning model are aggregated, and acquires a ranking of indicators preferred by each of the stakeholder groups when annotating; and a proposal unit that, when the annotation results differ among the stakeholder groups, proposes a change to the labels for the stakeholder groups that satisfy a predetermined condition, based on the ranking of the preferred indicators for each stakeholder group.

Citation Information

Patent Citations

  • Computer-implemented system and method for determining user attention

    JP2021525424A

  • Data processing method and apparatus, electronic device, and storage medium

    JP2022077969A

  • Machine learning-based visual device selection system

    JP2022531413A

  • Scoring autonomous vehicle trajectories using reasonable crowd data

    US20220063666A1

  • Information processing method, control program, information processing apparatus, and information processing system

    JP2023104050A