Crowdsourcing point allocation method and device, electronic equipment and readable storage medium
By prepaying fixed or dynamic points on the crowdsourcing platform and combining truth inference and Shapley value calculation, crowdsourcing points are reasonably allocated, solving the problems of low task accuracy and low participation, and achieving higher quality labeling task completion and user enthusiasm.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NETEASE (HANGZHOU) NETWORK CO LTD
- Filing Date
- 2024-10-14
- Publication Date
- 2026-04-14
AI Technical Summary
On large-scale and highly standardized crowdsourcing platforms, the allocation of task points by participants is not reasonable enough, resulting in low accuracy of task results and low participation of labeling users. Existing technologies that allocate points solely based on ability levels cannot effectively reflect the completion quality of different labeling tasks.
By acquiring multiple annotation results and corresponding annotation users for annotation tasks, fixed or dynamic points are prepaid, and crowdsourcing points are reasonably allocated through truth inference, feature function and Shapley value calculation to reflect the contribution of annotation users.
It improved the accuracy of annotation task results and the participation of annotation users, prevented low-quality task completion, and achieved fairness and effectiveness in points allocation.
Smart Images

Figure CN121860460A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a crowdsourcing points allocation method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] The development of internet technology has driven the evolution of task allocation behavior. Tasks are no longer simply assigned to professional teams by the initiator, but are published on a voluntary basis through crowdsourcing platforms and handed over to the general public to complete.
[0003] Crowdsourcing, as a mechanism for solving problems in a distributed manner via the internet, uses a points system as an incentive mechanism on crowdsourcing platforms to encourage participants to actively answer questions. Participants are awarded points based on the accuracy of their answers after each submission.
[0004] In related technologies, for large-scale and highly standardized crowdsourcing platforms, the task recipients often divide the task points equally. The incentive points obtained by participants for completing the labeled tasks with high quality are consistent, resulting in insufficient incentive and thus low accuracy of task results and low participation of labeled users.
[0005] To address the aforementioned issues, related technologies can allocate task points to participants by determining their ability levels. However, a participant's ability level cannot fully represent the quality of their completion of different annotation tasks. Simply allocating task points based on the participant's ability level may result in unreasonable task point allocation, leading to low accuracy of task results and low participation from annotation users. Summary of the Invention
[0006] In view of this, the purpose of this application is to provide a crowdsourcing points allocation method, device, electronic device and readable storage medium, which requires users who undertake target annotation tasks to prepay a certain number of points for the target annotation task to prevent users from completing the annotation task with low quality; at the same time, when allocating points, the points are reasonably allocated by referring to the contribution of each user to the target annotation task, which helps to improve the accuracy of the annotation task results and the participation of users.
[0007] In a first aspect, embodiments of this application provide a crowdsourcing points allocation method, the crowdsourcing points allocation method comprising:
[0008] Obtain multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user's ability level;
[0009] The truth labels for the target annotation task are obtained by performing truth inference on the multiple annotation results.
[0010] Based on the truth labels, a feature function is determined, and based on the feature function, the Shapley value of the labeling user corresponding to each labeling result is determined;
[0011] Crowdsourcing points are allocated to each annotation user corresponding to each annotation result based on the Shapley value, first score, and overall score of the target annotation task.
[0012] Secondly, embodiments of this application also provide a crowdsourcing points allocation device, the crowdsourcing points allocation device comprising:
[0013] The information acquisition module is used to acquire multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user's ability level;
[0014] The truth label inference module is used to perform truth inference on the multiple annotation results to obtain the truth label of the target annotation task;
[0015] The numerical determination module is used to determine a feature function based on the truth label, and to determine the Shapley value of the labeling user corresponding to each labeling result based on the feature function;
[0016] The points allocation module is used to allocate crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, first points, and the overall points of the target annotation task for each annotation user corresponding to each annotation result.
[0017] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the crowdsourcing points allocation method as described in any of the first aspects.
[0018] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the crowdsourcing points allocation method as described in any of the first aspects.
[0019] The crowdsourcing points allocation method, apparatus, electronic device, and readable storage medium provided in this application acquire multiple annotation results of a target annotation task, the annotation user corresponding to each annotation result, and the first points corresponding to the target annotation task; perform truth inference on the multiple annotation results to obtain the truth labels of the target annotation task; determine a feature function based on the truth labels, and determine the Shapley value of the annotation user corresponding to each annotation result based on the feature function; allocate crowdsourcing points to the annotation user corresponding to each annotation result based on the Shapley value of the annotation user corresponding to each annotation result, the first points, and the overall points of the target annotation task. In this way, annotation users undertaking the target annotation task need to prepay a certain amount of first points for the target annotation task to prevent them from completing the annotation task with low quality; at the same time, the allocation of points takes into account the contribution of each annotation user to the target annotation task for reasonable distribution, which helps to improve the accuracy of the annotation task results and the participation of annotation users.
[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a crowdsourcing points allocation method provided in this application embodiment;
[0023] Figure 2 This is a schematic diagram of the crowdsourcing points allocation process provided in the embodiments of this application;
[0024] Figure 3 This is one of the structural schematic diagrams of a crowdsourcing points allocation device provided in the embodiments of this application;
[0025] Figure 4 This is a second schematic diagram of a crowdsourcing points allocation device provided in an embodiment of this application;
[0026] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. Based on the embodiments of this application, every other embodiment obtained by those skilled in the art without inventive effort falls within the scope of protection of this application.
[0028] First, let me introduce the background of this application:
[0029] The development of internet technology has driven the evolution of task allocation behavior. Tasks are no longer simply assigned to professional teams by the initiator, but are published on a voluntary basis through crowdsourcing platforms and handed over to the general public to complete.
[0030] Crowdsourcing, as a mechanism for solving problems in a distributed manner via the internet, uses a points system as an incentive mechanism on crowdsourcing platforms to encourage participants to actively answer questions. Participants are awarded points based on the accuracy of their answers after each submission.
[0031] In related technologies, for large-scale and highly standardized crowdsourcing platforms, the task recipients often divide the task points equally. The incentive points obtained by participants for completing the labeled tasks with high quality are consistent, resulting in insufficient incentive and thus low accuracy of task results and low participation of labeled users.
[0032] To address the aforementioned issues, related technologies can allocate task points to participants by determining their ability levels. However, a participant's ability level cannot fully represent the quality of their completion of different annotation tasks. Simply allocating task points based on the participant's ability level may result in unreasonable task point allocation, leading to low accuracy of task results and low participation from annotation users.
[0033] Based on this, this application provides a crowdsourcing points allocation method to improve the accuracy of annotation task results and the participation of annotation users.
[0034] Please see Figure 1 , Figure 1 This is a flowchart illustrating a crowdsourcing points allocation method provided in an embodiment of this application. Figure 1 As shown in the embodiments of this application, the crowdsourcing points allocation method includes:
[0035] S101. Obtain multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user ability level of the annotation user.
[0036] S102. Perform truth inference on the multiple annotation results to obtain the truth label of the target annotation task.
[0037] S103. Determine the feature function based on the truth label, and determine the Shapley value of the labeling user corresponding to each labeling result based on the feature function.
[0038] S104. Based on the Shapley value, first score, and overall score of the target annotation task of each annotation user corresponding to the annotation result, crowdsourcing points are allocated to each annotation user corresponding to the annotation result.
[0039] The crowdsourcing points allocation method provided in this application requires users who undertake target annotation tasks to prepay a certain amount of initial points to prevent them from completing the annotation tasks with low quality. At the same time, the points are allocated reasonably by taking into account the contribution of each user to the target annotation task, which helps to improve the accuracy of the annotation task results and the participation of users.
[0040] The exemplary steps of the embodiments of this application are described below:
[0041] S101. Obtain multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user ability level of the annotation user.
[0042] Crowdsourcing, as a mechanism for distributed problem-solving via the internet, uses a points system as an incentive on crowdsourcing platforms to encourage participants to actively contribute. After each submission, participants earn points based on the accuracy of their answer and their contribution to the final correct solution. These points can be redeemed for gifts and bonuses and also serve as a basis for assessing a user's skill level.
[0043] In one possible implementation, after the target labeling task enters the crowdsourcing platform, multiple labeling users who have not labeled the target labeling task will choose to receive the target labeling task and label the target labeling task, thus obtaining multiple labeling results for the target labeling task.
[0044] If the target annotation task includes more than one annotation issue, different annotation users can receive different numbers of annotation issues to annotate. That is, different annotation users may have different numbers of annotation issues to annotate for the same target annotation task.
[0045] For example, suppose the target annotation task contains 100 questions, and the annotation time for the target annotation task is X minutes. When the target annotation task is accepted, annotation user A takes 20 annotation questions from the target annotation task, and annotation user B also takes 20 annotation questions from the target annotation task. After annotation user A has annotated 20 annotation questions, he takes 10 more annotation questions. Therefore, after the annotation time ends, for the target annotation task, annotation user A has annotated 30 annotation questions, and annotation user B has annotated 20 annotation questions.
[0046] In one possible implementation, in order to ensure the quality of each annotation user's annotation for the target annotation task and to encourage them to participate more actively in improving their own ability level, after determining that the annotation user has accepted the target annotation task, the annotation user needs to prepay a certain amount of first points for the target annotation task, and subsequently consider whether to return the first points and the amount of first points to be returned based on each annotation user's contribution.
[0047] In one possible implementation, the first integral is determined based on the target annotation task and / or the user's ability level of the annotation user.
[0048] Specifically, the first point prepaid by each annotation user for participating in the annotation activity of the target annotation task is determined through the following steps:
[0049] a1: Determine the preset fixed points corresponding to the target annotation task as the first points prepaid by each annotation user for participating in the annotation behavior of the target annotation task; or,
[0050] a2: Based on the user ability level of each annotation user and the preset point mapping relationship, determine the first point prepaid by each annotation user for participating in the annotation behavior of the target annotation task.
[0051] In this embodiment of the application, a preset fixed score may be set for different target annotation tasks. As long as the annotation user undertakes the target annotation task, the corresponding preset fixed score needs to be paid in advance.
[0052] In one possible implementation, the preset fixed integral for different target annotation tasks can be determined by the crowdsourcing platform through observation of historical annotation task data. It can be set according to the number of annotation questions contained in the target annotation task, the annotation deadline of the target annotation task, or the importance of the target annotation task, without being specifically limited here.
[0053] In another possible implementation, in order to improve the matching between the first score and the annotation user, and further enhance the annotation user's enthusiasm for undertaking the target annotation task and the annotation quality of the annotation user for the target annotation task, the first score of different annotation users for the same target annotation task can be dynamically set according to the point mapping relationship between the user ability level of different annotation users and the first score.
[0054] Specifically, the step "based on each annotation user's ability level and a preset points mapping relationship, determine the first points prepaid by each annotation user for participating in the annotation behavior of the target annotation task" includes:
[0055] b1: For each labeled user, based on the user's ability level and a preset ability level range, determine the target ability level range to which the user's ability level belongs, and determine the interval points corresponding to the target ability level range as the first points prepaid by the user for participating in the labeling behavior of the target labeling task; or...
[0056] b2: Based on the user ability level of each annotation user and the preset negative income threshold, determine the linear mapping relationship between the user ability level and the first point, and determine the first point prepaid by each annotation user for participating in the annotation behavior of the target annotation task based on the linear mapping relationship of the first point; wherein, the linear mapping relationship of the first point represents that the user ability level of the annotation user is negatively correlated with the first point.
[0057] In one possible implementation, user ability levels can be pre-divided to obtain multiple user ability level intervals. Each user ability level interval is assigned a preset fixed score. Based on the user ability level interval to which the labeled user belongs, the first score that the labeled user needs to prepay for the target labeling task is determined.
[0058] For example, user ability levels are divided into 1-10 levels, which can be further divided into three user ability level intervals: [1,4), [4,7), and [7,10]. The first score for the [1,4) user ability level interval is 20 points, the first score for the [4,7) user ability level interval is 15 points, and the first score for the [7,10] user ability level interval is 10 points. User A and user B are determined to receive the target annotation task. User A's user ability level is determined to be level 8, belonging to the [7,10] user ability level interval. Therefore, user A's first score for the target annotation task is 10 points. User B's user ability level is determined to be level 2, belonging to the [1,4) user ability level interval. Therefore, user B's first score for the target annotation task is 20 points.
[0059] In another possible implementation, it is also possible to determine the negative benefit threshold for the labeled user at different user ability levels, and then determine the linear mapping relationship between the user ability level and the first integral based on the different negative benefit thresholds. Finally, the first integral of the labeled user for the target labeling task can be determined based on the linear mapping relationship between the user ability level and the first integral.
[0060] The linear mapping relationship of the first integral represents that the user's ability level is negatively correlated with the first integral. That is, the higher the user's ability level, the lower the first integral required for the same target annotation task. This incentivizes the user to increase their enthusiasm for undertaking target annotation tasks and improve the annotation quality of the target annotation tasks, thereby improving the user's own ability level.
[0061] In this embodiment of the application, in order to ensure that different annotation users can prepay the corresponding first points, thereby improving the accuracy of the first point prepayment and the enthusiasm of annotation users to undertake the target annotation task, it is preferred to use different dynamic first point methods for different annotation users to determine the first points.
[0062] Furthermore, after confirming that each annotation user who accepted the target annotation task has prepaid the corresponding first point, different annotation users can annotate the target annotation task. At the same time, the crowdsourcing platform will collect the annotation results of each annotation user and calculate the truth label of the target annotation task.
[0063] S102. Perform truth inference on the multiple annotation results to obtain the truth label of the target annotation task.
[0064] In this embodiment of the application, when calculating the truth value of the target annotation task, it is necessary to determine that after a certain number of annotations are collected for the target annotation task, the truth value label is calculated. This can be done by inferring the truth value of the target annotation task based on the user ability level of each annotation user, the annotation results of each annotation user for the target annotation task, and the candidate answers of the target annotation task.
[0065] Specifically, the step "performing truth inference on the multiple annotation results to obtain the truth labels for the target annotation task" includes:
[0066] c1: Based on the user ability level of the annotating user corresponding to each annotation result, the annotation result, and the candidate answers of the target annotation task, determine the target annotation result with the largest proportion among multiple annotation results.
[0067] c2: Based on the target annotation results, obtain the ground truth label for the target annotation task.
[0068] In this embodiment of the application, the target annotation result with the largest proportion among multiple annotation results can be determined based on the user ability level of the annotation user corresponding to each annotation result, the annotation result, and the candidate answers of the target annotation task, and the true value label can be determined.
[0069] Here, the truth label can be set to 1 to annotate the annotation result a for user i. i It belongs to {0,1}, where 0 indicates that user i answered incorrectly and 1 indicates that user i answered correctly.
[0070] In one possible implementation, the ground truth label for the target annotation task can be determined using the following formula:
[0071]
[0072] Among them, t j represents the truth label for the target annotation task j; argmax represents a subset of the domain where any element in that subset is a function that maximizes the value of the function; y j Y represents the candidate answer for target annotation task j; N represents the set of candidate answers for target annotation task j; q represents the total number of annotation users; i This indicates the user's ability level, labeled as user i; a ij This represents the annotation result of user i for target annotation task j.
[0073] Furthermore, after obtaining the truth labels through truth inference, the feature function can be determined through the determined truth labels, and then the Shapley value of the labeled user corresponding to each labeling result can be calculated based on the feature function.
[0074] S103. Determine the feature function based on the truth label, and determine the Shapley value of the labeling user corresponding to each labeling result based on the feature function.
[0075] In this embodiment, a feature function can be obtained by weighting the annotation results of multiple table users who have annotated for the target annotation task and the user ability level of each annotation user, and then the Shapley value can be calculated using the feature function.
[0076] Specifically, the characteristic function is determined through the following steps:
[0077] d1: Determine the set of multiple annotations completed by multiple annotation users corresponding to multiple annotation results.
[0078] d2: Determine the feature function based on the user ability level and annotation results of each annotated user included in each annotation permutation set.
[0079] In this embodiment of the application, multiple labeling users who label the same target labeling task can be regarded as an alliance. In this alliance, multiple labeling users who have completed labeling can be arranged and combined in different ways to obtain multiple labeling arrangement sets.
[0080] For example, there are three annotation users who are annotating the target annotation task: annotation user A, annotation user B, and annotation user C. The annotation permutation sets that can be obtained for the above annotation users are: (annotation user A), (annotation user B), (annotation user B), (annotation user A, annotation user B), (annotation user A, annotation user C), (annotation user C, annotation user C) and (annotation user A, annotation user B, annotation user C).
[0081] Furthermore, based on determining the user ability level and annotation results of each annotated user included in multiple annotation permutation sets, the feature function is determined.
[0082] To reduce the computational steps and complexity, the truth label can be set to 1, and the annotation result a for user i can be labeled. i The range is {0,1}, where 0 indicates that user i answered incorrectly and 1 indicates that user i answered correctly. At the same time, the user ability level q of different labeled users is also set in the range of (0,1).
[0083] Specifically, the characteristic function can be calculated using the following formula:
[0084]
[0085] Where v(S) is the feature function; S is the set of labeled permutations; q i This indicates the user's ability level, labeled as user i; a i This represents the annotation result of user i for the target annotation task.
[0086] Furthermore, after determining the feature function, the average marginal contribution of each annotation user to the feature function v(S) can be calculated under all possible permutations of annotations completed by other annotationers, and then the Shapley value of the annotation user can be calculated.
[0087] Specifically, the step "determining the Shapley value of the annotation user corresponding to each annotation result based on the feature function" includes:
[0088] e1: For each labeled user, the marginal contribution of that user to the feature function in each set of labeled permutations.
[0089] e2: For each labeled user, based on the set weights corresponding to each labeled set, calculate the average marginal contribution of the labeled user to the feature function in each labeled set, and obtain the Shapley value of the labeled user.
[0090] In this application example, for each annotation user, it is necessary to calculate the marginal contribution of the annotation user under all possible annotation permutations of other annotationers, and then calculate the average marginal contribution of the annotation user to the feature function in each annotation permutation based on the weight value of each annotation permutation, thereby determining the Shapley value of the annotation user.
[0091] In one possible implementation, when the number of annotating users is large, calculating the Shapley value for all annotating users together would lead to a complex calculation process, thereby increasing the consumption of computer performance. Therefore, this embodiment first divides the annotating users into several groups (annotation permutation sets), and then calculates the set Shapley value within each group. It should be noted that the set Shapley value refers to the permutation of annotating users within each annotating set when calculating the Shapley value. Since the number of users in each annotating set is much smaller than the total number of annotating users, the efficiency of calculating the Shapley value can be further improved.
[0092] In one possible implementation, the set weight corresponding to the annotation permutation set can be determined based on the accuracy of the historical annotation results of the annotation users in each annotation permutation set.
[0093] Specifically, the Shapley value of a user can be calculated using the following formula:
[0094]
[0095] Where, φ i (V) represents the Shapley value of labeled user i; v(S) is the feature function; S is the set of labeled permutations; N represents the total number of labeled users.
[0096] In one possible implementation, the Shapley value can be positive or negative. A positive Shapley value indicates that the labeling user has made a positive contribution to the target labeling task and can obtain crowdsourcing points; while a negative Shapley value indicates that the labeling user has made a negative contribution to the target labeling task and needs to deduct crowdsourcing points. The allocation method of crowdsourcing points for each labeling user will be described below.
[0097] S104. Based on the Shapley value, first score, and overall score of the target annotation task of each annotation user corresponding to the annotation result, crowdsourcing points are allocated to each annotation user corresponding to the annotation result.
[0098] In this embodiment of the application, when determining the crowdsourcing points to be allocated to each labeled user, it is necessary to first determine whether the labeled user has made a positive or negative Shapley value contribution based on the Shapley value of each labeled user, and then calculate the crowdsourcing points that can be allocated to the labeled users.
[0099] Specifically, the step "assigning crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, first score, and overall score of the target annotation task for each annotation result" includes:
[0100] f1: Identify at least one negative contributing annotation user and at least one positive contributing annotation user among multiple annotation users who have made negative Shapley value contributions for the target annotation task.
[0101] f2: Based on the preset negative contribution ratio and the overall score of the target labeling task, determine the total deduction score for at least one negative contribution labeling user.
[0102] f3: For each negative contribution labeled user, based on the proportion of the negative contribution labeled user's Shapley value in the total Shapley value of the at least one negative contribution labeled user and the total deducted points, determine the deducted points for the negative contribution labeled user, deduct the deducted points from the negative contribution labeled user's first points, and return them to the labeled user.
[0103] f4: For each positive contribution labeling user, based on the Shapley value of the positive contribution labeling user, the overall score of the target labeling task, and the total deducted score, determine the crowdsourcing points allocated to the positive contribution labeling user.
[0104] A positive Shapley value indicates that the labeling user made a positive contribution to the target labeling task, while a negative Shapley value indicates that the labeling user made a negative contribution to the target labeling task.
[0105] In one possible implementation, a preset negative contribution ratio is set for each target annotation task. Based on the negative contribution ratio and the overall score that the entire target annotation task can provide, multiple negative contribution annotation users who generate negative Shapley value contributions need to be identified, and the total deduction points need to be deducted. Then, the total deduction points are distributed to each negative contribution annotation user according to the Shapley value contribution for point deduction.
[0106] The overall score and preset negative contribution ratio for different target annotation tasks can be different, and can be set according to the urgency, importance and number of annotation problems in the target annotation task.
[0107] For example, if the total score for the target annotation task is 50 points, and the preset negative contribution ratio is set to 20%, then the total number of points to be deducted for multiple negative contribution annotation users who generate negative Shapley values is 10 points.
[0108] Furthermore, for each negative contribution labeled user who generates a negative Shapley value contribution, the deduction points for that negative contribution labeled user are determined based on the proportion of the negative contribution labeled user's Shapley value in the total Shapley value of the at least one negative contribution labeled user and the total deduction points.
[0109] Specifically, the deduction of points for users with negative contributions can be determined using the following formula:
[0110]
[0111] Where, d ij α represents the deduction of points by user i for target annotation task j; α represents the preset negative contribution ratio; P represents the overall points of target annotation task j; This represents the negative Shapley value contribution of user i to the target annotation task j.
[0112] In another possible implementation, for at least one positive contribution labeling user who generates a positive Shapley value contribution, the total deducted points from at least one negative contribution labeling user need to be added to the overall points of the target labeling task before the points are allocated to each positive contribution labeling user.
[0113] Specifically, the step "For each positive contribution labeling user, based on the Shapley value of the positive contribution labeling user, the overall score of the target labeling task, and the total deducted score, determine the crowdsourcing points allocated to the positive contribution labeling user" includes:
[0114] g1: Based on the preset platform allocation ratio and the total deduction points, deduct the second points allocated to the crowdsourcing platform from the total deduction points to determine the additional reward points allocated to the at least one positive contribution user.
[0115] g2: Sum the additional reward points and the total points to determine the total reward points to be allocated to the at least one positive contribution user.
[0116] g3: For each positive contribution labeled user, the crowdsourcing points for that positive contribution labeled user are determined based on the proportion of the positive contribution labeled user's Shapley value in the total Shapley value of the at least one positive contribution labeled user and the total reward points.
[0117] In one possible implementation, the total reward points for users who make positive contributions include not only the overall points of the target labeling task, but also the total deducted points from users who make negative contributions. Before adding the total deducted points to the total reward points, the crowdsourcing platform needs to obtain a portion of the second points according to a certain preset platform allocation ratio.
[0118] Specifically, the second score is determined based on the preset platform allocation ratio and the total deducted score. The overall score and the preset platform allocation ratio for different target labeling tasks can be different. They can be set according to the urgency, importance, and number of labeling issues in the target labeling task. Alternatively, the preset platform allocation ratio can be the same, that is, the crowdsourcing platform is set so that the platform allocation ratio is consistent for different target labeling tasks.
[0119] For example, if the preset platform allocation ratio is set to β, and the total deduction points for at least one negative contribution labeling user who generates a negative Shapley value contribution are M, then the additional reward points that at least one positive contribution labeling user who generates a positive Shapley value contribution can obtain, excluding the overall points of the target labeling task, are (1-β)·M.
[0120] Furthermore, for each negative contribution user who generates a positive Shapley value contribution, the crowdsourcing points allocated to that positive contribution user are determined based on the proportion of the positive contribution user's Shapley value in the total Shapley value of the at least one positive contribution user and the total reward points.
[0121] Specifically, the crowdsourcing points allocated to a user with a positive contribution can be determined using the following formula:
[0122]
[0123] Where, r kj β represents the crowdsourcing points allocated to user k for target annotation task j; β represents the preset platform allocation ratio; α represents the preset negative contribution ratio; P represents the total points of target annotation task j. This represents the positive Shapley value contribution of annotation user k to the target annotation task j.
[0124] In one possible implementation, if the target annotation task contains multiple annotation problems, it is necessary to calculate the corresponding deduction score and reward score for each annotation problem, and then determine the total deduction score or total reward score for the target annotation task.
[0125] For example, if the target annotation task includes three annotation problems, annotation problem 1, annotation problem 2, and annotation problem 3; and annotation user A completes annotation problems 1, 2, and 3; after calculation, annotation user A's Shapley value for annotation problem 1 is negative, meaning annotation user A made a negative contribution to annotation problem 1 and needs to deduct 5 points; annotation user A's Shapley values for annotation problems 2 and 3 are positive, meaning annotation user A made positive contributions to annotation problems 2 and 3 and can be rewarded with 6 points and 3 points respectively, then for the target annotation task, annotation user A can obtain 6 + 3 - 5 = 4 points.
[0126] In one possible implementation, for at least one negative contribution labeling user who generates a negative Shapley value and whose points need to be deducted, the points to be deducted are first deducted from the negative contribution labeling user's first points for the target labeling task. If the first points cannot be deducted, then the points are deducted from the crowdsourcing points held by the negative contribution user. If there are still some points remaining after the first points are deducted, then the points need to be returned to the negative contribution labeling user.
[0127] For example, for user X who makes a negative contribution to the labeling task, the first point for the target labeling task is 10 points, and after calculation, the deduction points that user X needs to offset are 12 points. In this case, the 10 points prepaid by user X are deducted first, and then 2 points are deducted from user X's crowdsourcing points. For user G who makes a negative contribution to the labeling task, the first point for the target labeling task is 8 points, and after calculation, the deduction points that user G needs to offset are 5 points. In this case, 5 points are deducted from the first point and 3 points are returned to user G.
[0128] In another possible implementation, for at least one positive contribution user who generates a positive Shapley value contribution, after the reward points are distributed, the first point also needs to be returned to the positive contribution user.
[0129] Specifically, the crowdsourcing points allocation method further includes:
[0130] h1: For each positive contribution user, after the crowdsourcing points are allocated to the positive contribution user, all the first points prepaid by the positive contribution user are returned to the positive contribution user.
[0131] For example, for a positive contribution labeling user W, the first point for the target labeling task is 5 points, and after calculation, the reward points awarded to the positive contribution labeling user W are 10 points. Therefore, after completing the target labeling task, the positive contribution labeling user W can obtain 10 reward points and at the same time get back the first point of 5 points.
[0132] In one possible implementation, if for multiple negative contribution labeling users, if the first points of many negative contribution labeling users cannot completely cover the deducted points, then a preset negative contribution ratio is needed to ensure that the first points of all negative contribution labeling users participating in the target labeling task are sufficient to cover their negative contributions to the greatest extent.
[0133] Specifically, the crowdsourcing points allocation method further includes:
[0134] i1: Detect whether there are any users among the negative contribution labelers whose first score is less than the deducted score.
[0135] i2: If it exists, and the proportion of the number of users whose first score is less than the number of users whose score is deducted in the total number of users with negative contributions is greater than a preset proportion threshold, adjust the preset negative contribution proportion.
[0136] In this embodiment of the application, it is necessary to determine the proportion of users whose first points cannot cover the deducted points among the negative contribution users. If the proportion of the number of negative contribution users whose first points are less than the number of users whose points are deducted is greater than the total number of negative contribution users, it means that the first points of most negative contribution users cannot cover their negative contributions. Therefore, it is necessary to adjust the preset negative contribution ratio to reduce the deducted points of each negative contribution user.
[0137] The adjustment method for the preset negative contribution ratio is to reduce the preset negative contribution ratio, thereby reducing the deduction points required for each negative contribution labeling user, so as to ensure that the first points of all negative contribution labeling users participating in the target labeling task are sufficient to cover their negative contributions.
[0138] In one possible implementation, the preset quantity ratio threshold may be different for different target annotation tasks.
[0139] The crowdsourcing points allocation method in this application embodiment will be illustrated below with specific examples:
[0140] Please see Figure 2 , Figure 2 This is a schematic diagram of the crowdsourcing points allocation process provided in the embodiments of this application, such as... Figure 2As shown, there are multiple online tasks and multiple users annotating tasks to be received on the crowdsourcing platform. The crowdsourcing platform determines the tasks to be annotated from the multiple online tasks. The annotating users who need to receive the tasks to be annotated need to prepay points as required. After the points prepayment is completed, each annotating user annotates the tasks to be annotated, then performs truth inference, outputs the data labels corresponding to the truth labels, evaluates the contribution of Shapley values based on the data labels, and then allocates points to the annotating users who receive the tasks to be annotated.
[0141] In one possible implementation, through verification on a crowdsourcing platform, the crowdsourcing points allocation method provided in this application embodiment shows significant improvement in the quality of collected data, and can improve the overall quality of crowdsourcing tasks and the ability level of labeled users.
[0142] For example, in an online A / B test conducted on a crowdsourcing platform, as shown in Table 1, the crowdsourcing points allocation method provided in this application embodiment shows the best performance in multiple metrics, including the number of completed tasks (NCT), the mean (ANW) and median (MNW) of the number of participants in completed tasks, the number of labeled users (workers) with a skill level below 0.5 (QWAB), the number of workers with a skill level between 0.5 and 0.6 (NWAB), and the average skill level of workers (AAW).
[0143] Table 1. Results of Online A / B Testing of Crowdsourcing Points Allocation Method on the Crowdsourcing Platform
[0144]
[0145] The crowdsourcing points allocation method provided in this application involves obtaining multiple annotation results for a target annotation task, the annotation user corresponding to each annotation result, and the first points corresponding to the target annotation task; performing truth inference on the multiple annotation results to obtain the truth labels for the target annotation task; determining a feature function based on the truth labels, and determining the Shapley value for the annotation user corresponding to each annotation result based on the feature function; and allocating crowdsourcing points to the annotation user corresponding to each annotation result based on the Shapley value of the annotation user corresponding to each annotation result, the first points, and the overall points of the target annotation task. In this way, annotation users undertaking the target annotation task need to prepay a certain amount of first points for the target annotation task to prevent them from completing the annotation task with low quality; at the same time, the allocation of points takes into account the contribution of each annotation user to the target annotation task for reasonable distribution, which helps to improve the accuracy of the annotation task results and the participation of annotation users.
[0146] Based on the same inventive concept, this application also provides a crowdsourcing points allocation device corresponding to the crowdsourcing points allocation method. Since the principle of the device in this application is similar to the crowdsourcing points allocation method described above in the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0147] Please see Figure 3 , Figure 4 , Figure 3 This is one of the structural schematic diagrams of a crowdsourcing points allocation device provided in the embodiments of this application. Figure 4 This is a second schematic diagram of a crowdsourcing points allocation device provided in an embodiment of this application. Figure 3 As shown, the crowdsourcing points allocation device 300 includes:
[0148] The information acquisition module 310 is used to acquire multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user ability level of the annotation user;
[0149] The truth label inference module 320 is used to perform truth inference on the multiple annotation results to obtain the truth label of the target annotation task;
[0150] The numerical determination module 330 is used to determine a feature function based on the truth label, and to determine the Shapley value of the labeling user corresponding to each labeling result based on the feature function;
[0151] The points allocation module 340 is used to allocate crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, the first point, and the overall points of the target annotation task for each annotation user corresponding to each annotation result.
[0152] In one possible implementation, the crowdsourcing points allocation device 300 further includes a points determination module 350, which is used to determine the first points prepaid by each annotation user for participating in the annotation behavior of the target annotation task through the following steps:
[0153] The preset fixed points corresponding to the target annotation task are determined as the first points prepaid by each annotation user for participating in the annotation behavior of the target annotation task; or...
[0154] Based on each user's ability level and a preset points mapping relationship, the first points prepaid by each user for participating in the target annotation task are determined.
[0155] In one possible implementation, when the point determination module 350 determines the first point prepaid by each annotation user for participating in the target annotation task based on the user ability level of each annotation user and a preset point mapping relationship, the point determination module 350 is used to:
[0156] For each labeled user, based on the user's ability level and a preset ability level range, determine the target ability level range to which the user's ability level belongs, and determine the interval points corresponding to the target ability level range as the first points prepaid for the labeling behavior of participating in the target labeling task; or;
[0157] Based on the user ability level of each annotation user and a preset negative benefit threshold, a linear mapping relationship between the user ability level and the first point is determined, and based on the linear mapping relationship of the first point, the first point prepaid by each annotation user for participating in the annotation behavior of the target annotation task is determined; wherein, the linear mapping relationship of the first point indicates that the user ability level of the annotation user is negatively correlated with the first point.
[0158] In one possible implementation, when the truth label inference module 320 performs truth inference on the plurality of annotation results to obtain the truth label of the target annotation task, the truth label inference module 320 is used to:
[0159] Based on the user ability level of the annotating user corresponding to each annotation result, the annotation result, and the candidate answers of the target annotation task, determine the target annotation result with the largest proportion among multiple annotation results;
[0160] Based on the target annotation results, the ground truth labels for the target annotation task are obtained.
[0161] In one possible implementation, the numerical determination module 330 is used to determine the characteristic function through the following steps:
[0162] Determine the set of multiple annotations completed by multiple annotation users corresponding to multiple annotation results;
[0163] The feature function is determined based on the user ability level and annotation results of each annotated user included in each annotation permutation set.
[0164] In one possible implementation, when the numerical determination module 330 is used to determine the Shapley value of the annotation user corresponding to each annotation result based on the feature function, the numerical determination module 330 is used to:
[0165] For each labeled user, the marginal contribution of that user to the feature function in each set of labeled permutations is calculated.
[0166] For each labeled user, based on the set weights corresponding to each labeled set, the average marginal contribution of the labeled user to the feature function in each labeled set is calculated to obtain the Shapley value of the labeled user.
[0167] In one possible implementation, when the points allocation module 340 allocates crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, the first point, and the overall points of the target annotation task for each annotation user corresponding to the annotation result, the points allocation module 340 is used to:
[0168] Identify at least one negative contributing annotation user and at least one positive contributing annotation user among multiple annotation users who have made negative Shapley value contributions for the target annotation task.
[0169] Based on the preset negative contribution ratio and the overall score of the target labeling task, the total deduction points for at least one negative contribution labeling user are determined.
[0170] For each negative contribution labeled user, the deduction points for the negative contribution labeled user are determined based on the proportion of the negative contribution labeled user's Shapley value in the total Shapley value of the at least one negative contribution labeled user and the total deduction points. The first point of the negative contribution labeled user is deducted from the deduction points and then returned to the labeled user.
[0171] For each positive contribution labeling user, the crowdsourcing points allocated to that positive contribution labeling user are determined based on the user's Shapley value, the overall score of the target labeling task, and the total deducted points.
[0172] In one possible implementation, when the points allocation module 340 determines the crowdsourcing points allocated to each positive contribution labeling user based on the user's Shapley value, the overall points of the target labeling task, and the total deducted points, the points allocation module 340 is configured to:
[0173] Based on the preset platform allocation ratio and the total deduction points, the second points allocated to the crowdsourcing platform are deducted from the total deduction points to determine the additional reward points allocated to the at least one positive contribution user.
[0174] Add the additional reward points and the total points to determine the total reward points to be allocated to the at least one positive contribution user.
[0175] For each positive contribution user, the crowdsourcing points for that positive contribution user are determined based on the proportion of that positive contribution user's Shapley value in the total Shapley value of the at least one positive contribution user and the total reward points.
[0176] In one possible implementation, the crowdsourcing points allocation device 300 further includes a points return module 360, which is used for:
[0177] For each user with a positive contribution, after the crowdsourcing points are allocated to that user, all the first points prepaid by that user are returned to that user.
[0178] In one possible implementation, the crowdsourcing points allocation device 300 further includes a ratio adjustment module 370, the ratio adjustment module 370 being used for:
[0179] Detect whether there are any users among those who have negatively contributed whose first score is less than the score deducted from their score.
[0180] If such a user exists, and the proportion of the number of users whose first score is less than the number of users whose score is deducted is greater than a preset proportion threshold in the total number of users with negative contributions, then the preset negative contribution proportion is adjusted.
[0181] The crowdsourcing points allocation device provided in this application obtains multiple annotation results for a target annotation task, the annotation user corresponding to each annotation result, and the first points corresponding to the target annotation task; performs truth inference on the multiple annotation results to obtain the truth labels of the target annotation task; determines a feature function based on the truth labels, and determines the Shapley value of the annotation user corresponding to each annotation result based on the feature function; and allocates crowdsourcing points to the annotation user corresponding to each annotation result based on the Shapley value of the annotation user corresponding to each annotation result, the first points, and the overall points of the target annotation task. In this way, annotation users undertaking the target annotation task need to prepay a certain amount of first points for the target annotation task to prevent them from completing the annotation task with low quality; at the same time, the allocation of points takes into account the contribution of each annotation user to the target annotation task for reasonable distribution, which helps to improve the accuracy of the annotation task results and the participation of annotation users.
[0182] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 500 includes a processor 510, a memory 520, and a bus 530.
[0183] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 is running, the processor 510 communicates with the memory 520 via the bus 530, causing the processor 510 to execute the following instructions during operation:
[0184] Obtain multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user's ability level;
[0185] The truth labels for the target annotation task are obtained by performing truth inference on the multiple annotation results.
[0186] Based on the truth labels, a feature function is determined, and based on the feature function, the Shapley value of the labeling user corresponding to each labeling result is determined;
[0187] Crowdsourcing points are allocated to each annotation user corresponding to each annotation result based on the Shapley value, first score, and overall score of the target annotation task.
[0188] In one optional implementation, the instructions executed by the processor 510 determine the first point prepaid by each annotation user for participating in the annotation activity of the target annotation task through the following steps:
[0189] The preset fixed points corresponding to the target annotation task are determined as the first points prepaid by each annotation user for participating in the annotation behavior of the target annotation task; or...
[0190] Based on each user's ability level and a preset points mapping relationship, the first points prepaid by each user for participating in the target annotation task are determined.
[0191] In an optional implementation, the instructions executed by the processor 510, wherein determining the first points prepaid by each annotation user for participating in the annotation task based on the user ability level of each annotation user and a preset points mapping relationship, includes:
[0192] For each labeled user, based on the user's ability level and a preset ability level range, determine the target ability level range to which the user's ability level belongs, and determine the interval points corresponding to the target ability level range as the first points prepaid by the labeled user for participating in the labeling behavior of the target labeling task; or;
[0193] Based on the user ability level of each annotation user and a preset negative benefit threshold, a linear mapping relationship between the user ability level and the first point is determined, and based on the linear mapping relationship of the first point, the first point prepaid by each annotation user for participating in the annotation behavior of the target annotation task is determined; wherein, the linear mapping relationship of the first point indicates that the user ability level of the annotation user is negatively correlated with the first point.
[0194] In an optional implementation, the instructions executed by the processor 510, wherein performing truth inference on the plurality of annotation results to obtain the truth label of the target annotation task, includes:
[0195] Based on the user ability level of the annotating user corresponding to each annotation result, the annotation result, and the candidate answers of the target annotation task, determine the target annotation result with the largest proportion among multiple annotation results;
[0196] Based on the target annotation results, the ground truth labels for the target annotation task are obtained.
[0197] In one optional implementation, the characteristic function is determined in the instructions executed by the processor 510 through the following steps:
[0198] Determine the set of multiple annotations completed by multiple annotation users corresponding to multiple annotation results;
[0199] The feature function is determined based on the user ability level and annotation results of each annotated user included in each annotation permutation set.
[0200] In an optional implementation, the instructions executed by the processor 510, wherein determining the Shapley value of the annotation user corresponding to each annotation result based on the feature function, includes:
[0201] For each labeled user, the marginal contribution of that user to the feature function in each set of labeled permutations is calculated.
[0202] For each labeled user, based on the set weights corresponding to each labeled set, the average marginal contribution of the labeled user to the feature function in each labeled set is calculated to obtain the Shapley value of the labeled user.
[0203] In an optional implementation, the instructions executed by the processor 510, which include allocating crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, the first score, and the overall score of the target annotation task, include:
[0204] Identify at least one negative contributing annotation user and at least one positive contributing annotation user among multiple annotation users who have made negative Shapley value contributions for the target annotation task.
[0205] Based on the preset negative contribution ratio and the overall score of the target labeling task, the total deduction points for at least one negative contribution labeling user are determined.
[0206] For each negative contribution labeled user, the deduction points for the negative contribution labeled user are determined based on the proportion of the negative contribution labeled user's Shapley value in the total Shapley value of the at least one negative contribution labeled user and the total deduction points. The first point of the negative contribution labeled user is deducted from the deduction points and then returned to the labeled user.
[0207] For each positive contribution labeling user, the crowdsourcing points allocated to that positive contribution labeling user are determined based on the user's Shapley value, the overall score of the target labeling task, and the total deducted points.
[0208] In an optional implementation, the instructions executed by the processor 510, wherein determining the crowdsourcing points allocated to each positive contribution labeling user based on the user's Shapley value, the overall score of the target labeling task, and the total deducted points, includes:
[0209] Based on the preset platform allocation ratio and the total deduction points, the second points allocated to the crowdsourcing platform are deducted from the total deduction points to determine the additional reward points allocated to the at least one positive contribution user.
[0210] Add the additional reward points and the total points to determine the total reward points to be allocated to the at least one positive contribution user.
[0211] For each positive contribution user, the crowdsourcing points for that positive contribution user are determined based on the proportion of that positive contribution user's Shapley value in the total Shapley value of the at least one positive contribution user and the total reward points.
[0212] In one optional implementation, the instructions executed by the processor 510 further include:
[0213] For each user with a positive contribution, after the crowdsourcing points are allocated to that user, all the first points prepaid by that user are returned to that user.
[0214] In one optional implementation, the instructions executed by the processor 510 further include:
[0215] Detect whether there are any users among those who have negatively contributed whose first score is less than the score deducted from their score.
[0216] If such a user exists, and the proportion of the number of users whose first score is less than the number of users whose score is deducted is greater than a preset proportion threshold in the total number of users with negative contributions, then the preset negative contribution proportion is adjusted.
[0217] Through the above methods, annotation users undertaking target annotation tasks need to prepay a certain amount of initial points for the task to prevent them from completing the task with low quality. Simultaneously, the allocation of points takes into account each annotation user's contribution to the target annotation task, which helps improve the accuracy of the annotation results and user participation. Furthermore, the higher the user's ability level, the lower the initial points required for the same target annotation task, incentivizing users to increase their enthusiasm for undertaking target annotation tasks and improve the quality of their annotations, thereby raising their own ability level. Further, a different dynamic initial point allocation method is adopted for different annotation users to ensure that all annotation users can prepay the corresponding initial points, improving the accuracy of initial point prepayment and user enthusiasm for undertaking target annotation tasks. Finally, the preset negative contribution ratio is adjusted to reduce the points deducted from each negative contribution user, thus minimizing the points deducted from each negative contribution annotation user, ensuring that the initial points of all negative contribution annotation users participating in the target annotation task are sufficient to cover their negative contributions.
[0218] This application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the following instructions:
[0219] Obtain multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user's ability level;
[0220] The truth labels for the target annotation task are obtained by performing truth inference on the multiple annotation results.
[0221] Based on the truth labels, a feature function is determined, and based on the feature function, the Shapley value of the labeling user corresponding to each labeling result is determined;
[0222] Crowdsourcing points are allocated to each annotation user corresponding to each annotation result based on the Shapley value, first score, and overall score of the target annotation task.
[0223] In one alternative implementation, the instructions executable by the computer-readable storage medium determine the first point prepaid by each annotation user for participating in the annotation activity of the target annotation task through the following steps:
[0224] The preset fixed points corresponding to the target annotation task are determined as the first points prepaid by each annotation user for participating in the annotation behavior of the target annotation task; or...
[0225] Based on each user's ability level and a preset points mapping relationship, the first points prepaid by each user for participating in the target annotation task are determined.
[0226] In one optional implementation, the instructions executed by the computer-readable storage medium, wherein determining the first point prepaid by each annotation user for participating in the annotation task based on the user ability level of each annotation user and a preset point mapping relationship, includes:
[0227] For each labeled user, based on the user's ability level and a preset ability level range, determine the target ability level range to which the user's ability level belongs, and determine the interval points corresponding to the target ability level range as the first points prepaid by the labeled user for participating in the labeling behavior of the target labeling task; or;
[0228] Based on the user ability level of each annotation user and a preset negative benefit threshold, a linear mapping relationship between the user ability level and the first point is determined, and based on the linear mapping relationship of the first point, the first point prepaid by each annotation user for participating in the annotation behavior of the target annotation task is determined; wherein, the linear mapping relationship of the first point indicates that the user ability level of the annotation user is negatively correlated with the first point.
[0229] In one optional implementation, the instructions executable by the computer-readable storage medium, wherein the step of performing truth inference on the plurality of annotation results to obtain the truth label of the target annotation task, includes:
[0230] Based on the user ability level of the annotating user corresponding to each annotation result, the annotation result, and the candidate answers of the target annotation task, determine the target annotation result with the largest proportion among multiple annotation results;
[0231] Based on the target annotation results, the ground truth labels for the target annotation task are obtained.
[0232] In one optional implementation, the characteristic function is determined in the instructions executable by the computer-readable storage medium through the following steps:
[0233] Determine the set of multiple annotations completed by multiple annotation users corresponding to multiple annotation results;
[0234] The feature function is determined based on the user ability level and annotation results of each annotated user included in each annotation permutation set.
[0235] In one optional implementation, the instructions executable by the computer-readable storage medium, wherein determining the Shapley value of the annotation user corresponding to each annotation result based on the feature function, includes:
[0236] For each labeled user, the marginal contribution of that user to the feature function in each set of labeled permutations is calculated.
[0237] For each labeled user, based on the set weights corresponding to each labeled set, the average marginal contribution of the labeled user to the feature function in each labeled set is calculated to obtain the Shapley value of the labeled user.
[0238] In one optional implementation, the instructions executed by the computer-readable storage medium, wherein allocating crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, the first score, and the overall score of the target annotation task for each annotation result, includes:
[0239] Identify at least one negative contributing annotation user and at least one positive contributing annotation user among multiple annotation users who have made negative Shapley value contributions for the target annotation task.
[0240] Based on the preset negative contribution ratio and the overall score of the target labeling task, the total deduction points for at least one negative contribution labeling user are determined.
[0241] For each negative contribution labeled user, the deduction points for the negative contribution labeled user are determined based on the proportion of the negative contribution labeled user's Shapley value in the total Shapley value of the at least one negative contribution labeled user and the total deduction points. The first point of the negative contribution labeled user is deducted from the deduction points and then returned to the labeled user.
[0242] For each positive contribution labeling user, the crowdsourcing points allocated to that positive contribution labeling user are determined based on the user's Shapley value, the overall score of the target labeling task, and the total deducted points.
[0243] In one optional implementation, the instructions executable by the computer-readable storage medium, wherein determining the crowdsourcing points allocated to each positive contribution labeling user based on the user's Shapley value, the overall score of the target labeling task, and the total deducted points, includes:
[0244] Based on the preset platform allocation ratio and the total deduction points, the second points allocated to the crowdsourcing platform are deducted from the total deduction points to determine the additional reward points allocated to the at least one positive contribution user.
[0245] Add the additional reward points and the total points to determine the total reward points to be allocated to the at least one positive contribution user.
[0246] For each positive contribution user, the crowdsourcing points for that positive contribution user are determined based on the proportion of that positive contribution user's Shapley value in the total Shapley value of the at least one positive contribution user and the total reward points.
[0247] In one optional implementation, the instructions executed by the computer-readable storage medium further include:
[0248] For each user with a positive contribution, after the crowdsourcing points are allocated to that user, all the first points prepaid by that user are returned to that user.
[0249] In one optional implementation, the instructions executed by the computer-readable storage medium further include:
[0250] Detect whether there are any users among those who have negatively contributed whose first score is less than the score deducted from their score.
[0251] If such a user exists, and the proportion of the number of users whose first score is less than the number of users whose score is deducted is greater than a preset proportion threshold in the total number of users with negative contributions, then the preset negative contribution proportion is adjusted.
[0252] Through the above methods, annotation users undertaking target annotation tasks need to prepay a certain amount of initial points for the task to prevent them from completing the task with low quality. Simultaneously, the allocation of points takes into account each annotation user's contribution to the target annotation task, which helps improve the accuracy of the annotation results and user participation. Furthermore, the higher the user's ability level, the lower the initial points required for the same target annotation task, incentivizing users to increase their enthusiasm for undertaking target annotation tasks and improve the quality of their annotations, thereby raising their own ability level. Further, a different dynamic initial point allocation method is adopted for different annotation users to ensure that all annotation users can prepay the corresponding initial points, improving the accuracy of initial point prepayment and user enthusiasm for undertaking target annotation tasks. Finally, the preset negative contribution ratio is adjusted to reduce the points deducted from each negative contribution user, thus minimizing the points deducted from each negative contribution annotation user, ensuring that the initial points of all negative contribution annotation users participating in the target annotation task are sufficient to cover their negative contributions.
[0253] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0254] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0255] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0256] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0257] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0258] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A crowdsourcing points allocation method, characterized in that, The crowdsourcing points allocation method includes: Obtain multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user's ability level; The truth labels for the target annotation task are obtained by performing truth inference on the multiple annotation results. Based on the truth labels, a feature function is determined, and based on the feature function, the Shapley value of the labeling user corresponding to each labeling result is determined; Crowdsourcing points are allocated to each annotation user corresponding to each annotation result based on the Shapley value, first score, and overall score of the target annotation task.
2. The crowdsourcing points allocation method according to claim 1, characterized in that, The first point prepaid by each annotation user for participating in the annotation task is determined through the following steps: The preset fixed points corresponding to the target annotation task are determined as the first points prepaid by each annotation user for participating in the annotation behavior of the target annotation task; or... Based on each user's ability level and a preset points mapping relationship, the first points prepaid by each user for participating in the target annotation task are determined.
3. The crowdsourcing points allocation method according to claim 2, characterized in that, The determination of the first points prepaid by each annotation user for participating in the target annotation task, based on each annotation user's ability level and a preset point mapping relationship, includes: For each labeled user, based on the user's ability level and a preset ability level range, determine the target ability level range to which the user's ability level belongs, and determine the interval points corresponding to the target ability level range as the first points prepaid by the labeled user for participating in the labeling behavior of the target labeling task; or; Based on the user ability level of each annotation user and a preset negative benefit threshold, a linear mapping relationship between the user ability level and the first point is determined, and based on the linear mapping relationship of the first point, the first point prepaid by each annotation user for participating in the annotation behavior of the target annotation task is determined; wherein, the linear mapping relationship of the first point indicates that the user ability level of the annotation user is negatively correlated with the first point.
4. The crowdsourcing points allocation method according to claim 1, characterized in that, The step of performing truth inference on the multiple annotation results to obtain the truth label for the target annotation task includes: Based on the user ability level of the annotating user corresponding to each annotation result, the annotation result, and the candidate answers of the target annotation task, determine the target annotation result with the largest proportion among multiple annotation results; Based on the target annotation results, the ground truth labels for the target annotation task are obtained.
5. The crowdsourcing points allocation method according to claim 1, characterized in that, The characteristic function is determined through the following steps: Determine the set of multiple annotations completed by multiple annotation users corresponding to multiple annotation results; The feature function is determined based on the user ability level and annotation results of each annotated user included in each annotation permutation set.
6. The crowdsourcing points allocation method according to claim 5, characterized in that, The step of determining the Shapley value of the annotation user corresponding to each annotation result based on the feature function includes: For each labeled user, the marginal contribution of that user to the feature function in each set of labeled permutations is calculated. For each labeled user, based on the set weights corresponding to each labeled set, the average marginal contribution of the labeled user to the feature function in each labeled set is calculated to obtain the Shapley value of the labeled user.
7. The crowdsourcing points allocation method according to claim 1, characterized in that, The process of allocating crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, first score, and overall score of the target annotation task for each annotation result includes: Identify at least one negative contributing annotation user and at least one positive contributing annotation user among multiple annotation users who have made negative Shapley value contributions for the target annotation task. Based on the preset negative contribution ratio and the overall score of the target labeling task, the total deduction points for at least one negative contribution labeling user are determined. For each negative contribution labeled user, the deduction points for the negative contribution labeled user are determined based on the proportion of the negative contribution labeled user's Shapley value in the total Shapley value of the at least one negative contribution labeled user and the total deduction points. The first point of the negative contribution labeled user is deducted from the deduction points and then returned to the labeled user. For each positive contribution labeling user, the crowdsourcing points allocated to that positive contribution labeling user are determined based on the user's Shapley value, the overall score of the target labeling task, and the total deducted points.
8. The crowdsourcing points allocation method according to claim 7, characterized in that, For each positive contribution labeling user, the crowdsourcing points allocated to that user are determined based on their Shapley value, the overall score of the target labeling task, and the total deducted points. This includes: Based on the preset platform allocation ratio and the total deduction points, the second points allocated to the crowdsourcing platform are deducted from the total deduction points to determine the additional reward points allocated to the at least one positive contribution user. Add the additional reward points and the total points to determine the total reward points to be allocated to the at least one positive contribution user. For each positive contribution user, the crowdsourcing points for that positive contribution user are determined based on the proportion of that positive contribution user's Shapley value in the total Shapley value of the at least one positive contribution user and the total reward points.
9. The crowdsourcing points allocation method according to claim 8, characterized in that, The crowdsourcing points allocation method also includes: For each user with a positive contribution, after the crowdsourcing points are allocated to that user, all the first points prepaid by that user are returned to that user.
10. The crowdsourcing points allocation method according to claim 7, characterized in that, The crowdsourcing points allocation method also includes: Detect whether there are any users among those who have negatively contributed whose first score is less than the score deducted from their score. If such a user exists, and the proportion of the number of users whose first score is less than the number of users whose score is deducted is greater than a preset proportion threshold in the total number of users with negative contributions, then the preset negative contribution proportion is adjusted.
11. A crowdsourcing points allocation device, characterized in that, The crowdsourcing points allocation device includes: The information acquisition module is used to acquire multiple annotation results of the target annotation task, the annotation user corresponding to each annotation result, and the first score corresponding to the target annotation task; the first score is the score prepaid by each annotation user for participating in the annotation behavior of the target annotation task, and the first score includes the fixed score corresponding to the target annotation task and / or the dynamic score determined based on the user's ability level; The truth label inference module is used to perform truth inference on the multiple annotation results to obtain the truth label of the target annotation task; The numerical determination module is used to determine a feature function based on the truth label, and to determine the Shapley value of the labeling user corresponding to each labeling result based on the feature function; The points allocation module is used to allocate crowdsourcing points to each annotation user corresponding to each annotation result based on the Shapley value, first points, and the overall points of the target annotation task for each annotation user corresponding to each annotation result.
12. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the crowdsourcing points allocation method as described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the crowdsourcing points allocation method as described in any one of claims 1 to 10.