Labeling task allocation method and device, equipment and medium
By acquiring task characteristics and annotator capability profiles, and using predictive models to evaluate and dynamically adjust annotation task allocation, the problem of unreasonable annotation task allocation in existing technologies is solved, thereby improving annotation efficiency and accuracy.
Patent Information
- Application Number
- CN202510855952.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, annotation task allocation strategies ignore individual differences among annotators and the complexity of tasks, resulting in low efficiency and high error rates. They also lack real-time feedback mechanisms and cannot dynamically adjust allocation strategies.
By acquiring task characteristics and annotator capability profiles, a predictive model is used to assess annotators' task completion capabilities, dynamically adjust task allocation, and monitor task status in real time to optimize allocation strategies.
It improves the efficiency and accuracy of annotation tasks, and the dynamic adjustment of the allocation strategy makes task allocation more reasonable, thereby improving the annotation quality.
Smart Images

Figure CN120875320A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, device and medium for assigning annotation tasks. Background Technology
[0002] Currently, with the increasing demand for training artificial intelligence models, data annotation, as a crucial step in building high-quality AI models, faces significant challenges in efficiency and quality, becoming a bottleneck restricting AI development. Traditional task allocation strategies often employ random or simple round-robin methods, neglecting the differences between individual annotators and the diversity of task complexity. This extensive allocation approach leads to complex tasks being assigned to annotators with less expertise; for example, in medical image annotation, assigning tasks to non-medical professionals can reduce efficiency by over 60%, demonstrating a significant decrease in processing efficiency. Furthermore, high-difficulty tasks handled by low-skilled annotators increase the annotation error rate, potentially reaching 15%-20%, impacting the training effect of AI models. In addition, existing allocation methods lack real-time feedback mechanisms, failing to dynamically adjust allocation strategies based on annotator skill improvements or changes in task characteristics.
[0003] Therefore, how to reasonably allocate annotation tasks in order to improve the accuracy and efficiency of annotation has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method, apparatus, device, and medium for allocating annotation tasks to solve the problem of how to reasonably allocate annotation tasks in order to improve the accuracy and efficiency of annotation.
[0005] In a first aspect, embodiments of this application provide a method for assigning annotation tasks, including: Obtain the task characteristics of the task to be labeled, and obtain the capability profile characteristics of each labeler in the candidate labeler set; Input the task features and the ability profile features of all annotators into the trained prediction model, and output the prediction evaluation result of each annotator's completion of the task to be annotated; Based on the prediction and evaluation results, a target number of annotators are selected as target annotators, and the tasks to be annotated are sent to the target annotators according to a preset ratio. Obtain the real-time task status of each target annotator for the task to be labeled, adjust the parameters in the trained prediction model according to the real-time task status, and return to the process of inputting the task features and the ability profile features of all annotators into the trained prediction model until the task to be labeled is completed.
[0006] Secondly, embodiments of this application provide a labeling task allocation device, comprising: The feature acquisition module is used to acquire the task features of the task to be labeled, and to acquire the capability profile features of each labeler in the candidate labeler set; The prediction and evaluation module is used to input the task features and the ability profile features of all annotators into the trained prediction model and output the prediction and evaluation results of each annotator in completing the task to be labeled. The task allocation module is used to select a target number of annotators as target annotators based on the prediction and evaluation results, and send the tasks to be annotated to the target annotators according to a preset ratio; The dynamic adjustment module is used to obtain the real-time task status of each target annotator for the task to be labeled, adjust the parameters in the trained prediction model according to the real-time task status, and return to the process of inputting the task features and the ability profile features of all annotators into the trained prediction model until the task to be labeled is completed.
[0007] Thirdly, embodiments of this application provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the annotation task allocation method as described in the first aspect.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the annotation task allocation method as described in the first aspect.
[0009] The beneficial effects of the embodiments in this application compared with the prior art are: The process involves acquiring the task features of the task to be labeled and the capability profile features of each annotator in the candidate annotator set. These features are then input into a trained prediction model, which outputs a predicted evaluation result for each annotator's completion of the task. Based on this evaluation result, a target number of annotators are selected as target annotators. The task to be labeled is then sent to the target annotators according to a preset ratio. The real-time task status of each target annotator is obtained, and the parameters in the trained prediction model are adjusted based on this status. The process continues until the task to be labeled is completed. Specifically, using the annotator's capability profile features and the task features of the task to be labeled for task execution evaluation and prediction helps select target annotators, assign tasks to them, and monitor their task status to dynamically adjust the task allocation, improving the rationality of the allocation and effectively increasing the efficiency and accuracy of task labeling. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for a labeling task allocation method provided in Embodiment 1 of this application; Figure 2 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 2 of this application; Figure 3 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 3 of this application; Figure 4 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 4 of this application; Figure 5 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 5 of this application; Figure 6 This is an overall schematic diagram of a labeling task allocation method provided in Embodiment 5 of this application; Figure 7 This is a schematic diagram of the structure of a labeling task allocation device provided in Embodiment Six of this application; Figure 8 This is a schematic diagram of the structure of a computer device provided in Embodiment 7 of this application. Detailed Implementation
[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0013] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0014] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0015] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0016] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0017] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0018] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0019] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0020] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0021] To illustrate the technical solution of this application, specific embodiments are described below.
[0022] The annotation task allocation method provided in Embodiment 1 of this application can be applied to, for example, Figure 1 In this application environment, the client communicates with the server. The client includes, but is not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0023] See Figure 2 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 2 of this application. The above-described labeling task allocation method can be applied to... Figure 1The server-side component consists of computer devices that connect to clients and databases. Clients can correspond to task publishers and task receivers (i.e., annotators). The task publisher sends the task to be annotated to the server, which processes it accordingly and then distributes the task to the annotator on the corresponding client. In addition, the database stores the data information of the corresponding annotators and the data information corresponding to the task to be annotated. The server can directly obtain this data from the database.
[0024] like Figure 2 As shown, the annotation task assignment method may include the following steps: Step S201: Obtain the task characteristics of the task to be labeled, and obtain the capability profile characteristics of each labeler in the candidate labeler set.
[0025] The task to be labeled can refer to a task used to label a set of data or a single data, including but not limited to labeling image data, text data, and speech data.
[0026] Task features are used to characterize the task type, task requirements, task difficulty, data volume, and other features of the task to be labeled. These task features can be encoded by a corresponding feature extractor to form feature data that can be recognized by subsequent models, such as feature vectors. Of course, the above-mentioned task features can be formed by concatenating multiple features or by fusing multiple features to form a fused feature.
[0027] For example, the annotation task for medical images specifically involves annotating the types of lesions (such as tumors, nodules, etc.) in medical images. These medical images can include CT images, MRI images, etc. The difficulty of the annotation task can be divided into simple, medium, and complex levels. The aforementioned lesion types, image modes, and annotation difficulty can form the task feature expression for this annotation task.
[0028] The candidate annotator set is either a collection of selected candidate annotators or a collection of all annotators. Each annotator corresponds to a human entity. Therefore, each annotator has its own capability profile features, which are profile features that characterize the annotator's annotation capabilities, including but not limited to annotation efficiency, annotation skills, annotation accuracy, and task completion rate.
[0029] This capability profile feature can be obtained by analyzing the user information of the annotators, which includes basic information (e.g., education, major, years of work experience), skill information (e.g., skill assessment results, such as expert level, novice level, etc.), and historical behavior information (e.g., average annotation duration, modification frequency, annotation accuracy rate, etc.).
[0030] Optionally, this annotation task allocation method also includes: Get the task type of the task to be labeled, and get the labeling skill information of all labelers to be assigned tasks; The task type of the task to be labeled is matched one by one with the labeling skill information of all labelers to be assigned tasks, and the matched labelers are formed into a candidate labeler set.
[0031] For a large number of annotators, a coarse matching process can be added before performing step S201 above. This involves obtaining the annotation skill information of all annotators and matching this information with the task type to filter out annotators who can complete the task type.
[0032] This annotation skill information includes, but is not limited to, generating tags such as "Natural Language Processing - Text Classification: Expert Level" and "Medical Image Annotation Skills" through regularly organized skills assessments. When the task type is medical image annotation, annotators with "Medical Image Annotation Skills" are required. In this case, it is necessary to match all annotators with "Medical Image Annotation Skills" from all annotators to form a candidate annotator set.
[0033] Using coarse matching can effectively reduce the number of subsequent matches for annotators, thus improving matching efficiency.
[0034] Optionally, the task features of the task to be labeled can be obtained, including: Obtain the tasks to be labeled, extract the task type, labeling difficulty, and labeling objects for the corresponding tasks; Feature extraction is performed on task type, annotation difficulty, and annotation object to obtain the task features of the corresponding task to be annotated.
[0035] In this process, information is extracted from the task to be labeled to determine the task type, labeling difficulty, and labeling object. If the task type is the same as the task type directly obtained above, there is no need to extract it again; it can be obtained directly.
[0036] The labeling difficulty is a label representing the difficulty level of the task to be labeled. The labeling object is the data object in the task to be labeled, such as image data, text data, and voice data.
[0037] Step S202: Input the task features and the capability profile features of all annotators into the trained prediction model, and output the prediction evaluation results of each annotator's completion of the annotation task.
[0038] The input to the trained prediction model consists of two features: task features and capability profile features. The trained prediction model is used to calculate the prediction and evaluation results of the task to be labeled by performing the corresponding task based on the capability profile features.
[0039] For each annotator in the candidate annotation set, a prediction evaluation is performed. Each annotator corresponds to a prediction evaluation result, which is used to represent the efficiency, accuracy, and completion rate of the corresponding annotator in completing the annotation task.
[0040] The trained prediction model can employ reinforcement learning models such as neural networks and machine learning. The model takes task features and competency profile features as input and outputs evaluation results such as the quality and efficiency of completing the task using the competency profile features. During training, it collects the features of individuals who have completed the task and the competency profile features of the annotators who completed the task, and labels them with human evaluations of the completed task (including quality and efficiency). These two types of features and labels constitute the training set. Since annotators need to evaluate each annotation task, the training set can be obtained relatively easily.
[0041] Step S203: Based on the prediction and evaluation results, select the target number of annotators as target annotators, and send the annotated tasks to the target annotators according to a preset ratio.
[0042] The target number is a quantity parameter that is dynamically adjusted according to demand, used to limit the number of people who need to complete the task to be labeled. For example, the target number is 3.
[0043] If the predicted evaluation result is a quantitative value, then the predicted evaluation result is numerically sorted to obtain the predicted evaluation results for the top-ranked targets. The annotator corresponding to the top-ranked predicted evaluation result is the target annotator. If the predicted evaluation result is not a quantitative value, then methods such as matching calculations can be used to filter out the predicted evaluation results for the target targets. The annotator corresponding to the filtered predicted evaluation results is the target annotator.
[0044] The preset ratio can be dynamically adjusted according to needs. This preset ratio represents the allocation ratio of tasks, that is, a task is divided into corresponding scores according to the target quantity. The task can be allocated in an equal-division preset ratio. Of course, the preset ratio can also be determined based on the prediction evaluation results. If the prediction evaluation results are higher, the allocation ratio of tasks will be higher. For example, if the data volume of the task to be labeled is N, the target quantity is 3, and the prediction ratio is 6:3:1, then the first target labeler will be allocated 3N / 5 of the data volume, and so on.
[0045] Step S204: Obtain the real-time task status of each target annotator in the task to be labeled. Based on the real-time task status, adjust the parameters in the trained prediction model. Return to the execution and input the task features and the ability profile features of all annotators into the trained prediction model until the task to be labeled is completed.
[0046] After assigning the task to be labeled to the corresponding target labeler, the task execution status of the target labeler will be monitored in real time, i.e., real-time task status. This real-time task status is used to characterize the labeler's execution status of the task, including but not limited to workload, quality indicators, task progress, etc. Workload includes the number of currently unfinished tasks, the estimated time to complete the task to be labeled, the average labeling time, etc. Quality indicators include the labeler's recent accuracy rate, modification frequency, second review pass rate, etc. Task progress includes the progress of executing the task to be labeled, etc.
[0047] Based on the real-time task status, adjust the parameters in the trained prediction model. For example, when an annotator's current number of uncompleted tasks exceeds a threshold (e.g., 5) or the estimated completion time exceeds the task deadline, trigger the load balancing mechanism to assign new tasks to other candidates or recalculate the matching process. When an annotator's recent accuracy or second-round approval rate drops significantly, reduce their weight in the matching algorithm, decrease their chances of being assigned new tasks, and assign quality monitoring personnel to provide focused attention and guidance. When task progress lags behind or the data volume suddenly increases, adjust the matching strategy promptly, increase the number of annotators, or adjust the task allocation ratio to ensure timely task completion. For high-difficulty or urgent tasks, prioritize assigning them to annotators with rich experience and high accuracy, and set higher weights and priorities.
[0048] If the assigned task cannot be completed by the target annotator on time and to the required quality, the unannotated portion needs to be reassigned. The reassignment should be carried out according to the steps S202 to S203 above to determine the target annotator, and then the unannotated portion should be reassigned to a new target annotator.
[0049] Optionally, the parameters in the trained prediction model can be adjusted based on the real-time task status, including: For any target annotator, if the number of tasks in the target annotator's real-time task status exceeds the quantity threshold or the corresponding real-time task status does not meet the task progress requirements, then the annotating tasks assigned to the corresponding target annotator will be assigned to other target annotators. If the number of tasks in the real-time task status of other target annotators exceeds the quantity threshold or the corresponding real-time task status does not meet the task progress requirements, then increase the number of targets in the trained prediction model and / or adjust the preset ratio.
[0050] In cases where the workload of target annotators is too heavy, it is necessary to redistribute the data. In this case, the redistributed data volume can be allocated to other target annotators as originally determined. If other target annotators are unable to complete the task, the number of targets needs to be increased or the preset ratio needs to be adjusted so that the redistributed data volume can be completed as required.
[0051] This application embodiment obtains the task characteristics of the task to be labeled and the capability profile characteristics of each annotator in the candidate annotator set. The task characteristics and the capability profile characteristics of all annotators are input into a trained prediction model, which outputs a predicted evaluation result for each annotator's completion of the task. Based on the prediction evaluation results, a target number of annotators are selected as target annotators. The task to be labeled is sent to the target annotators according to a preset ratio. The real-time task status of each target annotator is obtained. Based on the real-time task status, the parameters in the trained prediction model are adjusted. The process returns to inputting the task characteristics and the capability profile characteristics of all annotators into the trained prediction model until the task to be labeled is completed. Specifically, using the capability profile characteristics of the annotators and the task characteristics of the task to be labeled for task execution evaluation and prediction, target annotators are selected, and the task to be labeled is assigned to them. Furthermore, by monitoring the task status of the target annotators, the allocation of labeling tasks is dynamically adjusted, improving the rationality of the allocation and effectively improving the completion efficiency and accuracy of labeling tasks.
[0052] See Figure 3 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 3 of this application. Figure 3 As shown, obtaining the capability profile features of each annotator in the candidate annotator set in step S201 above includes the following steps: Step S301: Obtain historical annotation behavior information for each annotator in the candidate annotator set.
[0053] Step S302: For any annotator, determine the annotation efficiency factor and quality stability factor of the corresponding annotator based on the annotator's historical annotation behavior information.
[0054] Step S303: Determine the capability profile characteristics of the corresponding annotator based on the annotation efficiency factor and the quality stability factor.
[0055] Among them, historical annotation behavior information is used to characterize the status of annotators during the historical execution of tasks, including but not limited to various features such as average annotation duration and modification frequency extracted from annotation operation logs.
[0056] When modeling the profile, factors are used as the basis for profile features. Factor analysis is used to process historical annotation behavior information. For example, key behavioral data such as annotation speed, accuracy, and task completion rate are selected as variables for factor analysis. Statistical software (such as SPSS, R language, etc.) is used to perform factor analysis, extract common factors (such as "annotation efficiency factor", "quality stability factor", etc.), and calculate the score of each common factor as a comprehensive indicator of the annotator profile.
[0057] This application embodiment constructs a profile of the annotator through factor analysis, thereby effectively providing data such as the annotator's annotation efficiency and annotation stability, which helps to improve the prediction accuracy and efficiency in the prediction model.
[0058] See Figure 4 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 4 of this application. Figure 4 As shown, obtaining the capability profile features of each annotator in the candidate annotator set in step S201 above may include the following steps: Step S401: Obtain the basic attribute information and annotation skill information of each annotator in the candidate annotator set, as well as the task requirements of the task to be annotated.
[0059] Step S402: For any annotator, based on the annotator's basic attribute information, annotation skill information, and task requirements, determine the initial efficiency weight of the annotation efficiency factor and the initial stability weight of the quality stability factor for the corresponding annotator.
[0060] The step S303 above, which determines the corresponding annotator's capability profile based on the annotation efficiency factor and quality stability factor, may include the following steps: Step S403: Based on the annotation efficiency factor and initial efficiency weight, combined with the quality stability factor and initial stability weight, determine the corresponding annotation personnel's capability profile characteristics.
[0061] Among them, the basic attribute information is used to characterize the basic information of the annotators. Information such as the annotators' education, major, and years of work experience are collected through the onboarding questionnaire. The annotation skill information can be referred to the description of the annotation skill information in the above embodiment 2. The annotation skill information obtained here is the information extraction of the annotators in the candidate annotator set.
[0062] This embodiment constructs dynamic weights, setting initial weights for each comprehensive indicator based on factors such as task urgency and annotator skill level. For example, a dynamic adjustment mechanism is introduced: when task urgency increases, the weight of the "annotation speed factor" is appropriately increased; when the task has high quality requirements, the weight of the "quality stability factor" is appropriately increased.
[0063] By combining weights and corresponding factors, the ability profile of an annotator can be constructed. This feature can be adaptively optimized through dynamic weight adjustment. By continuously optimizing the weight settings through historical annotation behavior information and experimental verification, the accuracy and practicality of the annotator profile can be ensured.
[0064] See Figure 5 This is a flowchart illustrating a labeling task allocation method provided in Embodiment 5 of this application. Figure 5As shown, after the annotation task in step S204 above is completed, the following steps may be included: Step S501: Obtain the annotation operation data of each target annotator for the assigned annotation task, identify suspicious annotations in the annotation operation data, and obtain suspicious annotations.
[0065] Step S502: Review the suspicious annotations, obtain the review results, and update the capability profile features of the target annotator corresponding to the suspicious annotations based on the review results.
[0066] After each annotation task is completed, the annotation results will be reviewed for quality to provide feedback and optimize the prediction model and the annotator profile model.
[0067] By analyzing data such as annotation trajectory, operation heatmap, annotation duration, and modification frequency in real time, suspicious annotations (such as excessive annotation frame adjustments, excessively short annotation duration, and excessively high modification frequency) are identified. Suspicious annotations undergo a second review, and the annotator profile (such as adjustment accuracy and quality stability indicators) and matching algorithm parameters are updated based on the review results.
[0068] In addition, by statistically analyzing indicators such as task completion rate, accuracy rate, and annotation speed of each annotator on a weekly or daily basis, a "skill-task" matching degree matrix is generated.
[0069] The matching algorithm parameters were optimized using a genetic algorithm, with the objective function being to maximize both the overall annotation efficiency and the quality-weighted score. The specific adjustment process is as follows: Analyze the matching matrix: identify the strengths and weaknesses of annotators in different task types, and discover mismatches; Adjusting factor weights: Based on the matching degree matrix analysis results, adjust the weights of factors such as annotation speed, accuracy, and task completion rate to more accurately reflect the actual capabilities of the annotators and the task requirements; Optimize reinforcement learning model parameters: If the matching algorithm uses a reinforcement learning model, retrain the model based on the matching degree matrix feedback, and adjust parameters such as learning rate, discount factor, and neural network structure to improve the accuracy and efficiency of the model in task allocation; Introduce new features or rules: Based on the matching degree matrix analysis results, introduce new features (such as the annotator's emotional state, work environment, etc.) or formulate new rules (such as limiting the number of times an annotator completes the same type of task consecutively) to enrich the annotator profile and optimize the task allocation strategy.
[0070] See Figure 6This is a schematic diagram of an annotation task allocation method provided in Embodiment 5 of this application. After dynamic task matching, quality monitoring and strategy optimization are performed to adjust the weights of factors for each annotator during annotator profile modeling. The prediction model and its related parameters in dynamic task matching can also be adjusted. This implementation effectively improves the efficiency and quality of data annotation, and also enhances the adaptability and flexibility of the entire matching process.
[0071] Corresponding to the annotation task allocation method in the above embodiment, Figure 7 This diagram shows a structural block diagram of the annotation task allocation device provided in Embodiment Six of this application. The annotation task allocation device is applied to... Figure 1 The server-side component connects to the client and database via a corresponding computer device. The client can correspond to a task publisher and a task receiver (i.e., annotator). The task publisher sends the task to be annotated to the server, which processes it and then distributes the task to the corresponding annotator on the client. Additionally, the database stores data information for the annotators and the tasks to be annotated, which the server can directly retrieve. For ease of explanation, only the parts relevant to the embodiments of this application are shown.
[0072] See Figure 7 The labeling task allocation device includes: The feature acquisition module 71 is used to acquire the task features of the task to be labeled, and to acquire the capability profile features of each labeler in the candidate labeler set; The prediction and evaluation module 72 is used to input the task features and the ability profile features of all annotators into the trained prediction model and output the prediction and evaluation results of each annotator's completion of the task to be annotated. The task allocation module 73 is used to select a target number of annotators as target annotators based on the prediction and evaluation results, and send the tasks to be annotated to the target annotators according to a preset ratio. The dynamic adjustment module 74 is used to obtain the real-time task status of each target annotator in the task to be labeled, adjust the parameters in the trained prediction model according to the real-time task status, and return to input the task features and the ability profile features of all annotators into the trained prediction model until the task to be labeled is completed.
[0073] Optionally, the feature acquisition module 71 includes: The first information acquisition unit is used to acquire historical annotation behavior information of each annotator in the candidate annotator set; The factor determination unit is used to determine the annotation efficiency factor and quality stability factor for any given annotator based on the annotator's historical annotation behavior information. The profile feature determination unit is used to determine the capability profile features of the corresponding annotator based on the annotation efficiency factor and the quality stability factor.
[0074] Optionally, the feature acquisition module 71 also includes: The second information acquisition unit is used to acquire the basic attribute information and annotation skill information of each annotator in the candidate annotator set, as well as the task requirements of the task to be annotated; The initial weight determination unit is used to determine the initial efficiency weight of the annotation efficiency factor and the initial stability weight of the quality stability factor for any annotator, based on the annotator's basic attribute information, annotation skill information, and task requirements. The portrait feature determination unit includes: The profile feature determination subunit is used to determine the capability profile features of the corresponding annotator based on the annotation efficiency factor and initial efficiency weight, combined with the quality stability factor and initial stability weight.
[0075] Optionally, the annotation task allocation device also includes: The task information acquisition module is used to obtain the task type of the task to be labeled, as well as the labeling skill information of all labelers to be assigned tasks; The candidate annotator determination module is used to match the task type of the task to be annotated with the annotation skill information of all annotators to be assigned tasks, determine the matching annotators, and form a candidate annotator set by all matching annotators.
[0076] Optionally, the feature acquisition module 71 includes: The task acquisition unit is used to acquire tasks to be labeled, extract tasks to be labeled, and obtain the task type, labeling difficulty, and labeling object of the corresponding tasks to be labeled. The task feature analysis unit is used to extract features from task type, annotation difficulty, and annotation object to obtain the task features of the corresponding task to be annotated.
[0077] Optionally, the dynamic adjustment module 74 includes: The task redistribution unit is used to redistribute the tasks assigned to the target annotator to other target annotators if the number of tasks in the target annotator's real-time task status exceeds the quantity threshold or the corresponding real-time task status does not meet the task progress requirements. The parameter condition unit is used to increase the number of targets in the trained prediction model and / or adjust the preset ratio if the number of tasks in the real-time task status of other target annotators exceeds the quantity threshold or the corresponding real-time task status does not meet the task progress requirements.
[0078] Optionally, the annotation task allocation device also includes: The suspicious annotation identification module is used to obtain the annotation operation data of each target annotator on the assigned annotation task after the task to be annotated is completed, and to identify suspicious annotations from the annotation operation data. The profile feature update module is used to review suspicious annotations, obtain review results, and update the capability profile features of the target annotator corresponding to the suspicious annotations based on the review results.
[0079] It should be noted that the information interaction and execution process between the above modules, units, and sub-units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0080] Figure 8 This is a schematic diagram of the structure of a computer device provided in Embodiment Seven of this application. Figure 8 As shown, the computer device of this embodiment includes: at least one processor ( Figure 8 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, wherein the processor executes the computer program to implement the steps in any of the above-described annotation task allocation method embodiments.
[0081] This computer device may include, but is not limited to, a processor, memory, and a database. Those skilled in the art will understand that... Figure 8 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components. For example, they may also include databases, network interfaces, displays, and input devices, which can be connected via a data bus.
[0082] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0083] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of the computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0084] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0085] The implementation of all or part of the processes in the methods of the above embodiments can also be accomplished by a computer program product. When the computer program product is run on a computer device, it enables the computer device to execute the steps in the above method embodiments.
[0086] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0087] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0088] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0089] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0090] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for assigning annotation tasks, characterized in that, include: Obtain the task characteristics of the task to be labeled, and obtain the capability profile characteristics of each labeler in the candidate labeler set; Input the task features and the ability profile features of all annotators into the trained prediction model, and output the prediction evaluation result of each annotator's completion of the task to be annotated; Based on the prediction and evaluation results, a target number of annotators are selected as target annotators, and the tasks to be annotated are sent to the target annotators according to a preset ratio. Obtain the real-time task status of each target annotator for the task to be labeled, adjust the parameters in the trained prediction model according to the real-time task status, and return to the process of inputting the task features and the capability profile features of all annotators into the trained prediction model until the task to be labeled is completed.
2. The annotation task allocation method according to claim 1, characterized in that, The acquisition of the capability profile features of each annotator in the candidate annotator set includes: Obtain historical annotation behavior information for each annotator in the candidate annotator set; For any annotator, the annotation efficiency factor and quality stability factor are determined based on the annotator's historical annotation behavior information. Based on the annotation efficiency factor and the quality stability factor, determine the corresponding ability profile characteristics of the annotator.
3. The annotation task allocation method according to claim 2, characterized in that, The process of obtaining the capability profile features of each annotator in the candidate annotator set also includes: Obtain the basic attribute information and annotation skill information of each annotator in the candidate annotator set, as well as the task requirements of the task to be annotated; For any annotator, based on the annotator's basic attribute information and annotation skill information, as well as the task requirements, determine the initial efficiency weight of the annotation efficiency factor and the initial stability weight of the quality stability factor corresponding to the annotator. The step of determining the capability profile characteristics corresponding to the annotator based on the annotation efficiency factor and the quality stability factor includes: Based on the annotation efficiency factor and the initial efficiency weight, combined with the quality stability factor and the initial stability weight, the capability profile characteristics corresponding to the annotator are determined.
4. The annotation task allocation method according to claim 1, characterized in that, Also includes: Get the task type of the task to be labeled, and get the labeling skill information of all labelers to be assigned tasks; The task type of the task to be labeled is matched one by one with the labeling skill information of all labelers to be assigned tasks to determine the matching labelers, and all the matching labelers are formed into a candidate labeler set.
5. The annotation task allocation method according to claim 1, characterized in that, The process of obtaining the task features of the task to be labeled includes: Obtain the task to be labeled, extract the task type, labeling difficulty, and labeling object corresponding to the task to be labeled; Feature extraction is performed on the task type, the annotation difficulty, and the annotation object to obtain the task features corresponding to the task to be annotated.
6. The annotation task allocation method according to claim 1, characterized in that, The step of adjusting the parameters in the trained prediction model according to the real-time task status includes: For any target annotator, if the number of tasks in the real-time task status of the target annotator exceeds the quantity threshold or the corresponding real-time task status does not meet the task progress requirements, then the task to be annotated assigned to the corresponding target annotator will be assigned to other target annotators. If the number of tasks in the real-time task status of other target annotators exceeds the quantity threshold or the corresponding real-time task status does not meet the task progress requirements, then the number of targets in the trained prediction model is increased and / or the preset ratio is adjusted.
7. The annotation task allocation method according to any one of claims 1 to 6, characterized in that, After the task to be labeled is completed, the following is also included: Obtain the annotation operation data of each target annotator for the assigned annotation task, and identify suspicious annotations from the annotation operation data to obtain suspicious annotations; The suspicious annotations are reviewed to obtain the review results. Based on the review results, the capability profile features of the target annotator corresponding to the suspicious annotations are updated.
8. A task assignment device, characterized in that, include: The feature acquisition module is used to acquire the task features of the task to be labeled, and to acquire the capability profile features of each labeler in the candidate labeler set; The prediction and evaluation module is used to input the task features and the ability profile features of all annotators into the trained prediction model and output the prediction and evaluation results of each annotator in completing the task to be labeled. The task allocation module is used to select a target number of annotators as target annotators based on the prediction and evaluation results, and send the tasks to be annotated to the target annotators according to a preset ratio; The dynamic adjustment module is used to obtain the real-time task status of each target annotator for the task to be labeled, adjust the parameters in the trained prediction model according to the real-time task status, and return to the process of inputting the task features and the ability profile features of all annotators into the trained prediction model until the task to be labeled is completed.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the annotation task allocation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the annotation task allocation method as described in any one of claims 1 to 7.