A data labeling task distribution method, device, equipment and medium

By dynamically allocating data annotation tasks and combining real-time status and permission information, the limitations of data annotation task distribution have been solved, achieving efficient and reasonable task distribution and resource allocation, and improving project progress and delivery efficiency.

CN121010181BActive Publication Date: 2026-02-06HANG ZHOU MINDFLOW TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511536203.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-02-06
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing data annotation task distribution methods have limited applicability to various scenarios, making it difficult to control project progress, resulting in task flow bottlenecks, processing timeouts, uneven resource allocation, and new teams having difficulty getting opportunities.

Method used

By acquiring the annotation task data package and project configuration information, the task is broken down into independent work units. Combined with the real-time status and permission information of the operators, the task is dynamically allocated, and an integrated processing mechanism for annotation, review, and rework is established to achieve multi-dimensional matching and dynamic resource distribution.

Benefits of technology

It improved project processing efficiency, shortened the average project cycle, ensured reasonable task distribution, avoided resource monopolies and task flow delays, and improved delivery efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010181B_ABST
    Figure CN121010181B_ABST
Patent Text Reader

Abstract

The application discloses a data labeling task distribution method and device, equipment and medium, and relates to the technical field of task distribution. The method comprises the following steps: obtaining labeling task data packets and project configuration information, splitting the labeling task data packets according to project distribution characteristics to obtain at least one batch of first to-be-distributed task sets; obtaining labeling work orders, and generating second to-be-distributed tasks according to the labeling work orders; obtaining repair work orders fed back after the second to-be-distributed tasks are audited, and generating third or fourth to-be-distributed tasks according to the repair work orders; and packaging the first, second, third and / or fourth to-be-distributed tasks into a current batch of work packages according to real-time team states, real-time individual states and permission information, and distributing the work packages to workers. The application considers multiple factors during task distribution, efficiently matches tasks and workers, realizes multi-dimensional matching and dynamic resource distribution, shortens the average project cycle, and significantly improves delivery efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of task distribution, and in particular relates to a data labeling task distribution method, device, equipment and medium. BACKGROUND

[0002] With the rapid development of deep learning, multi-modal technology and big data processing technology, data labeling has become one of the core technologies in the field of computer vision. Data labeling refers to marking, classifying, annotating and other operations on a large amount of initial data, so that machine learning models and subsequent personnel can better understand and process the data.

[0003] At present, data labeling needs to be performed by professional labeling personnel in the field, and in particular, the data labeling in the field of autonomous driving has high standards, so the labeled data needs to undergo quality inspection processing at various inspection nodes such as auditing, quality inspection and acceptance. If the results of auditing and acceptance do not meet the standards, the labeling personnel need to perform corresponding rework on the data to correct errors.

[0004] The distribution of data labeling work is usually based on simple distribution rules or manual allocation by project managers according to experience, which is limited in adaptation to scenarios and difficult to control project progress, which undoubtedly brings great challenges to the overall progress of data labeling projects. In order to better control the project progress, some distribution platforms also consider the business capabilities of workers and the delivery difficulty of projects to distribute tasks, but when facing large-scale labeling tasks, the business capabilities of workers cannot reflect the real-time work status, resulting in poor dynamic adaptability of actual task distribution, which easily causes task flow to be stuck, task processing to be overdue, etc. Moreover, relying on the business capabilities of workers to distribute tasks will make it difficult for new teams to get opportunities and the allocation of resources is not balanced enough.

[0005] Therefore, how to perform more efficient and reasonable labeling task distribution is an important topic to be solved in the industry at present. SUMMARY

[0006] Therefore, the embodiments of the present application provide a data labeling task distribution method, device, equipment and medium, so as to solve the problem that the current distribution of data labeling work is limited in adaptation to scenarios and difficult to control project progress.

[0007] According to a first aspect, the embodiments of the present application provide a data labeling task distribution method, which comprises:

[0008] Obtaining a labeling task data packet and project configuration information, splitting the labeling task data packet according to project distribution characteristics determined from the project configuration information to obtain at least one batch of first to-be-distributed task sets; each batch of first to-be-distributed task sets comprises at least one first to-be-distributed task;

[0009] obtain a labeling work order generated according to the first to-be-distributed task submitted by the operator, and generate a second to-be-distributed task according to the labeling work order;

[0010] obtain a repair work order fed back after the second to-be-distributed task is processed, and generate a third to-be-distributed task or a fourth to-be-distributed task according to the repair work order; the third to-be-distributed task corresponds to a repair work order that is rejected by the audit, and the fourth to-be-distributed task corresponds to a repair work order that passes the audit and needs to be modified;

[0011] determine real-time team states of the teams, real-time personal states of the operators, and permission information, pack the first, second, third, and / or fourth to-be-distributed tasks into a current batch of work packages according to the real-time team states, the real-time personal states, and the permission information, and distribute the work packages to the operators; the real-time team states are determined by a team task upper limit and the real-time personal states of all the operators in the team, the real-time personal states are determined by a number of to-be-processed labeling tasks in the work package and whether the work package is distributed, each operator can only process one work package at the same time, and the number of to-be-processed labeling tasks in the work package does not exceed a personal task upper limit.

[0012] In a first implementation manner of the first aspect, the determining the real-time team states of the teams, the real-time personal states of the operators, and the permission information, the packing the first, second, third, and / or fourth to-be-distributed tasks into the current batch of work packages according to the real-time team states, the real-time personal states, and the permission information, and the distributing the work packages to the operators specifically include:

[0013] determining the real-time team states of the teams;

[0014] in a case where the real-time team state of the team is determined to be a team distributable state, determining the real-time personal states of the operators in the team; the team distributable state is that the real-time personal state of at least one operator is a personal distributable state, and a total number of to-be-processed labeling task numbers in the work packages of all the operators in the team is lower than a team task upper limit.

[0015] in a case where it is determined that the real-time personal state of at least one operator is the personal distributable state, determining whether the operator whose real-time personal state is the personal distributable state has distributed the work package; the personal distributable state includes that the work package is not distributed, and a to-be-processed labeling task number in the work package is less than a personal work upper limit;

[0016] in a case where it is determined that there is at least one operator who has not distributed the work package, packing the first, second, third, and / or fourth to-be-distributed tasks into the current batch of work packages according to the permission information of the operators, and distributing the work packages to the operators who have not distributed the work package;

[0017] In a case where it is determined that all the workers distribute the work packages and there is at least one worker whose number of to-be-processed annotation tasks in the work package is less than the personal work upper limit, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch according to the permission information of the worker, and the task package of the current batch is merged into the work package already possessed by the worker.

[0018] In combination with the first aspect and the first implementation, in a second implementation of the first aspect, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch according to the permission information of the worker, and the work package is distributed to the worker who has not distributed the work package, specifically including:

[0019] In a case where the permission information is determined to be an annotator, the first, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the annotator;

[0020] In a case where the permission information is determined to be an auditor, the second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the auditor;

[0021] In a case where the permission information is determined to be an administrator, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the administrator.

[0022] In combination with the first aspect and the first implementation, in a third implementation of the first aspect, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch according to the permission information of the worker, and the task package of the current batch is merged into the work package already possessed by the worker, specifically including:

[0023] In a case where the permission information is determined to be an annotator, the first, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package of the current batch is merged into the work package of the annotator;

[0024] In a case where the permission information is determined to be an auditor, the second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package of the current batch is merged into the work package of the auditor;

[0025] In a case where the permission information is determined to be an administrator, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package of the current batch is merged into the work package of the administrator.

[0026] In combination with the first aspect and the second implementation, in a fourth implementation of the first aspect, the first, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the annotator, specifically including:

[0027] correlate the first to fourth to-be-distributed tasks to form a traceable chain among all the labeling tasks;

[0028] obtain the third and fourth to-be-distributed tasks, determine the first to-be-distributed task to which the third and fourth to-be-distributed tasks are associated according to the traceable chain, and determine the labeler corresponding to the first to-be-distributed task;

[0029] determine the batch to which each of the third and fourth to-be-distributed tasks belongs, and determine the generation time of each of the third and fourth to-be-distributed tasks;

[0030] determine whether the batch belongs to the current batch being processed by the labeler;

[0031] in the case where the batch belongs to the current batch, pack the third and / or fourth to-be-distributed tasks into a job package of the current batch, and distribute the job package to the labeler;

[0032] in the case where the batch is located before the current batch, determine the packing order between the third and fourth tasks according to the generation time, pack the first, third and / or fourth to-be-distributed tasks into a job package of the current batch according to the packing order, and pack the third and fourth tasks of the same batch into the same job package, and distribute the job package to the labeler;

[0033] in the case where there are no third and fourth to-be-distributed tasks, pack the first to-be-distributed task into a job package of the current batch, and distribute the job package to the labeler.

[0034] In combination with the first aspect, in a fifth implementation manner of the first aspect, before the step of determining the real-time team state of each team, the real-time personal state of each labeler, and the permission information, and packing the first, second, third and / or fourth to-be-distributed tasks into a job package of the current batch according to the real-time team state, the real-time personal state and the permission information, and distributing the job package to the labeler, the method further comprises:

[0035] obtain the first to fourth to-be-distributed tasks released by the labeler;

[0036] The labeler can only process one labeling task in a job package at the same time, and the labeling task is the first, second, third or fourth to-be-distributed task. When the operation time of the labeling task exceeds the preset maximum working time or the flow time of the labeling task exceeds the preset maximum natural cycle, the task is released passively;

[0037] The operation duration is obtained according to a current time and a first starting time of processing the labeling task, the flow duration is obtained according to the current time and a second starting time when the labeling task is distributed, the preset maximum working duration is obtained according to the first preset coefficient multiplied by an average historical processing duration of historical tasks, the average historical processing duration is an average value of historical processing durations of all historical tasks, the historical processing duration is obtained according to a third starting time when an operator processes a historical task and an ending time when the historical task is processed, the preset maximum natural cycle is obtained according to the second preset coefficient multiplied by an average historical flow duration of historical tasks, the average historical flow duration is an average value of historical flow durations of all historical tasks, and the historical flow duration is obtained according to a fourth starting time when an operator is distributed a historical task and an ending time.

[0038] According to a second aspect, the embodiments of the present application further provide a data labeling task distribution device, the device comprises:

[0039] A project splitting module is configured to obtain a labeling task data packet and project configuration information, split the labeling task data packet according to project distribution characteristics determined from the project configuration information, and obtain at least one batch of first to-be-distributed task sets; each batch of first to-be-distributed task sets comprises at least one first to-be-distributed task.

[0040] A first generation module is configured to obtain a labeling work order generated according to a first to-be-distributed task submitted by an operator, and generate a second to-be-distributed task according to the labeling work order.

[0041] A second generation module is configured to obtain a repair work order fed back after the second to-be-distributed task is audited, and generate a third to-be-distributed task or a fourth to-be-distributed task according to the repair work order; the third to-be-distributed task corresponds to an audited rejected repair work order, and the fourth to-be-distributed task corresponds to an audited passed repair work order that needs to be modified.

[0042] A task distribution module is configured to determine real-time team states of each team, real-time individual states of each operator, and permission information, pack the first, second, third, and / or fourth to-be-distributed tasks into a current batch of work packages according to the real-time team states, real-time individual states, and permission information, and distribute the work packages to the operators; the real-time team state is determined by a team task upper limit and real-time individual states of all operators in the team, the real-time individual state is determined by a number of to-be-processed labeling tasks in the work package and whether the work package is distributed, each operator can only process one work package at the same time, and the number of to-be-processed labeling tasks in the work package does not exceed an individual task upper limit.

[0043] According to a third aspect, the embodiments of the present application further provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the data labeling task distribution method according to any one of the above embodiments when executing the program.

[0044] According to a fourth aspect, the embodiments of the present application further provide a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the steps of the data labeling task distribution method according to any one of the above embodiments.

[0045] The data labeling task distribution method, device, equipment and medium of the present application split the labeling task data packet by the project distribution characteristics determined from the project configuration information, and then split the labeling task data packet into independent job units, which are convenient for distribution to each job worker for parallel processing, thereby improving the project processing efficiency. The second to-be-distributed task is generated according to the labeling work order, and the third and fourth to-be-distributed tasks are generated according to the repair work order. By associating each labeling task, an integrated processing mechanism of labeling, auditing and repair can be established, and the labeling task quality responsibility can be traced back to the specific labeler or auditor, thereby better controlling the project quality. According to the real-time team state, real-time individual state and permission information, the first, second, third and / or fourth to-be-distributed tasks are packaged into the current batch of job packets, and the job packets are distributed to each job worker. The current batch can be distributed to multiple job workers at the same time for batch processing. Each job worker can only process one job packet at the same time, and the number of to-be-processed labeling tasks in the job packet does not exceed the individual task upper limit. In this way, the job worker can more efficiently process the labeling task and adjust the resource allocation in real time according to the project progress and the state of the job worker. By setting a team task upper limit for each team, it is ensured that each team, especially a new team, can be distributed a certain number of labeling tasks and the tasks can be smoothly transferred, avoiding resource monopoly. At the same time, the real-time team state is determined by the team task upper limit and the real-time individual state of all job workers in the team, ensuring that the tasks are more reasonably distributed to the job workers of the corresponding team. The individual task upper limit can be determined according to the predicted productivity and response time of the job worker. By setting an individual task upper limit for each job worker, not only can the labeling task be prevented from being maliciously occupied, but also the impact of individual transfer lag on the entire project progress can be avoided. At the same time, the real-time individual state is determined by the number of to-be-processed labeling tasks in the job packet and whether the job packet is distributed, ensuring that the labeling task is more reasonably distributed to the corresponding job worker, and achieving efficient matching of labeling tasks and job workers. Considering multiple dimensions when distributing tasks, efficient matching of tasks and job workers is achieved, multi-dimensional matching and dynamic resource distribution are achieved, the average project cycle is shortened, the delivery efficiency is significantly improved, and automatic task distribution can reduce manual operation time, adapt to different labeling scenarios, and be horizontally expanded to millions of tasks. Attached Figure Description

[0046] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the invention in any way. In the drawings:

[0047] Figure 1 One of the flowcharts of the data annotation task distribution method provided by the present invention is shown;

[0048] Figure 2 This diagram illustrates the process of distributing annotation tasks to annotators in the data annotation task distribution method provided by the present invention.

[0049] Figure 3 The second flowchart of the data annotation task distribution method provided by the present invention is shown;

[0050] Figure 4 A schematic diagram of the data annotation task distribution device provided by the present invention is shown;

[0051] Figure 5 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] With the rapid development of deep learning, multimodal technologies, and big data processing technologies, data annotation has become one of the core technologies in the field of computer vision. Data annotation refers to the labeling, classification, and annotation of large amounts of initial data so that machine learning models and subsequent operators can better understand and process this data.

[0054] Currently, data annotation requires professional annotators in the field. In particular, the field of autonomous driving has high standards for data annotation. Therefore, the annotated data also needs to undergo quality inspection at various checkpoints such as review, quality control, and acceptance. If the review and acceptance results are not up to standard, the annotators need to rework the data to correct the errors.

[0055] The distribution of data labeling work is usually based on simple distribution rules or manually assigned by project managers according to experience, which is limited in adapting to scenarios and difficult to control project progress, which undoubtedly brings great challenges to the overall progress of the data labeling project. In order to better control the project progress, some distribution platforms also consider the business capabilities of workers and the delivery difficulty of projects when distributing tasks, but when faced with large-scale labeling tasks, the business capabilities of workers cannot reflect their real-time work status, resulting in poor dynamic adaptability of actual task distribution, which can easily cause task flow to be stuck, task processing to be overdue, and so on. Moreover, relying on the business capabilities of workers to distribute tasks can make it difficult for new teams to get opportunities and the allocation of resources is not balanced enough.

[0056] In summary, how to perform more efficient and reasonable labeling task distribution is an important issue that needs to be solved in the industry at present.

[0057] Due to the above technical problems, a data labeling task distribution method is provided in the embodiments of the present application, which aims to realize efficient matching of labeling tasks and workers through dynamic resource distribution, multi-dimensional matching rules and real-time task monitoring, and provide a more efficient flow mode for labeling tasks. The data labeling task distribution method of the embodiments of the present application can be used in electronic devices, including but not limited to computers, mobile terminals, etc. Figure 1 is a flowchart of the data labeling task distribution method according to the embodiments of the present application, as Figure 1 shown, the method can include the following steps:

[0058] S101, obtain labeling task data packets and project configuration information, and split the labeling task data packets according to the project distribution characteristics determined from the project configuration information to obtain at least one batch of first to-be-distributed task sets, wherein each batch of first to-be-distributed task sets contains at least one first to-be-distributed task.

[0059] The labeling task data packets are split by the project distribution characteristics, so that the labeling task data packets are split into independent work units, and the first to-be-distributed tasks obtained are the work units described above, which facilitates distribution to each worker for parallel processing and improves project processing efficiency. The first to-be-distributed task set is a task set composed of at least one first to-be-distributed task. Considering that some labeling task data packets are relatively complex and difficult to process, the labeling task data packets are split into at least one batch of first to-be-distributed task sets according to the project distribution characteristics. It can be understood that the more complex and difficult the labeling task data packets are to process, the more batches they will be split into.

[0060] When a new data labeling project needs to be carried out, the labeling task data packet of the data labeling project and the project configuration information configured for the labeling task data packet are obtained accordingly. Then, feature extraction is performed on the project configuration information to obtain project distribution features, and the corresponding splitting of the labeling task data packet is performed according to the project distribution features to obtain a first batch of first to-be-distributed task sets.

[0061] In this embodiment, the first to-be-distributed task is a labeling task that needs to be labeled by an operator, and is also a new task to be distributed, which accounts for a large proportion of all to-be-distributed tasks.

[0062] S102, obtaining a labeling work order generated according to the first to-be-distributed task submitted by the operator, and generating a second to-be-distributed task according to the labeling work order. Each second to-be-distributed task corresponds to a labeling work order.

[0063] In this embodiment, each first to-be-distributed task will be distributed to an operator for labeling processing. After the operator finishes processing the first to-be-distributed task, the operator will submit a labeling work order that has been labeled. The labeling work order is obtained after the first to-be-distributed task is labeled by the operator. In most cases, the labeling work order cannot be directly accepted but needs to go through the processing of various nodes such as review and acceptance and pass through them before being successfully accepted, so as to ensure the quality of the data delivered to the customer. In this embodiment, a second to-be-distributed task is generated according to the labeling work order. The second to-be-distributed task is a labeling task that needs to be audited by an operator.

[0064] The second to-be-distributed task also accounts for a large proportion of all to-be-distributed tasks. In a theoretical case, each first to-be-distributed task will generate a second to-be-distributed task. In an actual case, the number of second to-be-distributed tasks generated from the labeling work order of the first to-be-distributed task does not exceed the first to-be-distributed task.

[0065] Considering the development of intelligent labeling task auditing technology, an intelligent labeling task auditing system can be embedded in the distribution platform. Some or all of the auditing work of the labeling task can be directly intervened and automatically audited by the system. Therefore, the labeling work order can select some or all of them to generate corresponding second to-be-distributed tasks according to the actual situation. At the same time, a mapping relationship is established between the labeling work order corresponding to the first to-be-distributed task, the second to-be-distributed task corresponding to the labeling work order, and the three, so as to better trace the task flow and distribute the task.

[0066] It should be noted that even if the review work of all labeling tasks is directly intervened and automatically reviewed by the system, part of the labeling work orders submitted by all workers can still be selected for manual review by other workers, that is, in this case, the second distribution task can still be generated for part of the labeling work orders.

[0067] S103, obtaining the repair work order fed back after the second distribution task is reviewed, and generating a third distribution task or a fourth distribution task according to the repair work order.

[0068] In the embodiment, whether it is manual review or automatic review, the second distribution task can be divided into three cases after review: review is rejected, review is passed and no modification is needed, and review is passed and modification is needed. The first and third cases will generate repair work orders, and the third distribution task or the fourth distribution task will be generated according to the repair work order. The third distribution task and the fourth distribution task are both labeling tasks that need to be repaired, and the difference between them is that the degree of repair is different. One repair work order will only generate one labeling task. The third distribution task corresponds to the repair work order rejected by the review, and the fourth distribution task corresponds to the repair work order passed by the review and needing to be modified.

[0069] After the worker completes the third distribution task or the fourth distribution task, he will submit the repair work order after repair processing. In some cases, the repair work order after repair processing can also generate a second distribution task, that is, the repair work order after repair processing still needs to be reviewed again to ensure the overall quality of the project. Of course, the repair work order after repair processing can also be processed by other review platforms or other reviewers. In these cases, the second distribution task will not be generated.

[0070] The first to fourth distribution tasks constitute all the distribution tasks and participate in the subsequent automatic distribution of tasks to each worker. Based on the first to fourth distribution tasks, a task pool can be constructed, and all undistributed distribution tasks are stored in the task pool.

[0071] S104, determining the real-time team state of each team, the real-time personal state of each worker, and the permission information of each worker, packaging the first distribution task, the second distribution task, the third distribution task and / or the fourth distribution task into a current batch of work package according to the real-time team state, the real-time personal state and the permission information, and distributing the work package to the worker. Each worker's work package has at least one labeling task, and the labeling task is composed of the first distribution task, the second distribution task, the third distribution task and / or the fourth distribution task.

[0072] It should be noted that the current batch, for example, batch 1, can be distributed to multiple operators for batch, that is, one batch can generate multiple job packages and be handed over to the corresponding number of operators for processing. After the operator submits the job package, if there is no new labeling task in the batch, the labeling task in the batch with the nearest target delivery time will be automatically assigned. For batches that have submitted job packages, operators can also switch batches independently, process a particular batch with high priority, and keep up with the delivery schedule of urgent data to avoid the backlog of urgent tasks.

[0073] In the present embodiment, the real-time team state is determined by the team task upper limit and the real-time personal state of all operators in the team, and the real-time personal state is determined by the number of pending labeling tasks in the job package and whether the job package is distributed. Each operator can only handle one job package at the same time, and the number of pending labeling tasks in the job package does not exceed the personal task upper limit. Among them, the real-time team state includes team distributable state and team non-distributable state, and the real-time personal state includes personal distributable state and personal non-distributable state.

[0074] The team task upper limit can be set by the user, for example, the project general manager. By setting the team task upper limit for each team, it is ensured that each team, especially new teams, can be assigned a certain number of labeling tasks and the tasks can be smoothly transferred, avoiding resource monopoly. At the same time, the real-time team state is determined by the team task upper limit and the real-time personal state of all operators in the team, ensuring that the tasks are more reasonably distributed to the operators of the corresponding team.

[0075] The personal task upper limit can be set by the user, for example, the project general manager and the project manager of each team. For example, the user can determine the personal task upper limit according to the predicted productivity and response time of the operator. By setting the personal task upper limit for each operator, it not only avoids malicious occupation of labeling tasks, but also avoids affecting the overall project progress due to personal transfer lag. At the same time, the real-time personal state is determined by the number of pending labeling tasks in the job package and whether the job package is distributed, ensuring that the labeling tasks are more reasonably distributed to the corresponding operators, and achieving efficient matching of labeling tasks and operators. Each operator can only handle one job package at the same time, and the number of pending labeling tasks in the job package does not exceed the personal task upper limit, which can ensure that the operator processes the labeling task more efficiently and adjusts the resource allocation in real time according to the project progress and the state of the operator.

[0076] The data labeling task distribution method of the present application splits the labeling task data package by the project distribution characteristics determined from the project configuration information, and then splits the labeling task data package into individual independent job units, which are convenient for distribution to each job worker for parallel processing, improving project processing efficiency; according to the labeling work order, a second to-be-distributed task is generated, and according to the repair work order, a third and fourth to-be-distributed task is generated, and by associating each labeling task, an integrated labeling, auditing and repair processing mechanism can be established, and the labeling task quality responsibility can be traced back to the specific labeler or auditor, and then the project quality can be better controlled; according to the real-time team state, real-time individual state and permission information, the first, second, third and / or fourth to-be-distributed tasks are packaged into the current batch of job packages, and the job packages are distributed to each job worker, and the current batch can be distributed to multiple job workers for batch processing, each job worker can only process one job package at the same time, and the number of to-be-processed labeling tasks in the job package does not exceed the individual task upper limit, which can ensure that the job worker can more efficiently process the labeling task and adjust the resource allocation in real time according to the project progress and the state of the job worker; by setting a team task upper limit for each team, it is ensured that each team, especially a new team, can be distributed a certain number of labeling tasks and the tasks can be smoothly transferred, avoiding resource monopoly, while the real-time team state is determined by the team task upper limit and the real-time individual state of all job workers in the team, ensuring that the tasks are more reasonably distributed to the job workers of the corresponding team, the individual task upper limit can be determined according to the predicted productivity and response time of the job worker, by setting an individual task upper limit for each job worker, not only can the labeling task be prevented from being maliciously occupied, but also the impact of individual transfer lag on the entire project progress can be avoided, while the real-time individual state is determined by the number of to-be-processed labeling tasks in the job package and whether the job package has been distributed, ensuring that the labeling task is more reasonably distributed to the corresponding job worker, and efficient matching of labeling tasks and job workers is achieved; considering multiple dimensions when distributing tasks, efficient matching of tasks and job workers is achieved, multi-dimensional matching and dynamic resource distribution are achieved, which shortens the average project cycle, significantly improves delivery efficiency, and automatic task distribution can reduce manual operation time, adapt to different labeling scenarios, and can be horizontally expanded to millions of tasks.

[0077] In the present embodiment, step S101 specifically comprises:

[0078] S1011, acquire labeling task data package and project configuration information. Each new data labeling project will generate a corresponding labeling task data package, and each labeling task data package will also have corresponding project configuration information.

[0079] S1012, according to the preset feature extraction rule, the project configuration information is extracted, the project configuration feature is obtained, and the project distribution feature is obtained by assembling the project configuration feature. Wherein, the preset feature extraction rule determines which feature is extracted and the number of extracted features, such as the preset feature extraction rule stipulates that the business scene, the continuity annotation requirement of scene, the number of required annotation objects, the project delivery period, the project importance and other project configuration features are extracted from the project configuration information. The specific process of feature assembly is:

[0080]

[0081] Among them, indicates the project distribution feature; indicates the business scene feature in the project configuration feature; indicates the continuity annotation requirement feature of scene in the project configuration feature; indicates the number (how many) of required annotation objects feature in the project configuration feature; indicates the project delivery period feature in the project configuration feature; indicates the project importance feature in the project configuration feature.

[0082] Among them, the delivery period feature determines that the project can be split into several batches for processing, and the importance feature determines whether there is a batch that can be inserted to be processed in advance in the batch that the project is split into.

[0083] Suppose the current annotation task data packet is divided into K batches, so that Among them, indicates the first batch of first distribution task set, contains at least one first distribution task, the priority of is higher than that of , that is, The priority of each first distribution task in is higher than that of each first distribution task in , and there is also a priority between each first distribution task in . When the new data annotation project access system is started, according to the project distribution feature, the annotation task data packet corresponding to the new data annotation project is split, and at least one batch of first distribution task set is also obtained. The importance feature extracted from the project configuration information of the new data annotation project determines whether there is a queue in the batch that the project is split into, for example, the first first distribution task split out in the annotation task data packet corresponding to the new data annotation project becomes the new , will replace the original , the original is updated to , accordingly, each batch after the original is also postponed by one batch, and the total number of batches becomes K+1. The first to-be-distributed task determines the specific priority according to the batch order and the continuity requirement of the scene, and determines the first to-be-distributed task packaged in the current batch according to the priority.

[0084] It should be noted that the preset feature extraction rule can be configured by the user.

[0085] S1013, according to the project distribution characteristics, the annotation task data packet is split to obtain at least one batch of first to-be-distributed task set.

[0086] In this embodiment, step S104 specifically includes:

[0087] S1041, determine the real-time team state of each team.

[0088] S1042, in the case where the real-time team state of the team is determined to be a team distributable state, determine the real-time personal state of each worker in the team. The team distributable state is specifically: at least one worker's real-time personal state is in a personal distributable state, and the total number of the number of to-be-processed annotation tasks in the work package of all workers in the team is less than the team task upper limit. Other cases are team non-distributable state.

[0089] S1043, in the case where at least one worker's real-time personal state is determined to be in a personal distributable state, determine whether the worker whose real-time personal state is in a personal distributable state has distributed a work package. The personal distributable state specifically includes the following cases: no work package is distributed; the number of to-be-processed annotation tasks in the work package is less than the personal work upper limit, that is, there is a number of empty tasks in the work package, and accordingly, new to-be-distributed tasks can be added to the work package.

[0090] S1044, in the case where at least one worker is determined not to have distributed a work package, according to the permission information of the worker, the first, second, third and / or fourth to-be-distributed task is packaged into a work package of the current batch, and the work package is distributed to the worker who has not distributed a work package.

[0091] In this embodiment, the permission information contains three categories, one is an annotator, another is an auditor, and one is an administrator. The permission information can be determined by the identity information of the worker and bound to the work account of the worker.

[0092] S1045, in a case where it is determined that all the workers distribute the work packages and there is at least one worker whose number of to-be-processed labeling tasks in the work package is less than the personal work upper limit, according to the permission information of the worker, the first, second, third and / or fourth to-be-distributed tasks are packaged into a current batch of work package, and the current batch of task package is merged into the work package already owned by the worker. It can be understood that the total number of to-be-distributed tasks in the current batch of work package plus the total number of to-be-processed labeling tasks does not exceed the personal work upper limit, and each worker can still only process one work package.

[0093] According to the real-time team state of the team, the real-time personal state of the worker, and the permission information, the distribution of each type of task is performed, multiple dimensions are considered when distributing tasks, efficient matching of tasks and workers is performed, multi-dimensional matching and dynamic resource distribution are realized, so as to shorten the average project cycle, significantly improve the delivery efficiency, and automatic task distribution can reduce manual operation time, adapt to different labeling scenarios, and be horizontally expanded to millions of task quantities.

[0094] When a new worker accesses the system or all to-be-distributed tasks have been distributed, the worker can not distribute the work package, more specifically, the process of packaging the current batch of work package and distributing the work package to the worker who does not distribute the work package is as follows:

[0095] In a case where it is determined that the permission information is a labeler, the first, third and / or fourth to-be-distributed tasks are packaged into a current batch of work package, and the work package is distributed to the labeler.

[0096] In a case where it is determined that the permission information is an auditor, the second, third and / or fourth to-be-distributed tasks are packaged into a current batch of work package, and the work package is distributed to the auditor.

[0097] In a case where it is determined that the permission information is an administrator, the first, second, third and / or fourth to-be-distributed tasks are packaged into a current batch of work package, and the work package is distributed to the administrator.

[0098] It can be understood that the distribution of to-be-distributed tasks is performed only when there are to-be-distributed tasks.

[0099] The work package of the worker whose permission information is a labeler includes the following cases: composed of the first to-be-distributed task, i.e., all new labeling tasks; composed of the first and third to-be-distributed tasks; composed of the first, third and fourth to-be-distributed tasks; composed of the third or fourth to-be-distributed task.

[0100] The work package of the auditor includes the following cases: composed of the second to-be-distributed task; composed of the second and fourth to-be-distributed tasks; composed of the second and third to-be-distributed tasks; composed of the second, third and fourth to-be-distributed tasks; composed of the third and / or fourth to-be-distributed task.

[0101] The work package of the administrator includes the following cases: composed of the first to-be-distributed task; composed of the second to-be-distributed task; composed of the third and / or fourth to-be-distributed task; composed of the first, second and third to-be-distributed tasks; composed of the first, second and fourth to-be-distributed tasks; composed of the second and third to-be-distributed tasks; composed of the second and fourth to-be-distributed tasks; composed of the second, third and fourth to-be-distributed tasks; composed of the first, second, third and fourth to-be-distributed tasks.

[0102] In the case of normal task flow, the workman will continuously process the various types of labeling tasks in the distributed work package, and the workman can only perform one labeling task at the same time. After completing a labeling task, the next labeling task will be automatically distributed. More specifically, the process of packing the current batch of tasks into the workman's existing work package is as follows:

[0103] In the case of determining that the permission information is the labeling person, the first, third and / or fourth to-be-distributed tasks are packed into the work package of the current batch, and the work package of the current batch is merged into the work package of the labeling person;

[0104] In the case of determining that the permission information is the auditor, the second, third and / or fourth to-be-distributed tasks are packed into the work package of the current batch, and the work package of the current batch is merged into the work package of the auditor;

[0105] In the case of determining that the permission information is the administrator, the first, second, third and / or fourth to-be-distributed tasks are packed into the work package of the current batch, and the work package of the current batch is merged into the work package of the administrator.

[0106] The way of packing the work package of the current batch according to the permission information is as described in step S1044, which will not be repeated here. The work package of the current batch will be merged with the work package already distributed to the workman as an added package. It can be understood that the total number of to-be-distributed tasks in the work package of the current batch plus the total number of to-be-processed labeling tasks does not exceed the personal work upper limit, and each workman can still only process one work package at the same time.

[0107] The team task upper limit of the A team is set as m. First, it is determined whether the team distribution state of the A team is a team distributable state. In the case of determining the team distributable state, if the real-time personal state of all workers in the A team is a personal non-distributable state, no work package will be distributed to any worker in the A team in this case. If there is a worker in the A team whose real-time personal state is a personal distributable state, it is then determined whether all workers in the A team have distributed work packages at present. If there are workers who have not distributed work packages, new work packages can be distributed to these workers according to the permission information of the workers. If all workers have distributed work packages and there are workers whose number of to-be-processed labeling tasks in the work package is less than the personal work upper limit, these workers can be packaged into the current batch of work packages according to the permission information of the workers, and the work packages corresponding to the workers are merged.

[0108] In the embodiment, in order to avoid the accumulation of labeling tasks that need to be repaired, task scheduling is also performed when distributing labeling tasks to optimize dynamic resource distribution. Taking a labeler as an example, the process of distributing labeling tasks also includes:

[0109] The first to fourth to-be-distributed tasks are associated to form a trace chain between all labeling tasks.

[0110] For example, each first to-be-distributed task can be set with a unique number, and the number of the first to-be-distributed task can be used as a unique identifier of the first to-be-distributed task. Each labeling work order can also be set with a unique number, and the number of the labeling work order can be used as a unique identifier of the labeling work order. Each repair work order can also be set with a unique number, and the number of the repair work order can be used as a unique identifier of the labeling work order. By associating the number of the labeling work order with the number of the first to-be-distributed task that generates the labeling work order, a mapping relationship between the first and second to-be-distributed tasks can be established. Then, by associating the number of the repair work order with the number of the labeling work order that generates the repair work order, a mapping relationship between the second and third / fourth to-be-distributed tasks can be established. Thus, a complete mapping relationship between the first to fourth to-be-distributed tasks is established, and a trace chain between all labeling tasks is formed. The trace chain can be: the first to-be-distributed task with number A→the second to-be-distributed task with number B→the third to-be-distributed task with number C. In this way, an integrated processing mechanism of labeling, auditing and repair is established, and the labeling task quality responsibility can be traced back to specific labelers or auditors, thereby better controlling the project quality.

[0111] obtaining a third and a fourth to-be-distributed task, determining a first to-be-distributed task associated with the third and the fourth to-be-distributed task according to the traceability chain, and processing the first to-be-distributed task corresponding to the labeling personnel. Each worker can associate his identity information with the corresponding labeling personnel when processing the labeling task, and then the corresponding worker can be traced, including the labeling personnel, the auditing personnel, and the administrator.

[0112] determining a batch to which each of the third and the fourth to-be-distributed task belongs, and determining a generation time point of generating each of the third and the fourth to-be-distributed task.

[0113] determining whether the batch belongs to a current batch being processed by the labeling personnel, and in the case of determining that the batch belongs to the current batch, packing the third and / or the fourth to-be-distributed task into a work package of the current batch, and directly distributing the work package of the current batch to the worker or merging the work package of the current batch into a work package of the labeling personnel according to a specific condition of the personal distributable state of the labeling personnel. The third or the fourth to-be-distributed task generated by the current batch has the highest priority and is preferentially distributed. The specific number of the third and the fourth to-be-distributed task generated by the repair order that can be packed is determined by the upper limit of the personal task and the number of to-be-processed labeling tasks. Assuming that the number of the third and the fourth to-be-distributed task of the current batch in a certain time period exceeds the difference between the upper limit of the personal task and the number of to-be-processed labeling tasks, the labeling personnel continues to process the labeling task in hand, so that the number of to-be-processed labeling tasks in the work package of the labeling personnel will continue to decrease until the difference can pack all the third and the fourth to-be-distributed task of the current batch.

[0114] determining that the batch is located before the current batch, determining a packing order between the third and the fourth task according to the generation time point, packing the first, the third and / or the fourth to-be-distributed task into a work package of the current batch according to the packing order, and packing the third and the fourth task of the same batch into the same work package, and directly distributing the work package of the current batch to the worker or merging the work package of the current batch into a work package of the labeling personnel according to a specific condition of the personal distributable state of the labeling personnel. The third or the fourth to-be-distributed task generated by the non-current batch has the second highest priority, and the first to-be-distributed task has the third priority. The third or the fourth to-be-distributed task of the second highest priority is preferentially packed when packing, and the first to-be-distributed task is added for packing if the labeling task can still be packed. The specific number of the labeling task that can be packed is determined by the upper limit of the personal task and the number of to-be-processed labeling tasks.

[0115] In the case where no third and fourth to-be-distributed tasks are determined, the first to-be-distributed task is packaged into a current batch job package, the first to-be-distributed task determines a specific priority according to a batch order and a continuity marking requirement of a scene, determines the first to-be-distributed task packaged in the current batch according to the priority, and according to the specific condition of the personal distributable state of the marker, directly distributes the current batch job package to the marker or merges it into the job package of the marker.

[0116] Take a marker as an example for description, please refer to Figure 2 , a marker participates in the marking task processing of batches 1, 2, 3 and 4 in chronological order, and continuously processes the job package of batch 4. In the processing process, a third to-be-distributed task of batch 2 is generated at T1, a third to-be-distributed task of batch 1 is generated at T2, another third to-be-distributed task of batch 2 is generated at T3, and a third to-be-distributed task of batch 4 is generated at T3, in chronological order. The above T1 to T3 are the generation time of each marking task. Since the number of to-be-processed marking tasks in the job package of batch 4 has not reached the personal task upper limit of the marker, the third to-be-distributed task of batch 4 generated at T3 can be directly packaged into a job package and merged into the job package being processed. The marker submits all the marking jobs in the batch 4 job package at T4. Since T1 is the earliest one among all time points, the third to-be-distributed task of batch 2 at T1 and the other third to-be-distributed task of batch 2 at T3 are packaged into a batch 5 job package together and distributed to the marker for processing. During the processing of the batch 5 job package, a third to-be-distributed task of batch 3 is generated at T5. Since T2 is before T5, the third to-be-distributed task of batch 1 at T2 is packaged first, and then the original batch 5 job package is merged. The third to-be-distributed task of batch 3 at T6 is packaged, and the original batch 5 job package is merged. During the processing of the batch 5 job package, another third to-be-distributed task of batch 4 is generated at T6. The marker submits all the marking jobs in the batch 5 job package at T7. Thus, the third to-be-distributed task of batch 4 at T6 can be packaged together with the first to-be-distributed task obtained by splitting, and distributed to the marker as a batch 6 job package.

[0117] Different from the labeler, the highest and the second highest priority of the reviewer are unchanged, and the third priority is the second to-be-distributed task. Generally, the administrator does not directly participate in the processing of the labeling task. In some special cases such as the labeler, the reviewer leaves the team or logs out of the account, the task flow may be stuck, and the administrator will process any type of labeling personnel. For the administrator, the priority can not be set, and the tasks are processed in the order of generation.

[0118] In this embodiment, the task progress is also monitored in real time, and the overdue task is automatically released. For details, please refer to Figure 3 The method can further include the following steps:

[0119] S201, obtain the labeling task data packet and the project configuration information, split the labeling task data packet according to the project distribution characteristics determined from the project configuration information, and obtain at least one batch of first to-be-distributed task set. Specifically as Figure 1 The step S101 is described above, and will not be repeated here.

[0120] S202, obtain the labeling work order generated according to the first to-be-distributed task submitted by the operator, and generate the second to-be-distributed task according to the labeling work order. Specifically as Figure 1 The step S102 is described above, and will not be repeated here.

[0121] S203, obtain the repair work order fed back after the second to-be-distributed task is processed by the reviewer, and generate the third to-be-distributed task or the fourth to-be-distributed task according to the repair work order. Specifically as Figure 1 The step S103 is described above, and will not be repeated here.

[0122] S204, obtain the first to fourth to-be-distributed tasks released by the operator. Wherein, the release is divided into two cases of active release and passive release. When the operator considers that he cannot complete a labeling task in the work package, he can actively release the labeling task; within the same time, the operator can only process one labeling task in the work package, and when the operation time of the labeling task exceeds the preset maximum working time or the flow time of the task exceeds the preset maximum natural cycle, the labeling task is passively released.

[0123] In the embodiment, the operation duration is obtained according to a current time and a first starting time of processing the labeling task, the time node at which the operator just starts to process a certain labeling task in the work package is the first starting time, and the difference between the first starting time and the current time can obtain the operation duration of the labeling task; the circulation duration is obtained according to the current time and a second starting time at which the labeling task is distributed, when the operator takes the work package, the second starting time at which all the labeling tasks in the work package are distributed, and the difference between the second starting time and the current time can obtain the circulation duration of the labeling task.

[0124] By setting the two time thresholds of the preset maximum working duration and the preset maximum natural period, the labeling task of the operator can be actively released by using the double time thresholds, the released labeling task is re-joined in the task pool and is re-distributed to other operators in the future, so that the influence of the individual circulation lag on the whole project progress is avoided.

[0125] The preset maximum working duration is obtained by multiplying a first preset coefficient (for example, 3) by an average historical processing duration of historical tasks, the average historical processing duration is an average value of the historical processing durations of all the historical tasks, the historical processing duration of a certain historical task is obtained according to a third starting time at which the operator just starts to process the historical task and an ending time at which the historical task is processed, the time node at which a certain historical task is just started to be processed by the operator is the third starting time, and the time node at which the historical task is processed (submitted by the operator) is the ending time, the difference between the first ending time and the third starting time can obtain the historical processing duration.

[0126] The preset maximum natural period is obtained by multiplying a second preset coefficient (for example, 2) by an average historical circulation duration of historical tasks, the average historical circulation duration is an average value of the historical circulation durations of all the historical tasks, the historical circulation duration of a certain historical task is obtained according to a fourth starting time at which the operator is distributed the historical task and the ending time at which the historical task is processed, the time node at which a certain historical task is distributed to the operator is the fourth starting time, and the difference between the ending time and the fourth starting time can obtain the historical circulation duration.

[0127] The historical tasks include the first to fourth to-be-distributed tasks that have been processed in the task pool, the four types of tasks can be used as historical tasks, and the preset maximum working duration and the preset natural period are determined.

[0128] It should be noted that the third and fourth to-be-distributed tasks are distributed to the operators, and the corresponding flow time is counted. The third and fourth to-be-distributed tasks are distributed to the operators and started by the operators, and the corresponding processing time is counted. That is, the third and fourth to-be-distributed tasks generated according to the repair order are isolated from the original task and are independently timed.

[0129] In order to avoid the error of abnormal data from causing the task to be misreleased, in the case that the historical processing time or the historical flow time of the historical task in the task pool that has been processed is not more than a preset time length (for example, 1 minute), the corresponding historical processing time or historical flow time is not counted in the calculation of the subsequent preset maximum working time or preset maximum natural cycle.

[0130] S205, determine the real-time team state of each team, the real-time personal state of each operator, and the permission information of each operator, package the first to-be-distributed task, the second to-be-distributed task, the third to-be-distributed task and / or the fourth to-be-distributed task into a current batch of work package according to the real-time team state, the real-time personal state and the permission information, and distribute the work package to the operators. The distribution method of the non-released labeling task is specifically as follows Figure 1 The step S104 is described above and will not be repeated here.

[0131] The released annotation task is distributed to the workers in the same team for processing, that is, the released annotation task is transferred to the workers in the same team, the released annotation task is used to accumulate project processing experience for the workers in the same team, and the accuracy of subsequent annotation task distribution is improved. Specifically, first, the release source of the released annotation task and the permission information of the release source are determined, that is, the annotation task is released at which worker and the permission information of the corresponding worker is confirmed; next, it is determined that the worker who has the same team and the same permission information as the release source is in a personal distributable state, if there is, it is further determined whether the worker or the workers have processed the same batch of annotation tasks as the release source, if there is, the released annotation task or the released annotation task and other annotation tasks are packaged into the current batch of work package, and then distributed to the corresponding worker or merged into the work package of the worker. It should be noted that if there are multiple workers who have processed the same batch of annotation tasks as the release source, the released annotation task can be randomly distributed to one of the workers; if there is no worker who has processed the same batch of annotation tasks as the release source, it is further determined whether the worker or the workers have processed the same project annotation task as the release source, if there is, the released annotation task or the released annotation task and other annotation tasks are packaged into the current batch of work package, and then distributed to the corresponding worker or merged into the work package of the worker. It should be noted that if there are multiple workers who have processed the same project annotation task as the release source, the released annotation task can be randomly distributed to one of the workers; if there is no worker who has processed the same batch of annotation tasks or the same project annotation task, the released annotation task or the released annotation task and other annotation tasks are packaged into the current batch of work package, and then distributed to the pipeline worker in the team for processing. That is, the released annotation task is preferentially transferred to the worker who has the same team, the same permission information and has processed the same batch of annotation tasks for processing, and then preferentially transferred to the worker who has the same team, the same permission information and has processed the same project annotation task for processing, and if none of them, the released annotation task is transferred to the administrator for processing.

[0132] The data annotation task distribution device provided in the embodiment of the present application is described below. The data annotation task distribution device described below can be referred to in correspondence with the data annotation task distribution method described above.

[0133] Due to the above technical problems, in the embodiment of the present application, a data annotation task distribution device is also provided, which aims to realize efficient matching of annotation tasks and workers through dynamic resource distribution, multi-dimensional matching rules and real-time task monitoring, and provide a more efficient transfer mode of annotation tasks. Figure 4 is a structural schematic diagram of the data annotation task distribution method according to the embodiment of the present application, like Figure 4As shown, the apparatus can include:

[0134] A project splitting module 10 is configured to acquire the labeling task data package and the project configuration information, split the labeling task data package according to the project distribution features determined from the project configuration information, and obtain at least one batch of first to-be-distributed task sets, wherein each batch of first to-be-distributed task sets contains at least one first to-be-distributed task.

[0135] A first generating module 20 is configured to acquire labeling work orders generated according to the first to-be-distributed tasks submitted by the workers, and generate second to-be-distributed tasks according to the labeling work orders. Each second to-be-distributed task corresponds to one labeling work order.

[0136] A second generating module 30 is configured to acquire repair work orders fed back after the second to-be-distributed tasks are audited, and generate third to-be-distributed tasks or fourth to-be-distributed tasks according to the repair work orders.

[0137] A task distribution module 40 is configured to determine real-time team states of each team, real-time individual states of each worker, and permission information of each worker, pack the first to-be-distributed tasks, the second to-be-distributed tasks, the third to-be-distributed tasks, and / or the fourth to-be-distributed tasks into a current batch of work packages according to the real-time team states, the real-time individual states, and the permission information, and distribute the work packages to the workers. Each worker's work package contains at least one labeling task, and the labeling task is composed of the first to-be-distributed tasks, the second to-be-distributed tasks, the third to-be-distributed tasks, and / or the fourth to-be-distributed tasks.

[0138] The data labeling task distribution device of the application splits the labeling task data packet according to the project distribution characteristics determined from the project configuration information, and then splits the labeling task data packet into independent job units, which are convenient for distribution to each job worker for parallel processing, thereby improving the project processing efficiency; the second to-be-distributed task is generated according to the labeling work order, and the third and fourth to-be-distributed tasks are generated according to the repair work order, and an integrated labeling, auditing and repair processing mechanism can be established by associating each labeling task, the labeling task quality responsibility can be traced back to the specific labeler or auditor, and the project quality can be better controlled; according to the real-time team state, real-time individual state and permission information, the first, second, third and / or fourth to-be-distributed tasks are packaged into the current batch of job packets, and the job packets are distributed to each job worker, and the current batch can be distributed to multiple job workers at the same time for batch processing, each job worker can only process one job packet at the same time, and the number of to-be-processed labeling tasks in the job packet does not exceed the individual task upper limit, which can ensure that the job worker processes the labeling task more efficiently and adjusts the resource allocation in real time according to the project progress and the state of the job worker; by setting a team task upper limit for each team, it is ensured that each team, especially a new team, can be distributed a certain number of labeling tasks and the tasks can be smoothly transferred, avoiding resource monopoly, and the real-time team state is determined by the team task upper limit and the real-time individual state of all job workers in the team, ensuring that the tasks are more reasonably distributed to the job workers of the corresponding team, the individual task upper limit can be determined according to the predicted productivity and response time of the job worker, by setting an individual task upper limit for each job worker, not only can the labeling task be prevented from being maliciously occupied, but also the impact of individual transfer lag on the entire project progress can be avoided, and the real-time individual state is determined by the number of to-be-processed labeling tasks in the job packet and whether the job packet is distributed, ensuring that the labeling task is more reasonably distributed to the corresponding job worker, and efficient matching of labeling tasks and job workers is achieved; considering multiple dimensions when distributing tasks, efficient matching of tasks and job workers is achieved, multi-dimensional matching and dynamic resource distribution are achieved, thereby shortening the average project cycle, significantly improving delivery efficiency, and automatic task distribution can reduce manual operation time, adapt to different labeling scenarios, and be horizontally expanded to millions of tasks.

[0139] Figure 5 An example of an entity structure diagram of an electronic device is shown in Figure 5 The electronic device can include a processor 510, a communication interface 520, a memory 330, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can invoke the logic command in the memory 530 to execute the data labeling task distribution method, which comprises:

[0140] Obtain the labeling task data packet and the project configuration information, split the labeling task data packet according to the project distribution characteristics determined from the project configuration information, and obtain at least one batch of first to-be-distributed task sets; each batch of first to-be-distributed task sets contains at least one first to-be-distributed task;

[0141] Obtain the labeling work order generated according to the first to-be-distributed task submitted by the operator, and generate a second to-be-distributed task according to the labeling work order;

[0142] Obtain the repair work order fed back after the second to-be-distributed task is audited, and generate a third to-be-distributed task or a fourth to-be-distributed task according to the repair work order; the third to-be-distributed task corresponds to the repair work order that is audited and rejected, and the fourth to-be-distributed task corresponds to the repair work order that is audited and passed and needs to be modified;

[0143] Determine the real-time team state of each team, the real-time personal state of each operator, and the permission information, pack the first, second, third and / or fourth to-be-distributed tasks into a current batch of work packages according to the real-time team state, the real-time personal state and the permission information, and distribute the work packages to the operators; the real-time team state is determined by the team task upper limit and the real-time personal state of all operators in the team, the real-time personal state is determined by the number of to-be-processed labeling tasks in the work package and whether the work package is distributed, each operator can only process one work package at the same time, and the number of to-be-processed labeling tasks in the work package does not exceed the personal task upper limit.

[0144] In addition, the logical instructions in the memory 530 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0145] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the above-mentioned provided data labeling task distribution method for executing the data labeling task, and the method comprises:

[0146] Obtain the labeling task data packet and the project configuration information, split the labeling task data packet according to the project distribution characteristics determined from the project configuration information, and obtain at least one batch of first to-be-distributed task sets; each batch of first to-be-distributed task sets contains at least one first to-be-distributed task;

[0147] Obtain the labeling work order generated according to the first to-be-distributed task submitted by the operator, and generate a second to-be-distributed task according to the labeling work order;

[0148] Obtain the repair work order fed back after the second to-be-distributed task is audited, and generate a third to-be-distributed task or a fourth to-be-distributed task according to the repair work order; the third to-be-distributed task corresponds to the repair work order that is rejected by the audit, and the fourth to-be-distributed task corresponds to the repair work order that is passed by the audit and needs to be modified;

[0149] Determine the real-time team state of each team, the real-time personal state of each operator, and the permission information, pack the first, second, third and / or fourth to-be-distributed tasks into a current batch of work packages according to the real-time team state, real-time personal state and permission information, and distribute the work packages to the operators; the real-time team state is determined by the team task upper limit and the real-time personal state of all operators in the team, the real-time personal state is determined by the number of to-be-processed labeling tasks in the work package and whether the work package is distributed, each operator can only process one work package at the same time, and the number of to-be-processed labeling tasks in the work package does not exceed the personal task upper limit.

[0150] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for distributing data annotation tasks, characterized in that, The method comprises: Obtaining an annotation task data packet and project configuration information, splitting the annotation task data packet according to project distribution characteristics determined from the project configuration information to obtain at least one batch of first to-be-distributed task sets; each batch of first to-be-distributed task sets comprises at least one first to-be-distributed task, and the project distribution characteristics comprise business scenario characteristics, scenario continuity annotation requirement characteristics, required annotation object quantity characteristics, project delivery period characteristics, and project importance degree characteristics; Obtaining an annotation work order generated according to a first to-be-distributed task submitted by a worker, and generating a second to-be-distributed task according to the annotation work order; Obtaining a repair work order fed back after the second to-be-distributed task is audited, and generating a third to-be-distributed task or a fourth to-be-distributed task according to the repair work order; the third to-be-distributed task corresponds to a repair work order that is audited as rejected, and the fourth to-be-distributed task corresponds to a repair work order that is audited as passed and needs to be modified; Determining real-time team states of each team, real-time personal states of each worker, and permission information, packaging the first, second, third, and / or fourth to-be-distributed tasks into a current batch of work packages according to the real-time team states, real-time personal states, and permission information, and distributing the work packages to the workers; the real-time team state is determined by a team task upper limit and real-time personal states of all workers in the team, the real-time personal state is determined by a number of to-be-processed annotation tasks in a work package and whether the work package is distributed, each worker can only process one work package at the same time, and the number of to-be-processed annotation tasks in the work package does not exceed a personal task upper limit; The determination of the real-time team states of each team, the real-time personal states of each worker, and the permission information, the packaging of the first, second, third, and / or fourth to-be-distributed tasks into a current batch of work packages according to the real-time team states, real-time personal states, and permission information, and the distribution of the work packages to the workers specifically comprise: Determining real-time team states of each team; In a case where the real-time team state of a team is determined to be a team distributable state, determining real-time personal states of each worker in the team; the team distributable state is that the real-time personal state of at least one worker is a personal distributable state, and the total number of to-be-processed annotation task numbers in the work packages of all workers in the team is lower than a team task upper limit; In a case where it is determined that the real-time personal state of at least one worker is a personal distributable state, determining whether the worker whose real-time personal state is the personal distributable state has distributed a work package; the personal distributable state comprises not distributing a work package and the number of to-be-processed annotation tasks in the work package being less than a personal work upper limit; In a case where it is determined that at least one worker has not distributed a work package, packaging the first, second, third, and / or fourth to-be-distributed tasks into a current batch of work packages according to the permission information of the worker, and distributing the work packages to the worker who has not distributed a work package. In a case where it is determined that all workers distribute the work packages and there is at least one worker whose work package has a number of to-be-processed annotation tasks less than the individual work upper limit, according to the permission information of the worker, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the task package of the current batch is merged into the work package already possessed by the worker. The packaging of the first, second, third, and / or fourth to-be-distributed tasks into a work package of the current batch according to the permission information of the worker, and the distribution of the work package to the worker who has not distributed the work package, specifically includes: In a case where it is determined that the permission information is that of an annotator, the first, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the annotator. In a case where it is determined that the permission information is that of an auditor, the second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the auditor. In a case where it is determined that the permission information is that of an administrator, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the administrator.

2. The data labeling task distribution method of claim 1, wherein, The packaging of the first, second, third, and / or fourth to-be-distributed tasks into a work package of the current batch according to the permission information of the worker, and the merging of the task package of the current batch into the work package already possessed by the worker, specifically includes: In a case where it is determined that the permission information is that of an annotator, the first, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package of the current batch is merged into the work package of the annotator. In a case where it is determined that the permission information is that of an auditor, the second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package of the current batch is merged into the work package of the auditor. In a case where it is determined that the permission information is that of an administrator, the first, second, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package of the current batch is merged into the work package of the administrator.

3. The method of claim 1, wherein, The packaging of the first, third, and / or fourth to-be-distributed tasks into a work package of the current batch, and the distribution of the work package to the annotator, specifically includes: The first to fourth to-be-distributed tasks are associated to form a traceable chain among all annotation tasks. The third and fourth to-be-distributed tasks are obtained, and the first to-be-distributed task associated with the third and fourth to-be-distributed tasks and the annotator processing the first to-be-distributed task are determined according to the traceable chain. The batch to which each of the third and fourth to-be-distributed tasks belongs is determined, and the generation time of each of the third and fourth to-be-distributed tasks is determined. It is determined whether the batch belongs to the current batch being processed by the annotator. In a case where it is determined that the batch belongs to the current batch, the third and / or fourth to-be-distributed tasks are packaged into a work package of the current batch, and the work package is distributed to the annotator. In a case where it is determined that the batch is located before the current batch, the packaging order between the third and fourth tasks is determined according to the order of the generation time, the first, third, and / or fourth to-be-distributed tasks are packaged into a work package of the current batch according to the packaging order, the third and fourth tasks in the same batch are packaged into the same work package, and the work package is distributed to the annotator. In a case where it is determined that there is no third and fourth to-be-distributed task, the first to-be-distributed task is packaged into a job package of a current batch, and the job package is distributed to the annotator.

4. The method of claim 1, wherein, Before the step of determining the real-time team state of each team, the real-time individual state of each job worker, and the permission information, and packaging the first, second, third, and / or fourth to-be-distributed task into a job package of a current batch according to the real-time team state, the real-time individual state, and the permission information, and distributing the job package to the job worker, the method further comprises: acquiring the first to fourth to-be-distributed tasks released by the job worker; In a case where the operation duration of the annotation task exceeds a preset maximum working duration or the flow duration of the annotation task exceeds a preset maximum natural cycle, the task is released passively; The operation duration is obtained according to a current time and a first start time of processing the annotation task, the flow duration is obtained according to a current time and a second start time when the annotation task is distributed, the preset maximum working duration is obtained according to a first preset coefficient multiplied by an average historical processing duration of historical tasks, the average historical processing duration is an average value of historical processing durations of all historical tasks, the historical processing duration is obtained according to a third start time of processing a historical task by the job worker and an end time when the historical task is processed, the preset maximum natural cycle is obtained according to a second preset coefficient multiplied by an average historical flow duration of historical tasks, the average historical flow duration is an average value of historical flow durations of all historical tasks, and the historical flow duration is obtained according to a fourth start time when a historical task is distributed to the job worker and an end time.

5. The method of claim 1, wherein, The method comprises: acquiring the annotation task data packet and the project configuration information; extracting features from the project configuration information according to a preset feature extraction rule to obtain project configuration features, assembling the project configuration features to obtain project distribution features, and splitting the annotation task data packet according to the project distribution features to obtain at least one batch of first to-be-distributed task sets. The method comprises:

6. A data labeling task distribution apparatus, characterized by, The device comprises: a project splitting module configured to acquire an annotation task data packet and project configuration information, split the annotation task data packet according to project distribution features determined from the project configuration information to obtain at least one batch of first to-be-distributed task sets, and each batch of first to-be-distributed task sets comprises at least one first to-be-distributed task, wherein the project distribution features comprise a business scenario feature, a continuity annotation requirement feature of a scenario, a quantity feature of required annotation objects, a project delivery cycle feature, and a project importance feature. The first generation module is configured to obtain a labeling work order generated according to a first to-be-distributed task submitted by an operator, and generate a second to-be-distributed task according to the labeling work order; The second generation module is configured to obtain a repair work order fed back after the second to-be-distributed task is audited, and generate a third to-be-distributed task or a fourth to-be-distributed task according to the repair work order; the third to-be-distributed task corresponds to a repair work order whose audit is rejected, and the fourth to-be-distributed task corresponds to a repair work order whose audit is passed and needs to be modified; The task distribution module is configured to determine real-time team states of each team, real-time individual states of each operator, and permission information, pack the first, second, third and / or fourth to-be-distributed tasks into a current batch of work packages according to the real-time team states, the real-time individual states and the permission information, and distribute the work packages to the operators; the real-time team state is determined by a team task upper limit and the real-time individual states of all operators in the team, and the real-time individual state is determined by the number of to-be-processed labeling tasks in the work package and whether the work package is distributed; each operator can only process one work package at the same time, and the number of to-be-processed labeling tasks in the work package does not exceed an individual task upper limit; The task distribution module specifically includes: determining real-time team states of each team; In a case where the real-time team state of the team is determined to be a team distributable state, determining real-time individual states of each operator in the team; the team distributable state is that at least one real-time individual state of the operator is in an individual distributable state, and the total number of to-be-processed labeling task numbers in the work packages of all operators in the team is lower than a team task upper limit; In a case where it is determined that at least one real-time individual state of the operator is in an individual distributable state, determining whether the operator whose real-time individual state is in the individual distributable state has distributed the work package; the individual distributable state includes: not distributing the work package; and the number of to-be-processed labeling tasks in the work package is less than an individual work upper limit; In a case where it is determined that at least one operator has not distributed the work package, packing the first, second, third and / or fourth to-be-distributed tasks into a current batch of work packages according to the permission information of the operator, and distributing the work packages to the operators who have not distributed the work packages; In a case where it is determined that all operators have distributed the work packages and that at least one operator has a number of to-be-processed labeling tasks in the work package that is less than an individual work upper limit, packing the first, second, third and / or fourth to-be-distributed tasks into a current batch of work packages according to the permission information of the operator, and merging the current batch of task packages into the work package already possessed by the operator; The method specifically includes: In a case where it is determined that the permission information is that of a labeling operator, packing the first, third and / or fourth to-be-distributed tasks into a current batch of work packages, and distributing the work packages to the labeling operator; In a case where it is determined that the permission information is that of an auditor, packing the second, third and / or fourth to-be-distributed tasks into a current batch of work packages, and distributing the work packages to the auditor; In a case where the authority information is determined to be an administrator, the first, second, third, and / or fourth to-be-distributed task is packaged into a job package of a current batch, and the job package is distributed to the administrator.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the data labeling task distribution method of any one of claims 1-5 when executing the program.

8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the steps of the data labeling task distribution method of any one of claims 1-5 when executed by the processor.

Citation Information

Patent Citations

  • Data collaborative labeling method and device

    CN114169428A

  • Data processing method and device, electronic equipment and storage medium

    CN115829305A