Corpus tagging task processing method and device, storage medium and electronic equipment
By introducing task control strategies and a large-scale task allocation model into the corpus annotation platform, dynamically calculating the number of tasks, and combining user capabilities and corpus feature profiles, the problems of unfair task assignment and low efficiency in existing technologies are solved, achieving efficient corpus annotation task allocation and exclusive distribution.
Patent Information
- Application Number
- CN202511621379.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-02-03
AI Technical Summary
Existing corpus annotation platforms suffer from unfair task assignment, duplicate annotation, and low annotation efficiency when handling large-scale annotation tasks. They also lack effective task slicing and scheduling capabilities, resulting in a poor user experience.
By determining the task request from the user, a task control strategy is adopted to schedule user tasks in the corpus task pool, thereby achieving automatic allocation and exclusivity of target corpus annotation tasks. The task allocation processing model dynamically calculates the number of tasks, and personalized task allocation is performed by combining user ability profiles and corpus feature profiles.
It has achieved automated and reasonable distribution of corpus classification tasks, ensuring the uniqueness and exclusivity of task allocation, significantly improving annotation efficiency, avoiding repetitive operations, and optimizing the overall annotation throughput and quality of the platform.
Smart Images

Figure CN121455641A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of computer technology, and particularly relates to a corpus annotation task processing method and device, a storage medium and an electronic device. BACKGROUND
[0002] In the natural language processing scene, high-quality task annotation corpus is the basis for training and optimizing algorithm models. At present, on the corpus annotation platform, artificial selection of tasks or linear queuing distribution mode is generally used to manage and distribute the corpus to be classified.
[0003] However, the existing technical solutions have deficiencies in processing large-scale annotation tasks. For example, using the mode of artificial selection or simple queue distribution, it is difficult to guarantee the fairness of task taking, which can easily lead to repeated annotation of the same corpus by different users, or cause the problem of missing annotation of part of the corpus. In addition, the existing distribution mechanism also lacks effective task slicing and scheduling capabilities, resulting in low overall annotation efficiency and poor user experience. SUMMARY
[0004] The embodiment of the present specification provides a corpus annotation task processing method, device, storage medium and electronic device, and the technical solution is as follows: In a first aspect, the embodiment of the present specification provides a corpus annotation task processing method, and the method comprises: determining a task taking request of a user end for a corpus annotation task; adopting a task control strategy to perform user task scheduling processing on a corpus task pool based on the task taking request to obtain a target corpus annotation task with a task distribution quantity; sending each target corpus annotation task to the client end, so that the client end performs corpus data classification and annotation processing based on the target corpus annotation task.
[0005] In a second aspect, the embodiment of the present specification provides a corpus annotation task processing device, and the device comprises: A request determination module is configured to determine a task taking request of a user end for a corpus annotation task. A task scheduling module is configured to adopt a task control strategy to perform user task scheduling processing on a corpus task pool based on the task taking request to obtain a target corpus annotation task with a task distribution quantity. A task sending module is configured to send each target corpus annotation task to the client end, so that the client end performs corpus data classification and annotation processing based on the target corpus annotation task.
[0006] In a third aspect, the embodiments of the present specification provide a computer storage medium, which stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and performing the method steps described above.
[0007] In a fourth aspect, the embodiments of the present specification provide an electronic device, which can include a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and performing the method steps described above.
[0008] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects: In one or more embodiments of the present specification, the service platform determines a task taking request of a user end for a corpus annotation task, adopts a task control strategy to perform user task scheduling processing on a corpus task pool based on the task taking request to obtain a target corpus annotation task of a task allocation quantity, sends each target corpus annotation task to the client end, and enables the client end to perform corpus data classification annotation processing based on the target corpus annotation task, thereby realizing the automation and rationality of corpus classification task distribution, reducing human intervention, guaranteeing the uniqueness and exclusivity of corpus task allocation, effectively avoiding the problem of repeated operation of the same corpus by multiple users, and significantly improving the annotation efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present specification, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0010] Figure 1 is a flowchart of a corpus annotation task processing method provided by an embodiment of the present specification; Figure 2 is a flowchart of user task scheduling processing provided by an embodiment of the present specification; Figure 3 is a flowchart of task allocation processing provided by an embodiment of the present specification; Figure 4 is a flowchart of annotation behavior monitoring provided by an embodiment of the present specification; Figure 5 is a structural diagram of a corpus annotation task processing device provided by an embodiment of the present specification; Figure 6 is a structural diagram of an electronic device. DETAILED DESCRIPTION
[0011] With reference to the drawings of the embodiments in the specification, the technical solutions in the embodiments of the specification will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the specification, rather than all the embodiments of the specification. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the specification.
[0012] In the description of the specification, it should be understood that the terms "first", "second" and the like are used only for descriptive purposes, and cannot be construed as indicating or implying relative importance. In the description of the specification, it should be noted that, unless otherwise explicitly specified and limited, "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device. The specific meaning of the above terms in the specification can be understood by the person of ordinary skill in the art. In addition, in the description of the specification, "a plurality of" means two or more, unless otherwise specified. "And / or", which describes the relationship between the associated objects, means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents a "or" relationship between the associated objects.
[0013] The specification will be described in detail below with reference to specific embodiments.
[0014] In one embodiment, as shown in Figure 1 A corpus annotation task processing method is specifically proposed, which can be implemented by relying on a computer program and can run on a corpus annotation task processing device based on the von Neumann system. The computer program can be integrated in an application or run as a stand-alone tool application. The corpus annotation task processing device can be a service platform, including but not limited to: a personal computer, a tablet computer, a handheld device, a vehicle-mounted device, a server device, a computing device, or other processing devices connected to a wireless modem, etc. In different networks, the terminal device can be called by different names, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, device in 5G network or future evolution network, etc.
[0015] Specifically, the corpus annotation task processing method comprises: S102: determining a task receiving request of a user terminal for a corpus annotation task; Corpus annotation task: can refer to the operation of classifying, labeling or marking one or more corpus data (such as a piece of text, a picture or a piece of audio) according to a preset label system (such as emotion classification, entity recognition, topic classification, etc.).
[0016] Specifically, the logged-in user (for example, an annotator) triggers a task request operation on the user interface operated by the user through an interactive software (for example, clicking the “take task” button or executing a specific shortcut key).
[0017] After the user terminal receives the operation instruction, it sends a task taking request to the service platform (that is, the execution subject of the method of the application). The task taking request, optionally, carries at least a user ID for uniquely identifying the user. Optionally, the request can also carry other auxiliary information, such as a request timestamp, a user current state, or a user desired task type, etc.
[0018] After receiving the request, the service platform parses it, determines the legitimacy of the user ID, and triggers the subsequent task scheduling step.
[0019] S104: Based on the task taking request, the task control strategy is used to perform user task scheduling processing on the corpus task pool to obtain a target corpus annotation task with a task allocation quantity; Corpus task pool: refers to a logical or physical storage unit for storing a large amount of corpus data to be annotated, such as a database table, a data warehouse or a distributed file system. Each piece of corpus in the corpus task pool can have a state identifier, such as ‘to be allocated’ (or ‘unprocessed’), ‘locked’ (or ‘processing’), ‘completed’, etc.
[0020] Task control strategy: refers to a set of pre-set rules, algorithms or models used for decision-making when performing task scheduling. The strategy is the core basis for realizing the “user task scheduling processing” in S104. In an embodiment, the strategy includes, for example, a concurrency control strategy (for preventing tasks from being taken repeatedly) and a quantity allocation strategy (for determining the quantity of tasks allocated this time).
[0021] Illustratively, the service platform can start to perform this step after determining that the request of S102 is valid, such as through a user task scheduling module (for example, a microservice or a processing process).
[0022] In an embodiment, the task control strategy is specifically executed as follows based on the task taking request: Performing the concurrency control and task screening step: the service platform (also known as the server side) first accesses the corpus task pool (for example, a query database). In order to prevent different users from taking the same task at the same time, the server side excludes the corpus in the state of 'locked' or 'completed' when retrieving. In other words, the server side only selects from the set of corpus in the state of 'to be distributed'.
[0023] Performing the step of determining the number of task assignments: the server side needs to determine the number of task assignments N assigned to the user side this time.
[0024] In an optional embodiment, the "number of task assignments" can be a fixed value pre-configured by the system administrator (for example, N = 10).
[0025] In another optional embodiment, the "number of task assignments" can be a dynamically calculated value. For example, the server side can query the historical labeling efficiency, accuracy and other information (i.e. user ability portrait) of the user according to the user ID carried in S102 request, and combine the estimated difficulty of the corpus to be distributed (i.e. task feature portrait), and use a large model to dynamically calculate a reasonable N value.
[0026] Performing the task locking and generating step: the server side selects N corpus from the set of corpus to be distributed according to the above strategy. After the selection is completed, the server side immediately (or in an atomic operation) updates the state of the N corpus in the corpus task pool to locked, and binds the locked state with the user ID requested in S102. This "locking and binding" operation is the key to concurrency control, ensuring that any other user (such as user B) cannot retrieve the N locked corpus when performing S104 before the current user A completes and submits the N tasks. Finally, the server side packages the N locked corpus (including its content, ID, etc.) to generate one or more target corpus labeling tasks.
[0027] S106: sending each of the target corpus labeling tasks to the client to enable the client to perform corpus data classification and labeling processing based on the target corpus labeling tasks.
[0028] Illustratively, after successfully generating target corpus labeling tasks in S104, the service platform responds to and sends the task data (for example, a JSON data package containing N corpus contents) to the user side that initiated the request through the network (such as wired or wireless network).
[0029] After the user side receives the data package, it parses it and presents the content (such as text) of the target corpus and the corresponding labeling tools (such as classification label options, labeling boxes, etc.) to the user on the user interface (UI).
[0030] The user (annotator) analyzes and judges each of the N pieces of corpus on the client side, and performs the corpus data classification and annotation processing (for example, selects one or more classification labels for each piece of corpus).
[0031] In a complete business closed loop, when the user completes all (or part) of the target corpus annotation tasks on the client side and clicks the "submit" button, the client sends the annotation result data (including corpus ID and annotation label) back to the service platform as the server side. After receiving the annotation result, the server side verifies and stores the result, and updates the corpus task pool, updates the state of the corresponding corpus ID from 'locked' to 'completed' (or 'to be reviewed'), thereby releasing the lock, so that the system can continue to flow.
[0032] In a feasible implementation, before specifically performing the step of sending each target corpus annotation task to the client, the method further comprises: determining the target corpus corresponding to the target corpus annotation task, and configuring the target corpus state of the target corpus as a locked state; Illustratively, when the "locked state" configuration is completed, the target corpus becomes "unassignable" for subsequent S102 (task request) initiated by all other concurrent users (such as user B, user C). When other users also initiate requests, the S104 (task scheduling) performed by them will automatically skip all corpus in the "locked" state when searching the "corpus task pool", thereby ensuring that other users can only obtain other unassigned corpus. It can be understood that by "locking and binding the user in real time", it is ensured that any target corpus can only be held and processed by one user at the same time.
[0033] monitoring the corpus annotation submission request of the user end for the target corpus annotation task, and configuring the target corpus state of the target corpus as an annotation completed state based on the corpus annotation submission request.
[0034] The corpus annotation submission request refers to the data packet or API (application program interface) call sent by the user end to the server end after the user completes the annotation operation. The purpose of this request is to inform the server end that the previously issued target corpus annotation task has been processed by the user. In an optional embodiment, the request should at least include: user identification, task identification, annotation result.
[0035] Corpus state: refers to a field or mark stored in the server-side database (i.e. corpus task pool) for representing the stage of a specific corpus in the current business process. In the complete process of the present application, the state should at least include: 'to be allocated' (or 'unprocessed'), 'locked state' (or 'processing'), and 'annotation completion state' introduced in this step.
[0036] When the user completes the batch of target corpus annotation tasks issued in S106 (for example, selects a classification label for 10 corpora) on the user end (for example, a web browser), the user will trigger a submission action (for example, click the "Submit" button). After receiving the submission action, the user end encapsulates the "corpus annotation submission request" (as defined in "Noun explanation", containing UserID, Corpus IDs and annotation result data), and sends it asynchronously to the above-mentioned communication interface listened by the server end.
[0037] After receiving the request, the server end performs request verification. After verification, the server end extracts the annotation result data from the request and stores it persistently (for example, writes to an "annotation result table" or updates the main data table). After successfully storing the annotation result, the server end performs a database state update operation. All target corpora involved in the "corpus annotation submission request" and verified are updated from the locked state to the annotation completion state in the "target corpus state" field of the "corpus task pool".
[0038] In the embodiment of the present application, the service platform determines the task taking request of the user end for the corpus annotation task, adopts a task control strategy based on the task taking request to perform user task scheduling processing on the corpus task pool to obtain a task allocation number of target corpus annotation tasks, sends each target corpus annotation task to the client, and enables the client to perform corpus data classification annotation processing based on the target corpus annotation task. The automation and rationality of corpus classification task distribution are realized, human intervention is reduced, the uniqueness and exclusivity of corpus task allocation are guaranteed, the problem of repeated operation of the same corpus by multiple users is effectively avoided, and the annotation efficiency is significantly improved.
[0039] Optionally, please refer to Figure 2 , Figure 2 is a flowchart of a user task scheduling process proposed in the present application. The specific implementation of the user task scheduling processing based on the task taking request and adopting the task control strategy to obtain a task allocation number of target corpus annotation tasks can refer to the following ways: S202: determine the user ability portrait of the user end; User capability profile refers to a data structure or feature vector used to describe the comprehensive ability of the user end (i.e. the annotator) in the corpus annotation task. The profile is generated by the server side based on the statistical analysis and modeling of the user's historical annotation behavior data (e.g. historical annotation accuracy, average annotation speed, preferred task field, historical task completion rate, etc.). The profile is dynamically updated to reflect the user's current ability level in real time.
[0040] Illustratively, the server side first parses the user's unique identifier (UserID) from the request. Then, the server side uses the UserID to query and extract the latest "user capability profile" corresponding to the UserID from its internally maintained "user feature library" or "profile database".
[0041] In an optional embodiment, the "user capability profile" is continuously updated by a background process (or model) based on the user's newly submitted annotation results (such as after the submission in S204), ensuring the timeliness and accuracy of the profile. For example, if a user's recent annotation accuracy has significantly decreased, the "ability score" in the profile will also be adjusted accordingly.
[0042] S204: Based on the user capability profile, a task allocation processing large model is used to determine the number of tasks allocated to the user end, and based on the number of tasks allocated, the user task scheduling process is performed on the corpus task pool to obtain the target corpus annotation task for the user end.
[0043] Task allocation processing large model: refers to a pre-trained large language model (LLM) adapted in the task allocation scenario; in this embodiment, the core function of the model is to take the "user capability profile" determined in S202 as one of the inputs, and through its internal complex nonlinear processing logic, output a decision result.
[0044] Task allocation quantity: refers to a numerical value (N), which is calculated by the "task allocation processing large model" based on a specific "user capability profile". The value (N) is no longer a fixed configuration parameter in the system (e.g. everyone gets 10 every time), but is dynamically changing. It represents the number of corpus entries that the system believes the user is most suitable for in the current batch based on the evaluation of the user's ability.
[0045] Illustratively, the server side takes the "user ability profile" as input, feeds it to the pre-deployed "task allocation processing large model", performs a forward inference calculation through the large model, and its internal logic (e.g., a trained regression network) determines an optimal "task allocation number" (N) based on the input profile features.
[0046] For example: For the "high accuracy, high speed" profile, a larger N value (e.g., N = 30) should be output to maximize the throughput of this efficient user; for the "low accuracy, low speed" (e.g., novice) profile, a smaller N value (e.g., N = 5) should be output to allow the user to focus on a small number of tasks, ensuring quality and reducing the risk of task timeout.
[0047] In addition, through the task allocation processing large model, the subsequent scheduling process is performed with N as the core constraint, and the task allocation processing large model initiates a query to the "corpus task pool" (e.g., a database) to request N pieces of corpus with a state of 'to be classified'. After obtaining the N target corpora, the target corpus annotation tasks corresponding to the target corpora are generated.
[0048] In the embodiments of the present specification, by the above-mentioned manner, instead of using the "one-size-fits-all" fixed number allocation mode, the task allocation processing large model is used to deeply understand the user's ability, realizing personalized and intelligent "dynamic task package size" allocation, which can fully exert its efficiency, and novice users can also work under controllable pressure, thereby optimizing the platform's annotation throughput, task flow efficiency, and final annotation quality as a whole.
[0049] Optionally, the following exemplifies a model training process of a task allocation processing large model: In some embodiments, a trained basic large language model can be obtained, and the basic large language model is adapted to the task allocation processing scenario to obtain a task allocation processing large model. Generally, directly applying the basic large language model to the task allocation processing scenario is difficult to adapt to the new task allocation processing scenario. Therefore, the initial task allocation processing large model is created by obtaining the basic large language model, and sample data such as "sample user ability profile" in the new task allocation processing scenario (which can also include sample task feature profiles of sample corpora to be classified in the sample corpus task pool) are obtained. Since the basic large language model is usually an open source AIGC model that has been trained and has content generation capability, in the present specification, only the adaptation of the task allocation processing scenario to the basic large language model is required. Specifically, the sample data can be used to fine-tune the initial task allocation processing large model, and after the model fine-tuning training is completed, the task allocation processing large model adapted to the task allocation processing scenario is obtained.
[0050] Model creation: obtain a basic large language model, create an initial task allocation processing scene plug-in model for a task allocation processing scene, and compose an initial task allocation processing large model based on the basic large language model and the initial task allocation processing scene plug-in model; the basic large language model (MLLM) includes but is not limited to DeepSeek large model, GPT series large model, etc.
[0051] Sample data acquisition: acquire sample data under a new task allocation processing scene, which is sample user ability portrait and other sample data (also including sample task feature portrait of sample corpus to be classified in a sample corpus task pool) under the task allocation processing scene.
[0052] Sample data labeling: label the corresponding sample corpus labeling label based on the task allocation processing demand of the task allocation processing scene.
[0053] Model training process: input the sample data into the initial task allocation processing large model for at least one round of model training; in the model forward training process: based on the sample data, the initial task allocation processing large model is used to obtain a predicted corpus labeling task; In the model backward training process, the model loss value is determined based on the predicted corpus labeling task and the corpus labeling task label using a model loss function (such as Euclidean distance loss, hinge loss, and cross-entropy loss), and the model parameter adjustment of the initial task allocation processing scene plug-in model in the initial task allocation processing large model is performed based on the model loss value. The model structure of the basic large language model can be kept unchanged until the model training end condition is met to obtain the basic large language model and the task allocation processing scene plug-in model, complete the model fusion of the basic large language model and the task allocation processing scene plug-in model, and obtain the trained task allocation processing large model.
[0054] Illustratively, the initial task allocation processing scene plug-in model can be created based on a machine learning model.
[0055] Model fusion of the basic large language model and the task allocation processing scene plug-in model: the model structure layer weight of the task allocation processing scene plug-in model is fused with the basic large language model, the model structure layer parameters of the target model structure layer corresponding to the model structure layer weight in the basic large language model are fused with the model structure layer weight, and the model structure layer weight of the task allocation processing scene plug-in model can only exist in part of the corresponding model structure layer weight among all the model structure layers in the basic large language model. The model structure layer parameters are updated based on the model structure layer weight for this part of the target model structure layer. In this way, the reference update process of all model structure layer weights is completed, and the task allocation processing large model is obtained.
[0056] Optionally, the model end training condition of the model can include, for example, a value of a loss function being less than or equal to a preset loss function threshold, a number of iterations reaching a preset number threshold, and the like. The specific model end training condition can be determined based on actual conditions, and is not specifically limited here.
[0057] It should be noted that the machine learning model involved in one or more embodiments of the present specification includes, but is not limited to, fitting of one or more of a convolutional neural network (CNN) model, a deep neural network (DNN) model, a recurrent neural network (RNN), an embedding model, a gradient boosting decision tree (GBDT) model, a logistic regression (LR) model, and the like.
[0058] Optionally, please refer to Figure 3 , Figure 3 is a flowchart of a task allocation process proposed in the present specification. Specifically, the task allocation quantity for the user end is determined based on the user capability portrait by using a task allocation processing large model, and the target corpus labeling task for the user end is obtained by performing user task scheduling processing on the classified corpus in the corpus task pool based on the task allocation quantity. The following methods can be referred to: S302: inputting the user capability portrait into a task allocation processing large model; S304: obtaining a task feature portrait of the classified corpus in the corpus task pool by using the task allocation processing large model, the task feature portrait being constructed based on text complexity, professional field attribution, and estimated labeling time of the classified corpus; The task feature portrait refers to a data structure or feature vector for describing the attributes of a single classified corpus in the corpus task pool. The portrait is constructed by the task allocation processing large model after natural language processing (NLP) analysis of the corpus content.
[0059] Illustratively, one or more dedicated natural language processing (NLP) sub-modules (for example, a text classifier, a regression model, and the like) scheduled and integrated by the task allocation processing large model convert the original, unstructured “classified corpus” (for example, a piece of text) in the “corpus task pool” into a structured “task feature portrait” that can be used for matching operations in subsequent S306. The "task feature profile" is constructed based on the following key dimensions: 1. Text complexity: a quantitative evaluation of the difficulty level of the text in terms of understanding and processing.
[0060] Construction: In one embodiment, the task assignment processing large model can calculate a complexity score by analyzing a plurality of linguistic features of the text to be classified. These features include, but are not limited to, syntactic structure, lexical difficulty, readability indicators: for example, using industry-recognized readability algorithms (such as Flesch-Kincaid index, Gunning Fog index, etc.) to derive a comprehensive score.
[0061] 2. Professional field attribution: a classification label of the professional category to which the text to be classified belongs in terms of content.
[0062] Construction: In one embodiment, the task assignment processing large model can perform topic modeling or text classification on the text to be classified based on its powerful domain knowledge base and context understanding capabilities. For example, the model can determine that a piece of text belongs to a pre-defined professional field such as "financial reports", "medical diagnosis", "sports event review", or "legal contract terms".
[0063] 3. Estimated annotation time: a predictive value of the time required for a "standard ability" user (or based on a specific user profile) to complete the annotation of the text.
[0064] Construction: In one embodiment, the task assignment processing large model uses a plurality of features such as "text complexity", absolute length of the text (e.g. number of characters or words), and "professional field attribution" as references, and outputs an estimated time value (e.g. 0.5 minutes, 3.0 minutes, 10.0 minutes, etc.) through a pre-trained regression model.
[0065] In an optional implementation, the execution timing of S304 can be asynchronous. Specifically, the server-side background process can call the large model to complete S304 in advance when the "text to be classified" enters the "text task pool", and store (or cache) the generated "task feature profile" (including complexity, field, ETC, etc.) as a structured data field in the database entry of the text.
[0066] It can be understood that when S302 (user request) occurs, there is no need to perform real-time (Real-time) NLP calculation of S304, but the generated portraits can be directly and quickly retrieved from the database. This reduces the calculation delay of S306 (matching) step, significantly improves the response speed of task taking, and optimizes the user experience.
[0067] S306: calculating the corpus matching degree between the user ability portrait and each task feature portrait of the corpus to be classified in the corpus task pool through the task allocation processing large model, and determining the number of task allocations for the user end based on the task feature portrait and the user ability portrait; The corpus matching degree refers to a value (for example, a normalized score between 0 and 1) calculated by the task allocation processing large model in real time to quantify the degree of fit between the user ability portrait and the task feature portrait. The higher the score, the more suitable it is to assign the piece of corpus to be classified (task) to the user end for processing.
[0068] The detailed process of S306 step can be divided into the following two sub-processes executed by the task allocation processing large model: Sub-process A: calculating the corpus matching degree Input: query vector: the user ability portrait determined in S302 (for example, containing features such as {expertise: finance, accuracy: 98%}). Candidate vector set: the task feature portraits of all "corpora to be classified" in the "corpus task pool" obtained in S304 (for example, T1={field: finance, complexity: high}, T2={field: sports, complexity: low}, etc.).
[0069] Sub-process A calculation mechanism: the task allocation processing large model performs similarity or correlation calculation on the "query vector" and each "task feature portrait" in the "candidate vector set". The calculation can be based on the cosine similarity of specific dimension vectors, for example, calculating the similarity of the user's "expertise" vector and the task's "professional field attribution" vector.
[0070] Output: The output of this sub-process is the "priority allocation list". In this list, all "corpora to be classified" have been sorted in descending order of their "corpus matching degree" scores with the current user.
[0071] Sub-process B: determining the number of task allocations Input: user ability portrait (characteristics representing user efficiency, such as "personalized efficiency coefficient"), task feature portrait; (optionally) system preset target task time window (for example, expecting the user to complete a batch of tasks within 30 minutes).
[0072] Computer mechanism: A dynamic quantitative calculation is performed by the task allocation processing large model, which can adopt a "target time window-based" strategy: Step i: The task allocation processing large model extracts the "personalized efficiency coefficient" (for example, Efficiency = 1.5x) from the "user ability portrait".
[0073] Step ii: The task allocation processing large model extracts the "estimated labeling time consumption" from the "task feature portrait". Preferably, the large model will refer to the average estimated labeling time consumption of the "high matching degree" corpus screened in sub-process A (for example, Avg_ETC = 5 minutes per item).
[0074] Step iii: The task allocation processing large model performs a calculation based on the pre-set "target task time window" (for example, Window = 30 minutes), for example: Task allocation quantity (N) = Window / (Avg_ETC / Efficiency).
[0075] Output: The output of this sub-process is a dynamic integer value N, for example, N = 30 / (5 / 1.5) ≈ 9. The N value is the "task allocation quantity" for this allocation.
[0076] After the execution of step S306, the large model has obtained both the "priority allocation list" (from sub-process A) and the "task allocation quantity N" (from sub-process B). These two calculation results will be used as inputs for step S308 to perform the final "selection" operation.
[0077] In one possible implementation, the determination of the task allocation quantity for the user end based on the task feature portrait and the user ability portrait can be performed in the following manner: Step A2: Obtain a pre-set target task time window, based on the personalized efficiency coefficient in the user ability portrait that represents the historical labeling efficiency of the user end, the target task time window represents the pre-set duration that the user end is expected to complete a single batch of tasks; The target task time window refers to a benchmark duration parameter (for example: 15 minutes, 30 minutes or 60 minutes) that is pre-set by the system (for example, platform administrator) to guide task allocation. As defined in step A2 of this embodiment, it represents the pre-set duration that the system expects a user end (labeler) to complete a batch of tasks after receiving the batch of tasks. This parameter is the core anchor point for calculating the "task allocation quantity" in this embodiment, and its purpose is to unify task packages of different difficulties and different users to a relatively fixed expected completion time.
[0078] The personalized efficiency coefficient refers to a quantitative value (e.g., 0.8, 1.0, 1.5) extracted from the user capability profile. The coefficient is derived by the server based on long-term analysis of the user's "historical labeling efficiency" and is used to represent the user's labeling speed ratio compared to a "standard user" (e.g., coefficient 1.0). For example, a coefficient of 1.5 represents that the user's labeling speed is 1.5 times the standard speed (high efficiency); a coefficient of 0.8 represents that the speed is 80% of the standard speed (slower).
[0079] Illustratively, through the task allocation process, the large model reads the preset value of the target task time window from the system global configuration, and at the same time, from the input "user capability profile", analyzes and extracts the "personalized efficiency coefficient" representing the user's historical efficiency.
[0080] Step A4: From the task feature profile, obtain the average estimated labeling time consumption of the to-be-classified corpus. Based on the average estimated labeling time consumption and the personalized efficiency coefficient, determine the personalized estimated single-time consumption of the user end; The average estimated labeling time consumption refers to the reference time consumption value obtained from the "task feature profile". This value can be the global average time consumption of all to-be-classified corpora in the "corpus task pool", or the local average value (e.g., 5.5 minutes per piece) calculated by the large model extracting the "estimated labeling time consumption" (ETC) feature of those "to-be-classified corpora" judged as "high matching degree" (understandable as the to-be-assigned part).
[0081] The personalized estimated single-time consumption refers to an intermediate value calculated in real time, representing the actual time required by the system to estimate the completion of a single corpus by a specific user (defined by the "personalized efficiency coefficient") to process a specific task (defined by the "average estimated labeling time consumption").
[0082] Illustratively, the large model refers to the results of S306 matching degree calculation to circumscribe the high matching degree to-be-classified corpus set with matching degree greater than the threshold value. The model extracts the estimated labeling time consumption from the task feature profile of the set and calculates the average value to obtain the average estimated labeling time consumption.
[0083] The core calculation of the large model in this step is to adjust the reference time consumption of this step based on the coefficient of step A2 to determine the "personalized estimated single-time consumption". The calculation method is: personalized estimated single-time consumption = average estimated labeling time consumption / personalized efficiency coefficient.
[0084] Step A6: Based on the target task time window and the personalized estimated single-time consumption, determine the task allocation quantity.
[0085] The large model performs the final calculation of this embodiment to determine the task allocation number (N) in the following manner: task allocation number (N) = target task time window / personalized estimated single time consumption.
[0086] This calculation method shows that, under the condition of fixed total time budget (time window), the shorter the personalized time consumption of a user in processing a single task, the more tasks N he can be allocated.
[0087] In an optional embodiment, since the calculation result N can be a floating point number, an integer number of the result is also obtained by performing an integer operation (e.g., floor or round) on the result.
[0088] S308: The task allocation processing large model selects target corpus indicated by the task allocation number from the to-be-classified corpus according to the corpus matching degree, and generates a target corpus annotation task corresponding to the target corpus.
[0089] The corpus matching degree refers to a quantitative score representing the degree of fit between the user and the task.
[0090] The task allocation number refers to an integer value N calculated dynamically.
[0091] The target corpus refers to a specific corpus subset.
[0092] The target corpus annotation task refers to a final generated data package or object. The object is formed after the target corpus is locked and packaged, and can be directly sent to the user end.
[0093] Illustratively, the task allocation processing large model selects N target corpora from the to-be-classified corpus from high to low according to the corpus matching degree, the status field of the N corpora in the corpus task pool is updated from 'to be classified' to 'locked state', and the locked state is bound to the current user ID initiating the request, which can realize the concurrent control (avoiding repeated collection) of the present application. Further, the contents (e.g., text ID, text content, annotation requirements, etc.) of the N target corpora are extracted and encapsulated into one (or more) structured data package (e.g., a JSON object) to obtain the target corpus annotation task.
[0094] In the present specification, by introducing a task allocation processing large model, the task scheduling is upgraded from "passive response" to "active intelligent matching". The settings of S304 (acquiring task portrait) and S306 (calculating matching degree) provide protection for the system, ensuring that (for example) high-difficulty or specific field tasks can be accurately allocated to (S302) users with corresponding ability or professional background, thereby solving the problem of low labeling quality caused by task mismatch. Further, the settings of S306 (determining the number of task allocations) and S308 (selecting by quantity) provide intelligent regulation for the system. This mechanism can dynamically calculate the size of the personalized task package based on the dual portrait of user ability and task difficulty, rather than a "one-size-fits-all" fixed number, ensuring the optimal allocation of platform resources (human and task), while ensuring quality, maximizing platform labeling efficiency and throughput.
[0095] Optionally, as shown in Figure 4 Figure 4 is a process schematic diagram of annotation behavior monitoring. In the task processing method of the corpus annotation task of one or more embodiments of the present specification, the following steps are further included: S402: Monitor the annotation behavior data of the user end; Annotation behavior data refers to the process data generated by the user when performing annotation operations on the user end, rather than only referring to the final submitted annotation results. In one embodiment, the data at least includes: Time data: for example, the dwell time of a single corpus, the average annotation time consumption, the hesitation time from the first click to the final submission.
[0096] Interaction data: for example, the click frequency of the label, the number of label modifications, the mouse movement track, and the usage mode of the shortcut key.
[0097] Distribution data: for example, in the current batch (N tasks issued by S106), the distribution of the labels selected by the user (for example, whether 99% all select the same label).
[0098] Illustratively, the server end (or through the script embedded in the user end) starts a behavior listener, which collects the "annotation behavior data" (as defined in "Noun explanation") on the client interface in real time (or in micro-batches, Micro-batch).
[0099] S404: Determine the real-time annotation quality of the user based on the annotation behavior data using a quality evaluation model; Quality Assessment Model refers to a pre-trained AI model (e.g. anomaly detection model, classifier or a regression model) that can be a sub-module of the Task Assignment Processing Large Model (mentioned in S302) or an independent model. The input of this model is the "labeled behavior data" monitored in S402, and the output is the "instant labeling quality" in S404. This model learns the behavior data paradigm of "high-quality users" and "low-quality / cheating users" from a large amount of historical data, thereby acquiring the ability to predict the current user's labeling quality.
[0100] Instant labeling quality refers to the evaluation signal or score output by the "Quality Assessment Model" in S404 after real-time (or quasi-real-time) analysis of S402 data. The signal can be: a classification label: e.g. high quality, medium trust, suspicious (suspected low quality), cheating (suspected data brushing); a quantitative score: e.g. Quality_Score = 0.95 (high quality) or Suspicion_Score = 0.8 (highly suspicious).
[0101] Illustratively, by performing S402, its data stream is continuously or periodically (e.g. every 10 seconds) fed into this step. The execution method is: Input the behavior data collected in S402 (e.g. {avg_time: 1.5s / item, label_distribution: {"A": 100%}}) into the "Quality Assessment Model", and the model compares this behavior data with the baseline. The baseline can come from: the "task feature portrait" constructed in S304 (e.g. the estimated labeling time ETC of this batch of tasks is 15 seconds per item).
[0102] the "user ability portrait" of the user (S302) (e.g. the user's historical average time consumption is 12 seconds per item).
[0103] For example, the quality assessment model determines that the current user's avg_time: 1.5s is much lower than the ETC: 15s, and its label_distribution is extremely unbalanced.
[0104] Output: The "Quality Assessment Model" performs inference and outputs an "instant labeling quality" signal, e.g. Quality_Flag = "Suspicious_Low_Effort" (suspected low quality, data brushing).
[0105] S406: Perform task labeling intervention processing on the user end based on the instant labeling quality.
[0106] The execution manner of the task labeling intervention processing: Dynamic insertion of a "quality inspection task": the system selects a quality inspection task (for example, QC_001, whose known answer is 'label B') that is highly similar to the current task (issued by S106) from a preset "standard answer library" (quality inspection task pool). The system dynamically inserts this QC_001 into the user's current locked task queue (for example, as the next task to be labeled).
[0107] Instant verification: the system waits for the user to label QC_001. When the user submits the labeling result of QC_001 (for example, the user still selects 'label A' by inertia), the system immediately compares it with the "standard answer" ('label B').
[0108] Intervention execution: if the verification fails: it proves that the judgment of S404 is accurate. The system performs strong intervention, for example, one or more of the following: freeze permissions: immediately "freeze" the user's "submission permission" (S202) for the batch (remaining) tasks issued by S106; recycle tasks: release all (remaining) tasks locked by the user from the 'locked state' to the 'to be allocated' state for other users to take; send a warning: send a "quality warning" pop-up window to the user end.
[0109] Real-time update of user portrait: based on the "quality inspection failure" event, the "user ability portrait" of the user is dynamically downgraded.
[0110] Further, if the verification succeeds: it proves that S404 may be a false judgment (or the user has been alerted and started to label seriously). The system removes the "suspicious" state and allows the user to continue.
[0111] In one possible implementation, if the instant labeling quality is less than a first quality threshold, the user ability portrait is updated. If the instant labeling quality is less than a second quality threshold, the labeling submission permission of the client is frozen until the labeling result of the quality inspection task meets the preset standard.
[0112] In this specification, a set of "in-process" real-time quality monitoring and intervention loops are achieved by the above-mentioned manner, rather than traditional "after-the-fact" quality inspection. This mechanism can achieve "instant stop loss", that is, before the user submits (such as S202) low-quality labeling data and pollutes the database, it is identified by the evaluation model of S404 and intercepted by the intervention means (such as freezing or inserting a quality inspection task) of S406, thereby greatly saving the subsequent quality inspection and rework costs.
[0113] Further, the scheme feeds back the "in-process" monitoring results to the "pre-process" task allocation model (such as S302-S308) in real time through S406 (e.g., real-time updating of user portrait). This enables the entire system to have real-time, self-adaptive error correction and optimization capabilities: once the real-time labeling quality of a user is determined to be low, the user's ability portrait will be immediately down-weighted, causing the task matching model of S306 to automatically avoid assigning similar high-value or high-difficulty tasks to the user in the next round, thus realizing an intelligent closed loop from "allocation" to "monitoring" to "allocation".
[0114] The embodiments of the present specification will be described below in conjunction with Figure 5 The corpus annotation task processing device provided by the embodiments of the present specification will be described in detail. It should be noted that, Figure 5 The corpus annotation task processing device shown in the embodiment is used to execute the method of the present specification Figures 1-4 The method of the embodiment shown is only shown in relation to the embodiments of the present specification for ease of description, and specific technical details not disclosed are described with reference to the method of the present specification Figures 1-4 The embodiment shown.
[0115] Please refer to Figure 5 which shows a structural schematic diagram of the corpus annotation task processing device according to the embodiments of the present specification. The corpus annotation task processing device 1 can be realized by software, hardware, or a combination of the two to become all or part of a device. According to some embodiments, the corpus annotation task processing device 1 includes a request determination module 11, a task scheduling module 12, and a task sending module 13, which are specifically used for: The request determination module 11 is configured to determine a task receiving request of a user end for a corpus annotation task; The task scheduling module 12 is configured to perform user task scheduling processing on a corpus task pool based on the task receiving request to obtain a target corpus annotation task in a task allocation quantity by using a task control strategy; The task sending module 13 is configured to send each of the target corpus annotation tasks to the client end, so that the client end performs corpus data classification and annotation processing based on the target corpus annotation tasks.
[0116] In a possible implementation, the user task scheduling processing on the corpus task pool based on the task receiving request to obtain a target corpus annotation task in a task allocation quantity by using a task control strategy includes: determining a user ability portrait for the user end; determining a task allocation quantity for the user end by using a task allocation processing large model based on the user ability portrait, and performing user task scheduling processing on the corpus task pool in the task allocation quantity to obtain a target corpus annotation task for the user end.
[0117] In an implementable embodiment, based on the user capability profile, a task allocation processing large model is adopted to determine a task allocation quantity for the user end, and based on the task allocation quantity, user task scheduling processing is performed on the to-be-classified corpus in the corpus task pool to obtain a target corpus annotation task for the user end, comprising: inputting the user capability profile into a task allocation processing large model; obtaining, by the task allocation processing large model, a task feature profile of the to-be-classified corpus in the corpus task pool, the task feature profile being constructed based on text complexity, professional field attribution, and estimated annotation time consumption of the to-be-classified corpus; calculating, by the task allocation processing large model, a corpus matching degree between the user capability profile and each of the task feature profiles of the to-be-classified corpus in the corpus task pool, and determining a task allocation quantity for the user end based on the task feature profiles and the user capability profile; selecting, by the task allocation processing large model, target corpora indicated by the task allocation quantity from the to-be-classified corpora according to the corpus matching degree, and generating target corpus annotation tasks corresponding to the target corpora.
[0118] In an implementable embodiment, determining a task allocation quantity for the user end based on the task feature profile and the user capability profile comprises: obtaining a preset target task time window, determining a personalized efficiency coefficient representing historical annotation efficiency of the user end in the user capability profile, and the target task time window representing a preset time length expected for the user end to complete a single batch of tasks; from the task feature profile, obtaining an average estimated annotation time consumption of the to-be-classified corpus, and based on the average estimated annotation time consumption and the personalized efficiency coefficient, determining a personalized estimated single-time consumption of the user end; based on the target task time window and the personalized estimated single-time consumption, determining a task allocation quantity.
[0119] In an implementable embodiment, further comprising: monitoring annotation behavior data of the user end; based on the annotation behavior data, adopting a quality evaluation model to determine instant annotation quality of the user; based on the instant annotation quality, performing task annotation intervention processing on the user end.
[0120] In an implementable embodiment, based on the instant annotation quality, performing task annotation intervention processing on the user end, comprising: if the instant annotation quality is less than a first quality threshold, performing profile updating processing on the user capability profile; If the instant labeling quality is less than a second quality threshold, the labeling submission permission of the client is frozen until the labeling result of the quality inspection task meets a preset standard.
[0121] In an implementation, before the target corpus labeling tasks are sent to the client, the method further includes: determining a target corpus corresponding to the target corpus labeling task, and configuring a target corpus state of the target corpus as a locked state; and / or, monitoring a corpus labeling submission request of the user end for the target corpus labeling task, and configuring the target corpus state of the target corpus as a labeling completed state based on the corpus labeling submission request.
[0122] It should be noted that the corpus labeling task processing apparatus provided in the above embodiments is used to execute the corpus labeling task processing method, and only the division of the above functional modules is used as an example for illustration. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the corpus labeling task processing apparatus and the corpus labeling task processing method provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0123] The above embodiment numbers of the present specification are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0124] The present specification also provides a computer storage medium, which can store a plurality of instructions, the instructions being suitable for being loaded and executed by a processor to implement the corpus labeling task processing method of the above Figures 1-4 embodiments. The specific implementation process can be referred to the specific description of the above Figures 1-4 embodiments, which will not be repeated here.
[0125] The present specification also provides a computer program product, which stores at least one instruction, the at least one instruction being loaded and executed by the processor to implement the corpus labeling task processing method of the above Figures 1-4 embodiments. The specific implementation process can be referred to the specific description of the above Figures 1-4 embodiments, which will not be repeated here.
[0126] Please refer to Figure 6A structural block diagram of an electronic device according to an embodiment of the present disclosure is provided. The electronic device according to the present disclosure can include one or more of the following components: a processor 1010, a memory 1020, an input device 1030, an output device 1040, and a bus 1050. The processor 1010, the memory 1020, the input device 1030, and the output device 1040 can be connected to each other through the bus 1050.
[0127] The processor 1010 can include one or more processing cores. The processor 1010 connects various parts within the entire electronic device using various interfaces and lines, and performs various functions of the electronic device and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1020, and calling data stored in the memory 1020. Alternatively, the processor 1010 can be implemented in at least one of a hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 1010 can integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes an operating system, a user interface, and an application program, etc.; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 1010, but can be implemented by a separate communication chip.
[0128] The memory 1020 can include a random access memory (RAM) and can also include a read-only memory (ROM). Alternatively, the memory 1020 includes a non-transitory computer-readable storage medium. The memory 1020 can be used to store instructions, programs, codes, code sets, or instruction sets.
[0129] The input device 1030 is configured to receive input instructions or data, and the input device 1030 includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 1040 is configured to output instructions or data, and the output device 1040 includes, but is not limited to, a display device and a speaker. In the embodiments of the present specification, the input device 1030 can be a temperature sensor configured to obtain the operating temperature of the electronic device. The output device 1040 can be a speaker configured to output an audio signal.
[0130] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above-described drawings does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than those shown in the drawings, or combine certain components, or different component arrangements. For example, the electronic device further includes a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WIFI) module, a power supply, a Bluetooth module, and the like, which are not described herein.
[0131] In the embodiments of the present specification, the execution subject of each step can be the electronic device described above. Alternatively, the execution subject of each step is an operating system of the electronic device. The operating system can be an Android system, an IOS system, or other operating systems, and the embodiments of the present specification do not limit the operating system.
[0132] In the electronic device, Figure 6 In the electronic device, the processor 1010 can be configured to call a program stored in the memory 1020 and perform to implement the corpus annotation task processing method as described in the various method embodiments of the present specification.
[0133] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiments can be completed by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory, or a random access memory, etc.
[0134] The above only describes the preferred embodiments of the present specification, and of course cannot limit the scope of the rights of the present specification, so equivalent changes made according to the claims of the present specification still fall within the scope of the present specification.
Claims
1. A method for processing corpus annotation tasks, characterized in that, The method includes: Determine the user's task claiming request for the corpus annotation task; Based on the task request, a task control strategy is used to perform user task scheduling on the corpus task pool to obtain the target corpus annotation task with the assigned task quantity. Each of the target corpus annotation tasks is sent to the client, so that the client can perform corpus data classification and annotation processing based on the target corpus annotation tasks.
2. The method according to claim 1, characterized in that, The target corpus annotation task, which uses a task control strategy to perform user task scheduling on the corpus task pool based on task request to obtain the number of task allocations, includes: Determine the user capability profile for the aforementioned user terminal; Based on the user capability profile, a task allocation processing model is used to determine the number of tasks to be allocated to the user. Based on the number of tasks allocated, user task scheduling processing is performed on the corpus to be classified in the corpus task pool to obtain the target corpus annotation task for the user.
3. The method according to claim 2, characterized in that, Based on the user capability profile, a task allocation processing model is used to determine the number of tasks to be allocated to the user. Based on the number of tasks allocated, user task scheduling processing is performed on the unclassified corpus in the corpus task pool to obtain the target corpus annotation task for the user, including: The user capability profile is input into the task allocation and processing model. The task feature profile of the corpus to be classified in the task pool is obtained through the task allocation and processing model. The task feature profile is constructed based on the text complexity, professional domain affiliation, and estimated annotation time of the corpus to be classified. The task allocation processing model calculates the corpus matching degree between the user capability profile and each task feature profile of the corpus to be classified in the corpus task pool, and determines the number of tasks to be allocated to the user based on the task feature profile and the user capability profile. The task allocation processing model selects the target corpus indicated by the task allocation quantity from the corpus to be classified based on the corpus matching degree, and generates the target corpus annotation task corresponding to the target corpus.
4. The method according to claim 3, characterized in that, The step of determining the number of tasks to be allocated to the user based on the task feature profile and the user capability profile includes: Obtain a preset target task time window, and determine a personalized efficiency coefficient representing the historical annotation efficiency of the user terminal based on the user capability profile. The target task time window represents the preset time expected for the user terminal to complete a single batch of tasks. From the task feature profile, the average estimated annotation time of the corpus to be classified is obtained, and based on the average estimated annotation time and the personalized efficiency coefficient, the personalized estimated single-item time of the user terminal is determined. The number of tasks to be assigned is determined based on the target task time window and the personalized estimated time for a single task.
5. The method according to claim 1, characterized in that, The method further includes: Monitor the annotation behavior data of the user terminal; The user's real-time annotation quality is determined using a quality assessment model based on the annotation behavior data. Based on the real-time annotation quality, the user terminal is subjected to task annotation intervention processing.
6. The method according to claim 5, characterized in that, The task annotation intervention process performed on the user terminal based on the real-time annotation quality includes: If the quality of the instant annotation is less than the first quality threshold, then the user capability profile is updated. If the quality of the instant annotation is less than the second quality threshold, the annotation submission permission of the client is frozen until the annotation result of the quality inspection task meets the preset standard.
7. The method according to claim 1, characterized in that, Before sending each of the target corpus annotation tasks to the client, the method further includes: Determine the target corpus corresponding to the target corpus annotation task, and configure the target corpus state of the target corpus to a locked state; and / or, Monitor the user's submission request for the target corpus annotation task, and configure the target corpus status to annotation completed based on the submission request.
8. A corpus annotation task processing device, characterized in that, The device includes: The request determination module is used to determine the user's task retrieval request for the corpus annotation task; The task scheduling module is used to perform user task scheduling on the corpus task pool based on task retrieval requests and to obtain the target corpus annotation tasks with the assigned task quantity by adopting task control strategies. The task sending module is used to send each of the target corpus annotation tasks to the client, so that the client can perform corpus data classification and annotation processing based on the target corpus annotation tasks.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the steps of the method as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as described in any one of claims 1 to 7.