A man-machine collaborative data labeling method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610693948.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本申请提供了一种人机协同的数据标注方法、装置、电子设备及存储介质,以解决现有技术中传统人工标注成本与标注效率失衡,难以确保标注数据的可靠性和一致性;纯模型标注模型自身存在偏差和错误模式,缺乏对标签的解释与校验能力,标注策略缺乏动态适应性的问题
[0016]本申请实施例提供的上述技术方案与现有技术相比具有如下优点:本申请实施例提供的该方法,通过构建具有可变参数区域的提示词模板,将目标标注任务的需求信息嵌入所述提示词模板,得到目标提示词;将所述目标提示词并行输入多个大语言模型中,输出多组候选标注结果及所述候选标注结果对应的推理逻辑链;从多组所述候选标注结果及所述候选标注结果对应的推理逻辑链筛选出符合预设条件的候选标注结果,以及对应的推理逻辑链;采用预设的评估方式对所述候选标注结果,以及对应的推理逻辑链进行处理得到标准标注数据;将所述标准标注数据作为所述目标标注任务的标注结果,并存储于标注数据库。实现使标注者能够在每次交互中针对特定任务或文本特征调整提示词,确保标注结果更加精准和符合实际需求。多维自动化评估与人工验证,确保标注数据在多轮交互中的稳定性和准确性,并确保最终采纳的标注数据具备高可靠性和任务贴合度。
Smart Images

Figure CN122819174A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data annotation technology, and in particular to a human-computer collaborative data annotation method, apparatus, electronic device and storage medium. Background Technology
[0002] In machine learning tasks such as natural language processing and computer vision, labeled data is the foundation for model training. Early methods relied primarily on traditional manual annotation. However, with the development of pre-trained models, pure model annotation has gradually become more common, utilizing existing models to automatically generate pseudo-labels or directly output annotation results to replace the manual annotation process.
[0003] Traditional manual annotation involves establishing annotation standards, having annotators review each data point manually, and then having professionals conduct random sampling reviews. Pure model annotation, on the other hand, uses initial annotation data to fine-tune the model, and then uses the optimized model to re-annotate the data.
[0004] Existing technologies typically suffer from the following problems: traditional manual annotation is slow, resulting in an imbalance between labor costs and annotation efficiency; subjective differences exist among different annotators, making it difficult to ensure the reliability and consistency of the annotated data; pure model annotation inherently contains biases and error patterns; there is a lack of ability to interpret and verify labels, making it difficult to detect annotation errors; and annotation strategies lack dynamic adaptability. Summary of the Invention
[0005] This application provides a human-machine collaborative data annotation method, device, electronic device, and storage medium to solve the problems in the prior art where the cost and efficiency of traditional manual annotation are unbalanced, making it difficult to ensure the reliability and consistency of annotated data; pure model annotation models themselves have biases and error patterns, lack the ability to interpret and verify labels, and the annotation strategy lacks dynamic adaptability.
[0006] Firstly, this application provides a human-machine collaborative data annotation method, including: Construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word; The target prompt words are input into multiple large language models in parallel, and multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results are output. Candidate annotation results that meet preset conditions and corresponding inference logic chains are selected from multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results. The candidate annotation results and the corresponding inference logic chain are processed using a preset evaluation method to obtain standard annotation data; The standard annotation data is used as the annotation result of the target annotation task and stored in the annotation database.
[0007] In one possible implementation, the step of processing the candidate annotation results and the corresponding inference logic chain using a preset evaluation method to obtain standard annotation data includes: The candidate annotation results and the corresponding inference logic chain are manually verified according to the preset annotation requirements. Candidate annotation results that meet all the preset annotation requirements, along with their corresponding inference logic chains, are selected as standard annotation data.
[0008] In one possible implementation, the step of processing the candidate annotation results and the corresponding inference logic chain using a preset evaluation method to obtain standard annotation data includes: The candidate annotation results and the corresponding inference logic chain are manually verified according to the preset annotation requirements. Select candidate annotation results that meet the preset annotation requirements, update the candidate annotation results, and use the updated candidate annotation results and the corresponding inference logic chain as standard annotation data.
[0009] In one possible implementation, the step of processing the candidate annotation results and the corresponding inference logic chain using a preset evaluation method to obtain standard annotation data includes: The candidate annotation results and the corresponding inference logic chain are manually verified according to the preset annotation requirements. When all the candidate annotation results do not meet the preset annotation requirements, the input standard annotation results are received and used as standard annotation data.
[0010] In one possible implementation, the method further includes: When some or all of the candidate annotation results meet the preset annotation requirements, non-standard information is extracted from the inference logic chain corresponding to the candidate annotation results. The prompt word template is updated based on the non-standard information.
[0011] In one possible implementation, the step of filtering candidate annotation results that meet preset conditions from multiple sets of candidate annotation results and the corresponding inference logic chains of the candidate annotation results, and the corresponding inference logic chains, includes: Automated quality assessment is performed on multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; Candidate annotation results that meet preset conditions and their corresponding reasoning logic chains are selected and retained. The preset conditions include that the perplexity value of the candidate annotation results is less than a first threshold and the semantic similarity value is greater than a second threshold.
[0012] In one possible implementation, the method further includes: When there are candidate annotation results that do not meet the preset conditions, a regeneration mechanism is triggered to re-execute the steps of inputting the target prompt word into multiple large language models in parallel and outputting multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
[0013] Secondly, this application provides a human-machine collaborative data annotation device, comprising: The acquisition module is used to construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word; The output module is used to input the target prompt words into multiple large language models in parallel and output multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results; The filtering module is used to filter out candidate annotation results that meet preset conditions and corresponding inference logic chains from multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; The evaluation module is used to process the candidate annotation results and the corresponding inference logic chain using a preset evaluation method to obtain standard annotation data; The storage module is used to store the standard annotation data as the annotation result of the target annotation task in the annotation database.
[0014] Thirdly, this application provides an electronic device, including: a processor and a memory, wherein the processor is configured to execute a human-machine collaborative data annotation program stored in the memory to implement the human-machine collaborative data annotation method described in any one of the first aspects.
[0015] Fourthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the human-machine collaborative data annotation method described in the first aspect.
[0016] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application constructs a prompt word template with a variable parameter region, embeds the requirement information of the target annotation task into the prompt word template, and obtains the target prompt word; the target prompt word is input into multiple large language models in parallel, and multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results are output; candidate annotation results that meet preset conditions and the corresponding inference logic chains are selected from the multiple sets of candidate annotation results and the corresponding inference logic chains; the candidate annotation results and the corresponding inference logic chains are processed using a preset evaluation method to obtain standard annotation data; the standard annotation data is used as the annotation result of the target annotation task and stored in the annotation database. This enables annotators to adjust the prompt words for specific tasks or text features in each interaction, ensuring that the annotation results are more accurate and meet actual needs. Multi-dimensional automated evaluation and manual verification ensure the stability and accuracy of the annotation data in multiple rounds of interaction, and ensure that the finally adopted annotation data has high reliability and task fit. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0020] Figure 1 A flowchart illustrating a human-machine collaborative data annotation method provided in this application embodiment; Figure 2 A flowchart illustrating another human-machine collaborative data annotation method provided in this application embodiment; Figure 3 A schematic diagram of the structure of a human-machine collaborative data annotation device provided in this application embodiment; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0023] To address the technical problems in existing technologies, such as the imbalance between cost and efficiency in traditional manual annotation, the difficulty in ensuring the reliability and consistency of labeled data, the inherent biases and error patterns in pure model annotation models, the lack of interpretation and verification capabilities for labels, and the lack of dynamic adaptability in annotation strategies, this application provides a human-machine collaborative data annotation method that can achieve high-quality, controllable data annotation, improve data annotation efficiency, and ensure the reliability and consistency of labeled data.
[0024] Figure 1 This is a flowchart illustrating a human-computer collaborative data annotation method provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps: S101. Construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word.
[0025] The human-machine collaborative data annotation method provided in this application is applied to data annotation scenarios, specifically to human-machine collaborative data annotation systems. Through the co-evolution of model learning and human cognition within the system, high-quality and controllable data annotation is achieved, improving data annotation efficiency and building a solid, high-quality data foundation for subsequent model training.
[0026] In this embodiment, before the data annotation workflow starts, a prompt word template with a variable parameter area needs to be pre-built and injected into the system for operation. This prompt word template can be flexibly expanded according to the specific needs of the annotation task. When handling different types of annotation tasks, the variable parameters in the template can be modified or new parameters can be added according to the actual needs of the current annotation task, and the content of the prompt word template can be adjusted accordingly. Therefore, this template can adapt to various annotation scenarios without rebuilding the entire workflow, thereby improving the task adaptability and configuration efficiency of the data annotation system.
[0027] Specifically, based on the annotation requirements of the target annotation task, the corresponding requirement information is obtained, and this requirement information is embedded into the variable parameter area of the aforementioned prompt word template to generate target prompt words. By driving the annotation process through configurable prompt word templates, the model in the subsequent system can flexibly respond to different annotation requirements. When the annotation task changes, only the parameter configuration in the template needs to be adjusted to achieve task adaptation, significantly improving the system's reusability and scalability.
[0028] S102. Input the target prompt words into multiple large language models in parallel, and output multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
[0029] In this embodiment, the generated target prompt word is input in parallel into multiple large language models in the system. Each model, upon receiving the target prompt word, performs natural language understanding and generation based on its content, and independently completes the annotation task indicated by the target prompt word. Each model outputs a set of candidate annotation results, along with a corresponding inference logic chain. The inference logic chain describes the intermediate reasoning process from understanding the prompt word to generating the annotation result. By calling multiple large language models in parallel, the system can simultaneously obtain multiple sets of annotation outputs based on the same prompt word but from different model perspectives, providing diverse candidate bases for subsequent result output and quality evaluation.
[0030] S103. Select candidate annotation results that meet the preset conditions and the corresponding inference logic chains from multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results.
[0031] In this embodiment, after multiple large models output multiple sets of candidate annotation results, the system enters an automated quality assessment stage to filter out low-quality model outputs. Specifically, the candidate results output by multiple models cannot be used as final annotation data and need to undergo further screening. Based on preset conditions, the system performs automated quality assessment on the multiple sets of candidate annotation results and the corresponding inference logic chains, initially screening and retaining candidate annotation results that meet the preset conditions, as well as the corresponding inference logic chains.
[0032] In an optional embodiment of this application, candidate annotation results that meet preset conditions, along with their corresponding inference logic chains, proceed to the next stage of manual verification; while candidate annotation results that do not meet the preset conditions, along with their corresponding inference logic chains, are marked as low-quality results by the system and removed from the current process. Through this process, the system can automatically remove model output results that clearly do not meet annotation requirements before manual intervention, thereby improving overall annotation efficiency and quality control.
[0033] S104. The candidate annotation results and the corresponding inference logic chain are processed using a preset evaluation method to obtain standard annotation data.
[0034] In this embodiment of the application, after the above-mentioned automated quality assessment process, the candidate annotation results that meet the preset conditions, as well as the corresponding inference logic chain, will enter the further verification and evaluation stage.
[0035] In this stage, annotators, based on the pre-defined annotation requirements and standards of the target annotation task, use pre-defined evaluation methods to judge the annotation results output by the model and the corresponding inference logic chain. Annotators need to determine whether the candidate annotation result truly meets the pre-defined annotation requirements and standards, and based on its corresponding inference logic chain, judge whether the generation process of the result is reasonable and logically coherent. Specifically, the verification results can include the following three situations: all content of the candidate annotation result meets the pre-defined annotation requirements; some content of the candidate annotation result meets the pre-defined annotation requirements; and all content of the candidate annotation result does not meet the pre-defined annotation requirements.
[0036] Furthermore, the candidate annotation results and their corresponding inference logic chains are processed based on the verification results to obtain standard annotation data. By further evaluating the quality of the aforementioned automated screening results, the limitations of automated indicators in semantic understanding and task fit judgment are compensated for, ensuring that the final standard data has high reliability and task fit.
[0037] S105. Use the standard annotation data as the annotation result of the target annotation task and store it in the annotation database.
[0038] In this embodiment, after both automated screening and manual verification, the system uses the final standard annotation data as the annotation result for the target annotation task and stores it in the annotation database. Simultaneously, the system can also structure and store various types of information related to the target annotation task, specifically including: the task identifier and requirement description of the target annotation task, the original annotation results output by the model, the corresponding inference logic chain, and the evaluation records and judgment criteria generated during the manual verification process. The above information is stored in association according to a preset data structure, forming high-quality samples that can be directly used for subsequent model training or fine-tuning.
[0039] By continuously accumulating the aforementioned structured data, the system can construct a high-quality, reusable supervised dataset. This provides a solid corpus foundation for the iterative optimization of large language models during subsequent annotation tasks, improving the model's understanding ability and output reliability in relevant annotation tasks.
[0040] The technical solution provided in this application involves constructing a prompt word template with a variable parameter region, embedding the requirement information of the target annotation task into the prompt word template to obtain target prompt words; inputting the target prompt words in parallel into multiple large language models, outputting multiple sets of candidate annotation results and the corresponding inference logic chains of the candidate annotation results; filtering out candidate annotation results that meet preset conditions and their corresponding inference logic chains from the multiple sets of candidate annotation results and their corresponding inference logic chains; processing the candidate annotation results and their corresponding inference logic chains using a preset evaluation method to obtain standard annotation data; and storing the standard annotation data as the annotation result of the target annotation task in the annotation database. This enables annotators to adjust the prompt words for specific tasks or text features in each interaction, ensuring that the annotation results are more accurate and meet actual needs. Multi-dimensional automated evaluation and manual verification ensure the stability and accuracy of the annotation data in multiple rounds of interaction, and ensure that the finally adopted annotation data has high reliability and task fit.
[0041] Figure 2 This is a flowchart illustrating another human-machine collaborative data annotation method provided in an embodiment of this application. Figure 2 The process shown includes the following steps: S201. Construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word.
[0042] The human-machine collaborative data annotation method provided in this application is applied to data annotation scenarios, specifically to human-machine collaborative data annotation systems. Through the co-evolution of model learning and human cognition within the system, high-quality and controllable data annotation is achieved, improving data annotation efficiency and building a solid, high-quality data foundation for subsequent model training.
[0043] In this embodiment, before the data annotation workflow starts, a prompt word template with a variable parameter area needs to be pre-built and injected into the system for operation. This prompt word template can be flexibly expanded according to the specific needs of the annotation task. When handling different types of annotation tasks, the variable parameters in the template can be modified or new parameters can be added according to the actual needs of the current annotation task, and the content of the prompt word template can be adjusted accordingly. Therefore, this template can adapt to various annotation scenarios without rebuilding the entire workflow, thereby improving the task adaptability and configuration efficiency of the data annotation system.
[0044] Specifically, based on the annotation requirements of the target annotation task, the system obtains the corresponding requirement information and embeds this information into the variable parameter area of the aforementioned prompt word template to generate target prompt words. That is, the system generates more specific target prompt words for a specific target annotation task using a configurable prompt word template. The system then uses these target prompt words to drive the subsequent annotation process and execute the target annotation task.
[0045] Through the above processing, the originally general prompt word template is converted into an executable specific instruction, ensuring that the prompt content input to the model is highly consistent with the actual needs of the target annotation task at both the semantic and task context levels. This avoids problems such as output content deviating from task requirements and semantic ambiguity caused by the use of generalized prompt words, thereby improving the accuracy of model response and task adaptability in the annotation process.
[0046] S202. Input the target prompt words into multiple large language models in parallel, and output multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
[0047] In this embodiment, the generated target prompt words are input in parallel to multiple large language models in the system. These multiple large language models can be various mainstream model architectures, such as DeepSeek, ChatGPT, and Qwen. Each model, upon receiving the target prompt word, performs natural language understanding and generation operations based on the prompt word's content and independently completes the annotation task indicated by the target prompt word. Each model outputs a set of candidate annotation results, along with a corresponding inference logic chain. This inference logic chain describes the intermediate reasoning process from understanding the prompt word to generating the annotation result, and may include a logical path for key information extraction, intermediate step derivation, and final conclusion generation.
[0048] By calling multiple large language models with different architectures in parallel, the system can obtain multiple sets of candidate annotation results from different model perspectives under the same annotation task, reducing the systematic bias that may exist in a single model. At the same time, it retains the reasoning logic chain corresponding to the candidate annotation results output by each model, providing diversified basis and traceable reasoning information for subsequent annotation result screening and quality evaluation.
[0049] S203. Perform automated quality assessment on multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results.
[0050] In this embodiment, after multiple large models output multiple sets of candidate annotation results, the system enters an automated quality assessment stage to filter out low-quality model outputs. Specifically, the candidate results output by multiple models cannot be used as final annotation data. The system needs to perform an automated quality assessment on the multiple sets of candidate annotation results and the corresponding inference logic chains according to preset conditions. This assessment stage performs a preliminary screening of each set of outputs, identifies and retains candidate annotation results that meet the preset conditions, as well as the corresponding inference logic chains.
[0051] S204. Filter and retain candidate annotation results that meet the preset conditions, as well as the corresponding reasoning logic chain.
[0052] In this embodiment, through automated quality assessment, the system filters and retains candidate annotation results that meet preset conditions, along with their corresponding inference logic chains. These preset conditions include a perplexity value less than a first threshold and a semantic similarity value greater than a second threshold for the candidate annotation results.
[0053] Specifically, the system automatically calculates two evaluation metrics: perplexity and semantic similarity. The perplexity value measures the fluency and certainty of the generated text; a lower perplexity value indicates more natural text and a more reasonable language structure. The system effectively eliminates low-quality outputs with confusing language, unclear logic, or deviation from the task's theme by filtering candidate annotations with perplexity values greater than or equal to a first threshold. Simultaneously, the system uses semantic similarity to measure the model's language modeling ability for text generation; a higher semantic similarity value indicates greater semantic consistency between generated texts, meaning the output is closer to the expected semantic distribution. The system effectively eliminates low-quality outputs with significant deviations from the expected semantics by filtering candidate annotations with semantic similarity values less than a second threshold.
[0054] When the above candidate annotation results simultaneously satisfy the condition that the perplexity value is less than the first threshold and the semantic similarity value is greater than the second threshold, it indicates that the output result meets the set quality requirements in terms of both text fluency and semantic consistency, and can proceed to the next stage of manual verification.
[0055] In an optional embodiment of this application, when there are candidate annotation results that do not meet the preset conditions, a regeneration mechanism is triggered to re-execute the steps of inputting the target prompt word into multiple large language models in parallel and outputting multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
[0056] Specifically, for candidate annotations and their corresponding inference chains that do not meet the preset conditions, the system will automatically discard these unqualified outputs to prevent them from entering the subsequent manual verification stage, thereby reducing unnecessary manual processing time. Simultaneously, the system triggers a regeneration mechanism, re-inputting the aforementioned target prompts into multiple large language models and requesting each model to generate new candidate annotations. Through this regeneration mechanism, the system can effectively avoid semantic misjudgments, logical jumps, or grammatical inconsistencies that may occur in a single random output by the model, improving the overall stability of the output. This automated quality assessment and regeneration process is repeated until candidate annotations that meet the preset conditions are obtained, or the preset maximum number of retries is reached. Thus, the system ensures that candidate annotations entering the manual verification stage have a basically acceptable quality threshold, thereby improving the efficiency and effectiveness of subsequent manual evaluation.
[0057] S205. Based on the preset annotation requirements, manually verify the candidate annotation results and the corresponding inference logic chain.
[0058] In this embodiment, after the automated quality assessment process described above, candidate annotation results that meet the preset conditions, along with their corresponding inference logic chains, will enter the manual verification stage. In this stage, annotators, based on the preset annotation requirements and standards of the target annotation task, will judge the annotation results output by the model and the corresponding inference logic chains. Annotators need to determine whether the candidate annotation results truly meet the preset annotation requirements and standards, and, based on their corresponding inference logic chains, whether the generation process of the results is reasonable and logically coherent.
[0059] S206. Select all candidate annotation results that meet the preset annotation requirements, and the corresponding reasoning logic chain as standard annotation data.
[0060] In this embodiment, when all content of the candidate annotation result meets the preset annotation requirements, the system uses the candidate annotation result and its corresponding inference logic chain as standard annotation data. Specifically, after careful verification, the annotation personnel confirm that the candidate annotation result and its corresponding inference logic chain meet the preset annotation requirements of the target annotation task in terms of content accuracy, format standardization, and task relevance, and no further modification or screening is required. Therefore, the candidate annotation result and its corresponding inference logic chain are used as standard annotation data.
[0061] S207. Select candidate annotation results that meet the preset annotation requirements, update the candidate annotation results, and use the updated candidate annotation results and the corresponding inference logic chain as standard annotation data.
[0062] In this embodiment, when a portion of the candidate annotation result meets preset annotation requirements—that is, after verification by the annotator according to the preset requirements—it is confirmed that the candidate annotation result has certain problems, such as unclear reasoning logic, conclusions inconsistent with semantics, or ignoring specific context. To address these problems, the annotator modifies and updates the problematic or unreasonable content, using the updated candidate annotation result and its corresponding reasoning logic chain as standard annotation data. Manual verification is used to precisely correct content that does not meet the requirements, thereby preserving valid information and improving the overall quality of the annotation data.
[0063] S208. When all the contents of the candidate annotation results do not meet the preset annotation requirements, receive the input standard annotation results and use the standard annotation results as standard annotation data.
[0064] In this embodiment, when the annotator verifies the candidate annotation result according to preset annotation requirements and confirms that none of the content meets the preset requirements, the annotator can directly input the standard annotation result that meets the preset requirements. The system receives the standard annotation result input by the annotator and uses it as the final standard annotation data. By having the annotator input directly, the system can complete the annotation task based on the annotator's professional knowledge even when the model output is completely unavailable, ensuring that the target annotation task can still obtain standard data that meets the requirements and avoiding task interruption due to low model output quality.
[0065] S209. Use the standard annotation data as the annotation result of the target annotation task and store it in the annotation database.
[0066] In this embodiment, after both automated screening and manual evaluation, the system uses the final standardized annotation data as the annotation result for the target annotation task and stores it in the annotation database. Simultaneously, the system can also structurally store various types of information related to the target annotation task, specifically including: the task identifier and requirement description of the target annotation task, the original annotation results output by the model, the corresponding inference logic chain, and the evaluation records and judgment criteria generated during the manual evaluation process. The above information is stored in association according to a preset data structure, forming high-quality samples that can be directly used for subsequent model training or fine-tuning.
[0067] S210 When some or all of the candidate annotation results meet the preset annotation requirements, extract non-standard information from the reasoning logic chain corresponding to the candidate annotation results; update the prompt word template based on the non-standard information.
[0068] In this embodiment, when the annotator verifies the candidate annotation result according to the preset annotation requirements and confirms that some or all of the content of the candidate annotation result meets the preset annotation requirements and standards, it indicates that the candidate annotation result has certain problems, such as unclear reasoning logic, conclusions that do not match the semantics, or ignoring specific context. At this time, the annotator can extract non-standard information from the reasoning logic chain corresponding to the candidate annotation result. This non-standard information may be the cause of the above-mentioned problems in the candidate annotation result, such as the model's deviation in understanding the task, omission of contextual information, or misinterpretation of specific constraints.
[0069] The annotators input correction instructions based on the non-standard information and embed these instructions into the prompt word template, thus updating the prompt word template. Through this feedback and correction process, the blind spots that originally caused model errors are transformed into explicit constraints in the prompt words, enabling the optimized prompt words to guide the model to avoid similar errors in subsequent generation.
[0070] Figure 2 The illustrated process provides a human-machine collaborative data annotation method. By constructing prompt word templates with variable parameter regions, the system can quickly adapt to feedback from the human verification stage. This allows the updated prompt words to guide the model to focus more on the user's concerns in the next generation task, optimizing its understanding path and reasoning strategy. This process not only achieves high-quality, controllable data annotation but also promotes the co-evolution of model learning and human cognition in the human-machine collaborative data annotation system. On the one hand, the model continuously corrects its reasoning logic and output expression with the help of human feedback, gradually improving its performance on specific annotation tasks. On the other hand, users, relying on the model's efficient generation capabilities, significantly reduce the workload of manual annotation and improve overall annotation efficiency.
[0071] Figure 3 This is a schematic diagram of the structure of a human-machine collaborative data annotation device provided in an embodiment of this application. Figure 3 As shown, the device includes: The acquisition module 301 is used to construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word; Output module 302 is used to input the target prompt words into multiple large language models in parallel and output multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results; The filtering module 303 is used to filter out candidate annotation results that meet preset conditions and corresponding inference logic chains from multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; Evaluation module 304 is used to process the candidate annotation results and the corresponding inference logic chain using a preset evaluation method to obtain standard annotation data; The storage module 305 is used to store the standard annotation data as the annotation result of the target annotation task in the annotation database.
[0072] In one possible implementation, the evaluation module 304 is specifically used to: manually verify the candidate annotation results and the corresponding inference logic chain according to preset annotation requirements; and select the candidate annotation results and the corresponding inference logic chain that all content meets the preset annotation requirements as standard annotation data.
[0073] In one possible implementation, the evaluation module 304 is specifically used to: manually verify the candidate annotation results and the corresponding inference logic chain according to preset annotation requirements; select candidate annotation results whose content meets the preset annotation requirements, update the candidate annotation results, and use the updated candidate annotation results and the corresponding inference logic chain as standard annotation data.
[0074] In one possible implementation, the evaluation module 304 is specifically used to: manually verify the candidate annotation results and the corresponding inference logic chain according to preset annotation requirements; when all contents of the candidate annotation results do not meet the preset annotation requirements, receive the input standard annotation results and use the standard annotation results as standard annotation data.
[0075] In one possible implementation, the device further includes: The update module 306 is specifically used to: extract non-standard information from the inference logic chain corresponding to the candidate annotation result when some or all of the candidate annotation result meets or does not meet the preset annotation requirements; and update the prompt word template according to the non-standard information.
[0076] In one possible implementation, the filtering module 303 is specifically used for: automatically evaluating the quality of multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; filtering and retaining candidate annotation results that meet preset conditions, as well as the corresponding inference logic chains, wherein the preset conditions include the perplexity value of the candidate annotation results being less than a first threshold and the semantic similarity value being greater than a second threshold.
[0077] In one possible implementation, the device further includes: The regeneration module 307 is specifically used to: trigger the regeneration mechanism when there are candidate annotation results that do not meet the preset conditions, and re-execute the steps of inputting the target prompt word into multiple large language models in parallel and outputting multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
[0078] like Figure 4 As shown in the figure, this application provides a device including a processor 411, a communication interface 412, a memory 413, and a communication bus 414, wherein the processor 411, the communication interface 412, and the memory 413 communicate with each other through the communication bus 414. Memory 413 is used to store computer programs; In one embodiment of this application, when the processor 411 executes the program stored in the memory 413, it implements the human-machine collaborative data annotation method provided in any of the foregoing method embodiments, including: Construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word; The target prompt words are input into multiple large language models in parallel, and multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results are output. Candidate annotation results that meet preset conditions and corresponding inference logic chains are selected from multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results. The candidate annotation results and the corresponding inference logic chain are processed using a preset evaluation method to obtain standard annotation data; The standard annotation data is used as the annotation result of the target annotation task and stored in the annotation database.
[0079] In one possible implementation, the candidate annotation results and the corresponding inference logic chain are manually verified according to preset annotation requirements; the candidate annotation results and the corresponding inference logic chain that all content meets the preset annotation requirements are selected as standard annotation data.
[0080] In one possible implementation, the candidate annotation results and the corresponding inference logic chain are manually verified according to preset annotation requirements; candidate annotation results that partially meet the preset annotation requirements are selected, and the candidate annotation results are updated. The updated candidate annotation results and the corresponding inference logic chain are used as standard annotation data.
[0081] In one possible implementation, the candidate annotation results and the corresponding inference logic chain are manually verified according to preset annotation requirements; when all contents of the candidate annotation results do not meet the preset annotation requirements, the input standard annotation results are received and the standard annotation results are used as standard annotation data.
[0082] In one possible implementation, when some or all of the candidate annotation results meet the preset annotation requirements, non-standard information is extracted from the reasoning logic chain corresponding to the candidate annotation results; and the prompt word template is updated based on the non-standard information.
[0083] In one possible implementation, an automated quality assessment is performed on multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; candidate annotation results that meet preset conditions and the corresponding inference logic chains are selected and retained, wherein the preset conditions include that the perplexity value of the candidate annotation results is less than a first threshold and the semantic similarity value is greater than a second threshold.
[0084] In one possible implementation, when there are candidate annotation results that do not meet the preset conditions, a regeneration mechanism is triggered to re-execute the steps of inputting the target prompt word into multiple large language models in parallel and outputting multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
[0085] This application embodiment also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; it may also include combinations of the above types of memory. When one or more programs in the storage medium can be executed by one or more processors, the aforementioned human-machine collaborative data annotation method executed on the human-machine collaborative data annotation device side is implemented. The processor is used to execute the human-machine collaborative data annotation program stored in the memory to implement the following steps of executing the human-machine collaborative data annotation method on the human-machine collaborative data annotation device side: A prompt word template with a variable parameter region is constructed, and the requirement information of the target annotation task is embedded into the prompt word template to obtain the target prompt word. The target prompt word is input into multiple large language models in parallel, and multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results are output. Candidate annotation results that meet preset conditions and their corresponding inference logic chains are selected from the multiple sets of candidate annotation results and their corresponding inference logic chains. The candidate annotation results and their corresponding inference logic chains are processed using a preset evaluation method to obtain standard annotation data. The standard annotation data is used as the annotation result of the target annotation task and stored in the annotation database.
[0086] In one possible implementation, the candidate annotation results and the corresponding inference logic chain are manually verified according to preset annotation requirements; the candidate annotation results and the corresponding inference logic chain that all content meets the preset annotation requirements are selected as standard annotation data.
[0087] In one possible implementation, the candidate annotation results and the corresponding inference logic chain are manually verified according to preset annotation requirements; candidate annotation results that partially meet the preset annotation requirements are selected, and the candidate annotation results are updated. The updated candidate annotation results and the corresponding inference logic chain are used as standard annotation data.
[0088] In one possible implementation, the candidate annotation results and the corresponding inference logic chain are manually verified according to preset annotation requirements; when all contents of the candidate annotation results do not meet the preset annotation requirements, the input standard annotation results are received and the standard annotation results are used as standard annotation data.
[0089] In one possible implementation, when some or all of the candidate annotation results meet the preset annotation requirements, non-standard information is extracted from the reasoning logic chain corresponding to the candidate annotation results; and the prompt word template is updated based on the non-standard information.
[0090] In one possible implementation, an automated quality assessment is performed on multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; candidate annotation results that meet preset conditions and the corresponding inference logic chains are selected and retained, wherein the preset conditions include that the perplexity value of the candidate annotation results is less than a first threshold and the semantic similarity value is greater than a second threshold.
[0091] In one possible implementation, when there are candidate annotation results that do not meet the preset conditions, a regeneration mechanism is triggered to re-execute the steps of inputting the target prompt word into multiple large language models in parallel and outputting multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0093] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0094] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0095] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A human-computer collaborative data annotation method, characterized in that, The method includes: Construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word; The target prompt words are input into multiple large language models in parallel, and multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results are output. Candidate annotation results that meet preset conditions and corresponding inference logic chains are selected from multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results. The candidate annotation results and the corresponding inference logic chain are processed using a preset evaluation method to obtain standard annotation data; The standard annotation data is used as the annotation result of the target annotation task and stored in the annotation database.
2. The method according to claim 1, characterized in that, The process of using a preset evaluation method to process the candidate annotation results and the corresponding inference logic chain to obtain standard annotation data includes: The candidate annotation results and the corresponding inference logic chain are manually verified according to the preset annotation requirements. Candidate annotation results that meet all the preset annotation requirements, along with their corresponding inference logic chains, are selected as standard annotation data.
3. The method according to claim 1, characterized in that, The process of using a preset evaluation method to process the candidate annotation results and the corresponding inference logic chain to obtain standard annotation data includes: The candidate annotation results and the corresponding inference logic chain are manually verified according to the preset annotation requirements. Select candidate annotation results that meet the preset annotation requirements, update the candidate annotation results, and use the updated candidate annotation results and the corresponding inference logic chain as standard annotation data.
4. The method according to claim 1, characterized in that, The process of using a preset evaluation method to process the candidate annotation results and the corresponding inference logic chain to obtain standard annotation data includes: The candidate annotation results and the corresponding inference logic chain are manually verified according to the preset annotation requirements. When all the candidate annotation results do not meet the preset annotation requirements, the input standard annotation results are received and used as standard annotation data.
5. The method according to claim 3 or 4, characterized in that, The method further includes: When some or all of the candidate annotation results meet the preset annotation requirements, non-standard information is extracted from the inference logic chain corresponding to the candidate annotation results. The prompt word template is updated based on the non-standard information.
6. The method according to claim 1, characterized in that, The step of filtering candidate annotation results that meet preset conditions from multiple sets of candidate annotation results and the corresponding inference logic chains of the candidate annotation results, and the corresponding inference logic chains, includes: Automated quality assessment is performed on multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; Candidate annotation results that meet preset conditions and their corresponding reasoning logic chains are selected and retained. The preset conditions include that the perplexity value of the candidate annotation results is less than a first threshold and the semantic similarity value is greater than a second threshold.
7. The method according to claim 6, characterized in that, The method further includes: When there are candidate annotation results that do not meet the preset conditions, a regeneration mechanism is triggered to re-execute the steps of inputting the target prompt word into multiple large language models in parallel and outputting multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results.
8. A human-machine collaborative data annotation device, characterized in that, The device includes: The acquisition module is used to construct a prompt word template with a variable parameter region, embed the requirement information of the target annotation task into the prompt word template, and obtain the target prompt word; The output module is used to input the target prompt words into multiple large language models in parallel and output multiple sets of candidate annotation results and the inference logic chain corresponding to the candidate annotation results; The filtering module is used to filter out candidate annotation results that meet preset conditions and corresponding inference logic chains from multiple sets of candidate annotation results and the inference logic chains corresponding to the candidate annotation results; The evaluation module is used to process the candidate annotation results and the corresponding inference logic chain using a preset evaluation method to obtain standard annotation data; The storage module is used to store the standard annotation data as the annotation result of the target annotation task in the annotation database.
9. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute a human-computer collaborative data annotation program stored in the memory to implement the human-computer collaborative data annotation method according to any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the human-machine collaborative data annotation method according to any one of claims 1-7.