Data labeling method and device

By converting the task flowchart created by users into original annotation tasks and using a large language model to generate sub-task execution results, the problem of low data annotation efficiency in the existing technology is solved, and an efficient and accurate data annotation process is achieved.

CN120196415APending Publication Date: 2025-06-24ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510301277.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

It is difficult to efficiently mark data in the prior art, especially when the demand for model usage is increasing, how to better mark data has become an important issue.

Method used

Obtain the annotation task by converting the user-created task flowchart into the original annotation task and performing task filling in the original annotation task. Analyze the annotated task to obtain the execution relationship between multiple task sequences and their subtasks. The data to be marked is used as task execution input, and subtasks within each task sequence are executed in parallel according to the execution relationship. The subtask execution results are generated using a large language model, and the target execution results are determined as data labels.

Benefits of technology

An efficient data labeling process is realized, the accuracy and efficiency of data labeling are improved, and the stability and robustness of the system are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196415A_ABST
    Figure CN120196415A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data annotation method and device.The data annotation method comprises the steps that in the data annotation process, a task flow chart created by a user is converted into an original annotation task, task filling is conducted in the original annotation task, and an annotation task is obtained; the method comprises the following steps: converting a user-defined task flow chart into an executable task, analyzing a labeling task to obtain a plurality of task sequences and an execution relationship among sub-tasks in the task sequences, and taking to-be-labeled data submitted by a user as task execution input; and task execution of the sub-tasks in each task sequence is performed in parallel according to the execution relationship to obtain a plurality of task execution results, a target execution result is determined in each task execution result and is used as a label of the to-be-labeled data, and data labeling is performed in a manner of creating a task flow chart in a user-defined manner to obtain the label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of data processing technologies, and in particular, to a data annotation method and apparatus. Background Art

[0002] With the continuous development of Internet technologies and the continuous improvement of information technologies, more and more data needs to be processed. In order to process data efficiently, various models have begun to be widely used in data processing. Users can use expert models to process specific tasks, or use large language models to process various tasks, or use hybrid expert models to process various tasks. As the demand for model usage continues to increase, the demand for data annotation of various models is also increasing. How to better perform data annotation has become an issue that all parties need to pay attention to. Summary of the Invention

[0003] One or more embodiments of this specification provide a data annotation method, including: converting a task flow chart created by a user into an original annotation task, and performing task filling in the original annotation task to obtain an annotation task. Analyzing the annotation task to obtain a plurality of task sequences and the execution relationships between the subtasks within the task sequences. Using the data to be annotated submitted by the user as a task execution input, and performing the task execution of the subtasks within each task sequence in parallel according to the execution relationships to obtain a plurality of task execution results. The task execution includes invoking the large language model corresponding to the task sequence to generate the execution result of the subtask. Determining a target execution result among the task execution results, and using the target execution result as the label of the data to be annotated.

[0004] One or more embodiments of this specification provide a data annotation apparatus, including: an annotation task generation module configured to convert a task flow chart created by a user into an original annotation task, and perform task filling in the original annotation task to obtain an annotation task. An annotation task analysis module configured to analyze the annotation task to obtain a plurality of task sequences and the execution relationships between the subtasks within the task sequences. A task execution module configured to use the data to be annotated submitted by the user as a task execution input, and perform the task execution of the subtasks within each task sequence in parallel according to the execution relationships to obtain a plurality of task execution results. The task execution includes invoking the large language model corresponding to the task sequence to generate the execution result of the subtask. A label obtaining module configured to determine a target execution result among the task execution results, and use the target execution result as the label of the data to be annotated.

[0005] One or more embodiments of this specification provide a data annotation device, including: a processor; and a memory configured to store computer-executable instructions, the processor executing the computer-executable instructions to implement the following process: converting a task flow chart created by a user into an original annotation task, and performing task filling in the original annotation task to obtain an annotation task. Parsing the annotation task to obtain a plurality of task sequences and the execution relationships between the subtasks within the task sequences. Using the data to be annotated submitted by the user as the task execution input, and performing the task execution of the subtasks within each task sequence in parallel according to the execution relationships to obtain a plurality of task execution results. The task execution includes calling the large language model corresponding to the task sequence to generate the execution result of the subtask. Determining a target execution result among the task execution results, and using the target execution result as the label of the data to be annotated.

[0006] One or more embodiments of this specification provide a computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions implementing the following process when executed: converting a task flow chart created by a user into an original annotation task, and performing task filling in the original annotation task to obtain an annotation task. Parsing the annotation task to obtain a plurality of task sequences and the execution relationships between the subtasks within the task sequences. Using the data to be annotated submitted by the user as the task execution input, and performing the task execution of the subtasks within each task sequence in parallel according to the execution relationships to obtain a plurality of task execution results. The task execution includes calling the large language model corresponding to the task sequence to generate the execution result of the subtask. Determining a target execution result among the task execution results, and using the target execution result as the label of the data to be annotated. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Figure 1 It is a schematic diagram of the implementation environment of a data annotation method provided by one or more embodiments of this specification; Figure 2 It is a processing flow chart of a data annotation method provided by one or more embodiments of this specification; Figure 3 It is a schematic diagram of a task flow chart provided by one or more embodiments of this specification; Figure 4A processing flow chart of a data annotation method applied to a data annotation scenario provided for one or more embodiments of this specification; Figure 5 A schematic diagram of an embodiment of a data annotation device provided for one or more embodiments of this specification; Figure 6 A schematic structural diagram of a data annotation device provided for one or more embodiments of this specification. Detailed implementation manners

[0008] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification with reference to the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.

[0009] The data annotation method provided by one or more embodiments of this embodiment is applicable to the implementation scenario of data annotation. Referring to Figure 1 , this implementation environment at least includes: a server 101. The server 101 can be one or more servers, a server cluster composed of several servers, or a cloud server of a cloud computing platform. The server 101 is used to perform the conversion of annotation tasks and execute the annotation tasks to obtain the labels of the data to be annotated. In addition, the server 101 can also cooperate with the user terminal 102 to create a task flow chart and receive the data to be annotated submitted by the user terminal 102.

[0010] This implementation environment may further include a graph execution engine 101-1. The graph execution engine 101-1 can be deployed inside the server 101 or independently deployed. The graph execution engine 101-1 can perform the conversion of annotation tasks and execute the annotation tasks to obtain the labels of the data to be annotated. It should be noted that the graph execution engine 101-1 can also cooperate with the server 101 to perform the above operations, that is: the server 101 can transmit the task flow chart and / or the data to be annotated to the graph execution engine 101-1, and the graph execution engine 101-1 performs the conversion of annotation tasks and executes the annotation tasks to obtain the labels of the data to be annotated.

[0011] The implementation environment may further include a user terminal 102, which may be a personal computer, a laptop computer, a smart phone, a tablet computer, an e-book reader, a wearable device, a device for information interaction based on AR (Augmented Reality) / VR (Virtual Reality), etc. The user terminal 102 may interact with the server 101 to perform the creation process and / or editing process of the task flow chart, and the user terminal 102 may also send the data to be annotated to the server 101. In addition, the user terminal 102 may receive the label of the data to be annotated sent by the server 101.

[0012] In this implementation environment, the user can interact with the server 101 through the user terminal 102 to create a task flow chart. The server 101 and the graph execution engine 101-1 cooperate to convert the task flow chart into an original annotation task. The graph execution engine 101-1 fills tasks in the original annotation task to obtain an annotation task, and parses the annotation task to obtain multiple task sequences and the execution relationship between the subtasks within each task sequence. The server 101 transmits the data to be annotated submitted by the user terminal 102 to the graph execution engine 101-1. The graph execution engine 101-1 uses the data to be annotated as the task execution input, and performs the task execution of the subtasks within each task sequence in parallel according to the execution relationship to obtain multiple task execution results, and determines the label of the data to be annotated in each task execution result. In this way, data annotation is obtained through the user-defined creation of a task flow chart.

[0013] One or more embodiments of the data annotation method provided in this embodiment are as follows: Refer to Figure 2 , the data annotation method provided in this embodiment specifically includes steps S202 to S208.

[0014] Step S202, convert the task flow chart created by the user into an original annotation task, and fill tasks in the original annotation task to obtain an annotation task.

[0015] In this embodiment, the task flow chart may be a visual flow chart created and configured by the user to describe the data annotation task flow. The user can create and configure the task flow chart in the annotation service. The task flow chart may include multiple task nodes, and the association relationship may be configured between each task node to represent the process sequence between each task node. The task flow chart may include multiple task nodes and / or the association relationship between each task node. Among them, the user includes an operating user and / or a management user.

[0016] A task node can be a node created and configured by a user to describe a specific task in a data annotation task flow. The user can create and configure task nodes in the annotation service. The user can configure multiple task nodes for selection when creating a task flow chart. Task nodes can represent different tasks through node configuration information. A task node can represent a large language model call task, or a annotation result query task, or an application programming interface call task, or a key information extraction task, or a function editing task, or a conditional judgment task. Optionally, a task node is created based on a node creation request submitted by the user; the task node is configured according to the node configuration information submitted by the user.

[0017] Specifically, during the node configuration process, the user can trigger the node creation interface on the service page of the annotation service to submit a node creation request. After the node creation is completed, the user can input node configuration information in the node configuration window and / or the node configuration page to configure the node. Multiple identical task nodes can be configured in the task flow chart; among them, the node configuration information includes operation configuration data and / or task description data. Optionally, the node creation request is generated after the node creation interface configured on the service page of the annotation service is triggered.

[0018] For example, when the user configures a task node representing a large language model call task, the node configuration information may include the data input interface of the target large language model, the prompt words, and / or the configuration data corresponding to the call operation. Optionally, the node configuration window is configured on the service page of the annotation service.

[0019] Specifically, during the creation process of the task flow chart, the task flow chart can be created based on the instructions submitted by the user. The user can submit a node selection instruction to select a task node, and the user can also submit a node configuration instruction to configure the association relationship between task nodes, that is: the task flow chart is constructed based on task nodes and the association relationship between task nodes. By submitting instructions by the user, the user-defined construction of the annotation process of data annotation is realized. The conversion process of the task flow chart can be performed after the task flow chart is created; in an optional implementation provided in this embodiment, the task flow chart is constructed in the following manner: Select multiple task nodes according to the node selection instruction submitted by the user; Configure the association relationship between each task node according to the node configuration instruction submitted by the user to obtain the task flow chart.

[0020] Specifically, after receiving a task flow chart creation request, the corresponding task nodes can be selected according to the node selection instructions submitted by the user through the selection control of the task nodes on the configuration page of the task flow chart, and the association relationships between the task nodes can be configured according to the node configuration instructions submitted by the user through the connection control of the task nodes on the configuration page. Based on the selected task nodes and the association relationships between the task nodes, a task flow chart is constructed. Optionally, the task flow chart is created in the annotation service.

[0021] For example, as Figure 3 shown, the task flow chart may include multiple task nodes, and there are association relationships between the task nodes to represent the execution order. A task node can establish association relationships with two or more task nodes; among them, node a is the start node or input node, and node h is the end node or output node.

[0022] In addition, a configuration page for the task flow chart can also be provided to the user, enabling the user to create the task flow chart conveniently and intuitively through page interaction, improving the user experience. In another optional implementation provided in this embodiment, the task flow chart is constructed in the following manner: Obtain the node selection instructions submitted by the user by triggering the selection control of the task nodes on the configuration page of the task flow chart; Respond to the node selection instructions to select multiple task nodes, and obtain the node configuration instructions submitted by the user by triggering the association control of the task nodes on the configuration page; Respond to the node configuration instructions to configure the association relationships between the task nodes, and construct a task flow chart according to the multiple task nodes and the association relationships between the task nodes.

[0023] Optionally, the node selection instructions are generated after the selection control of the task nodes configured on the configuration page of the task flow chart is triggered; the node configuration instructions are generated after the association control of the task nodes configured on the configuration page is triggered.

[0024] During specific implementation, after the task flow chart is created, the task flow chart is converted into an original annotation task according to the multiple task nodes configured in the task flow chart created by the user and the association relationships between the task nodes, and the task is filled in the original annotation task to obtain an annotation task. Specifically, a subtask configured to have the same output data as the input data can be filled in the original annotation task to obtain an annotation task.

[0025] In this embodiment, the original annotation task is an executable task converted from the task flow chart created by the user to represent the data annotation task flow; the annotation task is the executable task of the task filled in the original annotation task, which also represents the data annotation task flow. Multiple task sequences may be included in the original annotation task and / or the annotation task. For any task sequence, the task sequence includes multiple subtasks and the execution relationship and / or execution order between the subtasks. Optionally, the annotation task is constructed based on multiple subtasks and / or multiple operators and the execution relationship between the subtasks and / or the operators.

[0026] Among them, the task sequence may be an operator group, specifically a serial operator group, and the subtask may be an operator. Optionally, the task sequence includes an operator group, the subtask includes an operator, and the execution relationship includes the hierarchy of the operators within the operator group.

[0027] Specifically, during the conversion of the task flow chart, multiple task nodes in the task flow chart can be converted into multiple subtasks, the association relationship between the task nodes in the task flow chart can be converted into the execution relationship between the subtasks, and the original annotation task can be constructed based on the subtasks and the execution relationship between the subtasks. By converting the task flow chart into an executable task, it assists the operating user in data annotation and improves the user experience; in an optional implementation manner provided in this embodiment, converting the task flow chart into the original annotation task includes: Extract multiple task nodes in the task flow chart and the association relationship between the task nodes; Query the node configuration information corresponding to each task node, generate multiple subtasks according to the node configuration information, and determine the execution relationship between the subtasks according to the association relationship; Construct the original annotation task based on multiple subtasks and the determined execution relationship.

[0028] Specifically, analyze the task flow chart to obtain multiple task nodes in the task flow chart and the association relationship between the task nodes, query the node configuration information corresponding to each task node, generate subtasks based on the node configuration information, generate the execution relationship between the subtasks based on the association relationship between the task nodes, and construct the original annotation task based on the subtasks and the execution relationship between the subtasks.

[0029] Continuing with the above example, Figure 3After the shown task flow chart is converted into the original annotation task, the original annotation task includes eight operators a - h, where a is the starting operator and h is the ending operator. The original annotation task contains three task sequences and / or operator groups, namely (b→c), (e→f→g), and (d) respectively. The subtasks and / or operators in the three task sequences and / or operator groups are executed concurrently. Then the original annotation task can be expressed as {a→[(b→c), (e→f→g), (d)]→h}.

[0030] In this embodiment, the subtasks in the original annotation task and / or the subtasks of the annotation task can correspond to the task nodes in the task flow chart created by the user. Optionally, the subtasks include: large language model call task, annotation result query task, application programming interface call task, key information extraction task, function editing task, and / or conditional judgment task.

[0031] In practical applications, during the concurrent execution of the subtasks within each task sequence of the original annotation task, there may be a situation where improper concurrent control leads to inaccurate annotation results. For example: Since the execution times of each subtask are different, the input data of the subtasks with multiple input data is incomplete, affecting the input data and / or output data of the subsequent subtasks. Another example is that there are deadlock task nodes in the task flow chart created by the user. After being converted into the original annotation task, a deadlock problem will be encountered, affecting the task execution. In response to this, the subtasks of each task sequence in the original annotation task can be filled to the same level, so that the number of subtasks in each task sequence is the same, and then the subtasks at each level can be executed concurrently. It is also possible to fill empty tasks in the original annotation task to obtain the annotation task, improving the fault tolerance rate and enhancing the system robustness. In an optional implementation manner provided in this embodiment, obtaining the annotation task by filling tasks in the original annotation task includes: Analyze the original annotation task to obtain multiple original task sequences and the execution relationships between the subtasks within each original task sequence, and determine the target task sequence in each original task sequence according to the number of subtasks within each original task sequence; Determine the missing task levels of each original task sequence relative to the target task sequence according to the execution relationship; Fill empty tasks in the missing task levels of each original task sequence to obtain the annotation task.

[0032] Among them, the empty task includes a subtask with the same output data and input data.

[0033] Specifically, parsing the original annotation task obtains multiple original task sequences and / or operator groups, and obtains the execution relationships of each subtask within the original task sequences and / or operator groups. Determine the target task sequence as the original task sequence with the largest number of subtasks within the original task sequences and / or operator groups. Traverse the remaining task sequences according to the execution relationships to determine the missing task levels relative to the target task sequence. Determine the number of empty tasks to be filled in the remaining task sequences according to the missing task levels, and fill in the empty tasks at the missing task levels. Among them, the task levels of each subtask in the task sequence can represent the execution relationships and / or execution orders of each subtask; optionally, the execution relationships are determined according to the task levels of each subtask.

[0034] Continuing with the above original annotation task as an example, the original annotation task can be expressed as {a→[(b→c), (e→f→g), (d)]→h}. Count the number of subtasks in each task sequence, determine the target task sequence as (a→e→f→g→h), determine that the missing task level of (a→b→c→h) is the fourth level, and determine that the missing levels of (a→d→h) are the third and fourth levels. Then the annotated task after filling in the empty nodes can be expressed as {a→[(b→c→empty), (e→f→g), (d→empty→empty)]→h}; among them, each subtask is executed in sequence according to the task levels.

[0035] Specifically, during the task filling process, there may be subtasks with multiple input data and / or multi-input operators in the original annotation task, and the multi-input operators may exist in multiple task sequences at the same time; in response to this, the task level of the multi-input operator in the target task sequence can be used as the reference level and / or target level, and calculate the missing task levels of each original task sequence according to the execution levels of the multi-input operator in the remaining task sequences, that is: determine the filling level of the empty task according to the task level of the subtask with multiple input data in the original annotation task, and fill in the empty task at the filling level of the original annotation task to obtain the annotated task, which improves the accuracy of the task filling operation; in an optional implementation manner provided in this embodiment, obtaining the annotated task by filling tasks in the original annotation task includes: Parse the original annotation task to obtain multiple original task sequences and the execution relationships between each subtask within the original task sequences, and determine the target task sequence in each original task sequence according to the number of subtasks in each original task sequence; Filter out the target subtasks with multiple input data in the target task sequence according to the execution relationships, and determine the target level where the target subtasks are located in the target task sequence; Calculate the missing task levels of each original task sequence according to the execution levels and the target level of the target subtasks in each task sequence; Fill in the empty tasks at the missing task levels of each original task sequence to obtain the annotated task.

[0036] For example, if the original annotation task is expressed as {a→[(b→c→d→e→f),(g→d→h)]→i}, then the target task sequence is (a→b→c→d→e→f→i). Among them, there are target subtasks of multiple input data and / or multi-input operators d and i. d and i are at the fourth level and the seventh level respectively in the target task sequence. Then, it can be determined that the missing task levels in the task sequence (a→g→d→h→i) are the third level and the sixth level. Then, the task sequence after filling the empty tasks is (a→g→empty→d→h→empty→i), and the annotated task after filling can be expressed as {a→[(b→c→d→e→f),(g→empty→d→h→empty)]→i}.

[0037] In addition, the filling level of the empty task can also be determined according to the missing task level and the processing time of the remaining tasks at the missing task level, that is: determine the filling level of the empty task according to the execution time of the subtasks at each level in the original annotation task, and fill the empty task in the filling level to obtain the annotation task. By configuring the tasks with longer processing times to be executed in parallel at the same level, the execution efficiency of the annotation task is improved; in another alternative implementation provided in this embodiment, obtaining the annotation task by filling tasks in the original annotation task includes: parsing the original annotation task to obtain multiple original task sequences and the execution relationship between the subtasks within the original task sequences, determining the target task sequence in each original task sequence according to the number of subtasks within each original task sequence; determining the missing task levels of each original task sequence relative to the target task sequence according to the execution relationship, and screening out the subtasks corresponding to the missing task levels in the remaining original task sequences; querying and / or predicting the processing time of the subtasks; if the processing time is greater than the threshold, migrating the subtasks in this original task sequence to this level and filling the empty task at the original level of this subtask; if the processing time is less than the threshold, no processing is required.

[0038] For example, if the original annotation task is expressed as {a→[(b→c),(e→f→g),(d)]→h}, and it is detected that the processing times of tasks b, f, and d are relatively long, then migrate tasks b and d to the same level as task f and fill the empty tasks, and the obtained annotation task is expressed as {a→[(empty→b→c),(e→f→g),(empty→d→empty)]→h}.

[0039] In addition, after the task filling operation is completed in the original annotation task, multiple task sequences of the annotation task and the execution relationships between subtasks within each task sequence can be directly obtained without performing the operation of parsing the annotation task. Specifically, after converting the task flow chart into the original annotation task, the original annotation task can be parsed to obtain multiple original task sequences, and empty tasks can be filled in the original task sequences to obtain multiple task sequences of the annotation task and the execution relationships between subtasks within the task sequences. Optionally, the annotation task is obtained by performing task filling in the original annotation task.

[0040] During the specific execution process, the user can also pre-construct the data to be annotated into a dataset to be annotated, which is used to annotate multiple data to be annotated during the execution of the annotation task. The dataset to be annotated can be submitted to sequentially annotate multiple tasks to be annotated, improving the data annotation efficiency. In an optional implementation manner provided in this embodiment, the data to be annotated is obtained in the following way: obtain the dataset to be annotated submitted by calling the service interface of the annotation service, and extract multiple data to be annotated included in the dataset to be annotated.

[0041] In addition, the user can also annotate a single data to be annotated. The user can submit the data to be annotated on the service page of the annotation service, improving the convenience for the user. In another optional implementation manner provided in this embodiment, the data to be annotated is obtained in the following way: obtain the data to be annotated input after the user triggers the data input control configured on the service page of the annotation service.

[0042] In this embodiment, the data to be annotated can be the data that needs to be data-annotated during the training and / or application of the model. The data to be annotated can include text data, image data, and / or audio data. Specifically, the data to be annotated can be the training samples of the model, or the basic data of the baseline model and / or the general model. For example, the data to be annotated can be the question text of the answer label to be annotated required during the training of the question-and-answer model, or the image of the category label to be annotated required during the training of the image classification model. Also for example, the data to be annotated can be the dialogue text of the reply label to be annotated of the question-and-answer model.

[0043] It should be noted that the task nodes can be filled in the task flow chart created by the user first to obtain the target task flow chart, and the target task flow chart can be converted into an annotation task. Step S202 can be replaced by: filling the task nodes in the task flow chart created by the user to obtain the target task flow chart, converting the target task flow chart into an annotation task, and combining it with the remaining steps and / or optional implementation manners provided in this embodiment to form a new embodiment. Among them, filling the task nodes includes filling empty task nodes. The specific operation of filling the task nodes can be performed with reference to the way of performing task filling in the original annotation task above.

[0044] Step S204: Parse the annotation task to obtain multiple task sequences and the execution relationships between the subtasks within each task sequence.

[0045] In specific implementation, parse the annotation task to obtain multiple task sequences included in the annotation task and / or multiple subtasks within the task sequence, and parse each task sequence to obtain the execution relationships between the subtasks within the task sequence. Optionally, the execution relationship includes the execution order of each subtask.

[0046] During the specific execution process, it is also possible to parse the annotation task to obtain multiple subtasks included in the annotation task and the execution relationships between the subtasks, and traverse the annotation task based on the execution relationships to determine the subtasks within the same task sequence and / or the same serial operator group and construct the task sequence.

[0047] It should be noted that it is also possible to first obtain the execution relationship and then construct the task sequence according to the execution relationship; step S204 can be replaced by: parse the annotation task to obtain multiple subtasks and the execution relationships between the subtasks, select subtasks according to the execution relationship to construct multiple task sequences, and combine them with the remaining steps and / or optional implementation manners provided in this embodiment to form a new embodiment.

[0048] It should also be noted that, as an executable task, the annotation task can also be directly executed based on the data to be annotated; that is, step S204 may not be executed, and directly after step S202 is completed, use the data to be annotated submitted by the user as the task execution data, and perform the subtasks within the task sequence included in the annotation task in parallel according to the execution relationships between the subtasks in the annotation task to obtain multiple task execution results, and combine them with the remaining steps provided in this embodiment to form a new implementation manner.

[0049] Step S206: Use the data to be annotated submitted by the user as the task execution input, and perform the subtasks within each task sequence in parallel according to the execution relationship to obtain multiple task execution results.

[0050] In specific implementation, use the data to be annotated submitted by the user as the input data of the starting task and / or input operator in the annotation task, and concurrently execute the subtasks of each task sequence. For the subtasks within each task sequence, perform the task execution according to the execution relationship and / or task hierarchy to obtain the task execution results corresponding to each task sequence. Optionally, the subtasks within the task sequence perform the task execution according to the execution relationship and / or task hierarchy.

[0051] In the specific execution process, to improve the accuracy and availability of the data output by subtasks with multiple input data and / or multi-input operators, during the parallel execution of tasks, after the subtasks at the same task level are completed, the execution results can be input to the subtasks at the next level to improve the accuracy and reliability of the data annotation results output by the annotation tasks, and thus improve the accuracy and reliability of the labels.

[0052] Specifically, the data to be annotated can be concurrently input to the subtasks in the first execution order within each task sequence. After each subtask is completed, the execution result is used as the input data for the subtasks in the second execution order. The subtasks at the remaining task levels and / or task execution orders are executed with reference to the above example until each task sequence completes the execution of the subtasks within the task sequence to obtain multiple task execution results. Further, a thread can be allocated for each task sequence to execute the subtasks within the task sequence; optionally, the subtasks within each task sequence are executed concurrently through multiple threads.

[0053] In addition, the subtasks within each task sequence can also be executed according to the task levels; step S206 can be replaced with: using the data to be annotated submitted by the user as the input for task execution, and concurrently executing the subtasks within each task sequence according to the task levels of the subtasks to obtain multiple task execution results, and combining them with the remaining steps and / or optional implementation manners provided in this embodiment to form a new embodiment.

[0054] In practical applications, to improve the accuracy of data annotation, multiple dimensions of data can be collected in various ways for data annotation processing. For example: data annotation results and / or labels can be obtained by means of assisted question answering with large language models, assisted question answering with expert models, and / or collecting data to extract key data; in this regard, the subtasks within each task sequence can also be configured accordingly according to the data annotation method corresponding to the task sequence to improve the usability of the task execution results. For example, if multiple different large language models are used to assist in data annotation for the data to be annotated, the call tasks of different large language models can be included in each task sequence. Optionally, task execution includes calling the large language model corresponding to the task sequence to generate the task execution result.

[0055] Based on the above description of the subtasks, specifically during the task execution process, there are also differences in the execution of different types of subtasks. The following separately describes the task execution of each type of subtask.

[0056] (1) Task execution of the large language model call task The configuration information of the large language model call task may include the data input interface of the target large language model, prompt words, and / or operation configuration data. The input data and / or prompt words can be input into the large language model for data processing, and the output result of the large language model is used as the execution result of the large language model call task. By assisting data annotation through the large language model, the usability of the data annotation result is improved. In an optional implementation provided in this embodiment, task execution includes: Determine the target large language model according to the configuration information of the subtask corresponding to the large language model call in the task sequence, and input the execution result of the upper-level subtask and the prompt words in the configuration information into the target large language model for text generation; Obtain the text data generated by the target large language model, and use the text data as the execution result of the subtask.

[0057] In addition, the large language model call task can also be configured in the first execution order in the task sequence, and the large language model is called based on the data to be annotated. The large language model call task can also perform data processing according to the type of input data. For example: determine the target large language model according to the configuration information of the subtask in the task sequence, and input the data to be annotated and / or the execution result of the upper-level task and the prompt words in the configuration information into the target large language model for data processing; use the processing result output by the target large language model as the execution result of the subtask.

[0058] (2)Task execution of the annotation result query task In practical applications, the results of some basic model answers and / or model judgments may be the same. For example, the answers of various question-and-answer models when replying to users' greetings can be the same. In response to this, a feature library corresponding to the annotation result can be established, and the information of the corresponding feature library is written into the configuration information of the annotation result query task. During task execution, the features of the input data can be extracted, and the target feature matching the feature is queried in the feature library corresponding to the subtask, and the label associated with the target feature is used as the execution result, saving computing resources and improving data annotation efficiency. In an optional implementation provided in this embodiment, task execution includes: Extract the features of the data to be annotated, and query the feature library that matches the configuration information of the subtask; Screen out the target features in the feature library whose feature similarity to the feature is greater than the threshold, and use the label associated with the target feature as the execution result of the subtask.

[0059] Optionally, the configuration information includes: interface information of the feature library and / or operation configuration information.

[0060] In addition, the annotation result query task can also be configured at the intermediate task level; for example: extracting the features of the execution results of the subtasks at the upper level; querying the feature library that matches the configuration information of the subtasks, querying the target features that match the feature in the feature library, and using the labels associated with the target features as the execution results of the subtasks.

[0061] (3)Task execution of the application programming interface call task In practical applications, there are also some data with high timeliness. The data with high timeliness can be queried and annotated by calling the application programming interface. For example, it can be queried whether the numerical values of financial data are correct; in this regard, the data query interface corresponding to the subtask can be called to query the data, and the queried data is used as the execution result; in an optional implementation manner provided in this embodiment, the task execution includes: Calling the data query interface that matches the configuration information of the subtask to query the data, and using the data returned by the interface call as the execution result of the subtask.

[0062] (4)Task execution of the key information extraction task During the specific execution process, the key information of the input data can also be extracted and used as the execution result to enable the subtasks at the next level to perform data processing. For example, the keywords in the text to be annotated can be extracted; in this regard, the key information can be extracted through the key information extraction task, and the extracted key information is used as the execution result; in an optional implementation manner provided in this embodiment, the task execution includes: Extracting the key information that matches the configuration information of the subtask from the input data, and using the key information as the execution result of the subtask.

[0063] (5)Task execution of the function editing task During the specific execution process, some lightweight functions can also be edited and implemented through subtasks, thereby improving the data annotation efficiency. For example, the aggregation logic function can be implemented by writing engineering code; in this regard, the operation configuration data corresponding to various functions can be configured in the configuration information of the function editing task, and the data is processed according to the operation configuration information during the task execution process; in an optional implementation manner provided in this embodiment, the task execution includes: Parsing the operation configuration information in the configuration information of the subtask, and performing data processing on the input data according to the operation configuration information to obtain the execution result of the subtask.

[0064] (6)Task execution of the conditional judgment task During the specific execution process, the task sequence executed by the annotation task can be controlled by means of conditional judgment, and the execution efficiency of the annotation task can be improved by controlling the execution direction. For example, by configuring a conditional judgment task, the execution of subsequent subtasks within the task sequence can be aborted when the input data does not meet the conditions; in this regard, the judgment conditions can be configured in the configuration information to determine whether specific fields in the input data meet the judgment conditions; in an optional implementation manner provided in this embodiment, the task execution includes: Detect whether the input data meets the conditions corresponding to the configuration information of the subtask, and determine the subtask at the next level according to the detection result.

[0065] It should be noted that the subtasks in the annotation task can include any type of task provided above. For the same type of task, it can also be configured in the same and / or different task sequences after adjusting the configuration information; that is: during the execution of any subtask in the annotation task, the task execution can include any one of the execution methods provided above, or a combination of two or more methods. For example: the task execution includes: extracting the key information in the input data that matches the configuration information of the subtask, and inputting the key information and the prompt words included in the configuration information into the target large language model corresponding to the configuration information for data processing; using the data returned by the target large language model as the execution result of the subtask.

[0066] Step S208, determine the target execution result among the task execution results, and use the target execution result as the label of the data to be annotated.

[0067] In specific implementation, traverse the multiple task execution results, screen out the target execution result among the task execution results, and use the target execution result as the label of the data to be annotated, that is: determine the label of the data to be annotated among the execution results.

[0068] In this embodiment, the task execution result can be the result obtained after the subtasks in the task sequence are executed according to the execution relationship, and can represent the label obtained by the annotation method corresponding to the task sequence. Specifically, the task execution results corresponding to each task sequence can be the same or different; for example: for the data to be annotated corresponding to the classification model, the task execution results corresponding to each task sequence belong to the same classification, or, in the case where the other task execution results are the same, there is a task execution result corresponding to one task sequence that is different from the other task execution results.

[0069] In addition, there are also cases where the key information of the task execution results is the same. The task execution results with the same key information can be considered as the same task execution results. For example, for the data to be annotated of a question-and-answer model, the task execution result can be the answer text to the question text. There may be differences in word order and word usage for the same question. The answer texts with the same key information can be regarded as the same task execution results, and the answer texts with different key information can be regarded as different task execution results. It should be noted that the task execution results can exist in the form of parameter data, and whether the task execution results are the same can be judged by comparing the parameters.

[0070] Specifically, in the process of determining the target execution result, to improve the usability and reliability of the labels of the data to be annotated, the task execution result with a relatively large proportion can be selected from multiple task execution results as the target execution result. For example, the target execution result can be determined through a voting algorithm; or, multiple task execution results can be divided into multiple categories, the number of task execution results in each category in the classification results can be counted, and the task execution result with a larger number can be used as the target execution result. For example, the target execution result can be determined through a clustering algorithm. In addition, the target weight corresponding to each category can be calculated by combining the weights configured for the subtasks to which each task execution result belongs with clustering, and the task execution result corresponding to the target weight can be used as the target execution result.

[0071] In the specific execution process, clustering processing can be performed on each task execution result to obtain clustering groups. The target clustering group can be selected according to the number of task execution results in the clustering group, and the task execution results in the target clustering group can be used as the target execution result and / or label, improving the usability of the target execution result and / or label. In an optional implementation manner provided in this embodiment, determining the target execution result among each task execution result includes: Performing clustering processing on each task execution result to obtain clustering groups; Sorting the clustering groups according to the number of task execution results in the clustering groups, and using the task execution results in the group corresponding to the target order as the target execution result.

[0072] Specifically, clustering processing is performed on each task execution result to obtain multiple clustering groups, the number of task execution results in each clustering group is counted, and the clustering groups are sorted according to the number. The task execution results in the group corresponding to the target order are used as the target execution result, that is: the task execution results in the clustering group with the largest number are used as the target execution result and / or label.

[0073] Alternatively, the label can also be directly determined in the execution results of each task, that is: determine the label of the data to be labeled in the execution results of each task; in another alternative implementation provided in this embodiment, determine the target execution result in the execution results of each task, and use the target execution result as the label of the data to be labeled, including: Perform clustering processing on the execution results of each task to obtain clustering groups; Sort the clustering groups according to the number of results of the task execution results in the clustering groups, and use the task execution results in the group corresponding to the target order as the label of the data to be labeled.

[0074] Alternatively, the target execution result and / or label can also be determined in multiple task execution results by pre-configuring weights; for example: calculate the weights of each clustering group according to the weights configured for the subtasks to which the task execution results in each clustering group belong, select the target group corresponding to the target weight, and use the task execution results in the target group as the target execution result and / or the label of the data to be labeled.

[0075] In addition, the label of the data to be labeled can also be determined in the aggregated result after aggregation; in an alternative implementation provided in this embodiment, determine the target execution result in the execution results of each task, and use the target execution result as the label of the data to be labeled, including: select the task execution results for aggregation processing according to the preset configuration, obtain multiple aggregated results, and determine the label of the data to be labeled in each aggregated result.

[0076] Specifically, the label of the data to be labeled can be determined in each aggregated result by means of weights and / or voting.

[0077] In practical applications, the label of the data to be labeled can exist in the form of parameter data; in view of this, to improve the user's perception of the data annotation result, the annotation result corresponding to the data to be labeled can be generated according to the label and returned to the user for the user to view; in an alternative implementation provided in this embodiment, after obtaining the label of the data to be labeled, it further includes: If the data to be labeled is the text to be marked, perform intent marking on the text to be marked according to the label to obtain the marked text, and return the marked text to the user.

[0078] If the data to be labeled is the question text to be decided, generate the decision text of the question text according to the label, and return the decision text to the user.

[0079] Specifically, the intention of the data to be annotated can be marked according to the tags to obtain the marked text, and the marked text can be returned to the user so that the user can view the intention data of the data to be annotated. Alternatively, the decision text and / or response text of the data to be annotated can also be generated according to the tags, and the decision text and / or response text can be returned to the user so that the user can query the decision text of the data to be annotated.

[0080] It should be noted that the data annotation for the data to be annotated can also be multiple annotations. For example, while annotating the intention data and the response text for the data to be annotated, the labels of the data to be annotated can be obtained by aggregating multiple execution results; step S208 can be replaced by aggregating the execution results of each task to obtain the label of the data to be annotated, and combining it with the remaining steps and / or optional implementation manners of this embodiment to form a new embodiment.

[0081] In addition, during the data annotation process, the task flow chart created by the user can be converted into an annotation task, the annotation task can be parsed to obtain multiple task sequences, the data to be annotated submitted by the user can be used as the task execution input, and the subtasks within each task sequence can be executed in parallel according to the execution relationship between the tasks within the task sequence to obtain multiple task execution results. The label of the data to be annotated is determined from the task execution results, and combined with the remaining steps and / or optional implementation manners of this embodiment to form a new embodiment; the function of customizing the annotation task flow is provided to the user operating the data annotation. By converting the annotation task and filling in the empty tasks, the stability of the annotation task execution is improved, the system robustness is enhanced, and at the same time, the task execution results corresponding to multiple task sequences are obtained and selected, improving the usability of the data annotation results and / or labels. Optionally, the annotation task is obtained by filling in the empty tasks after being converted based on the task flow chart created by the user.

[0082] In summary, for one or more data annotation methods provided in this embodiment, the task flow chart created by the user is converted into an original annotation task, the annotation task is obtained by filling in tasks in the original task, the annotation task is parsed to obtain multiple task sequences and the execution relationship of each subtask within the task sequence, the data to be annotated submitted by the user is used as the task execution input, and the subtasks within each task sequence are executed in parallel according to the execution relationship to obtain multiple task execution results. The label of the data to be annotated is determined from the task execution results. By concurrently executing the subtasks in the annotation task, the annotation efficiency is improved. At the same time, by constructing the annotation task by filling in the empty tasks, the task levels within each task sequence are made consistent, improving the system stability and robustness, and improving the usability of the data annotation results; Further, after parsing the original annotation task, empty tasks can be filled in the task sequence of the original annotation task to obtain the task sequence of the annotation task and the execution relationship of each subtask within the task sequence, so that the levels of the task sequences of the annotation task are the same, improving the stability and robustness during the operation of the annotation task, and further improving the accuracy of the labels; In addition, the user can also trigger the creation interface configured on the service page of the annotation service to create a task flow chart, which represents the task flow of data annotation, provides a way for the operating user of data annotation to customize the task flow, improves the user perception, and at the same time decomposes the complex annotation task into multiple subtasks, improving the annotation efficiency.

[0083] The following takes the application of a data annotation method provided in this embodiment in a data annotation scenario as an example, and combines Figure 4 to further illustrate the data annotation method provided in this embodiment. The data annotation method applied to the data annotation scenario specifically includes the following steps.

[0084] Step S402, convert the task flow chart created by the user into an original annotation task.

[0085] Step S404, parse the original annotation task to obtain multiple original task sequences and the execution relationship between each subtask within the original task sequence.

[0086] Step S406, determine the target task sequence in each original task sequence according to the number of subtasks in each original task sequence.

[0087] Step S408, screen out the target subtasks with multiple input data in the target task sequence according to the execution relationship, and determine the target level where the target subtasks are located in the target task sequence.

[0088] Step S410, calculate the missing task levels of each original task sequence according to the execution level and the target level of the target subtasks within each original task sequence.

[0089] Step S412, fill empty tasks at the missing task levels of each original task sequence to obtain multiple task sequences of the annotation task.

[0090] Step S414, based on the data to be annotated submitted by the user, concurrently execute the subtasks within each task sequence according to the execution relationship to obtain multiple task execution results.

[0091] Previously, the dataset to be annotated submitted by calling the service interface of the annotation service can be obtained; or, the data to be annotated input by the user on the service page of the annotation service can be obtained. Optionally, the dataset to be annotated includes multiple data to be annotated.

[0092] Step S416: Determine the labels of the data to be labeled among the execution results of each task.

[0093] Step S418: Generate the annotation results of the data to be labeled based on the labels and the data to be labeled, and send them to the user.

[0094] It should be noted that before Step S402 is executed, it may further include: obtaining the task flow chart created by the user. Among them, the task flow chart can be created in the following way: select multiple task nodes according to the node selection instruction submitted by the user; configure the association relationship between each task node according to the node configuration instruction submitted by the user to obtain the task flow chart.

[0095] It should be noted that any one step or any combination of multiple steps among Step S402 to Step S418 can be combined with any one step or any combination of multiple steps among the above-mentioned Step S202 to Step S208 to form a new implementation manner according to the needs of implementation and deployment; in addition, according to the actual deployment needs, any one or any combination of technical features among Step S402 to Step S418 can be selected and combined with any one or more technical features provided by the above-mentioned Step S202 to Step S208 to form a new implementation manner; or, any one or any combination of technical features among Step S402 to Step S418 can also be replaced by any one or more technical feature combinations provided by the above-mentioned Step S202 to Step S208 according to the actual deployment needs to form a new implementation manner, which will not be elaborated here one by one.

[0096] An embodiment of a data annotation device provided in this specification is as follows: In the above embodiment, a data annotation method is provided. Correspondingly, a data annotation device is also provided, which will be described below with reference to the drawings.

[0097] Refer to Figure 5 , which shows a schematic diagram of an embodiment of a data annotation device provided in this embodiment.

[0098] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiments described below are only illustrative.

[0099] This embodiment provides a data annotation device, and the device includes: An annotation task generation module 502, configured to convert the task flow chart created by the user into an original annotation task, and perform task filling in the original annotation task to obtain an annotation task; An annotation task parsing module 504, configured to parse the annotation task to obtain multiple task sequences and the execution relationship between each subtask within the task sequence; The task execution module 506 is configured to use the data to be annotated submitted by the user as the task execution input, and perform the task execution of the subtasks within each task sequence in parallel according to the execution relationship to obtain multiple task execution results; the task execution includes calling the large language model corresponding to the task sequence to generate the execution result of the subtask. The label obtaining module 508 is configured to determine the target execution result among the task execution results, and use the target execution result as the label of the data to be annotated.

[0100] An embodiment of a data annotation device provided in this specification is as follows: Corresponding to the above-described data annotation method, based on the same technical concept, one or more embodiments of this specification also provide a data annotation device, which is used to execute the data annotation method provided above. Figure 6 It is a schematic structural diagram of a data annotation device provided by one or more embodiments of this specification.

[0101] A data annotation device provided in this embodiment includes: As Figure 6 shown, the data annotation device may have relatively large differences due to configuration or performance, and may include one or more processors 601 and a memory 602. One or more application programs or data may be stored in the memory 602. Among them, the memory 602 may be short-term storage or persistent storage. The application programs stored in the memory 602 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the data annotation device. Further, the processor 601 may be set to communicate with the memory 602 and execute a series of computer-executable instructions in the memory 602 on the data annotation device. The data annotation device may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input / output interfaces 605, and one or more keyboards 606.

[0102] In a specific embodiment, the data annotation device includes a memory and one or more programs, where one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions in the data annotation device, and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions: Convert the task flow chart created by the user into an original annotation task, and perform task filling in the original annotation task to obtain an annotation task; Parse the annotation task to obtain multiple task sequences and the execution relationships between the subtasks within each task sequence; Use the data to be annotated submitted by the user as the task execution input, and perform the task execution of the subtasks within each task sequence in parallel according to the execution relationships to obtain multiple task execution results; the task execution includes calling the large language model corresponding to the task sequence to generate the execution results of the subtasks; Determine the target execution result among the task execution results, and use the target execution result as the label of the data to be annotated.

[0103] An embodiment of a computer-readable storage medium provided in this specification is as follows: Corresponding to the data annotation method described above, based on the same inventive concept, one or more embodiments of this specification also provide a computer-readable storage medium.

[0104] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, and the computer-executable instructions, when executed, implement the following process: Convert the task flow chart created by the user into an original annotation task, and perform task filling in the original annotation task to obtain an annotation task; Parse the annotation task to obtain multiple task sequences and the execution relationships between the subtasks within each task sequence; Use the data to be annotated submitted by the user as the task execution input, and perform the task execution of the subtasks within each task sequence in parallel according to the execution relationships to obtain multiple task execution results; the task execution includes calling the large language model corresponding to the task sequence to generate the execution results of the subtasks; Determine the target execution result among the task execution results, and use the target execution result as the label of the data to be annotated.

[0105] It should be noted that the embodiment of the computer-readable storage medium in this specification and the embodiment of the data annotation method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding method described above, and the repeated parts will not be elaborated.

[0106] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. The key points of each embodiment are the differences from other embodiments. For example, the device embodiment, the equipment embodiment, and the computer-readable storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. Please refer to the relevant description of the method embodiment when reading the relevant content in the device embodiment, the equipment embodiment, and the computer-readable storage medium embodiment.

[0107] The foregoing describes particular embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In certain embodiments, multitasking and concurrent processing are also possible or may be advantageous.

[0108] In the 1930s, it was obvious to distinguish whether an improvement to a technology was a hardware improvement (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with hardware entity modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that as long as the method flow is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0109] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0110] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0111] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0112] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0113] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable network live processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable network live processing devices produce means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0114] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable network live processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0115] These computer program instructions can also be loaded onto a computer or other programmable network live processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.

[0116] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0117] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0118] A computer-readable medium includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0119] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising at least one..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0120] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0121] The above are only embodiments of this document and are not intended to limit this document. For those skilled in the art, this document may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this document shall be included within the scope of the claims of this document.

Claims

1. A data annotation method, comprising: Converting the task flow chart created by the user into an original annotation task, and performing task filling in the original annotation task to obtain the annotation task; Parsing the labeling task to obtain multiple task sequences and execution relationships between subtasks in the task sequences; The data to be annotated submitted by the user is used as task execution input, and task execution of subtasks in each task sequence is performed in parallel according to the execution relationship to obtain multiple task execution results; the task execution includes calling the large language model corresponding to the task sequence to generate the execution results of the subtasks; The target execution result is determined in each task execution result, and the target execution result is used as a label of the data to be labeled.

2. According to the data annotation method of claim 1, before the step of converting the task flow chart created by the user into an original annotation task and performing task filling in the original annotation task to obtain the annotation task is executed, it also includes: Selecting multiple task nodes according to the node selection instruction submitted by the user; According to the node configuration instruction submitted by the user, the association relationship between the task nodes is configured to obtain the task flow chart; The task node is configured according to the node configuration information submitted by the user.

3. According to the data labeling method of claim 1, the data to be labeled is obtained in the following manner: Acquire a data set to be annotated submitted by calling a service interface of an annotation service, and extract a plurality of data to be annotated contained in the data set to be annotated; Alternatively, the data to be annotated which is input by the user after triggering a data input control configured on a service page of the annotation service is obtained.

4. The data annotation method according to claim 1, wherein converting the task flow chart created by the user into the original annotation task comprises: Extract multiple task nodes and the association relationship between the task nodes in the task flow chart; Querying node configuration information corresponding to each of the task nodes, generating a plurality of subtasks according to the node configuration information, and determining an execution relationship between the subtasks according to the association relationship; The original labeling task is constructed based on the multiple subtasks and the determined execution relationship.

5. According to the data annotation method of claim 1, the step of performing task filling in the original annotation task to obtain the annotation task comprises: Parsing the original labeling task to obtain multiple original task sequences and execution relationships between subtasks in the original task sequences, and determining a target task sequence in each original task sequence according to the number of subtasks in each original task sequence; Determine the missing task levels of each original task sequence relative to the target task sequence according to the execution relationship obtained by analysis, and fill in empty tasks at the missing task levels of each original task sequence to obtain the labeled task.

6. The data labeling method according to claim 5, wherein determining the missing task levels of each original task sequence relative to the target task sequence according to the execution relationship obtained by parsing comprises: Filtering out target subtasks with multiple input data in the target task sequence according to the execution relationship obtained by parsing, and determining the target level of the target subtask in the target task sequence; The missing task level of each original task sequence is calculated according to the execution level of the target subtask in each original task sequence and the target level.

7. According to the data labeling method of claim 1, the execution result of the large language model generation subtask corresponding to the calling task sequence includes: Determine a target large language model according to configuration information of a subtask corresponding to the large language model call in the task sequence, and input the execution result of the subtask at the previous level and the prompt words in the configuration information into the target large language model for text generation; The text data generated by the target large language model is obtained, and the text data is used as the execution result of the subtask.

8. The data annotation method according to claim 1, wherein the task execution further comprises: Extracting features of the data to be annotated, and searching a feature library that matches the configuration information of the subtask; A target feature whose feature similarity with the feature is greater than a threshold is screened out in the feature library, and a label associated with the target feature is used as an execution result of the subtask.

9. The data annotation method according to claim 1, wherein the task execution further comprises: Call the data query interface that matches the configuration information of the subtask to query the data, and use the data returned by the interface call as the execution result of the subtask; Alternatively, extract key information in the input data that matches the configuration information of the subtask, and use the key information as the execution result of the subtask; Alternatively, parsing the operation configuration information in the configuration information of the subtask, and performing data processing on the input data according to the operation configuration information to obtain the execution result of the subtask; Alternatively, it is detected whether the input data meets the conditions corresponding to the configuration information of the subtask, and the subtask of the next level is determined according to the detection result.

10. The data labeling method according to claim 1, wherein determining the target execution result in each task execution result comprises: Performing clustering processing on the execution results of each task to obtain cluster groups; The cluster groups are sorted according to the number of task execution results in the cluster groups, and the task execution results in the cluster groups corresponding to the target order are used as the target execution results.

11. The data labeling method according to claim 1, after the step of determining a target execution result from each task execution result and using the target execution result as a label of the data to be labeled, further comprises: If the data to be annotated is text to be marked, marking the text to be marked according to the label to obtain marked text, and returning the marked text to the user; If the data to be labeled is a question text to be decided, a decision text of the question text is generated according to the label, and the decision text is returned to the user.

12. A data annotation device, comprising: A labeling task generating module is configured to convert the task flow chart created by the user into an original labeling task, and perform task filling in the original labeling task to obtain the labeling task; A labeling task parsing module is configured to parse the labeling task to obtain multiple task sequences and execution relationships between subtasks in the task sequences; The task execution module is configured to use the data to be annotated submitted by the user as task execution input, and to perform task execution of subtasks in each task sequence in parallel according to the execution relationship to obtain multiple task execution results; the task execution includes calling the large language model corresponding to the task sequence to generate the execution results of the subtasks; The label acquisition module is configured to determine the target execution result in each task execution result, and use the target execution result as the label of the data to be labeled.

13. A data annotation device, comprising: processor; and a memory configured to store computer executable instructions, wherein the processor executes the computer executable instructions to implement the following process: Converting the task flow chart created by the user into an original annotation task, and performing task filling in the original annotation task to obtain the annotation task; Parsing the labeling task to obtain multiple task sequences and execution relationships between subtasks in the task sequences; The data to be annotated submitted by the user is used as task execution input, and task execution of subtasks in each task sequence is performed in parallel according to the execution relationship to obtain multiple task execution results; the task execution includes calling the large language model corresponding to the task sequence to generate the execution results of the subtasks; The target execution result is determined in each task execution result, and the target execution result is used as a label of the data to be labeled.

14. A computer-readable storage medium for storing computer-executable instructions, wherein the computer-executable instructions implement the steps of the method of claim 1 when executed.