Data annotation method and device, equipment and medium
By utilizing a high-quality annotation library and customized models on the data annotation platform, and determining the processing module for data annotation based on the annotation type, the problems of low efficiency of manual annotation and inconsistent quality of general models are solved, efficient and high-quality data annotation is achieved, and the performance of the machine model is improved.
Patent Information
- Application Number
- CN202510856673.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-03
AI Technical Summary
In the existing technology, the quality of data annotation has a significant impact on the performance of the machine model, but manual annotation is labor-intensive and inefficient, and the annotation quality of universal base models varies.
Obtain the data to be annotated and the annotation type information through the data annotation platform, use high-quality annotation libraries and customized annotation models, determine the corresponding annotation processing module according to the annotation type to perform data annotation, and improve the pertinence and quality of annotation.
It improves the efficiency and quality of data labeling, ensures that the labeled data can be better used to train machine learning models, and improves the performance of machine models.
Smart Images

Figure CN120744713A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data annotation method, device, equipment and medium. Background Art
[0002] With the development of computer technology, artificial intelligence (AI) technologies, including speech recognition, image recognition, and natural language processing, have gained widespread application. Within the AI field, building machine models for data classification, object detection, intent recognition, and semantic understanding can effectively improve the processing efficiency and accuracy of corresponding tasks.
[0003] In practical applications, data labeling is often required to train machine models to achieve desired computer tasks. The quality of the labeled data has a significant impact on the performance of the machine models. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a data labeling method, apparatus, device, and medium to improve the data quality of labeled data.
[0005] To solve the above technical problems, the embodiments of this specification are implemented as follows:
[0006] The embodiments of this specification provide a data annotation method, including:
[0007] Obtain the data to be labeled;
[0008] Acquire annotation type information for the data to be annotated; the annotation type information indicates the type of annotation information that can be annotated for the data to be annotated;
[0009] Determining a label processing module corresponding to the label type information according to the label type information; the label processing module has a data labeling function of labeling the label information corresponding to the label type information;
[0010] The labeling processing module is used to perform labeling processing on the data to be labeled to obtain labeled data containing labeling information.
[0011] The embodiments of this specification provide a data annotation device, including:
[0012] The labeling data acquisition module is used to obtain the data to be labeled;
[0013] A labeling type acquisition module, configured to acquire labeling type information to be labeled for the data to be labeled; the labeling type information indicates the type of labeling information that can be labeled for the data to be labeled;
[0014] A determination module, configured to determine, based on the annotation type information, an annotation processing module corresponding to the annotation type information; the annotation processing module having a data annotation function of annotating the annotation information corresponding to the annotation type information;
[0015] The processing module is used to perform labeling processing on the data to be labeled using the labeling processing module to obtain labeled data containing labeling information.
[0016] The embodiments of this specification provide a data annotation device, including:
[0017] at least one processor; and,
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the above-mentioned data labeling method.
[0020] An embodiment of this specification provides a computer-readable medium having computer-readable instructions stored thereon, and the computer-readable instructions can be executed by a processor to implement the above-mentioned data labeling method.
[0021] At least one embodiment provided in this specification can achieve the following beneficial effects: This embodiment can determine, based on the annotation type information of the data to be annotated, an annotation processing module corresponding to the annotation type information, wherein the annotation processing module has a data annotation function for annotating the annotation information corresponding to the annotation type information, and annotate the data to be annotated using the determined annotation processing module. Because the determined annotation processing module is targeted at the data to be annotated, the data quality of the annotated data can be improved by using the annotation processing module corresponding to the annotation type information to perform annotation processing on the data to be annotated. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of this specification or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0023] Figure 1 This is a schematic diagram of an application scenario of a data annotation method provided in an embodiment of this specification;
[0024] Figure 2This is a flowchart of a data annotation method provided in one embodiment of this specification;
[0025] Figure 3 This is a flowchart of a data annotation method provided in one embodiment of this specification;
[0026] Figure 4 The embodiments of this specification provide corresponding Figure 2 A structural diagram of a data labeling device;
[0027] Figure 5 The embodiments of this specification provide corresponding Figure 2 A structural diagram of a data labeling device. DETAILED DESCRIPTION
[0028] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.
[0029] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.
[0030] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0032] In order to solve the defects in the prior art, this solution provides the following embodiments:
[0033] Figure 1 This is a schematic diagram of an application scenario of a data annotation method provided in an embodiment of this specification. Figure 1 As shown, in the embodiments of this specification, the data to be labeled can be labeled using the data labeling platform 100 to obtain labeled data containing labeling information. For example, the data to be labeled A is labeled using the data labeling platform 100 to obtain data to be labeled A-labeling information 1, the data to be labeled B is labeled to obtain data to be labeled B-labeling information 2, the data to be labeled C is labeled to obtain data to be labeled C-labeling information 3, and so on.
[0034] Among them, the server carried by the data annotation platform 100 can be an independent physical server, or a server cluster or distributed file system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.
[0035] Specifically, the data annotation platform 100 can obtain the data to be annotated and the annotation type information to be annotated for the data to be annotated. For example, the data annotation platform 100 can obtain data to be annotated A and the annotation type information of data to be annotated A. Subsequently, the data annotation platform 100 can determine the annotation processing module corresponding to the annotation type information based on the annotation type information of data to be annotated A, where the annotation processing module has the data annotation function of annotating the annotation information corresponding to the aforementioned annotation type information. The annotation processing module can then be used to annotate the data to be annotated A, obtaining annotated data such as data to be annotated A - annotation information 1.
[0036] The data annotation method provided in the embodiments of this specification is introduced below with reference to the accompanying drawings.
[0037] Figure 2 This is a flow chart of a data annotation method provided in one embodiment of this specification. From a program perspective, Figure 2 The execution subject of the method can be a program installed on a server. Optionally, the method can also be executed by any device, equipment, platform, or device cluster with computing and processing capabilities, such as a data annotation platform.
[0038] like Figure 2 As shown, the data annotation method may include the following steps.
[0039] Step 202: Obtain data to be labeled.
[0040] In the embodiment of this specification, the data to be labeled may be data that needs to be labeled.
[0041] Data labeling is the process of converting unlabeled data into a form that can be understood by machine learning algorithms. By converting unlabeled data into labeled data, machine learning algorithms can learn various data processing tasks such as classification, regression, and object detection.
[0042] In practical applications, one or more methods can be used to obtain data to be labeled. For example, data to be labeled can be obtained from public datasets on the internet, crawled from social media, online forums, and other platforms using web crawlers, collected through crowdsourcing platforms, or collected using image acquisition devices, audio acquisition devices, and various sensors.
[0043] Optionally, the data to be labeled can be data from various business scenarios. For example, the data to be labeled can be data from computer vision scenarios, such as facial image data for emotion recognition, traffic image data for object detection, product image data for defect location detection, and medical imaging data for assisted medical diagnosis. The data to be labeled can also be data from natural language processing scenarios, such as product review data for sentiment analysis, news article data for topic classification, and text data for named entity recognition. The data to be labeled can also be data from medical scenarios, such as medical imaging data and genomic data for assisted medical diagnosis. The data to be labeled can also be data from financial scenarios, such as transaction records for assisting in fraud detection, financial data for credit risk assessment, and market fluctuation data for market forecasting. The data to be labeled can also be data from autonomous driving scenarios, such as point cloud data detected by lidar, traffic image data for traffic sign detection, and traffic video data for object tracking. The data to be labeled can also be data from educational scenarios, such as test answer data for assessing user knowledge and learning log data for monitoring user learning behavior. The data to be labeled can be data for recommendation scenarios, such as product click data and product purchase data used to analyze user preferences for products.
[0044] Optionally, the data to be annotated may be in various formats, for example, the format of the data to be annotated may be at least one of text format, voice format, image format, and video format.
[0045] Step 204: Obtain annotation type information for the data to be annotated; the annotation type information indicates the type of annotation information that can be annotated for the data to be annotated.
[0046] In the embodiments of this specification, the annotation type information of the data to be annotated may be an annotation type set according to actual annotation requirements, such as sentiment analysis type information, target detection type information, correlation type information, etc.
[0047] In actual applications, the data annotation platform may be pre-set with one or more annotation type information. Users can select the annotation type information to be annotated in the data annotation platform's operation interface based on actual annotation requirements. Alternatively, users can enter the annotation type information to be annotated in the data annotation platform's operation interface based on actual annotation requirements.
[0048] Specifically, annotation information refers to the labels or annotations added during the process of annotating the data to be annotated, which is used to describe the content, characteristics or categories of the data so that subsequent machine learning models can understand and process the data.
[0049] In an embodiment of the present specification, the annotation type information of the data to be annotated may be classification type information, and the classification type information may be information used to classify the data to be annotated. For example, if the data to be annotated is image data, the annotation information of the classification type information may specifically be "cat", "dog", "car", etc.; for example, if the data to be annotated is product review data, the annotation information of the classification type information may specifically be "positive review", "negative review", "neutral review", etc. The annotation type information of the data to be annotated may be target detection type information, and the target detection type information may be information used to perform target detection on the data to be annotated. For example, if the data to be annotated is traffic image data, the annotation information of the target detection type information may specifically be bounding box information representing the positions of pedestrians and vehicles in the traffic image data. The annotation type information of the data to be annotated may be key point detection type information, and the target detection type information may be information used to detect key points in the data to be annotated. For example, if the data to be annotated is facial image data, the annotation information of the key point detection type information may specifically be position information representing the eyes, nose, and mouth in the facial image data, etc. The annotation type information for the data to be annotated may be named entity recognition information, which may be information used to identify entities in the data to be annotated. For example, if the data to be annotated is text data, the annotation information for the named entity recognition information may specifically include person names, place names, or organization names. The annotation type information for the data to be annotated may also be correlation information, which may be information used to perform correlation analysis on the data to be annotated. For example, if the data to be annotated is two paragraphs of text, the annotation information for the correlation information may specifically include information indicating the correlation between the two paragraphs, such as "related" or "not related." The annotation type information for the data to be annotated may also be voiceprint recognition information, which may be information used to identify the user of the data to be annotated. For example, if the data to be annotated is voice data of a particular user, the annotation information for the voiceprint recognition information may specifically include information indicating the user's identity. The annotation type information for the data to be annotated may also be action recognition information, which may be information used to identify the actions of the target in the data to be annotated. For example, if the data to be annotated is sports-related video data, the annotation information for the action recognition type information may specifically represent user movements such as running, jumping, or shooting. The annotation type information for the data to be annotated may be business scenario type information, which may be information used to represent the business scenario of the data to be annotated. For example, if the data to be annotated is text data related to a financial scenario, the annotation information for the business scenario type information for the data to be annotated may specifically represent a financial scenario.For example, if the data to be labeled is traffic image data related to an autonomous driving scenario, the annotation information for the business scenario type information of the data to be labeled can specifically be information indicating the autonomous driving scenario. The annotation type information of the data to be labeled can also be event type information, which can be information used to indicate an event in the data to be labeled. For example, if the data to be labeled is meteorological data, the annotation information for the event type information can specifically be information indicating weather conditions such as heavy rain or typhoons. In practical applications, the annotation type information of the data to be labeled can also be other types of information, which are not listed here.
[0050] As a specific implementation, data to be annotated in different formats may have different annotation type information. For example, data to be annotated in image format may have object detection type information, key point detection type information, etc. Data to be annotated in text format may have named entity recognition type information. Data to be annotated in speech format may have voiceprint recognition type information. Data to be annotated in video format may have action recognition type information, etc.
[0051] As a specific implementation, data to be annotated in different formats can have the same annotation type information. For example, data to be annotated in image format, text format, voice format, and video format can all have application scenario type information, event type information, etc.
[0052] Optionally, any format of the data to be annotated may have one or more annotation type information.
[0053] In practical applications, the annotation type information of the data to be annotated can be determined based on the data processing task of the machine learning model to be trained. For example, if the data processing task of the machine learning model to be trained is to classify image data, the annotation type information can be classification type information. For example, if the data processing task of the machine learning model to be trained is to analyze the correlation between two text segments, the annotation type information can be correlation type information.
[0054] Step 206: Determine a labeling processing module corresponding to the labeling type information according to the labeling type information; the labeling processing module has a data labeling function of labeling the labeling information corresponding to the labeling type information.
[0055] In the embodiments of this specification, the data annotation platform may include one or more annotation processing modules.
[0056] Optionally, the data annotation platform may include a first annotation processing module for annotating the data to be annotated using a high-quality annotation library. The high-quality annotation library may store annotated data with a confidence level higher than a first preset threshold, and the annotated data in the high-quality annotation library may be used to annotate the processed data.
[0057] Optionally, the data annotation platform can include a second annotation processing module for annotating the data to be annotated using a trained base model. The base model is a general model pre-trained on large-scale data, possessing extensive knowledge and general capabilities, which can be fine-tuned to adapt to various specific tasks.
[0058] Optionally, the data annotation platform may include a third annotation processing module for annotating the data to be annotated using a customized annotation model, wherein the customized annotation model may be an annotation model trained using training data of the same annotation type as the data to be annotated.
[0059] In the embodiments of this specification, different annotation processing modules can annotate and process data to be annotated with different annotation types. For example, the first annotation processing module can annotate and process data to be annotated for classification, action recognition, key point detection, and object detection; the second annotation processing module can annotate and process data to be annotated for named entity recognition and relevance; the third annotation processing module can annotate and process data to be annotated for voiceprint recognition, business scenario, and event types, etc. Optionally, any one annotation processing module can annotate and process data to be annotated with annotation information corresponding to one or more annotation types.
[0060] As a specific implementation method, different annotation processing modules can also annotate data to be annotated with the same annotation type information. For example, the first annotation processing module, the second annotation processing module, and the third annotation processing module can all annotate data to be annotated with business scenario type information, event type information, etc.
[0061] In the embodiments of this specification, the data annotation platform may store a correspondence between annotation type information and an annotation processing module, so that the data annotation platform may use the correspondence to determine the annotation processing module corresponding to the annotation type information.
[0062] Step 208: Utilize the annotation processing module to perform annotation processing on the data to be annotated to obtain annotated data containing annotation information.
[0063] In the embodiments of this specification, a labeling processing module may be used to perform labeling processing on the data to be labeled, determine labeling information corresponding to the data to be labeled, and thereby obtain labeled data containing the labeling information.
[0064] In practical applications, labeled data can be used to train machine learning models, such as neural network models, linear regression models, decision tree models, etc.
[0065] As a specific implementation method, the annotated data can also be used to build or update the knowledge graph. For example, the entities, relationships, and attributes in the data to be annotated can be annotated to obtain the annotated data to build or update the knowledge graph.
[0066] Compared with the related technologies that use manual data labeling, manual labeling not only consumes a lot of manpower, but also has low data labeling efficiency and is easily affected by personal cognitive ability, and the labeling quality cannot be guaranteed. In the embodiment of this specification, the user can set the labeling type information of the data to be labeled, and the data labeling platform can then label the data to be labeled, which can improve the efficiency of data labeling. In addition, the embodiment of this specification uses the labeling processing module in the data labeling platform to perform data labeling, which is not affected by personal cognitive ability, thereby ensuring the quality of data labeling.
[0067] In addition, some related technologies use a universal base model for data annotation. However, the annotation quality of data with various annotation types using this universal base model varies. In the embodiments of this specification, a labeling processing module corresponding to the annotation type information can be determined based on the annotation type information of the data to be annotated. The labeling processing module has the data annotation function of annotating the annotation information corresponding to the annotation type information. Because the determined labeling processing module is targeted at the data to be annotated, the labeling processing module corresponding to the annotation type information can be used to annotate the data to be annotated, thereby improving the data quality of the annotated data.
[0068] based on Figure 2 The present specification also provides some implementation methods of the method, which are described below.
[0069] In an embodiment of the present specification, the annotation processing module may include a module for annotating the data to be annotated using a high-quality annotation library. The high-quality annotation library stores annotated data with a confidence level higher than a first preset threshold. Optionally, the annotating module annotates the data to be annotated to obtain annotated data containing annotation information, which may specifically include:
[0070] Retrieve labeled data whose similarity with the data to be labeled is greater than or equal to a second preset threshold from the high-quality labeling library to obtain a plurality of target labeled data.
[0071] Based on the labeling information of the plurality of target labeled data, labeling processing is performed on the data to be labeled to obtain labeled data containing the labeling information.
[0072] In the embodiment of this specification, the high-quality annotation library may be a database used to perform annotation processing on data to be annotated.
[0073] In practical applications, a high-quality annotation library can be pre-built in one or more ways, such as based on expert experience, or based on a base model or large model.
[0074] Specifically, the high-quality annotation library may store annotated data.
[0075] Optionally, the labeled data stored in the high-quality labeling library may be obtained by labeling historical data to be labeled.
[0076] Among them, the historical data to be labeled can be obtained from public data sets on the Internet, captured from social media, online forums and other platforms using web crawler technology, collected using crowdsourcing platforms, collected using image acquisition equipment, audio acquisition equipment and various sensors, or can be set using expert experience, etc.
[0077] Furthermore, the historical data to be labeled can be labeled based on expert experience, or the data to be labeled can be labeled using a base model or a large model, such as by constructing prompt words that instruct the base model or the large model to label the historical data to be labeled, and inputting the prompt words into the base model or the large model to label the data to be labeled, or the intelligent labeling iTAG module in the Alibaba Cloud artificial intelligence platform PAI can be used to label the historical data to be labeled, etc., to obtain labeled data.
[0078] As a specific implementation, the high-quality annotation library may store confidence levels corresponding to the annotated data. Specifically, the confidence levels stored in the high-quality annotation library may be obtained based on expert experience or a base model or a large model.
[0079] In the embodiment of this specification, the labeled data in the high-quality labeling library can be used to label the data to be labeled.
[0080] Specifically, labeled data whose similarity with the data to be labeled is greater than or equal to a second preset threshold can be retrieved from the high-quality labeling library to obtain several target labeled data, and then the obtained several target labeled data can be used to label the data to be labeled.
[0081] For example, feature vectors can be extracted from the labeled data in a high-quality database, such as image feature vectors from labeled data in image format, text feature vectors from labeled data in text format, audio feature vectors from labeled data in speech format, and video feature vectors from labeled data in video format. The similarity between the feature vectors of the labeled data and the feature vectors of the data to be labeled is then calculated, for example, by calculating cosine similarity, Euclidean distance, Manhattan distance, etc. This allows for obtaining a number of target labeled data whose similarity to the data to be labeled is greater than or equal to a second preset threshold.
[0082] As a specific implementation method, in order to improve the retrieval efficiency of labeled data and thus improve the efficiency of labeling the data to be labeled, feature vectors can be extracted from the labeled data in advance, and the extracted feature vectors can be stored in a high-quality database. Optionally, the high-quality labeling library in the embodiment of this specification can store feature vectors corresponding to the labeled data. In this way, when the data to be labeled needs to be labeled later, the stored feature vectors can be directly used to calculate the similarity between the feature vectors of the data to be labeled, without the need to extract feature vectors from the labeled data, thereby improving the labeling efficiency.
[0083] In the embodiment of this specification, the obtained plurality of target labeled data may be one target labeled data or a plurality of target labeled data.
[0084] As a specific implementation, if several target labeled data are one target labeled data, the target labeled data can be used to label the data to be labeled. For example, the labeling information of the target labeled data is determined as the labeling information of the data to be labeled, thereby obtaining labeled data containing the labeling information.
[0085] As a specific implementation method, if the several target labeled data are multiple target labeled data, then one target labeled data can be selected from the multiple target labeled data, such as randomly selecting one target labeled data, or selecting the target labeled data with the highest confidence, or selecting the target labeled data with the highest similarity to the data to be labeled, and then determining the labeling information of the selected target labeled data as the labeling information of the data to be labeled to obtain the labeled data containing the labeling information.
[0086] Alternatively, the large model can be used to determine the labeling information for the data to be labeled based on the labeling information of several target labeled data. For example, several target labeled data can be used as reference examples for the large model, and prompt words can be constructed to instruct the large model to label the data to be labeled based on the reference examples. The prompt words are then input into the large model, and the large model outputs labeled data containing the labeling information.
[0087] In an embodiment of the present specification, the annotation processing module may include a module for annotating the data to be annotated using a high-quality annotation library, wherein the high-quality annotation library may be pre-set, thereby improving data annotation efficiency by annotating the data to be annotated using the pre-set high-quality annotation library.
[0088] In addition, the high-quality annotation library stores annotated data with a confidence level higher than a first preset threshold, where the first preset threshold can be 80%, 90%, 95%, etc., so that the confidence level of the annotated data in the high-quality annotation library is higher, and then the annotated data in the high-quality annotation library is used to annotate the data to be annotated, which can improve the data quality of the annotated data.
[0089] Furthermore, in an embodiment of the present specification, the similarity between the labeled data retrieved from the high-quality database and the data to be labeled is greater than or equal to a second preset threshold, so that the similarity between the retrieved labeled data and the data to be labeled is relatively high, thereby using the labeling information of the retrieved labeled data to label the data to be labeled, which can further improve the data quality of the labeled data.
[0090] In practical applications, in order to improve the data quality of the labeled data, it is also possible to calculate the acceptable values of several target labeled data, and then use the target labeled data with higher acceptable values as acceptable data, and use the acceptable data to label the data to be labeled. Optionally, the labeling process of the data to be labeled based on the labeling information of the several target labeled data to obtain the labeled data containing the labeling information may specifically include:
[0091] Based on the respective weight factors of the plurality of target labeled data, the acceptable values of the plurality of target labeled data are calculated; the weight factors include at least one of a time decay factor, a model iteration version factor, a labeling specification iteration version factor, and a confidence factor.
[0092] The target labeled data whose adoptable value is greater than or equal to the third preset threshold is determined as adoptable data; there is at least one piece of adoptable data.
[0093] Based on the labeling information of the adoptable data, labeling information of the data to be labeled is determined.
[0094] The data to be labeled is labeled using the labeling information to obtain labeled data containing the labeling information.
[0095] In the embodiments of this specification, the high-quality annotation library may also store the annotation time of the annotated data or the storage time of the annotated data in the high-quality annotation library, the annotation rules of the annotated data, the confidence level of the annotated data, etc. Optionally, if the annotation information of the annotated data in the high-quality annotation library is obtained through a machine model, the quality annotation library may also store an iterative version of the machine model that outputs the annotation information of the annotated data.
[0096] Optionally, the time decay factor may represent the degree of influence of the labeled data on the acceptable value as time passes, specifically the degree of influence of the labeling time or storage time on the acceptable value. The model iteration version factor may represent the degree of influence of the iterative version of the machine model that outputs the labeling information of the labeled data on the acceptable value. The labeling specification iteration version factor may represent the degree of influence of the labeling rules of the labeling information of the labeled data on the acceptable value. The confidence factor may represent the degree of influence of the confidence of the machine model on the acceptable value.
[0097] In the embodiments of this specification, the acceptable value of the target labeled data can be calculated by various weight factors. For example, the longer the labeling time or storage time is from the current time, such as the time for calculating the acceptable value, the smaller the time decay factor is, and the lower the calculated acceptable value can be when other weight factors remain unchanged. The longer the formulation time of the iterative version of the machine model that outputs the labeling information of the labeled data is from the current time, the smaller the model iteration version factor is, and the lower the calculated acceptable value can be when other weight factors remain unchanged. The longer the formulation time of the labeling rules of the labeling information of the labeled data is from the current time, the smaller the labeling specification iteration version factor is, and the lower the calculated acceptable value can be when other weight factors remain unchanged. The higher the confidence of the machine model, the larger the confidence factor is, and the higher the calculated acceptable value can be when other weight factors remain unchanged.
[0098] As a specific implementation, the following formula can be used to calculate the acceptable value of the target labeled data:
[0099] Score=e -λΔt ×C×A×M×L
[0100] Among them, Score can represent the acceptable value of the target labeled data. -λΔtIt can represent the time decay factor. e can represent a natural constant, and λ can represent the decay coefficient. Specifically, λ can take values such as 0.01, 0.02, or 0.1. The value of λ can be determined based on actual needs and is not limited here. Δt can represent the number of days between the annotation time or storage time of the annotated data and the current time. C can represent the confidence level of the machine model, C∈[0,1]. In practical applications, the target annotated data can be obtained through manual data annotation, and A can represent the accuracy of the manual annotation, such as A∈[0.9,1]. M can represent a model iteration version factor, M∈(0,1], where M can be determined based on the iterative version of the machine model of the annotation information of the labeled data and the latest iterative version of the machine model, for example, it can be determined based on the ratio of the version number of the iterative version of the machine model of the annotation information of the labeled data and the version number of the latest iterative version of the machine model. L can represent a labeling specification iteration version factor, L∈(0,1], where L can be determined based on the labeling rule of the labeling information of the labeled data and the latest labeling rule, for example, it can be determined based on the ratio of the version number of the labeling rule of the labeling information of the labeled data and the version number of the latest labeling rule.
[0101] In the embodiment of this specification, the adoptable value may indicate the extent to which the target labeled data can be adopted as labeling information for determining the data to be labeled.
[0102] In practical applications, the acceptable value of the target labeled data can be represented by a specific decimal or percentage value between 0 and 1. Alternatively, other intervals, such as a value between 0 and 100, can be used to represent the acceptable value of the target labeled data. The specific representation of the acceptable value is not specifically limited here.
[0103] In the embodiment of the present specification, the target labeled data whose adoptable value is greater than or equal to the third preset threshold value can be determined as adoptable data. Specifically, the adoptable data can be one piece of adoptable data or multiple pieces of adoptable data.
[0104] As a specific implementation, if the adoptable data is a piece of adoptable data, the annotation information of the piece of adoptable data may be determined as the annotation information of the data to be annotated.
[0105] As a specific implementation method, if there are multiple pieces of acceptable data, one piece of acceptable data can be selected from the multiple pieces of acceptable data, such as randomly selecting one piece of acceptable data, or selecting the acceptable data with the highest confidence, or selecting the acceptable data with the largest acceptability value, and then the annotation information of the selected piece of acceptable data can be determined as the annotation information of the data to be labeled.
[0106] Optionally, the large model can be used to determine the labeling information of the data to be labeled based on the labeling information of the acceptable data. For example, the labeling information of one or more pieces of acceptable data can be used as reference examples for the large model. A prompt word can be constructed to instruct the large model to determine the labeling information of the data to be labeled based on the reference examples. The prompt word can then be input into the large model to obtain the labeling information of the data to be labeled as output by the large model.
[0107] In an embodiment of this specification, the high-quality annotation library may include high-quality annotation libraries for one or more business scenarios. For example, high-quality annotation libraries for computer vision scenarios, natural language processing scenarios, medical scenarios, financial scenarios, autonomous driving scenarios, educational scenarios, and other scenarios. In an embodiment of this specification, a target high-quality annotation library corresponding to the business scenario of the data to be annotated can be first determined, and then from the target high-quality annotation library, annotated data whose similarity with the data to be annotated is greater than or equal to a second preset threshold can be quickly and accurately retrieved, and then the annotated data to be annotated can be annotated quickly and accurately. Optionally, the method may also include:
[0108] Based on the business scenario to which the data to be annotated belongs, a target high-quality annotation library corresponding to the business scenario is determined.
[0109] Retrieving the labeled data whose similarity with the data to be labeled is greater than or equal to a second preset threshold from the high-quality labeling library may specifically include:
[0110] Retrieve labeled data from the target high-quality labeling library, the labeled data having a similarity with the data to be labeled that is greater than or equal to a second preset threshold.
[0111] In the embodiments of this specification, the target high-quality annotation library corresponding to the business scenario can be a target high-quality annotation library with the same business scenario as the data to be annotated. For example, if the data to be annotated is data for a financial scenario, the target high-quality annotation library corresponding to the data to be annotated can be a high-quality annotation library for financial scenarios.
[0112] Optionally, the high-quality annotation library may store business scenario information of the annotated data. The high-quality annotation library may be pre-divided into target high-quality annotation libraries for various business scenarios based on the business scenario information of the annotated data.
[0113] Alternatively, the high-quality annotation library may not store the business scenario information of the annotated data. For example, the data annotation platform may store a target high-quality annotation library and the correspondence between business scenarios and the target high-quality annotation library. Based on this correspondence, the target high-quality annotation library corresponding to the business scenario of the data to be annotated can be determined.
[0114] In practical applications, after determining the target high-quality annotation library corresponding to the business scenario, the labeled data can be retrieved from the target high-quality annotation library. For example, the similarity between the feature vectors of the labeled data in the target high-quality annotation library and the feature vectors of the data to be labeled can be calculated to retrieve the labeled data.
[0115] As a specific embodiment, if no labeled data with a similarity greater than or equal to a second preset threshold to the data to be labeled can be retrieved from the high-quality labeling library, the data to be labeled can be provided to other labeling processing modules for labeling. Other labeling processing modules may include modules for labeling the data to be labeled using customized labeling models, or modules for labeling the data to be labeled using trained base models, etc. In the embodiments of this specification, data labeling can be performed using multiple labeling processing modules to ensure labeling quality.
[0116] Optionally, a module that annotates the data to be annotated can use a high-quality annotation library to annotate the data. After obtaining the annotated data, the confidence level of the annotation information can be determined through expert experience or large models. If the confidence level of the annotation information is greater than or equal to a preset confidence threshold, the data to be annotated can be considered annotated. If the confidence level of the annotation information is less than the preset confidence threshold, the data to be annotated can be provided to other annotation processing modules for further annotation processing.
[0117] In the embodiments of this specification, the labeled data in the high-quality labeling library can also be updated regularly or irregularly. For example, the labeled data can be updated when the labeling time or storage time of the labeled data is a certain distance away from a certain time, and the labeled data can be updated when the labeling rules of the labeling information of the labeled data change. Specifically, the labeled data can be updated by using expert experience, large models, intelligent labeling iTAG modules, etc. For example, the labeling information of the labeled data, the confidence level of the labeling information, the business scenario information of the labeled data, etc. can be updated.
[0118] The embodiments of this specification can ensure the data quality of the labeled data in the high-quality labeling library by updating the labeled data in the high-quality labeling library, thereby improving the data quality of the labeled data obtained by labeling the data to be labeled based on the high-quality labeling library.
[0119] In practical applications, the labeling processing module may include a module for labeling the data to be labeled using a customized labeling model. The customized labeling model includes a labeling model trained using training data of the same labeling type as the data to be labeled. Optionally, labeling the data to be labeled using the labeling processing module to obtain labeled data containing labeling information may specifically include:
[0120] The data to be annotated is input into the customized annotation model to obtain annotation information output by the customized annotation model.
[0121] The labeling information is determined as the labeling information of the data to be labeled, and the labeled data is obtained.
[0122] Among them, a customized annotation model can refer to a machine learning model specially trained according to specific business needs, specific business scenarios, or specific data. In the embodiments of this specification, the customized annotation model can be a machine learning model trained using training data that is consistent with the annotation type of the data to be annotated. Specifically, the customized annotation model can be used to output annotation information for the data to be annotated of the annotation type to which its training samples belong.
[0123] In practical applications, the training data used to train the customized annotation model can include labeled data. Specifically, the labeled data can be obtained using the aforementioned method for obtaining labeled data from the high-quality annotation library. Alternatively, the labeled data from the high-quality annotation library can be directly used as the labeled data for training the customized annotation model.
[0124] Furthermore, based on the labeled data of various labeling types, labeling models corresponding to the labeled data of various labeling types can be trained respectively. Alternatively, based on the labeled data of various labeling types, a labeling model capable of labeling multiple types of data can be trained.
[0125] Alternatively, the annotation model can be trained using the self-training (SFT) method in the semi-supervised learning method based on historical unlabeled data. The historical unlabeled data can be obtained from public datasets on the internet, scraped from social media, online forums, and other platforms using web crawlers, collected using crowdsourcing platforms, collected using image acquisition devices, audio acquisition devices, and various sensors, or set up using expert experience.
[0126] SFT self-training involves applying the model's predicted label information to unlabeled data to train the machine model. Specifically, during SFT self-training, the model predicts the label information for the unlabeled data and uses this predicted information as the label information for training the machine model. This approach allows model training using unlabeled data without the need for manual data labeling.
[0127] Optionally, the labeling model can be trained using the SFT self-training method based on the data to be labeled. For example, the labeling model can be trained by combining the data to be labeled with historical data to be labeled.
[0128] In the embodiments of this specification, the labeling model is trained through the SFT self-training method, and there is no need to perform data labeling, such as there is no need to use expert experience to label data, thereby saving the labeling model training time, improving the labeling model training efficiency, and further improving the data labeling efficiency.
[0129] In practical applications, the data to be labeled can be input into the customized labeling model to obtain the labeling information output by the customized labeling model. The labeling information output by the customized labeling model can then be determined as the labeling information of the data to be labeled, thereby obtaining the labeled data.
[0130] As a specific embodiment, the training data of the customized annotation model may also include the confidence level of the labeled data. The confidence level of the labeled data may be obtained based on expert experience or a base model or a large model. The customized annotation model may then output the confidence level of the labeled information of the labeled data. Optionally, the customized annotation model is further configured to output the confidence level corresponding to the labeled information. The method may further include:
[0131] Determine whether the confidence level corresponding to the annotation information is greater than a fourth preset threshold, and obtain a first judgment result.
[0132] Determining the annotation information as the annotation information of the data to be annotated may specifically include:
[0133] If the first judgment result indicates that the confidence level corresponding to the labeling information is greater than the fourth preset threshold, the labeling information is determined as the labeling information of the data to be labeled.
[0134] In the embodiment of this specification, the annotation information with a confidence level greater than a fourth preset threshold can be determined as the annotation information of the data to be annotated, thereby obtaining the annotated data, which can improve the data quality of the annotated data.
[0135] As a specific implementation manner, after determining whether the confidence level corresponding to the annotation information is greater than a fourth preset threshold and obtaining a first determination result, the method may further include:
[0136] If the first judgment result indicates that the confidence level corresponding to the labeling information is not greater than the fourth preset threshold, the data to be labeled is determined as unsuccessfully labeled data.
[0137] Optionally, the labeled information whose confidence level is not greater than a fourth preset threshold value may be determined as unsuccessfully labeled data. Unsuccessfully labeled data may indicate that the module for labeling the data to be labeled using the customized labeling model was unable to label the data, or may indicate that the module for labeling the data to be labeled using the customized labeling model had a poor labeling effect.
[0138] In the embodiment of this specification, in order to improve the data quality of the annotated data, the unsuccessfully annotated data can also be provided to other annotation processing modules for annotation processing. Optionally, the method can also include:
[0139] The unsuccessfully labeled data is provided to another labeling processing module. The other labeling processing module includes at least one of a first labeling processing module and a second labeling processing module. The first labeling processing module is configured to label the data to be labeled using a high-quality labeling library. The high-quality labeling library stores labeled data with a confidence level above a first preset threshold. The second labeling processing module is configured to label the data to be labeled using a trained base model.
[0140] In an embodiment of the present specification, a customized labeling model can be trained using training data that is consistent with the labeling type of the data to be labeled, so that the customized labeling model is targeted at the data to be labeled, and the customized labeling model can be used to accurately label the labeled data of the labeling type, thereby improving the data quality of the labeled data.
[0141] In the embodiment of this specification, the labeling processing module includes a module for labeling the data to be labeled using the trained base model. The labeling processing module is used to label the data to be labeled to obtain the labeled data containing the labeling information, which may specifically include:
[0142] Acquire a plurality of labeled data pieces whose labeling information belongs to the labeling type information.
[0143] The base model is trained using the labeled data to obtain a target model.
[0144] The data to be labeled is input into the target model to obtain the labeling information output by the target model.
[0145] The labeling information is determined as the labeling information of the data to be labeled, and the labeled data is obtained.
[0146] In related technologies, a base model is a general model pre-trained on large-scale data. The base model has extensive knowledge and general capabilities, and can be adapted to various specific tasks through fine-tuning.
[0147] The base model used in the embodiments of this specification can be any one of the GPT (Generative Pre-trained Transformer) model, the BERT (Bidirectional Encoder Representations from Transformers) model, the CLIP (Contrastive Language-Image Pre-training) model, the T5 (Text-to-Text Transfer Transformer) model, the ViT (Vision Transformer) model, the ChatGLM model, and the LLaMA model.
[0148] In the embodiments of this specification, a base model can be trained using a number of labeled data to obtain a trained base model, and the trained base model can then be used to output labeling information of the data to be labeled.
[0149] Optionally, the several pieces of labeled data used to train the base model can be obtained by the aforementioned method of obtaining labeled data from the high-quality labeled library.
[0150] Optionally, the several pieces of labeled data used to train the base model can also be directly obtained from a high-quality annotation library. Optionally, the obtained labeled information belongs to the several pieces of labeled data of the annotation type information, specifically including:
[0151] Based on the annotation type information, a plurality of annotated data having annotation information belonging to the annotation type information is retrieved from a high-quality annotation library, wherein the high-quality annotation library stores annotated data having a confidence level higher than a first preset threshold.
[0152] In practical applications, the high-quality annotation library may also store annotation type information corresponding to the annotated data, so that based on the annotation type information stored in the high-quality annotation library, several annotated data with annotation information belonging to the annotation type information can be quickly retrieved.
[0153] As a specific implementation, the several pieces of labeled data used to train the base model may also be obtained from the training data used to train the customized labeled model.
[0154] Optionally, the acquired pieces of annotated information may be one piece of annotated information or multiple pieces of annotated information, which is not specifically limited here.
[0155] In the embodiments of this specification, using a plurality of labeled data to train the base model may be a process of fine-tuning the base model.
[0156] Among them, fine-tuning is a secondary training of a pre-trained general model such as a base model to improve the model's processing performance for data in a specific field.
[0157] In the embodiment of this specification, the data to be labeled is input into the target model, specifically, a prompt word is constructed to instruct the target model to output the labeling information of the data to be labeled, and then the constructed prompt word is input into the target model.
[0158] In an embodiment of this specification, the prompt words constructed to instruct the target model to output the labeling information of the data to be labeled may include information indicating the confidence level of the target model output, wherein the confidence level may be the confidence level corresponding to the labeling information of the data to be labeled output by the target model. The target model may then output the confidence level corresponding to the labeling information. Optionally, the target model is further configured to output the confidence level corresponding to the labeling information. The method may further include:
[0159] It is determined whether the confidence level corresponding to the annotation information is greater than a fifth preset threshold to obtain a second determination result.
[0160] Determining the annotation information as the annotation information of the data to be annotated may specifically include:
[0161] If the second judgment result indicates that the confidence level corresponding to the labeling information is greater than the fifth preset threshold, the labeling information is determined as the labeling information of the data to be labeled.
[0162] In the embodiment of this specification, it is possible to determine whether the confidence level corresponding to the annotation information is greater than the fifth preset threshold, and then the annotation information with a confidence level greater than the fourth preset threshold can be determined as the annotation information of the data to be annotated, thereby obtaining the annotated data, which can improve the data quality of the annotated data.
[0163] As a specific implementation, if the second judgment result indicates that the confidence level corresponding to the annotation information is not greater than the fifth preset threshold, the data to be annotated may be provided to other annotation processing modules for annotation processing. For example, the data to be annotated may be provided to a module for annotating the data to be annotated using a high-quality annotation library, or to a module for annotating the data to be annotated using a customized annotation model.
[0164] In practical applications, to improve the performance of the target model, the base model can be trained multiple times using the labeled data to obtain multiple trained base models, where each training can produce a trained base model. The model with the better performance among the multiple trained base models can then be used as the target model. Optionally, the training of the base model using the labeled data to obtain the target model may specifically include:
[0165] The base model is trained using the labeled data to obtain a plurality of trained base models.
[0166] The labeled data is used to verify the accuracy of the plurality of trained base models to obtain verification results.
[0167] The trained base model with an accuracy greater than or equal to a sixth preset threshold in the verification result is determined as the target model.
[0168] In the embodiments of this specification, the base model is trained using labeled data to obtain several trained base models. This can be accomplished by performing multiple rounds of iterative training on a base model using labeled data to obtain multiple trained base models corresponding to the multiple rounds of iterative training.
[0169] As a specific implementation method, the base model is trained using labeled data to obtain several trained base models. Alternatively, the labeled data can be used to train multiple different base models to obtain multiple trained base data corresponding to the multiple base models. For example, the GPT model, BERT model, CLIP model, T5 model, ViT model, ChatGLM model, and LLaMA model can be trained using labeled data to obtain multiple trained base data corresponding to each of these base models.
[0170] In practical applications, the labeled data can be pre-divided into a training dataset and a validation dataset. The base model is trained using the labeled data, specifically using the training dataset. The accuracy of several trained base models is verified using the labeled data, specifically using the validation dataset.
[0171] Optionally, if the trained base model with an accuracy greater than or equal to the sixth preset threshold is a trained base model, this trained base model can be used as the target model.
[0172] Optionally, if there are multiple trained base models whose accuracy is greater than or equal to the sixth preset threshold, a model can be randomly selected from these multiple trained base models as the target model, or the model with the highest accuracy among these multiple trained base models can be used as the target model, and so on.
[0173] Optionally, if there is no trained base model with an accuracy greater than or equal to the sixth preset threshold, it can be considered that the module for labeling the data to be labeled using the trained base model has a poor labeling effect. In this case, the data to be labeled can be provided to other labeling processing modules for labeling. For example, the data to be labeled can be provided to a module for labeling the data to be labeled using a high-quality labeling library, or provided to a module for labeling the data to be labeled using a customized labeling model.
[0174] In an embodiment of the present specification, the target model can be trained using several pieces of labeled data that are consistent with the labeling type information of the data to be labeled, so that the target model is targeted at the data to be labeled, and the target model can be used to accurately label the labeled data of the labeling type information, thereby improving the data quality of the labeled data.
[0175] In the embodiment of this specification, if the confidence level of the annotation information annotated by the annotation processing module corresponding to the annotation type information of the data to be annotated is low, other modules can be used for annotation processing to ensure the data quality of the annotated data.
[0176] As a specific implementation method, if the confidence of the annotation information annotated by each annotation processing module on the data to be annotated is less than or equal to the preset threshold, or the confidence of the annotation information of part of the data to be annotated is less than or equal to the preset threshold, manual annotation processing can also be used to process the data to be annotated or part of the data, so as to ensure the data quality of the annotated data.
[0177] In practical applications, in order to increase the amount and quality of data in the high-quality annotation library, and thereby improve the quality of the annotated data, in an embodiment of this specification, the annotated data with a confidence level higher than a first preset threshold can be stored in the high-quality annotation library. Optionally, the annotation information of the annotated data corresponds to a confidence level. After the annotated data to be annotated is annotated using the annotation processing module to obtain the annotated data containing the annotation information, the method may further include:
[0178] The annotated data having a confidence level higher than a first preset threshold is written into a high-quality annotated database. The high-quality annotated database stores annotated data having a confidence level higher than the first preset threshold.
[0179] In the embodiment of this specification, the confidence level of the annotation information of the annotated data can be obtained through expert experience or a large model, etc. Thus, the annotated data with a confidence level higher than a first preset threshold can be written into the high-quality annotation library.
[0180] In practical applications, in order to improve the accuracy and efficiency of labeling the data to be labeled, the data to be labeled may also be pre-processed. Optionally, before the labeling processing module labels the data to be labeled and obtains the labeled data containing the labeling information, the following steps may also be performed:
[0181] The data to be labeled is preprocessed to obtain preprocessed data to be labeled; the data preprocessing includes at least one of data standardization processing, data cleaning processing and data enhancement processing.
[0182] The labeling processing module is used to label the data to be labeled to obtain labeled data containing labeling information, which may specifically include:
[0183] The labeling processing module is used to perform labeling processing on the pre-processed data to be labeled, so as to obtain labeled data containing labeling information.
[0184] In the embodiment of the present specification, data preprocessing is performed on the data to be labeled, which may be at least one of data standardization, data cleaning, and data enhancement.
[0185] Data standardization involves transforming data and mapping it to the same dimension to make different features comparable. Common data standardization methods include Z-Score and Min-Max methods.
[0186] Data cleaning refers to a data processing method that improves data quality by removing noise, erroneous data, or missing data. For example, it can fill missing data, delete erroneous data, remove duplicate data, and unify data formats.
[0187] Data augmentation refers to a data processing method that expands data to increase its volume or diversity. For example, image data can be rotated, scaled, or color-shifted; text data can be augmented with synonyms, random word insertions, deletions, or word swaps; and audio data can be augmented with background noise or pitch adjustments.
[0188] Optionally, the data to be labeled in the embodiment of this specification may be one piece of data to be labeled, or may be multiple pieces of data to be labeled, which is not limited here.
[0189] In the embodiments of this specification, by preprocessing the data to be labeled, the data quality of the data to be labeled can be effectively improved, and the labeling processing module is used to label the preprocessed data to be labeled, which can improve the data quality of the obtained labeled data.
[0190] Figure 3 This is a flow chart of a data labeling method provided in one embodiment of this specification. Figure 3 As shown, the data annotation method may include the following steps.
[0191] Step 302: Obtain data to be labeled.
[0192] In the embodiments of this specification, the data to be annotated may include data from one or more business scenarios, such as computer vision scenarios, natural language processing scenarios, medical scenarios, financial scenarios, autonomous driving scenarios, education scenarios, and recommendation scenarios.
[0193] Optionally, the data to be annotated may include data in one or more formats, such as data in one or more business scenarios in text format, image format, voice format, and video format.
[0194] In practical applications, the data to be annotated can be multiple pieces of data. For example, the data to be annotated may include 200 pieces of text-formatted data for computer vision scenarios, 200 pieces of speech-formatted data for computer vision scenarios, 100 pieces of text-formatted data for natural language processing scenarios, 100 pieces of image-formatted data for medical scenarios, 100 pieces of text-formatted data for financial scenarios, 300 pieces of video-formatted data for autonomous driving scenarios, and so on.
[0195] Step 304: Data preprocessing.
[0196] In the embodiments of this specification, data preprocessing may be performed on the data to be labeled, such as performing at least one of data standardization, data cleaning, and data enhancement on the data to be labeled.
[0197] By preprocessing the data to be labeled, the labeling accuracy and efficiency can be improved.
[0198] In practical applications, data preprocessing and other operations may not be performed. For example, after obtaining the data to be labeled, the data labeling platform may directly label the data to be labeled, such as directly executing subsequent steps 306, 308, or 310. Step 304 may be omitted.
[0199] Step 306: Use the first annotation processing module to perform annotation processing on the data to be annotated.
[0200] In the embodiment of the present specification, the annotation processing module corresponding to the annotation type information can be determined according to the annotation type information of the data to be annotated, so that the annotation processing module corresponding to the annotation type information can be used to perform annotation processing on the data to be annotated.
[0201] In the embodiment of this specification, since the determined annotation processing module is targeted at the data to be annotated, the data quality of the annotated data can be improved by using the annotation processing module corresponding to the annotation type information to perform annotation processing on the data to be annotated.
[0202] As an implementation manner, the first annotation processing module may be a module for performing annotation processing on the data to be annotated using a high-quality annotation library.
[0203] Optionally, if the high-quality annotation database contains annotated data with a confidence level higher than a first preset threshold, the data to be annotated can be annotated by retrieving from the high-quality annotation database a number of target annotated data with a similarity greater than or equal to a second preset threshold. The specific annotation process can be found in the previous section and will not be detailed here.
[0204] As a specific implementation method, in order to improve data labeling efficiency and the data quality of the labeled data, a target high-quality labeling library corresponding to the business scenario to which the data to be labeled belongs can also be used for data labeling processing.
[0205] As a specific implementation, to improve the quality of the annotated data, the acceptable values of the target annotated data can be calculated based on their respective weight factors. The target annotated data with higher acceptable values can then be used as acceptable data, and the unannotated data can be annotated using the acceptable data. The weight factors can include at least one of a time decay factor, a model iteration factor, a labeling specification iteration factor, and a confidence factor.
[0206] In the embodiments of this specification, high-quality annotation libraries for different business scenarios can be used to retrieve target annotated data, improving both the efficiency of data annotation and the quality of the annotated data. Furthermore, acceptable data can be determined based on weighting factors such as time decay, model iteration factor, annotation specification iteration factor, and confidence factor, ensuring the accuracy and recall of the target annotated data.
[0207] In actual applications, if the acquired data to be labeled is not completely processed by the first labeling processing module, for example, after the first labeling processing module is used to label multiple data to be labeled, the confidence of the labeling information of part of the data to be labeled is greater than a preset threshold, and the confidence of the labeling information of another part of the data to be labeled is less than or equal to the preset threshold, the other part of the data to be labeled can be called data that has not been successfully labeled by the first labeling processing module, and the second labeling processing module can be used to continue the labeling processing, and the data labeling platform can execute step 308.
[0208] Step 308: Use the second annotation processing module to perform annotation processing on the data to be annotated.
[0209] In the embodiments of this specification, the second annotation processing module is used to perform annotation processing on the data to be annotated. The second annotation processing module may be used to perform annotation processing on the data that was not successfully annotated by the first annotation processing module, or the second annotation processing module may be used to perform annotation processing on the data to be annotated whose annotation type information corresponds to the second annotation processing module.
[0210] In actual applications, there may be multiple annotation processing modules corresponding to the annotation type information of the data to be annotated. The embodiments of this specification can select any one of these multiple annotation processing modules to perform annotation processing on the data to be annotated, such as randomly selecting a annotation processing module, or selecting a relatively idle annotation processing module for annotation processing. If there is unsuccessfully annotated data after the selected annotation processing module performs annotation processing on the data to be annotated, the remaining annotation processing modules corresponding to the annotation type information of the data to be annotated can be used to perform annotation processing on the unsuccessfully annotated data. Optionally, the second annotation processing module can be a annotation processing module corresponding to the annotation type information of the data that was unsuccessfully annotated by the first annotation processing module.
[0211] As a specific implementation, if there is unsuccessfully labeled data, the unsuccessfully labeled data may also be provided to any other labeling processing module for processing. Optionally, the second labeling processing module may be a labeling processing module different from the first labeling processing module.
[0212] As an implementation manner, the second annotation processing module may be a module for performing annotation processing on the data to be annotated using a customized annotation model.
[0213] Optionally, the customized labeling model can be trained using training data that is consistent with the labeling type of the data to be labeled. The customized labeling model can be used to output labeling information of the data to be labeled of the labeling type to which its training samples belong.
[0214] In the embodiments of this specification, by training a customized annotation model using training data that is consistent with the annotation type of the data to be annotated, the generalization and accuracy of the customized annotation model can be improved.
[0215] The data to be annotated can be input into the customized annotation module to obtain the annotation information output by the customized annotation module, thereby performing annotation processing on the data to be annotated.
[0216] In actual applications, if the acquired data to be labeled is not completely processed by the second labeling processing module, for example, after the second labeling processing module is used to label multiple data to be labeled, the confidence of the labeling information of part of the data to be labeled is greater than a preset threshold, and the confidence of the labeling information of another part of the data to be labeled is less than or equal to the preset threshold, the other part of the data to be labeled can be called data that has not been successfully labeled by the second labeling processing module, and the third labeling processing module can be used to continue the labeling processing, and the data labeling platform can execute step 310.
[0217] Step 310: Utilize the third annotation processing module to perform annotation processing on the data to be annotated.
[0218] In the embodiments of this specification, the third annotation processing module is used to perform annotation processing on the data to be annotated. The third annotation processing module can be used to perform annotation processing on the data that was not successfully annotated by the second annotation processing module, or the third annotation processing module can be used to perform annotation processing on the data to be annotated whose annotation type information corresponds to the third annotation processing module.
[0219] As a specific implementation, the third annotation processing module may be an annotation processing module corresponding to the annotation type information of the data that was not successfully annotated by the second annotation processing module.
[0220] As a specific implementation, the third annotation processing module may be a different annotation processing module from the second annotation processing module.
[0221] In the embodiment of this specification, the third labeling processing module may be a module for labeling the data to be labeled using the trained base model.
[0222] Optionally, a third annotation processing module is used to perform annotation processing on the data to be annotated. Specifically, a prompt word is constructed to instruct the target model to output the annotation information of the data to be annotated, and then the constructed prompt word is input into the trained base model to obtain the annotation information output by the trained base model, thereby annotating the data to be annotated.
[0223] As a specific implementation, to improve the performance of the trained base model, the base model can be trained multiple times using labeled data to obtain multiple trained base models, where each training cycle can produce a trained base model. The model with the best performance among the multiple trained base models can then be used as the trained base model.
[0224] In the embodiments of this specification, during the training of the base model, the strong reasoning capabilities of the base model can be utilized to merge labeled information of the same type or business scenario to identify the evaluation standard features for that type or business scenario. Simultaneously, the historical labeled data in the training samples can be verified, and the base model can be iteratively trained multiple times to select the model with better performance as the trained base model. Furthermore, the base model can be trained using a small number of training samples, rather than a large number of training samples, and the trained base model can be applied to unlabeled data from multiple business scenarios.
[0225] In the embodiments of this specification, each annotation processing module assigns a confidence level to the annotated data obtained after processing the data to be annotated. Therefore, when the confidence level of the annotated data is less than or equal to a preset threshold, the data to be annotated can be submitted as unsuccessfully annotated data to another annotation processing module, such as another annotation processing module, another two annotation processing modules, or manual annotation processing, thereby improving the quality of the annotated data.
[0226] In practical applications, if the first labeling processing module completely processes the acquired data to be labeled, for example, if after labeling multiple pieces of data to be labeled using the first labeling processing module, no labeling information with a confidence level less than or equal to a preset threshold exists, then the second and third labeling processing modules may not be required to label the data to be labeled, and steps 308 and 310 may be omitted. Similarly, if after labeling multiple pieces of data to be labeled using the second labeling processing module, no labeling information with a confidence level less than or equal to a preset threshold exists, then the third labeling processing module may not be required to label the data to be labeled, and step 310 may be omitted.
[0227] As a specific implementation, if the data to be annotated includes data with multiple annotation types, multiple annotation processing modules corresponding to the data with multiple annotation types can be used to perform data annotation processing in parallel.
[0228] For example, if 300 pieces of data to be annotated are obtained, of which 200 pieces have annotation type information corresponding to the first annotation processing module and 100 pieces have annotation type information corresponding to the third annotation processing module, the first annotation processing module can be used to annotate the 200 pieces of data, while the third annotation processing module can be used to annotate the 100 pieces of data.
[0229] For example, if 500 pieces of data to be annotated are obtained, of which 200 pieces have annotation type information corresponding to the first annotation processing module, 120 pieces have annotation type information corresponding to the second annotation processing module, and 180 pieces have annotation type information corresponding to the third annotation processing module, the first annotation processing module can be used to annotate the 200 pieces of data, the second annotation processing module can be used to annotate the 120 pieces of data, and the third annotation processing module can be used to annotate the 180 pieces of data.
[0230] It can be understood that the above embodiment is described by taking the first labeling processing module as a module for labeling the labeled data using a high-quality labeling library, the second labeling processing module as a module for labeling the labeled data using a customized labeling model, and the third labeling processing module as a module for labeling the labeled data using a trained base model as an example. In actual applications, the first labeling processing module, the second labeling processing module, and the third labeling processing module can be any one of a module for labeling the labeled data using a high-quality labeling library, a module for labeling the labeled data using a customized labeling model, and a module for labeling the labeled data using a trained base model. For example, in addition to the above-described situation, as another embodiment, the first labeling processing module can be a module for labeling the labeled data using a customized labeling model, the second labeling processing module can be a module for labeling the labeled data using a high-quality labeling library, and the third labeling processing module can be a module for labeling the labeled data using a trained base model. Alternatively, as another embodiment, the first labeling processing module may be a module for labeling the labeled data using a customized labeling model, the second labeling processing module may be a module for labeling the labeled data using a trained base model, and the third labeling processing module may be a module for labeling the labeled data using a high-quality labeling library. Alternatively, as another embodiment, the first labeling processing module may be a module for labeling the labeled data using a trained base model, the second labeling processing module may be a module for labeling the labeled data using a high-quality labeling library, and the third labeling processing module may be a module for labeling the labeled data using a customized labeling model. Alternatively, as another embodiment, the first labeling processing module may be a module for labeling the labeled data using a trained base model, the second labeling processing module may be a module for labeling the labeled data using a customized labeling model, and the third labeling processing module may be a module for labeling the labeled data using a high-quality labeling library, and so on. The specific module functions can be set according to actual needs and are not specifically limited here.
[0231] In the embodiments of this specification, there is no limitation on the order in which steps 306, 308, and 310 are executed. For example, step 308 may be executed first, then step 306, and then step 310. Alternatively, step 310 may be executed first, then step 306, and then step 308. Alternatively, steps 306, 308, and 310 may be executed simultaneously.
[0232] In the embodiments of this specification, when the amount of data to be labeled is large, for example, when there are hundreds, thousands or even tens of thousands of data items to be labeled, multiple labeling processing modules can be combined for labeling processing, and each labeling processing module labels and processes the data to be processed with the labeling type information corresponding to each labeling processing module, thereby enabling efficient and accurate data labeling processing.
[0233] Step 312: Obtain labeled data.
[0234] The first data annotation module, the second annotation processing module, and the third annotation processing module perform annotation processing on the data to be annotated, thereby obtaining annotated data containing annotation information.
[0235] The obtained labeled data may be labeled data whose confidence level of the labeled information is greater than a certain confidence level threshold.
[0236] Step 314: Write the labeled data into a high-quality labeling library.
[0237] Since the obtained labeled data can be labeled data whose confidence in the labeled information is greater than a certain confidence threshold, the labeled data can be written into a high-quality labeling library, thereby increasing the data volume and data quality of the labeled data in the high-quality labeling library, thereby improving the data labeling quality.
[0238] In actual applications, after obtaining the labeled data, the process can be terminated directly. The above step 314 can be omitted.
[0239] In the examples of this specification, a method for data annotation using a high-quality annotation library is proposed. Acceptable data can be determined based on weighting factors such as time decay, model iteration, annotation specification iteration, and confidence. This ensures the accuracy and recall of the target annotated data, enabling accurate data annotation using a high-quality annotation library. Compared to manually annotating data or using a general-purpose base model, data annotation accuracy using a high-quality annotation library is improved by 20%.
[0240] In the embodiments of this specification, data annotation is performed through various annotation processing modules, such as a module for annotating data using a high-quality annotation library, a module for annotating data using a customized annotation model, and a module for annotating data using a trained base model. A large amount of data to be annotated in various business scenarios and formats can be annotated and processed, with dual guarantees in terms of annotation efficiency and annotation accuracy, breaking through the singleness of the industry's intelligent auxiliary annotation effect.
[0241] Among them, intelligent auxiliary labeling (intelligent assisted labeling) refers to the process of using artificial intelligence technology to assist or partially replace manual data labeling.
[0242] In the embodiments of this specification, the annotation information of the annotated data obtained after each annotation processing module annotates the data corresponds to a confidence level, and then when the confidence level is less than or equal to a preset threshold, the data to be annotated can be provided to other annotation processing modules or manually for annotation processing, thereby ensuring the data quality of the annotated data.
[0243] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.
[0244] Figure 4 The embodiments of this specification provide corresponding Figure 2 A structural diagram of a data annotation device. Figure 4 As shown, the device may include:
[0245] The labeled data acquisition module 402 is used to acquire the data to be labeled.
[0246] The annotation type acquisition module 404 is configured to acquire annotation type information for the data to be annotated. The annotation type information indicates the type of annotation information that can be annotated for the data to be annotated.
[0247] The determination module 406 is configured to determine, based on the annotation type information, an annotation processing module corresponding to the annotation type information. The annotation processing module has a data annotation function for annotating the annotation information corresponding to the annotation type information.
[0248] The processing module 408 is configured to perform labeling processing on the data to be labeled using the labeling processing module to obtain labeled data containing labeling information.
[0249] based on Figure 4 The present specification also provides some specific implementation plans of the device, which are described below.
[0250] Optionally, the annotation processing module includes a module for performing annotation processing on the data to be annotated using a high-quality annotation library, wherein the high-quality annotation library stores annotated data with a confidence level higher than a first preset threshold.
[0251] The processing module 408 may be specifically configured to:
[0252] Retrieve labeled data whose similarity with the data to be labeled is greater than or equal to a second preset threshold from the high-quality labeling library to obtain a plurality of target labeled data.
[0253] Based on the labeling information of the plurality of target labeled data, labeling processing is performed on the data to be labeled to obtain labeled data containing the labeling information.
[0254] Optionally, the device may further include:
[0255] The annotation library determination module is used to determine a target high-quality annotation library corresponding to the business scenario based on the business scenario to which the data to be annotated belongs.
[0256] Retrieving the labeled data whose similarity with the data to be labeled is greater than or equal to a second preset threshold from the high-quality labeling library may specifically include:
[0257] Retrieve labeled data from the target high-quality labeling library, the labeled data having a similarity with the data to be labeled that is greater than or equal to a second preset threshold.
[0258] Optionally, the labeling process of the data to be labeled based on the labeling information of the plurality of target labeled data to obtain labeled data containing the labeling information may specifically include:
[0259] Based on the respective weight factors of the plurality of target labeled data, the acceptable values of the plurality of target labeled data are calculated; the weight factors include at least one of a time decay factor, a model iteration version factor, a labeling specification iteration version factor, and a confidence factor.
[0260] The target labeled data whose adoptable value is greater than or equal to the third preset threshold is determined as adoptable data; there is at least one piece of adoptable data.
[0261] Based on the labeling information of the adoptable data, labeling information of the data to be labeled is determined.
[0262] The data to be labeled is labeled using the labeling information to obtain labeled data containing the labeling information.
[0263] Optionally, the labeling processing module includes a module for labeling the data to be labeled using a customized labeling model. The customized labeling model includes a labeling model trained using training data consistent with the labeling type of the data to be labeled.
[0264] The processing module 408 may be specifically configured to:
[0265] The data to be annotated is input into the customized annotation model to obtain annotation information output by the customized annotation model.
[0266] The labeling information is determined as the labeling information of the data to be labeled, and the labeled data is obtained.
[0267] Optionally, the customized annotation model is further configured to output a confidence level corresponding to the annotation information. The apparatus may further include:
[0268] The first judgment module is used to judge whether the confidence level corresponding to the annotation information is greater than a fourth preset threshold value, and obtain a first judgment result.
[0269] Determining the annotation information as the annotation information of the data to be annotated may specifically include:
[0270] If the first judgment result indicates that the confidence level corresponding to the labeling information is greater than the fourth preset threshold, the labeling information is determined as the labeling information of the data to be labeled.
[0271] Optionally, after obtaining the first judgment result, determining whether the confidence level corresponding to the annotation information is greater than a fourth preset threshold may further include:
[0272] If the first judgment result indicates that the confidence level corresponding to the labeling information is not greater than the fourth preset threshold, the data to be labeled is determined as unsuccessfully labeled data.
[0273] Optionally, the device may further include:
[0274] A data providing module is configured to provide the unsuccessfully labeled data to other labeling processing modules. The other labeling processing modules include at least one of a first labeling processing module and a second labeling processing module. The first labeling processing module is configured to label the data to be labeled using a high-quality labeling library. The high-quality labeling library stores labeled data with a confidence level above a first preset threshold. The second labeling processing module is configured to label the data to be labeled using a trained base model.
[0275] Optionally, the labeling processing module includes a module for labeling the data to be labeled using the trained base model. The processing module 408 can be specifically used to:
[0276] Acquire a plurality of labeled data pieces whose labeling information belongs to the labeling type information.
[0277] The base model is trained using the labeled data to obtain a target model.
[0278] The data to be labeled is input into the target model to obtain the labeling information output by the target model.
[0279] The labeling information is determined as the labeling information of the data to be labeled, and the labeled data is obtained.
[0280] Optionally, the target model is further configured to output a confidence level corresponding to the annotation information. The apparatus may further include:
[0281] The second judgment module is used to judge whether the confidence corresponding to the annotation information is greater than a fifth preset threshold, and obtain a second judgment result.
[0282] Determining the annotation information as the annotation information of the data to be annotated may specifically include:
[0283] If the second judgment result indicates that the confidence level corresponding to the labeling information is greater than the fifth preset threshold, the labeling information is determined as the labeling information of the data to be labeled.
[0284] Optionally, the acquiring of the annotation information belonging to the plurality of annotated data of the annotation type information may specifically include:
[0285] Based on the annotation type information, a plurality of annotated data whose annotation information belongs to the annotation type information is retrieved from a high-quality annotation library; the high-quality annotation library stores annotated data with a confidence level higher than a first preset threshold.
[0286] Optionally, the training of the base model using the labeled data to obtain a target model may specifically include:
[0287] The base model is trained using the labeled data to obtain a plurality of trained base models.
[0288] The labeled data is used to verify the accuracy of the plurality of trained base models to obtain verification results.
[0289] The trained base model with an accuracy greater than or equal to a sixth preset threshold in the verification result is determined as the target model.
[0290] Optionally, the device may further include:
[0291] The data preprocessing module is used to perform data preprocessing on the data to be labeled to obtain preprocessed data to be labeled. The data preprocessing includes at least one of data standardization, data cleaning, and data enhancement.
[0292] The processing module 408 may be specifically configured to:
[0293] The labeling processing module is used to perform labeling processing on the pre-processed data to be labeled, so as to obtain labeled data containing labeling information.
[0294] Optionally, the labeled data has a corresponding confidence level. The device may further include:
[0295] The data writing module is configured to write the annotated data having a confidence level higher than a first preset threshold into a high-quality annotated database, wherein the high-quality annotated database stores annotated data having a confidence level higher than the first preset threshold.
[0296] Based on the same idea, the embodiments of this specification also provide devices corresponding to the above methods.
[0297] Figure 5 The embodiments of this specification provide corresponding Figure 2 A structural diagram of a data annotation device. Figure 5 As shown, the device 500 may include:
[0298] at least one processor 510; and,
[0299] A memory 530 in communication with the at least one processor; wherein,
[0300] The memory 530 stores instructions 520 that can be executed by the at least one processor 510. The instructions are executed by the at least one processor 510 to enable the at least one processor 510 to implement the above-mentioned data labeling method.
[0301] Based on the same idea, the embodiments of this specification also provide a computer-readable medium corresponding to the above method. The computer-readable medium stores computer-readable instructions, which can be executed by a processor to implement the above data labeling method.
[0302] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The devices, equipment and methods provided in the embodiments of this specification correspond to each other, so the devices and equipment also have beneficial technical effects similar to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding devices and equipment will not be repeated here.
[0303] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures such as diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always program the improved process flow into the hardware circuit to obtain the corresponding hardware circuit structure. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is an integrated circuit whose logical function is determined by user programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0304] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0305] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0306] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0307] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0308] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0309] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0310] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0311] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0312] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0313] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0314] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0315] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0316] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A data annotation method, comprising: Obtain the data to be labeled; Obtaining annotation type information for the data to be annotated; The annotation type information indicates the type of annotation information that can be annotated with the data to be annotated; Determining, according to the annotation type information, an annotation processing module corresponding to the annotation type information; The annotation processing module has a data annotation function of annotating the annotation information corresponding to the annotation type information; The labeling processing module is used to perform labeling processing on the data to be labeled to obtain labeled data containing labeling information.
2. The method according to claim 1, wherein the annotation processing module comprises a module for annotating the data to be annotated using a high-quality annotation library; the high-quality annotation library stores annotated data with a confidence level higher than a first preset threshold; The processing module is used to perform labeling processing on the data to be labeled to obtain labeled data containing labeling information, specifically including: Retrieving labeled data whose similarity to the data to be labeled is greater than or equal to a second preset threshold from the high-quality labeling library to obtain a plurality of target labeled data; Based on the labeling information of the plurality of target labeled data, labeling processing is performed on the data to be labeled to obtain labeled data containing the labeling information.
3. The method of claim 2, further comprising: Based on the business scenario to which the data to be annotated belongs, determining a target high-quality annotation library corresponding to the business scenario; Retrieving the labeled data whose similarity with the data to be labeled is greater than or equal to a second preset threshold from the high-quality labeling library specifically includes: Retrieve labeled data from the target high-quality labeling library, the labeled data having a similarity with the data to be labeled that is greater than or equal to a second preset threshold.
4. The method according to claim 2, wherein the labeling process of the data to be labeled based on the labeling information of the plurality of target labeled data to obtain labeled data containing the labeling information specifically comprises: Calculating the adoptable values of the plurality of target labeled data based on respective weight factors of the plurality of target labeled data; The weight factor includes at least one of a time decay factor, a model iteration version factor, a labeling specification iteration version factor, and a confidence factor; Determining the target labeled data, wherein the adoptable value is greater than or equal to a third preset threshold, as adoptable data; wherein there is at least one piece of adoptable data; Determining the labeling information of the data to be labeled based on the labeling information of the adoptable data; The data to be labeled is labeled using the labeling information to obtain labeled data containing the labeling information.
5. The method of claim 1, wherein the labeling processing module comprises a module for labeling the data to be labeled using a customized labeling model; the customized labeling model comprises a labeling model trained using training data consistent with the labeling type of the data to be labeled; The tagging processing module is used to tag the data to be tagged to obtain tagged data containing tagging information, specifically including: Inputting the data to be annotated into the customized annotation model to obtain annotation information output by the customized annotation model; The labeling information is determined as the labeling information of the data to be labeled, and the labeled data is obtained.
6. The method of claim 5, wherein the customized annotation model is further configured to output a confidence level corresponding to the annotation information; the method further comprising: Determine whether the confidence level corresponding to the annotation information is greater than a fourth preset threshold, and obtain a first determination result; Determining the labeling information as the labeling information of the data to be labeled specifically includes: If the first judgment result indicates that the confidence level corresponding to the labeling information is greater than the fourth preset threshold, the labeling information is determined as the labeling information of the data to be labeled.
7. The method according to claim 6, wherein after determining whether the confidence level corresponding to the annotation information is greater than a fourth preset threshold, obtaining the first determination result, further comprising: If the first judgment result indicates that the confidence level corresponding to the labeling information is not greater than the fourth preset threshold, the data to be labeled is determined as unsuccessfully labeled data.
8. The method of claim 7, further comprising: Providing the unsuccessfully labeled data to other labeling processing modules; The other annotation processing modules include at least one of a first annotation processing module and a second annotation processing module; the first annotation processing module is used to perform annotation processing on the data to be annotated using a high-quality annotation library; the high-quality annotation library stores annotated data with a confidence level higher than a first preset threshold; The second labeling processing module is used to perform labeling processing on the data to be labeled using the trained base model.
9. The method according to claim 1, wherein the labeling processing module includes a module for labeling the data to be labeled using the trained base model; labeling the data to be labeled using the labeling processing module to obtain labeled data containing labeling information specifically includes: Acquire several pieces of annotated data whose annotation information belongs to the annotation type information; Using the labeled data to train the base model to obtain a target model; Inputting the data to be labeled into the target model to obtain labeling information output by the target model; The labeling information is determined as the labeling information of the data to be labeled, and the labeled data is obtained.
10. The method according to claim 9, wherein the target model is further configured to output a confidence level corresponding to the annotation information; the method further comprising: Determine whether the confidence level corresponding to the annotation information is greater than a fifth preset threshold, and obtain a second determination result; Determining the labeling information as the labeling information of the data to be labeled specifically includes: If the second judgment result indicates that the confidence level corresponding to the labeling information is greater than the fifth preset threshold, the labeling information is determined as the labeling information of the data to be labeled.
11. The method according to claim 9, wherein obtaining the labeled data of the plurality of labeled data items belonging to the labeled type information specifically comprises: Based on the annotation type information, a plurality of annotated data pieces whose annotation information belongs to the annotation type information are retrieved from a high-quality annotation library; The high-quality annotation library stores annotated data with a confidence level higher than a first preset threshold.
12. The method according to claim 9, wherein the training of the base model using the labeled data to obtain a target model comprises: Using the labeled data to train the base model to obtain a plurality of trained base models; Verifying the accuracy of the plurality of trained base models using the labeled data to obtain verification results; The trained base model with an accuracy greater than or equal to a sixth preset threshold in the verification result is determined as the target model.
13. The method according to claim 1, before using the annotation processing module to annotate the data to be annotated to obtain annotated data containing annotation information, further comprising: Performing data preprocessing on the data to be labeled to obtain preprocessed data to be labeled; The data preprocessing includes at least one of data standardization, data cleaning and data enhancement; The tagging processing module is used to tag the data to be tagged to obtain tagged data containing tagging information, specifically including: The labeling processing module is used to perform labeling processing on the pre-processed data to be labeled, so as to obtain labeled data containing labeling information.
14. The method according to any one of claims 1 to 13, wherein the labeling information of the labeled data corresponds to a confidence level; and after labeling the data to be labeled using the labeling processing module to obtain the labeled data containing the labeling information, the method further comprises: Writing the annotated data whose confidence level is higher than a first preset threshold into a high-quality annotation library; The high-quality annotation library stores annotated data with a confidence level higher than the first preset threshold.
15. A data annotation device, comprising: The labeling data acquisition module is used to obtain the data to be labeled; A labeling type acquisition module, configured to acquire labeling type information for the data to be labeled; The annotation type information indicates the type of annotation information that can be annotated with the data to be annotated; a determination module, configured to determine, based on the annotation type information, an annotation processing module corresponding to the annotation type information; The annotation processing module has a data annotation function of annotating the annotation information corresponding to the annotation type information; The processing module is used to perform labeling processing on the data to be labeled using the labeling processing module to obtain labeled data containing labeling information.
16. A data annotation device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the data labeling method according to any one of claims 1 to 14.
17. A computer-readable medium having computer-readable instructions stored thereon, wherein the computer-readable instructions can be executed by a processor to implement the data labeling method according to any one of claims 1 to 14.
Citation Information
Cited By
Electronic equipment and traffic data label generation method
CN121350255A
Electronic device and traffic data tag generation method
CN121350255B