Data labeling method and device, electronic equipment and storage medium
By using large language models for preliminary and target annotation in Internet advertising data labeling, the problem of insufficient number and accuracy of manual labeling data sets is solved, and the effectiveness of the advertising recommendation model and advertising conversion rate are improved.
Patent Information
- Application Number
- CN202311675623.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, relying on manual labeling data has led to the number and accuracy of labeling data sets that need to be improved, and it is difficult to meet the needs of understanding user behavior information in Internet advertising.
A data labeling method is proposed. By obtaining the data set to be marked and the preliminary prompt words, using a large language model for preliminary annotation, building the target prompt words based on the preliminary annotation results, further target annotation, and improving the accuracy of data labeling.
By building different prompt words at different marking stages, the accuracy of data annotation is improved, the effectiveness of the advertising recommendation model is improved, and the advertising conversion rate and platform revenue are improved.
Smart Images

Figure CN120125288A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a data annotation method, device, electronic device, and storage medium. Background Art
[0002] As the Internet increasingly penetrates into our lives, Internet advertising has gradually become a trend in the advertising industry. In Internet advertising, it is necessary to recommend advertisements based on various different user behavior information to achieve personalized recommendations for each user. Optionally, advertisement recommendations can be achieved by constructing models such as recall and ranking. Among them, the understanding of user behavior information requires the use of natural language processing technology to label different data sources, especially the understanding of advertisement data sources is particularly important. For advertisement data sources, they can be labeled according to information such as advertisement copywriting and material pictures, and then provided as features to the recall and ranking models, which can effectively improve the performance of the algorithm model. In related data annotation methods, data annotation generally relies on manual labor, resulting in room for improvement in the quantity and accuracy of the annotated data set. Summary of the Invention
[0003] This application proposes a data annotation method, device, electronic device, and storage medium.
[0004] In a first aspect, an embodiment of this application provides a data annotation method, the method including: obtaining a data set to be annotated and a first prompt word, where the data set to be annotated includes a plurality of advertisement copywriting data obtained after normalization processing, and the first prompt word is used to perform preliminary annotation on the plurality of advertisement copywriting data; inputting the data set to be annotated and the first prompt word into a annotation model to obtain a preliminary annotation result corresponding to the data set to be annotated output by the annotation model, where the annotation model is a large language model; based on the preliminary annotation result, constructing a target prompt word, where the target prompt word is used to indicate annotating the plurality of advertisement copywriting data according to the label meaning of the annotation label; inputting the target prompt word and the data set to be annotated into the annotation model to obtain a target annotation result corresponding to the data set to be annotated output by the annotation model.
[0005] Second aspect, an embodiment of the present application provides a data annotation device, the device includes: a data acquisition unit, configured to acquire a dataset to be annotated and a first prompt word, where the dataset to be annotated includes a plurality of advertisement copy data obtained after normalization processing, and the first prompt word is used to perform preliminary annotation on the plurality of advertisement copy data; a first output unit, configured to input the dataset to be annotated and the first prompt word into an annotation model, and obtain a preliminary annotation result corresponding to the dataset to be annotated output by the annotation model, where the annotation model is a large language model; a prompt word construction unit, configured to construct a target prompt word based on the preliminary annotation result, where the target prompt word is used to indicate annotating the plurality of advertisement copy data according to the label meaning of the annotation label; a second output unit, configured to input the target prompt word and the dataset to be annotated into the annotation model, and obtain a target annotation result corresponding to the dataset to be annotated output by the annotation model.
[0006] Third aspect, an embodiment of the present application provides an electronic device, including one or more processors and a memory; one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above-mentioned method.
[0007] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, and when the program code runs, the above-mentioned method is executed.
[0008] An embodiment of the present application provides a data annotation method, device, electronic device and storage medium. First, a dataset to be annotated and a first prompt word are acquired. The dataset to be annotated includes a plurality of advertisement copy data obtained after normalization processing, and the first prompt word is used to perform preliminary annotation on the plurality of advertisement copy data. Then, the dataset to be annotated and the first prompt word are input into an annotation model, and a preliminary annotation result corresponding to the dataset to be annotated output by the annotation model is obtained. The annotation model is a large language model. Then, a target prompt word is constructed based on the preliminary annotation result. The target prompt word is used to indicate annotating the plurality of advertisement copy data according to the label meaning of the annotation label. Finally, the target prompt word and the dataset to be annotated are input into the annotation model, and a target annotation result corresponding to the dataset to be annotated output by the annotation model is obtained. Through the above method, different prompt words are constructed at different marking stages to annotate the advertisement copy data in the dataset to be annotated, which can improve the accuracy of data annotation. Description of the Drawings
[0009] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0010] Figure 1 The flowchart of a data annotation method proposed in an embodiment of the present application is shown;
[0011] Figure 2 The flowchart of a data annotation method proposed in another embodiment of the present application is shown;
[0012] Figure 3 The flowchart of a data annotation method proposed in yet another embodiment of the present application is shown;
[0013] Figure 4 The structural block diagram of a data annotation device proposed in an embodiment of the present application is shown;
[0014] Figure 5 The structural block diagram of a data annotation device proposed in an embodiment of the present application is shown;
[0015] Figure 6 The structural block diagram of an electronic device for executing the data annotation method according to the embodiments of the present application in the embodiments of the present application is shown;
[0016] Figure 7 The storage unit for storing or carrying the program code for implementing the data annotation method according to the embodiments of the present application in the embodiments of the present application is shown. Detailed implementation manners
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0018] As the Internet increasingly penetrates into our lives, Internet advertising has gradually become a trend in the advertising industry. In Internet advertising, it is necessary to recommend advertisements based on various different user behavior information to achieve personalized recommendations for each user. Here, user behavior information may include search behavior logs, browsing news, clicking on apps (applications), or advertisements, etc. Optionally, advertisement recommendations can be achieved by constructing models such as recall and ranking. And the understanding of user behavior information requires the use of natural language processing technology to label different data sources, especially the understanding of advertisement data sources is particularly important. For advertisement data sources, they can be labeled according to information such as advertisement copywriting and material pictures, and then provided as features to the recall and ranking models, which can effectively improve the performance of the algorithm models. For advertisement copywriting data, text classification tasks are performed using natural language processing technology. The quality of the text classification results will directly affect the quality of the recall and ranking model results, and further affect the conversion effect of advertisements, which has a direct impact on the efficiency and return on investment of advertisers in placing advertisements. At the same time, the quality of the text classification results will also affect the revenue of the platform side. If the classification is inaccurate, it will affect the user's advertisement experience, reduce the stickiness between users and the platform side, and thus reduce the platform's revenue. Therefore, performing text classification tasks on different data sources such as user behavior data and logs is very important and will directly affect the benefits of the platform side, advertisers, and users.
[0019] In the research on relevant data annotation methods, the inventors found that generally relying on manual labor to annotate data, the quantity and accuracy of the annotated data sets still need to be improved.
[0020] Therefore, the inventors proposed the data annotation method, device, electronic device, and storage medium in this application. First, obtain the data set to be annotated and the first prompt word. The data set to be annotated includes multiple advertisement copywriting data obtained after normalization processing. The first prompt word is used to perform preliminary annotation on the multiple advertisement copywriting data. Then, input the data set to be annotated and the first prompt word into the annotation model to obtain the preliminary annotation result corresponding to the data set to be annotated output by the annotation model. The annotation model is a large language model. Then, based on the preliminary annotation result, construct a target prompt word. The target prompt word is used to indicate annotating the multiple advertisement copywriting data according to the label meaning of the annotation label. Finally, input the target prompt word and the data set to be annotated into the annotation model to obtain the target annotation result corresponding to the data set to be annotated output by the annotation model. Through the above method, at different labeling stages, different prompt words are constructed to annotate the advertisement copywriting data in the data set to be annotated, which can improve the accuracy of data annotation.
[0021] The following will specifically describe the embodiments of this application with reference to the accompanying drawings.
[0022] Please refer to Figure 1, a data annotation method provided by an embodiment of the present application, is applied to an electronic device. The method includes:
[0023] Step S110: Obtain a dataset to be annotated and a first prompt. The dataset to be annotated includes multiple advertisement copy data obtained after normalization processing. The first prompt is used to perform preliminary annotation on the multiple advertisement copy data.
[0024] In an embodiment of the present application, the dataset to be annotated is a dataset that needs to be annotated and obtained. The dataset to be annotated includes multiple different advertisement copy data. Among them, the dataset to be annotated can be obtained from a preset database. The multiple different advertisement copy data can be multiple advertisement copy data obtained after normalization processing. Among them, normalization processing refers to clustering multiple advertisement copy data with a semantic similarity greater than or equal to a preset similarity into one advertisement copy data. Among them, the advertisement copy data is mainly in text form, which is expressed in natural language and can be text in any language.
[0025] The first prompt is a pre-constructed prompt used to instruct the annotation model to perform preliminary annotation on the advertisement copy data in the dataset to be annotated based on the tagging process of keywords, inference process, and reflection. Among them, the first prompt clarifies the role and task of the annotation model, and places certain restrictions on the requirements for tagging. For example, it can restrict the annotation model to annotate the advertisement copy data in the dataset to be annotated in two dimensions: the technique routine and the advertisement copy content. Among them, the technique routine refers to the routine techniques used to catch the user's attention when generating advertisement copy data. Further, the first prompt also proposes a tagging process based on keywords, inference process, and reflection. Specifically, the tagging process of keywords, inference process, and reflection can be: first, extract keywords, and use the ability of the annotation model to extract keywords to obtain keyword information such as entity words included in the advertisement copy data as a judgment basis; second, perform an inference process according to the extracted keywords by simulating the thinking ability of humans; third, label the advertisement copy data according to the keywords and inference process; fourth, reflect on the tagging result, and optimize the tagging result by combining the previous keywords and inference process. Through this tagging process of keywords - inference process and reflection, the annotation model can be guided to perform the tagging process step by step.
[0026] Exemplarily, the constructed first prompt can be: "You are a professional data annotation expert, and your task is to assign a classification label to each given sentence.
[0027] These texts are all advertisement copy for promotion and have their particularities. Therefore, pay attention to tagging from the following two aspects:
[0028] 1. In terms of gimmicks and routines: To attract clicks, attention-grabbing gimmicks and routines are often deliberately used. Please identify and label them accordingly.
[0029] These labels have nothing to do with the specific content, but rather focus on inducing users to click on the advertisement from the perspective of the psychology of the advertising audience.
[0030] 2. In terms of specific content: Apply some labels to the specific content. These labels must be the mainstream classification labels for advertising copy.
[0031] Please think step by step, reason carefully, and pay attention to the following three points:
[0032] 1. All labels should have specific and clear meanings, about 4 characters in length. Try to avoid using broad 2-character labels as much as possible.
[0033] 2. For gimmick and routine labels, you can either leave them blank (output "None") or select only one. For content labels, you can either leave them blank (output "None") or select at most two.
[0034] Label the advertising copy step by step according to the following process.
[0035] Step 1: Find relevant clue words (i.e., keywords, phrases, context information, semantic meanings, semantic relationships, tones, reference materials).
[0036] Step 2: Deduce the diagnostic reasoning process from the existing clue words and the input.
[0037] Step 3: Apply at most the classification labels based on the input text, clues, and reasoning.
[0038] Step 4: Consider whether the labeled results are reasonable.
[0039] Now, please label the following advertising copy according to the above specifications:
[0040] ”.
[0041] In the first hint word above, the role and task of the annotation model are first determined, telling the annotation model that it is expected to act as a data annotation expert to label the advertising copy data and imposing certain restrictions on the labeling requirements. Here, the advertising copy data can be labeled in two major dimensions. The first dimension is the gimmick and routine labels. Because in the advertising scenario, different routine gimmicks are often used to capture users' attention. Here, it is hoped that the annotation model can come up with some gimmick and routine labels on its own. The second dimension is the advertising copy content labels. Here, corresponding labels are applied according to the specific content of the novel advertising copy. It is also hoped that the annotation model can use the vast amount of knowledge it has learned to apply these content labels.
[0042] Then, based on the chain of thought approach, it is proposed that the annotation model can perform the annotation process step by step according to the keyword - reasoning process and reflection. The first step is to extract keywords, using the ability of the annotation model to extract keywords to obtain keyword information such as entity words in the text as the judgment basis; the second step is to simulate the human thinking ability according to the extracted keywords to carry out the reasoning process; the third step is to assign labels according to the keywords and the reasoning process; the fourth step is to reflect on the annotation results and optimize the annotation results by combining the previous keywords and the reasoning process. Through this keyword - reasoning process and reflection annotation process, it is possible to guide GPT4 to perform the annotation process step by step.
[0043] Through the first prompt constructed above, the ability of the annotation model can be used to actually annotate the multiple advertising copy data obtained after normalization.
[0044] Step S120: Input the dataset to be annotated and the first prompt into the annotation model, and obtain the preliminary annotation result corresponding to the dataset to be annotated output by the annotation model. The annotation model is a large language model.
[0045] In the embodiment of the present application, the annotation model can be the GPT - 4 (Generative Pre - trained Transformer) model in the Large Language Model (LLM).
[0046] As a way, the GPT - 4 interface can be called to call the annotation model. After obtaining the dataset to be annotated and the first prompt, the dataset to be annotated and the first prompt can be input into the annotation model. Thus, the annotation model can output the annotation labels of each advertising copy data in the dataset to be annotated based on the input first prompt, that is, obtain the preliminary annotation result corresponding to the dataset to be annotated. Among them, each advertising copy data can correspond to at least one initial annotation label.
[0047] In the embodiment of the present application, a preliminary label system can be constructed based on the preliminary annotation result. The preliminary label system can be a label system composed of all the annotation labels in the preliminary annotation result; or it can also be a label system composed of some of the annotation labels in the preliminary annotation result, and no specific limitation is made here.
[0048] Step S130: Based on the preliminary annotation result, construct a target prompt, where the target prompt is used to indicate annotating the multiple advertising copy data according to the label meaning of the annotation labels.
[0049] In the embodiments of the present application, after obtaining the preliminary annotation result, a target prompt word can be constructed based on the preliminary annotation result. Among them, the target prompt word can include the label meaning of each annotation label. Further, the target prompt word can also include inference precautions and the output form of the annotation label, etc.
[0050] At the same time, some examples can also be sampled and put into the prompt word, so that the annotation model can further understand the task through few-shot, and finally improve the labeling result.
[0051] Step S140: Input the target prompt word and the dataset to be annotated into the annotation model, and obtain the target annotation result corresponding to the dataset to be annotated output by the annotation model.
[0052] In the embodiments of the present application, the target annotation result corresponding to the dataset to be annotated can be understood as a set of target annotation labels corresponding to each advertisement copy data in the dataset to be annotated. Among them, each advertisement copy data can correspond to at least one target annotation label. Among them, the target annotation label can be understood as the final classification label of the advertisement copy data, that is, the target annotation label is used to represent the category to which each advertisement copy data belongs.
[0053] After obtaining the target prompt word, the target prompt word and the dataset to be annotated can be input into the annotation model again, so that the annotation model can re-annotate multiple advertisement copy data in the dataset to be annotated according to the target prompt word, and obtain the target annotation result corresponding to each advertisement copy data in the dataset to be annotated, and obtain the annotated dataset. Among them, the annotated dataset is a set of multiple advertisement copy data annotated with target annotation labels.
[0054] A data annotation method provided by the present application first obtains a dataset to be annotated and a first prompt word. The dataset to be annotated includes multiple advertisement copy data obtained after normalization processing. The first prompt word is used to perform preliminary annotation on the multiple advertisement copy data. Then, the dataset to be annotated and the first prompt word are input into the annotation model, and the preliminary annotation result corresponding to the dataset to be annotated output by the annotation model is obtained. The annotation model is a large language model. Then, based on the preliminary annotation result, a target prompt word is constructed. The target prompt word is used to indicate the annotation of the multiple advertisement copy data according to the label meaning of the annotation label. Finally, the target prompt word and the dataset to be annotated are input into the annotation model, and the target annotation result corresponding to the dataset to be annotated output by the annotation model is obtained. Through the above method, different prompt words are constructed at different labeling stages to annotate the advertisement copy data in the dataset to be annotated, which can improve the accuracy of data annotation.
[0055] Please refer to Figure 2, a data annotation method provided by an embodiment of the present application is applied to an electronic device. The method includes:
[0056] Step S210: Obtain a dataset to be annotated and a first prompt. The dataset to be annotated includes multiple advertisement copy data obtained after normalization processing. The first prompt is used to perform preliminary annotation on the multiple advertisement copy data.
[0057] Step S220: Input the dataset to be annotated and the first prompt into a annotation model, and obtain a preliminary annotation result corresponding to the dataset to be annotated output by the annotation model. The annotation model is a large language model.
[0058] Step S230: Based on the preliminary annotation result, construct a second prompt. The second prompt is used to indicate clustering processing of the annotation labels in the preliminary annotation result.
[0059] In an embodiment of the present application, the second prompt can be a prompt constructed based on the preliminary annotation result, and the task of the annotation model is clearly defined in the second prompt.
[0060] Clustering processing refers to sorting and merging all the annotation labels included in the preliminary marking result.
[0061] When constructing the second prompt based on the preliminary annotation result, duplicate removal processing can be performed on the annotation labels included in the preliminary annotation result to obtain a preliminary annotation result after duplicate removal (i.e., a preliminary label system), and then the annotation labels included in the preliminary annotation result after duplicate removal can be added to the prompt to obtain the second prompt.
[0062] Exemplarily, taking content labels as an example, the obtained preliminary annotation result may include hundreds of content labels. Here, it is desired to control the label system within about twenty content labels, and there is no overlap between the content labels. At this time, a second prompt can be constructed to perform clustering operations on these hundreds of content labels through the annotation model. At this time, the second prompt can be: "The number of labels in the following label system is a bit large. Please merge similar labels:
[0063] Label 1, Label 2, Label 3, and Label 4......"
[0064] Step S240: Input the preliminary annotation result and the second prompt into the annotation model, and obtain a clustering annotation result corresponding to the preliminary annotation result output by the annotation model.
[0065] In an embodiment of the present application, after constructing the second prompt, the preliminary annotation result and the second prompt can be input into the annotation model again, so that the annotation model can perform a clustering operation on the annotation labels included in the preliminary annotation result to obtain a clustered annotation result.
[0066] Exemplarily, if the second prompt is the above-mentioned prompt, then after inputting the second prompt and the preliminary annotation result into the annotation model together, the obtained clustered annotation result output by the annotation model can be:
[0067] According to the given label system, after sorting and merging, the labels can be classified as follows:
[0068] 1. Category 1:
[0069] - Label A: Label 1, Label 2,... etc.
[0070] 2. Category 2:
[0071] - Label B: Label 6, Label 7,... etc.
[0072] 3. Category 3:
[0073] - Label C: Label 8, Label 9,... etc. ...
[0075] 18. Category 18:
[0076] - Label R: Label 19, Label 20,... etc.
[0077] After performing a clustering operation on the preliminary annotation result through the annotation model, hundreds of content labels can be initially clustered into eighteen content labels.
[0078] Step S250: Construct a target prompt based on the clustered annotation result.
[0079] As a method, the constructing a target prompt based on the clustered annotation result includes: performing an optimization process on the clustered annotation result based on annotation accuracy and label coverage rate to obtain an optimized clustered annotation result; inputting a third prompt and the optimized clustered annotation result into the annotation model to obtain the label meanings of each annotation label in the optimized clustered annotation result output by the annotation model, where the third prompt is used to indicate the label meanings of each annotation label; constructing a target prompt based on the label meanings of each annotation label in the optimized clustered annotation result.
[0080] At this time, the label system of the novel industry can be further determined from two aspects: coverage and accuracy. In terms of accuracy, mainly sample some examples to check the accuracy and stability of the labeling results of the labeling model; in terms of coverage, it is hoped that the labels given by the labeling model can cover a wide range of samples as much as possible, and focus on the rationality of those head labels (labels that appear frequently). For long-tail labels, such as the labels of the copywriting that only appears in a very small number of samples, they can be appropriately filtered to ensure that the number of labels in the label system meets the requirements and reduce the complexity.
[0081] Specifically, through the method of supervised learning, professional knowledge can be used to manually review the initially constructed eighteen content labels to optimize the results of the labeling model by the labeling model. Here, the optimization is mainly carried out from the following aspects:
[0082] (1) Delete labels with too large a scope. Some labels have too wide a scope, resulting in very blurred boundaries of the labels, and it is easy to have a very poor discrimination situation. Such labels need to be deleted;
[0083] (2) Appropriately merge labels with very fine granularity. Some labels focus on representing a certain type of theme label, and they have strong feature types, but because the granularity is very small, the labels are too small and need to be merged;
[0084] (3) Appropriately add some labels. Combining business knowledge, the labels missed by the labeling model need to be added.
[0085] Through the above methods, the clustering labeling results are optimized to obtain the optimized clustering labeling results, so that the final label system of the novel industry can be determined based on the optimized clustering labeling results. Among them, the final label system can include labels in two dimensions: advertising copy content labels and technique and routine labels. The advertising copy content labels can include fourteen labels such as label A, label B,...; the technique and routine labels can include five labels such as label H, label J,....
[0086] After determining the label system of the novel industry, the meaning of each labeled label in the label system can be given through the labeling model. Specifically, a third prompt word can be constructed, and this third prompt word is used to indicate the output of the meaning of each labeled label. Exemplarily, the third prompt word can be set as:
[0087] "According to the label system given below, please give the meaning of each labeled label in it:
[0088] Advertising copy content labels: label A, label B,...;
[0089] Technique and routine labels: label H, label J,...".
[0090] After inputting the aforementioned third prompt word and the optimized clustering annotation results into the annotation model, the output results obtained from the annotation model can be:
[0091] "According to the given label system, after sorting, the explanations of each annotation label are as follows:
[0092] Technique and routine label:
[0093] a) Label H: It is interpreted as Content 1.
[0094] b) Label J: It is interpreted as Content 2. ...
[0096] Advertising copy content label:
[0097] a) Label A: It is interpreted as Content 3.
[0098] b) Label B: It is interpreted as Content 4.
[0099] ..."
[0101] After obtaining the label meanings of each annotation label through the above method, the label meanings of each annotation label can also be optimized and corrected through manual review to obtain more accurate label meanings.
[0102] As a method, when constructing the target prompt word based on the label meanings of each annotation label in the optimized clustering annotation results, the label meanings of each annotation label can be added to the prompt word to facilitate the annotation model to understand the label meanings; at the same time, some examples are sampled and put into the prompt word, and the annotation model is allowed to further understand the task through the few-shot method to obtain the target prompt word.
[0103] Exemplarily, the constructed target prompt word can be:
[0104] "You are a professional data annotation expert, and your task is to assign classification labels to each given sentence.
[0105] These texts are all advertising copy for promoting the online novel APP, and they have their particularities, so pay attention to tagging from the following two aspects:
[0106] 1. In terms of techniques and routines: In order to attract clicks, attention-grabbing techniques and routines are often deliberately used. Please identify and assign the corresponding labels. The examples and explanations of this type of label are as follows:
[0107] a) Label H: It is interpreted as Content 1.
[0108] b) Label J: It is interpreted as Content 2. ...
[0110] In short, such tags have nothing to do with the specific content, but rather focus on inducing clicks on advertisements from the perspective of the psychology of the advertising audience.
[0111] 2. Regarding the specific content: Add some tags to the specific content. These tags must be the classification tags of mainstream online novels. Examples of such tags are as follows:
[0112] Tag A, Tag B, Tag C,.... etc.
[0113] Explanations of some of the tags are as follows:
[0114] a) Tag A: Interpreted as Content 3.
[0115] b) Tag B: Interpreted as Content 4. ...
[0117] Please think step by step, reason carefully, and at the same time pay attention to the following 3 points:
[0118] 1. The above two types of tags are just examples. You can add tags at any time, but do not add similar ones.
[0119] 2. All tags should have specific and clear meanings, about 4 characters in length, and try to avoid using broad 2-character words as much as possible.
[0120] 3. You can either not add (output "None") or add one tag of the technique and routine type, and you can either not add (output "None") or add at most two tags of the content type.
[0121] The following are several examples:
[0122] For the input
[0123] ```
[0124] 1. Text Content 1
[0125] 2. Text Content 2
[0126] 3. Text Content 3
[0127] ```
[0128] You need to output in the following table form (3 columns, separated by spaces between columns, and separated by commas between multiple tags):
[0129] ```
[0130] Serial Number Technique and Routine Tags Content Tags
[0131] 1 Tag H None
[0132] 2 Tag H None
[0133] 3 Tag J None
[0134] ```
[0135] Next, please tag the following advertising copy data according to the above specifications:
[0136] ```
[0137] %s
[0138] ```
[0139] ”.
[0140] Step S260: Input the target prompt word and the dataset to be labeled into the labeling model, and obtain the target labeling result corresponding to the dataset to be labeled output by the labeling model.
[0141] In the embodiment of the present application, the processes of steps S210 - S260 can also be used for the construction of label systems in other specific industries. In order to construct a label system for a specific industry and perform labeling classification, the rich knowledge of the GPT-4 large language model can be used, and appropriate prompt words can be constructed using prompt engineering to label the normalized advertising copy data, so as to determine the label system for the specific industry.
[0142] A data labeling method provided by the present application uses the rich knowledge of the large language model and prompt engineering technology to determine the label system, and uses the determined label system to label the dataset to be labeled, which can greatly improve the labeling efficiency, save labeling manpower, and reduce costs. At the same time, when constructing the target prompt word, letting the labeling model consider both the technique routine label and the advertising content label to label the advertising copy data in the dataset to be labeled can make the labeling model consider more comprehensively and the labeling result is more accurate.
[0143] Please refer to Figure 3 , a data labeling method provided by an embodiment of the present application is applied to an electronic device, and the method includes:
[0144] Step S310: Obtain a reference dataset, where the reference dataset includes multiple reference advertising copy data.
[0145] In the embodiment of the present application, it can be known that since the GPT-4 model needs to be called to label the advertising copy data, and the GPT-4 interface is charged according to tokens, that is to say, a certain cost will be generated for each call of the GPT-4 model. It is necessary to consider the cost issue. To improve the labeling efficiency under a limited number of calls, it is necessary to normalize the advertising copy data. For samples with particularly high semantic similarity, only one example needs to be selected for labeling, which can improve the labeling efficiency.
[0146] As a way, the reference data set is the obtained original advertisement copy data set. Among the multiple reference advertisement copy data included in the original advertisement copy data, there may be multiple similar reference advertisement copy data. Therefore, the reference data set can be normalized according to semantic similarity.
[0147] Step S320: Normalize the reference data set to obtain a data set to be labeled.
[0148] As a way, step S320 may specifically include: inputting the reference data set into a pre-trained text semantic matching model, and obtaining multiple semantic similarity data groups corresponding to the reference data set output by the text semantic matching model, where each of the semantic similarity data groups includes at least one reference advertisement copy data with semantic similarity; obtaining the data to be labeled corresponding to each of the multiple semantic similarity data groups, where the data to be labeled corresponding to each semantic similarity data group is any one reference advertisement copy data in each semantic similarity data group; and obtaining the data set to be labeled based on the data to be labeled corresponding to each of the multiple semantic similarity data groups.
[0149] In the embodiments of the present application, the text semantic matching model may be a Simbert model. The main principle of using the Simbert model to normalize the reference advertisement copy data in the reference data set is to use the Simbert model to convert the reference advertisement copy data into semantic vector embeddings, and these semantic vector embeddings have certain meanings in the vector space. When their semantic similarity is relatively high, the distance between the semantic vectors is relatively close, and when the similarity is relatively low, the distance between the semantic vectors is relatively far. According to this principle, multiple reference advertisement copy data in the reference data set can be clustered according to the vector similarity distance, so as to achieve the function of normalization.
[0150] Among them, the reference advertisement copy data included in the semantic similarity data group is at least one reference advertisement copy data with a semantic similarity greater than or equal to a preset similarity. Among them, the preset similarity is a similarity value preset to represent that the reference advertisement copy data is relatively similar.
[0151] In the embodiments of the present application, the Simbert model can be used to divide multiple reference advertisement copy data in the reference dataset into multiple semantically similar data groups. Among them, some semantically similar data groups may include multiple reference advertisement copy data. At this time, any one of the multiple reference advertisement copy data can be selected as the data to be labeled corresponding to the semantically similar data group. Exemplarily, a semantically similar data group includes 4 reference advertisement copy data, and these 4 reference advertisement copy data are very similar, with only individual punctuation marks and words that do not affect the semantics. If no normalization process is performed, all 4 reference advertisement copy data need to be labeled using the labeling model.
[0152] By using the Simbert model to perform a normalization operation on the above 4 reference advertisement copy data, these 4 reference advertisement copy data can be clustered into one advertisement copy according to semantic similarity (obtaining the data to be labeled corresponding to the semantically similar data group). Specifically, any one of the above 4 reference advertisement copy data can be selected to be labeled using the labeling model, which can greatly improve the labeling efficiency and also solve the cost.
[0153] In the embodiments of the present application, since there are a large number of cases where the multiple reference advertisement copy data in the reference dataset are semantically similar but have different text presentation forms, this normalization operation can effectively reduce the magnitude of the labeled data for advertisement copy. For example, there may originally be 100,000 advertisement copy, but after the normalization operation, only 60,000 advertisement copy may need to be labeled, effectively improving the labeling efficiency of the advertisement copy and reducing the cost.
[0154] Step S330: Obtain a dataset to be labeled and a first prompt. The dataset to be labeled includes multiple advertisement copy data obtained after normalization processing, and the first prompt is used to perform a preliminary label on the multiple advertisement copy data.
[0155] Step S340: Input the dataset to be labeled and the first prompt into the labeling model, and obtain the preliminary labeling result corresponding to the dataset to be labeled output by the labeling model. The labeling model is a large language model.
[0156] Step S350: Based on the preliminary labeling result, construct a target prompt, and the target prompt is used to indicate the labeling of the multiple advertisement copy data according to the label meaning of the labeling tag.
[0157] Step S360: Input the target prompt and the dataset to be labeled into the labeling model, and obtain the target labeling result corresponding to the dataset to be labeled output by the labeling model.
[0158] Step S370: Determine a training data set from the labeled data set based on the target labeling result.
[0159] In an embodiment of the present application, the labeled data set can be divided into a training data set and a test data set based on the target labeling result. Among them, when dividing the labeled data set into a training data set and a test data set based on the target labeling result, it can be divided according to a preset classification or according to a preset ratio, which is not specifically limited herein. Among them, the labeled data set at this time is a data set labeled with target labeling tags, and the advertising copy data included in the labeled data set is the same as the advertising copy data included in the data set to be labeled. The labeled data set is a data set obtained by labeling the data to be labeled.
[0160] Directly obtaining a training data set from the labeled data set labeled with target labeling tags can quickly obtain a high-quality training data set for training the model.
[0161] As a way, after step S370, it can also include: performing augmentation processing on the training data set to obtain a target training data set.
[0162] In an embodiment of the present application, the augmentation processing refers to generating advertising copy data similar to the advertising copy data in the training data set to obtain richer training data. Among them, the training data set can be subjected to anti-normalization processing to obtain a target training data set.
[0163] Step S380: Iteratively train the model to be trained based on the training data set until the training end condition is met, and obtain a target classification model.
[0164] In an embodiment of the present application, the model to be trained can be a pre-training + fine-tuning model such as BERT to construct a classification model. For example, the model to be trained can be a Roberta model.
[0165] The model to be trained can be iteratively trained through the training data set until the training end condition is met to obtain a target classification model, and then the target classification model can be verified through the test data set. Among them, the training end condition is a condition preset to represent the end of model training. For example, the training end condition can be set to the number of training times reaching a preset number; or, the training end condition can also be set to the loss function value being a preset value, etc., which is not specifically limited herein.
[0166] As a way, for the offline classification scenario, a pre-trained model with a relatively large number of Transformer layers can be used for classification. The more Transformer layers there are, the more complex the model is, and the better the model classification effect will be. For the online classification scenario, the model performance needs to be considered. To meet the real-time requirements, a 3-layer or 4-layer Transformer model can be used for classification. In this way, not only the classification effect of the model is relatively good, but also the online real-time performance can be satisfied.
[0167] As a way, step S380 may specifically include: iteratively training the to-be-trained model based on the target training data set until the training end condition is met, to obtain a target classification model.
[0168] A data annotation method provided by the present application iteratively trains a to-be-trained model through a training data set labeled with target annotation labels to obtain a target classification model. The target classification model can be used for the recall model and recommendation model in the downstream search and promotion scenarios, improving the advertising conversion effect and enhancing the stickiness between the platform and users.
[0169] Please refer to Figure 4 , a data annotation device 400 provided by an embodiment of the present application, the device 400 includes:
[0170] A data acquisition unit 410, configured to acquire a to-be-annotated data set and a first prompt word. The to-be-annotated data set includes a plurality of advertisement copy data obtained after normalization processing, and the first prompt word is used to perform preliminary annotation on the plurality of advertisement copy data.
[0171] As a way, the data acquisition unit 410 is specifically configured to acquire a reference data set, where the reference data set includes a plurality of reference advertisement copy data; perform normalization processing on the reference data set to obtain a to-be-annotated data set.
[0172] Further, the data acquisition unit 410 is specifically configured to input the reference data set into a pre-trained text semantic matching model, acquire a plurality of semantic similarity data groups corresponding to the reference data set output by the text semantic matching model, where each semantic similarity data group includes at least one reference advertisement copy data with semantic similarity; acquire the to-be-annotated data corresponding to each of the plurality of semantic similarity data groups, where the to-be-annotated data corresponding to each semantic similarity data group is any one reference advertisement copy data in each semantic similarity data group; and obtain the to-be-annotated data set based on the to-be-annotated data corresponding to each of the plurality of semantic similarity data groups.
[0173] The first output unit 420 is configured to input the dataset to be labeled and the first prompt into a labeling model, and obtain a preliminary labeling result corresponding to the dataset to be labeled output by the labeling model, where the labeling model is a large language model.
[0174] The prompt construction unit 430 is configured to construct a target prompt based on the preliminary labeling result, where the target prompt is used to indicate labeling the multiple advertising copy data according to the label meaning of the labeling label.
[0175] As a way, the prompt construction unit 430 is specifically configured to construct a second prompt based on the preliminary labeling result, where the second prompt is used to indicate clustering the labeling labels in the preliminary labeling result; input the preliminary labeling result and the second prompt into the labeling model, and obtain a clustering labeling result corresponding to the preliminary labeling result output by the labeling model; construct a target prompt based on the clustering labeling result.
[0176] Further, the prompt construction unit 430 is specifically configured to perform optimization processing on the clustering labeling result based on labeling accuracy and label coverage rate to obtain an optimized clustering labeling result; input a third prompt and the optimized clustering labeling result into the labeling model, and obtain the label meaning of each labeling label in the optimized clustering labeling result output by the labeling model, where the third prompt is used to indicate outputting the label meaning of each labeling label; construct a target prompt based on the label meaning of each labeling label in the optimized clustering labeling result.
[0177] The second output unit 440 is configured to input the target prompt and the dataset to be labeled into the labeling model, and obtain a target labeling result corresponding to the dataset to be labeled output by the labeling model.
[0178] Please refer to Figure 5 , the apparatus 400 further includes:
[0179] The model training unit 450 is configured to determine a training dataset from a labeled dataset based on the target labeling result; perform iterative training on a model to be trained based on the training dataset until a training end condition is met, and obtain a target classification model.
[0180] As a way, the model training unit 450 is specifically configured to perform enhancement processing on the training dataset to obtain a target training dataset; the performing iterative training on a model to be trained based on the training dataset until a training end condition is met, and obtaining a target classification model includes: performing iterative training on the model to be trained based on the target training dataset until a training end condition is met, and obtaining a target classification model.
[0181] It should be noted that the device embodiments in this application correspond to the foregoing method embodiments. For the specific principles in the device embodiments, reference may be made to the content in the foregoing method embodiments, which will not be elaborated here.
[0182] Next, Figure 6 an electronic device provided by this application will be described.
[0183] Please refer to Figure 6 , based on the above data annotation method and device, another electronic device 800 that can execute the foregoing data annotation method is further provided in the embodiments of this application. The electronic device 800 includes one or more (only one is shown in the figure) processors 802, a memory 804, and a network module 806 that are coupled to each other. Among them, a program that can execute the content in the foregoing embodiments is stored in the memory 804, and the processor 802 can execute the program stored in the memory 804.
[0184] Among them, the processor 802 may include one or more processing cores. The processor 802 connects various parts within the entire electronic device 800 through various interfaces and lines, and executes various functions of the electronic device 800 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 804, and by calling the data stored in the memory 804. Optionally, the processor 802 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 802 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the displayed content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 802 and may be implemented separately through a communication chip.
[0185] The memory 804 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. The memory 804 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 804 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the terminal 800 (such as a phone book, audio and video data, chat record data, etc.).
[0186] The network module 806 is used to receive and send electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an electronic device. The network module 806 may include various existing circuit elements for performing these functions. For example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a Subscriber Identity Module (SIM) card, a memory, and so on. The network module 806 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network, or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network, or a metropolitan area network. For example, the network module 806 can interact with a base station for information.
[0187] Please refer to Figure 7 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable storage medium 900, and the program code can be called by a processor to execute the method described in the above method embodiments.
[0188] The computer-readable storage medium 900 may be an electronic memory such as a flash memory, an Electrically Erasable Programmable Read-Only Memory (EEPROM), an EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 900 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 900 has a storage space for the program code 910 for executing any method step in the above method. These program codes can be read out from one or more computer program products or written into these one or more computer program products. The program code 910 can be compressed in an appropriate form, for example.
[0189] A data annotation method, device, electronic device, and storage medium provided by this application first obtain a dataset to be annotated and a first prompt. The dataset to be annotated includes multiple advertisement copy data obtained after normalization processing. The first prompt is used to perform preliminary annotation on the multiple advertisement copy data. Then, the dataset to be annotated and the first prompt are input into an annotation model, and a preliminary annotation result corresponding to the dataset to be annotated output by the annotation model is obtained. The annotation model is a large language model. Then, based on the preliminary annotation result, a target prompt is constructed. The target prompt is used to indicate annotating the multiple advertisement copy data according to the label meaning of the annotation label. Finally, the target prompt and the dataset to be annotated are input into the annotation model, and a target annotation result corresponding to the dataset to be annotated output by the annotation model is obtained. Through the above method, different prompts are constructed at different marking stages to annotate the advertisement copy data in the dataset to be annotated, which can improve the accuracy of data annotation.
[0190] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and all belong to the protection scope of the present invention.
Claims
1. A data annotation method, characterized in that, the method includes: Obtain a dataset to be annotated and a first prompt word. The dataset to be annotated includes multiple advertisement copy data obtained after normalization processing. The first prompt word is used to perform preliminary annotation on the multiple advertisement copy data; Input the dataset to be annotated and the first prompt word into an annotation model, and obtain the preliminary annotation result corresponding to the dataset to be annotated output by the annotation model. The annotation model is a large language model; Based on the preliminary annotation result, construct a target prompt word, where the target prompt word is used to indicate annotating the multiple advertisement copy data according to the label meaning of the annotation label; Input the target prompt word and the dataset to be annotated into the annotation model, and obtain the target annotation result corresponding to the dataset to be annotated output by the annotation model.
2. The method according to claim 1, characterized in that, the constructing a target prompt word based on the preliminary annotation result includes: Based on the preliminary annotation result, construct a second prompt word, where the second prompt word is used to indicate performing clustering processing on the annotation labels in the preliminary annotation result; Input the preliminary annotation result and the second prompt word into the annotation model, and obtain the clustering annotation result corresponding to the preliminary annotation result output by the annotation model; Based on the clustering annotation result, construct a target prompt word.
3. The method according to claim 2, characterized in that, the constructing a target prompt word based on the clustering annotation result includes: Based on annotation accuracy and label coverage, perform optimization processing on the clustering annotation result to obtain an optimized clustering annotation result; Input a third prompt word and the optimized clustering annotation result into the annotation model, and obtain the label meaning of each annotation label in the optimized clustering annotation result. Among them, the third prompt word is used to indicate outputting the label meaning of each annotation label; Based on the label meaning of each annotation label in the optimized clustering annotation result, construct a target prompt word.
4. The method according to claim 1, characterized in that, after inputting the target prompt word and the dataset to be annotated into the annotation model and obtaining the target annotation result corresponding to the dataset to be annotated output by the annotation model, it further includes: Based on the target annotation result, determine a training dataset from an annotation dataset; Iteratively train a model to be trained based on the training dataset until a training end condition is met, and obtain a target classification model.
5. The method according to claim 4, characterized in that, after determining the training dataset from the annotation dataset based on the target annotation result, it further includes: Perform augmentation processing on the training dataset to obtain a target training dataset; The iteratively training the model to be trained based on the training dataset until a training end condition is met and obtaining a target classification model includes: Iteratively train the model to be trained based on the target training dataset until a training end condition is met, and obtain a target classification model.
6. The method according to claim 1, wherein, the obtaining of the dataset to be labeled includes: obtaining a reference dataset, wherein the reference dataset includes a plurality of reference advertising copy data; performing normalization processing on the reference dataset to obtain the dataset to be labeled.
7. The method according to claim 6, wherein, the performing of normalization processing on the reference dataset to obtain the dataset to be labeled includes: inputting the reference dataset into a pre-trained text semantic matching model to obtain a plurality of semantic similarity data groups corresponding to the reference dataset output by the text semantic matching model, wherein each semantic similarity data group includes at least one reference advertising copy data with semantic similarity; obtaining the data to be labeled corresponding to each of the plurality of semantic similarity data groups, wherein the data to be labeled corresponding to each semantic similarity data group is any one of the reference advertising copy data in each semantic similarity data group; obtaining the dataset to be labeled based on the data to be labeled corresponding to each of the plurality of semantic similarity data groups.
8. A data annotation device, wherein, the device includes: a data acquisition unit for acquiring a dataset to be labeled and a first prompt word, wherein the dataset to be labeled includes a plurality of advertising copy data obtained after normalization processing, and the first prompt word is used for preliminarily annotating the plurality of advertising copy data; a first output unit for inputting the dataset to be labeled and the first prompt word into a labeling model to obtain a preliminary annotation result corresponding to the dataset to be labeled output by the labeling model, and the labeling model is a large language model; a prompt word construction unit for constructing a target prompt word based on the preliminary annotation result, and the target prompt word is used for instructing to annotate the plurality of advertising copy data according to the label meaning of the annotation label; a second output unit for inputting the target prompt word and the dataset to be labeled into the labeling model to obtain a target annotation result corresponding to the dataset to be labeled output by the labeling model.
9. An electronic device, wherein, it includes one or more processors; one or more programs are stored in the memory and configured to be executed by the one or more processors to perform the method according to any one of claims 1-7.
10. A computer-readable storage medium, wherein, program code is stored in the computer-readable storage medium, and when the program code is run by a processor, it performs the method according to any one of claims 1-7.
Citation Information
Cited By
Label labeling method and device, equipment, storage medium and computer program product
CN120850947A
Label marking method, device, equipment, storage medium and computer program product
CN120850947B