Prompt word processing method and device

Through the combination of agents and reinforcement learning, the annotation rules of large language models are identified and updated, and the target annotation prompt words are generated, which solves the problem of insufficient accuracy and convenience in the data generation of large language models and improves the competitiveness of the model.

CN120492871APending Publication Date: 2025-08-15ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510656826.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, large language models are insufficient in the data generation process, resulting in fierce competition in the data generation of model and difficult to meet high requirements.

Method used

Through the agent, the marked data and data labels are input into the large language model for labeling rules to identify the annotation rules, obtain the annotation rules, and update the annotation rules through reinforcement learning, and finally generate the target annotation prompt word to guide the large language model to perform data annotation processing.

Benefits of technology

It improves the accuracy and convenience of data generation of large language models and meets the high requirements of models in competition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492871A_ABST
    Figure CN120492871A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cue word processing method and device.The cue word processing method comprises the steps that in the cue word generating process, labeled data and data labels are input into a large language model through an intelligent agent for labeling rule recognition, and labeling rules are obtained; and inputting the labeled data and the labeling cue word containing the labeling rule into a large language model for data labeling processing to obtain a data labeling result, comparing the data labeling result with a data label, and updating the labeling rule according to the comparison result to obtain a target labeling rule after reinforcement learning is completed, and generating a target annotation prompt word based on the target annotation rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of data processing technology, and in particular to a prompt word processing method and device. Background Art

[0002] With the continuous development and promotion of the Internet and artificial intelligence technologies, more and more services can be automatically processed online, and many tasks can be automatically optimized and improved through advanced models such as large language models or intelligent agents. For example, during the training or parameter adjustment process of the first model, another model is used to generate data for model training or parameter adjustment. However, as model-related technologies become more and more complete, the competition faced by all parties is becoming more and more fierce. In this case, higher requirements are placed on the accuracy and convenience of model-generated data. Summary of the Invention

[0003] One or more embodiments of this specification provide a prompt word processing method, comprising: inputting, through an intelligent agent, labeled data and data labels into a large language model for labeling rule identification to obtain labeling rules; inputting the labeled data and labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain data labeling results; comparing the data labeling results with the data labels, and updating the labeling rules based on the comparison results to obtain target labeling rules for generating target labeling prompt words after reinforcement learning is completed.

[0004] One or more embodiments of the present specification provide a data labeling method based on prompt words, including: obtaining a target labeling rule obtained through reinforcement learning. Generating a target labeling prompt word based on the target labeling rule. Inputting the target labeling prompt word and the data to be labeled into the large language model for labeling processing to obtain a target labeling result; the reinforcement learning includes: obtaining labeling rules by calling the large language model by an intelligent agent to perform labeling rule identification, obtaining data labeling results based on the labeling prompt word containing the labeling rule, comparing the data labeling result with the data label, and updating the labeling rule according to the comparison result.

[0005] One or more embodiments of the present specification provide a prompt word processing device, comprising: a rule identification module, configured to input the labeled data and data labels into a large language model through an intelligent agent to identify the labeling rules and obtain the labeling rules. A labeling processing module, configured to input the labeled data and the labeling prompt words containing the labeling rules into the large language model to perform data labeling processing and obtain data labeling results. A rule updating module, configured to compare the data labeling results with the data labels and update the labeling rules based on the comparison results, so as to obtain a target labeling rule for generating a target labeling prompt word after reinforcement learning is completed.

[0006] One or more embodiments of the present specification provide a data labeling device based on prompt words, including: a rule acquisition module, configured to obtain a target labeling rule obtained through reinforcement learning. A prompt word generation module, configured to generate a target labeling prompt word based on the target labeling rule. A labeling processing module, configured to input the target labeling prompt word and the data to be labeled into the large language model for labeling processing to obtain a target labeling result; the reinforcement learning includes: obtaining labeling rules by calling the large language model for labeling rule identification by an intelligent agent, obtaining data labeling results by performing data labeling processing based on the labeling prompt word containing the labeling rule, comparing the data labeling results with the data label, and updating the labeling rules according to the comparison results.

[0007] One or more embodiments of this specification provide a prompt word processing device, comprising: a processor; and a memory configured to store computer-executable instructions, wherein when executed, the computer-executable instructions cause the processor to: input the labeled data and data labels into a large language model through an intelligent agent to identify labeling rules and obtain labeling rules. Input the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain data labeling results. Compare the data labeling results with the data labels, and update the labeling rules based on the comparison results to obtain target labeling rules for generating target labeling prompt words after reinforcement learning is completed.

[0008] One or more embodiments of the present specification provide a data labeling device based on prompt words, including: a processor; and a memory configured to store computer-executable instructions, wherein the computer-executable instructions, when executed, cause the processor to: obtain a target labeling rule obtained through reinforcement learning. Generate a target labeling prompt word based on the target labeling rule. Input the target labeling prompt word and the data to be labeled into the large language model for labeling processing to obtain a target labeling result; the reinforcement learning includes: obtaining labeling rules by calling the large language model for labeling rule identification through an intelligent agent, obtaining data labeling results by performing data labeling processing based on the labeling prompt word containing the labeling rule, comparing the data labeling results with the data label, and updating the labeling rules according to the comparison results.

[0009] One or more embodiments of this specification provide a computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: an intelligent agent inputs the labeled data and data labels into a large language model for labeling rule identification to obtain labeling rules. The labeled data and the labeling prompt words containing the labeling rules are input into the large language model for data labeling processing to obtain data labeling results. The data labeling results are compared with the data labels, and the labeling rules are updated based on the comparison results to obtain target labeling rules for generating target labeling prompt words after reinforcement learning is completed.

[0010] One or more embodiments of the present specification provide a computer-readable storage medium for storing computer-executable instructions, which implement the following process when executed: obtaining a target labeling rule obtained through reinforcement learning. Generate a target labeling prompt word based on the target labeling rule. Input the target labeling prompt word and the data to be labeled into the large language model for labeling processing to obtain a target labeling result; the reinforcement learning includes: obtaining a labeling rule by calling the large language model through an intelligent agent to perform labeling rule identification, obtaining a data labeling result by performing data labeling processing based on the labeling prompt word containing the labeling rule, comparing the data labeling result with the data label, and updating the labeling rule according to the comparison result. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate one or more embodiments of this specification or technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments described in this specification. Those skilled in the art can derive other drawings based on these drawings without inventive effort. Figure 1A schematic diagram of an implementation environment of a prompt word processing method provided in one or more embodiments of this specification; Figure 2 A processing flow chart of a prompt word processing method provided in one or more embodiments of this specification; Figure 3 A schematic diagram of a prompt word processing method for generating a marking prompt word provided in one or more embodiments of this specification; Figure 4 A flowchart of a prompt word processing method applied to a scenario of generating a marking prompt word provided in one or more embodiments of this specification; Figure 5 A processing flow chart of a data annotation method based on prompt words provided in one or more embodiments of this specification; Figure 6 A processing flow chart of a data annotation method based on prompt words applied to data annotation scenarios provided in one or more embodiments of this specification; Figure 7 A schematic diagram of an embodiment of a prompt word processing device provided in one or more embodiments of this specification; Figure 8 A schematic diagram of an embodiment of a data tagging device based on prompt words provided in one or more embodiments of this specification; Figure 9 A schematic diagram of the structure of a prompt word processing device provided in one or more embodiments of this specification; Figure 10 A schematic diagram of the structure of a data annotation device based on prompt words provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0012] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0013] The prompt word processing method provided in one or more embodiments of this specification can be applied to the implementation environment of the prompt word processing system. Figure 1 , the implementation environment includes at least: Server 101, reinforcement learning module 102, and large language model 103; wherein the reinforcement learning module 102 includes an agent 102-1 and an execution environment 102-2; Agent 102-1 and large language model 103 run on server 101. They can run on the same server or on different servers. Server 101 can be a single server, a server cluster consisting of multiple servers, or one or more cloud servers in a cloud computing platform. Agent 102-1 uses large language model 103 to perform labeling rule recognition and data labeling.

[0014] In this implementation environment, the intelligent agent 102-1 in the reinforcement learning module 102 inputs the labeled data and data labels stored in the execution environment 102-2 into the large language model 103 for labeling rule recognition to obtain the labeling rules, and inputs the labeled data and the labeling prompt words containing the labeling rules into the large language model 103 for data labeling processing to obtain the data labeling results. The intelligent agent 102-1 compares the data labeling results with the data labels and updates the labeling rules according to the comparison results to obtain the target labeling rules for generating the target labeling prompt words after the reinforcement learning is completed.

[0015] It should be noted that, considering that the labeled data, data labels and other related data involved in this specification may belong to the user's privacy to a certain extent, if you want to collect labeled data, data labels and other related data, you can obtain the user's authorization before collecting the data, so that the data collection operation complies with relevant data management regulations. For example, authorization can be performed when the user's labeled data or data labels are obtained for the first time, and authorization can also be performed any time the user's labeled data or data labels are obtained. The specific method of data authorization can be to send a user data authorization reminder to the user, and the user can obtain the user's data authorization after confirming the reminder through an instruction. Alternatively, the method of data authorization can also be to obtain the user's data authorization by signing a data authorization agreement.

[0016] One or more embodiments of a prompt word processing method provided in this specification are as follows: Reference Figure 2 The prompt word processing method provided in this embodiment specifically includes steps S202 to S206.

[0017] In step S202 , the intelligent agent inputs the labeled data and data labels into the large language model to identify the labeling rules and obtain the labeling rules.

[0018] The labeled data in this embodiment refers to data that has been labeled by machine or manually. The labeled data includes evaluation data for evaluating the effectiveness of the first model, wherein the evaluation data can be the input data and / or output data of the first model. In addition, the labeled data can also include the training data of the first model; the data label refers to the label of the labeled data, which is obtained by evaluating or judging the labeled data. Optionally, the data label includes at least one of the following types of labels: intention label, emotion label, relevance label, timeliness label, and quality label, wherein the emotion label can be positive, neutral, or negative, and the relevance label can be relevant or irrelevant. The first model can be a large financial insurance model. The large language model can be a Qwen2-14B model and / or a gpt4o model, or other models that can perform general tasks.

[0019] Here, the labeling process of the labeled data of the first model is explained by taking the correlation label as an example: after obtaining the input data and output data of the first model, the correlation of the input data and output data is judged to obtain the conclusion that the input data and output data of the first model are related, so the input data and output data of the first model are marked with a "correlated" label. In this example, the evaluation data composed of the input data and output data of the first model is used as the labeled data, and the obtained data label is the "correlated" sub-label in the correlation label.

[0020] The labeling rule refers to the rule for labeling the labeled data. The intelligent agent asks questions to the large language model so that the large language model can determine the labeling reasons of the labeled data being labeled with their respective data labels, and generates labeling rules based on the labeling reasons, that is, the labeling rule stipulates the reason why the data label of the labeled data is determined to be a specific label; using the above example, when the data label is a correlation label, the labeling rule can be a rule for the data label to be determined to be relevant, or a rule for the data label to be determined to be irrelevant, or a rule for the data label to be determined to be relevant and irrelevant, wherein the data label Rules that are considered relevant include: "1. The user input question is consistent with the topic of the question; 2. The model output directly answers the user input question; 3. The input question, question, and model output discuss the same topics, with a high degree of consistency. If any of these criteria are met, the data label is considered relevant." Rules that are considered irrelevant include: "1. The user input question is inconsistent with the topic of the question and model output; 2. The question and model output do not provide the specific information required by the user; 3. The question and model output lack specific information, making comparison and correlation impossible; 4. The input and output types do not match and fail to meet user needs. If any of these criteria are met, the data label is considered irrelevant." The agent is an entity capable of operating autonomously in a specific environment and making decisions based on the information it perceives. The agent can be a software program, robot, or other type of automated system that achieves a target task by interacting with the environment. Optionally, the agent includes a proxy agent, which is used to process tasks but does not directly participate in the final decision-making process. The agent in this embodiment invokes a large language model, making the large language model the core component of the agent. The agent is located within a reinforcement learning framework. In addition, the reinforcement learning framework may also include one or more of an environment module, an action module, and an observation module. The environment module can be used to store labeled data and data labels. The agent is used to observe the labeled data and data labels, self-heuristically discover and summarize labeling rules. The action module simulates the labeling process using labeling prompts. The observation module uses the labeled data and data labels stored in the environment to observe the results of the labeling process.

[0021] In a specific implementation, the process of obtaining the labeling rules by calling the large language model for labeling rule identification by the intelligent agent can be achieved by the intelligent agent inputting the labeled data and data labels into the large language model for labeling rule identification. In order to improve the accuracy of the large language model in labeling rule identification, the labeled data can also be classified before the intelligent agent inputs the labeled data and data labels into the large language model, so as to perform labeling rule identification for each classification result. In an optional implementation manner provided by this embodiment, the labeled data and data labels are input into the large language model by the intelligent agent for labeling rule identification. Before the labeling rule obtaining step is executed, the labeled data is classified in the following manner: Determining a label type of the data label; The labeled data are classified according to the label type to obtain various labeled data sets.

[0022] Specifically, the labeled data may be data stored in the execution environment of the agent. In the process of classifying the labeled data in the execution environment, the labeled data under the same label type may be used as a labeled data set.

[0023] For example, the labeled data stored in the execution environment of the intelligent agent is evaluation data composed of input data and output data of the first model, and the label types of the data labels are relevance labels and intention labels. Therefore, the data with data labels as relevance labels in the labeled data are taken as a labeled data set, and the data with data labels as intention labels in the labeled data are taken as another labeled data set. In addition, the label types of data labels can also include other label types, which are not limited here.

[0024] In addition, in the process of classifying the labeled data, the sub-data labels contained in the label type of the data label can be further determined, and each labeled data set can be classified according to the sub-data labels to obtain each labeled data subset, or the labeled data can be directly classified according to the sub-data labels contained in the data label to obtain each labeled data subset.

[0025] For example, the labeled data stored in the execution environment of the intelligent agent is evaluation data composed of input data and output data of the first model, and the label type of the data label is a correlation label. When the correlation label contains two sub-data labels, relevant and irrelevant, the labeled data with the sub-data label as relevant can be used as a labeled data subset, and the labeled data with the sub-data label as irrelevant can be used as another labeled data subset; in addition, the label type of the data label of the evaluation data can also include an intention label. When the sub-data labels under the intention label include n types, the labeled data can also be divided into n labeled data subsets.

[0026] During the specific implementation process, based on the classification of the labeled data according to the label type of the data label to obtain each labeled data set, when the labeled data and data labels are input into the large language model by the intelligent agent for labeling rule recognition, the intelligent agent can call the large language model to perform labeling rule recognition based on each labeled data set and data label to obtain the labeling rule. In an optional implementation manner provided by this embodiment, the labeled data and data labels are input into the large language model by the intelligent agent for labeling rule recognition to obtain the labeling rule in the following manner: The intelligent agent inputs the labeled data sets and the data labels corresponding to the labeled data sets into the large language model to perform labeling rule recognition to obtain the labeling rules.

[0027] Specifically, after classifying the labeled data to obtain various labeled data sets, the intelligent agent can input each labeled data set and the data labels corresponding to each labeled data set into the large language model in sequence or in batches. The large language model analyzes the reasons for the labeling and generates labeling rules based on the analysis.

[0028] For example, the first labeled data set with data labels as relevance labels, the second labeled data set with data labels as timeliness labels, the relevance labels and the timeliness labels can be input into the large language model through the intelligent agent to perform labeling rule recognition, and obtain the first labeling rule corresponding to the first labeled data set and the second labeling rule corresponding to the second labeled data set.

[0029] In addition, based on the above classification of the labeled data according to the sub-data labels to obtain each labeled data subset, when the labeled data and data labels are input into the large language model by the agent for labeling rule identification, the agent can call the large language model to perform labeling rule identification based on each labeled data subset and the sub-data labels to obtain the labeling rule. In an optional implementation provided by this embodiment, the labeled data and data labels are input into the large language model by the agent for labeling rule identification to obtain the labeling rule in the following manner: The intelligent agent inputs the labeled data subsets and the sub-data tags corresponding to the labeled data subsets into the large language model to perform labeling rule recognition, thereby obtaining sub-labeling rules; The sub-labeling rules are combined to obtain the labeling rule.

[0030] Specifically, after classifying the labeled data according to the sub-data labels to obtain each labeled data subset, the intelligent agent can input each labeled data subset and the sub-data labels corresponding to each labeled data set into the large language model. The large language model extracts the data features of the labeled data in each labeled data subset and the label features of the sub-data labels, analyzes the data features and label features, obtains sub-labeling rules, and then splices the sub-labeled data to obtain the labeling rules. In the process of identifying labeling rules, the large language model can analyze the labeling reasons and generate sub-labeling rules based on the labeling reasons obtained by analysis.

[0031] For example, the intelligent agent can input the labeled data subset with the data label "relevant", the labeled data subset with the data label "irrelevant", the "relevant" sub-data label and the "irrelevant" sub-data label into the large language model for labeling rule recognition, obtain the first sub-labeling rule and the second sub-labeling rule, and splice the first sub-labeling rule and the second sub-labeling rule to obtain the labeling rule.

[0032] During the specific implementation process, in order to improve data processing efficiency and reduce data volume, before the intelligent agent inputs the labeled data and data labels into the large language model for labeling rule identification to obtain the labeling rules, the labeled data can also be sampled, and the obtained sampling results and the data labels corresponding to the sampling results are input into the large language model for labeling rule identification to obtain the labeling rules. In an optional implementation provided by this embodiment, each labeled data set is sampled in the following manner: The labeled data in each labeled data set are sampled to obtain each sampled data set.

[0033] Specifically, during the sampling process, the labeled data in each labeled data set may be sampled by random sampling, or the labeled data in each data set may be sampled according to preset sampling parameters to obtain each sampled data set.

[0034] It should be noted that the timing of the sampling process can be set according to user needs. For example, the sampling process can also occur before the labeled data is classified according to the label type of the data label to obtain each labeled data set, that is, sampling is performed on the labeled data set. In addition, the sampling process can also occur after determining the sub-data tags contained in the label type of the data label, classifying each labeled data set according to the sub-data tags, and obtaining each labeled data subset. Or, it can also occur after classifying the labeled data according to the sub-data tags contained in the data label to obtain each labeled data subset. This embodiment is not limited to this.

[0035] In practical applications, there may be cases where at least two sub-data labels in a data label are relatively similar. In this case, based on the idea of comparative learning, at least two sub-data labels can be simultaneously input into the large language model for labeling rule recognition to improve the accuracy of the large language model in labeling rule recognition. In an optional implementation provided by this embodiment, the following method is used to input the labeled data and data labels into the large language model for labeling rule recognition through the intelligent agent to obtain the labeling rule: Determining associated sub-data tags having an associated relationship in the annotation tags; The agent inputs the associated sub-data tag and the annotated data subset corresponding to the associated sub-data tag into the large language model for annotating rule recognition to obtain an associated annotating rule; The associated labeling rules are spliced to obtain the labeling rules.

[0036] Specifically, in the process of calling the large language model through the intelligent agent to identify the labeling rules, at least two similar associated sub-data tags in the data tag are input into the large language model through the intelligent agent together for labeling rule identification, and the associated labeling rules corresponding to the associated sub-data tags are obtained. The associated labeling rules are spliced to obtain the labeling rules.

[0037] For example, for correlation labels, which may include associated sub-data labels such as "completely relevant", "partially relevant", and "irrelevant", the intelligent agent inputs the associated sub-data labels and the labeled data subset corresponding to the associated sub-data labels into the large language model for labeling rule recognition, and obtains the sub-labeling rules corresponding to "completely relevant", the sub-labeling rules corresponding to "partially relevant", and the sub-labeling rules corresponding to "irrelevant".

[0038] It should be noted that the above-mentioned determination of the associated sub-data tags with associated relationships in the annotation tags; the inputting of the associated sub-data tags and the annotation data subset corresponding to the associated sub-data tags into the large language model through the intelligent agent for annotation rule identification, and the annotation rule identification step of obtaining the associated annotation rules can be freely combined with other steps for annotation rule identification. Specifically, the annotated data can be first classified according to the tag type to obtain each annotated data set, and then each annotated data set can be classified according to the sub-data tags to obtain the annotated data subset, and then the annotated data subset is clustered according to the associated sub-data tags in the sub-data tags, and the obtained clustered data subset and the corresponding clustered data tags are input into the large language model through the intelligent agent for annotation rule identification to obtain the clustered annotation rules, and the clustered annotation rules are spliced to obtain the annotation rules.

[0039] For example, the data labels of the labeled data include relevance labels and intention labels. The labeled data are classified according to the label type of the data label to obtain the labeled data set corresponding to the relevance label and the labeled data set corresponding to the intention label. The labeled data set corresponding to the relevance label is then classified according to the sub-data labels of "completely relevant", "partially relevant" and "irrelevant" to obtain three data labeled subsets, and the labeled data set corresponding to the intention label is classified according to n word data pairs to obtain n labeled data subsets. The labeled data subsets are then clustered according to the "completely relevant", "partially relevant" and "irrelevant" in the sub-data labels to obtain clustered data subsets. The clustered data subsets and the corresponding clustered data labels are input into the large language model through the intelligent agent for labeling rule recognition to obtain clustered labeling rules, and the clustered labeling rules are spliced to obtain labeling rules.

[0040] In the specific implementation process, when the large language model is used to identify the labeling rules, the labeling reasons of the labeled data can be analyzed, and labeling rules can be generated based on the labeling reasons. In an optional implementation provided by this embodiment, the labeling rules are identified in the following manner: Identify the marking reason based on the marked data and the data label to obtain the marking reason; The marking rule is generated according to the marking reason.

[0041] The labeling reason refers to the reason why the labeled data was labeled with a certain data label, for example, the reason why the labeled data was labeled with the "related" sub-data label. Specifically, during the labeling rule identification process using the large language model, the labeling reason can be identified based on the labeled data and data labels to obtain the labeling reason, and then converted into a labeling rule.

[0042] Step S204 : Input the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain a data labeling result.

[0043] In the above, the intelligent agent inputs the labeled data and data labels into the large language model for labeling rule recognition. After obtaining the labeling rules, in this step, the intelligent agent generates labeling prompt words based on the labeling rules, and inputs the labeled data and labeling prompt words into the large language model for data labeling through the action (Action) in reinforcement learning to obtain the data labeling results. In this way, the data labeling results are subsequently compared with the data labels, and based on the comparison results, it is determined whether the labeling prompt words generated by the basic labeling rules can be used to guide the large language model to perform data labeling.

[0044] The annotation prompt refers to text or instructions, i.e., prompts, used to guide the large model to perform specific tasks or generate specific types of answers or outputs. Optionally, the annotation prompt includes annotation rules, which are used to guide the large model to annotate the annotated data according to the annotation rules. It should be noted that the annotated data has already been annotated to obtain data labels as benchmark labels. The data annotation results obtained by annotating the annotated data here are used to compare and confirm whether they are consistent with the data labels. If they are consistent, it indicates that the large language model was able to obtain correct data annotation results under the prompt of the annotation prompt containing the annotation rules. If they are inconsistent, it indicates that the annotation prompt containing the annotation rules may not achieve good results and needs to be optimized or updated.

[0045] In a specific implementation, after obtaining the labeling rules output by the large language model, the labeling rules can be written into the labeling prompt word template to obtain the labeling prompt words, which are used to guide the large language model. In an optional implementation provided by this embodiment, the labeling prompt words are obtained in the following manner: Identifying a labeled data set included in the labeled data, and determining a labeling prompt word template corresponding to the labeled data set; The marking rule is written into the marking prompt word template to obtain the marking prompt word.

[0046] For example, the labeled data set included in the labeled data is the labeled data set corresponding to the correlation label. The labeling prompt word template corresponding to the labeled data set can be "Please label the labeled data according to the following labeling rules" or "Please determine whether the output data of the first model is related to the input data according to the following labeling rules. Your answer can only be one of relevant or not relevant."

[0047] During the specific execution process, after generating annotation prompt words based on the annotation rules in the above manner, the annotated data and the annotation prompt words can be input into the large language model for data annotation processing to obtain the data annotation results. Specifically, during the data annotation processing of the large language model, the annotation prompt words can be semantically recognized, and the data annotation results of the annotated data can be determined based on the semantic recognition results.

[0048] In step S206 , the data annotation result is compared with the data label, and the annotation rule is updated according to the comparison result, so as to obtain a target annotation rule for generating a target annotation prompt word after the reinforcement learning is completed.

[0049] After the labeled data and the labeled prompt words are input into the large language model for data labeling processing to obtain the data labeling results, in this step, the data labeling results are compared with the data labels through observation to verify whether the large language model has outputted data labeling results that are consistent with the data labels or whose similarity is greater than the preset threshold under the guidance of the labeling prompt words. If the output data labeling results are consistent with the data labels or the similarity is greater than the preset threshold, it indicates that the labeling prompt words generated based on the labeling rules can produce a good guiding effect on the large language model and can be used for actual data labeling processing.

[0050] During the specific implementation process, in order to better evaluate the data annotation results output by the large language model, the data annotation results can be compared with the data labels to obtain a comparison result, and an annotation evaluation index can be calculated based on the comparison result. The annotation rules are updated according to the annotation evaluation index. In an optional implementation provided by this embodiment, the data annotation results are compared with the data labels in the following manner, and the annotation rules are updated according to the comparison result: Comparing the data annotation result with the data label using a regular matching method to obtain a comparison result; A labeling evaluation index is calculated based on the comparison result, and the labeling rule is updated according to the labeling evaluation index.

[0051] The annotation evaluation index may be an indicator for evaluating the data annotation results. Optionally, the annotation evaluation index includes the annotation accuracy. It should be noted that reinforcement learning is completed when the annotation evaluation index is greater than or equal to a preset evaluation index. The comparison result includes a first comparison result and / or a second comparison result. The first comparison result may be a successful comparison result, and the second comparison result may be a failed comparison result.

[0052] Specifically, the data annotation results can be compared with the data labels by using regular matching, and the annotation evaluation index can be calculated based on the comparison results. If the annotation evaluation index is less than the preset evaluation index, the annotation rule is updated; if the annotation evaluation index is greater than or equal to the preset evaluation index, the current annotation rule is used as the target annotation rule.

[0053] For example, regular matching can be used to compare c data annotation results with c data labels to obtain a first comparison results and b second comparison results (where a+b=c). The annotation evaluation index is calculated based on the number of first comparison results a and the number of data labels c. If the annotation evaluation index is less than the preset evaluation index, the annotation rule is updated.

[0054] In addition, in the process of comparing the data annotation results with the data labels, the semantic similarity between the data annotation results and the data labels can also be calculated. In an optional implementation provided by this embodiment, the data annotation results and the data labels are compared in the following manner, and the annotation rules are updated according to the comparison results: Calculating the semantic similarity between the data annotation result and the data label, and determining the comparison result of the first data annotation result whose semantic similarity is greater than the similarity threshold as a passed comparison; Calculating a labeling evaluation index based on the first data labeling result and the data label; If the marking evaluation index is less than the preset evaluation index, it is determined to update the marking rule.

[0055] The first data annotation result refers to a data annotation result in which the comparison result is a passing comparison. Furthermore, a data annotation result in which the comparison result is a failing comparison can be referred to as a second data annotation result. Optionally, the annotation evaluation metric includes an annotation accuracy rate. It should be noted that reinforcement learning is determined to be complete when the annotation evaluation metric is greater than or equal to a preset evaluation metric.

[0056] Specifically, in the process of comparing the data annotation results with the data labels, the semantic similarity between the data annotation results and the data labels can be calculated, and the comparison result of the first data annotation result whose semantic similarity is greater than the similarity threshold is determined to be a passed comparison. The proportion of the first data annotation result for the data label is used as the annotation evaluation index. If the annotation evaluation index is less than the preset evaluation index, the annotation rule is updated; if the annotation evaluation index is greater than or equal to the preset evaluation index, the current annotation rule is used as the target annotation rule and a target annotation prompt word is generated. In addition, after executing the sub-step of calculating the annotation evaluation index based on the first data annotation result and the data label, the annotation rule can also be directly updated according to the annotation evaluation index.

[0057] For example, the first feature of the data annotation result and the second feature of the extracted data label, the semantic similarity is calculated based on the first feature and the second feature, the comparison result of the first data annotation result with a semantic similarity greater than the similarity threshold is determined as a passed comparison, the ratio of the number of first data annotation results to the number of data labels is calculated, if the ratio is less than the preset value, the annotation rule is updated, if the ratio is greater than the preset value, the current annotation rule is used as the target annotation rule and the target annotation prompt word is generated.

[0058] It should be noted that, in the above calculation of the semantic similarity between the data annotation result and the data label, the comparison result of the first data annotation result whose semantic similarity is greater than the similarity threshold is determined as a passed comparison, and the comparison result of the second data annotation result whose semantic similarity is less than or equal to the similarity threshold is determined as a failed comparison, the annotation evaluation index can also be calculated based on the first data annotation result, the second data annotation result and the data label. The specific calculation process can be referred to the above calculation process of the annotation evaluation index, and this embodiment will not be repeated here.

[0059] In addition, the operation of comparing the data annotation results with the data labels can also be replaced by: inputting the data annotation results and the data labels into a classification function for classification processing to obtain a classification result; optionally, the classification result includes an annotation evaluation index.

[0060] For example, the data annotation results and data labels can be input into any classification function to obtain a classification report, where the classification report includes annotation evaluation indicators such as annotation accuracy.

[0061] Furthermore, if the calculated labeling evaluation index is less than the preset evaluation index, it indicates that the current reinforcement learning process has not ended and the labeling rules need to be updated to determine the target labeling rules. In an optional implementation provided in this embodiment, the labeling rules are updated in the following manner: If the labeling evaluation index is less than or equal to the preset evaluation index, sampling the labeled data, and inputting the obtained sampling results and the data labels corresponding to the sampling results into the large language model through the intelligent agent to perform labeling rule recognition to obtain an updated labeling rule; If the marking evaluation index is greater than the preset evaluation index, the current marking rule is used as the target marking rule.

[0062] Specifically, random sampling can be used to resample the labeled data. Since the sampling results obtained from each random sampling may be different, the sampling results and data labels are input into the large language model for labeling rule recognition to obtain updated labeling rules. It should be noted that the above process of updating labeling rules can be repeated until the reinforcement learning is completed.

[0063] The following takes an update process as an example to illustrate the update process of the labeling rules: sampling the labeled data to obtain the sampling results, inputting the sampling results and data labels into the large language model through the intelligent agent for labeling rule identification to obtain the labeling rules, inputting the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain the data labeling results; comparing the data labeling results with the data labels, and when the labeling evaluation index calculated based on the comparison result is less than the preset evaluation index, re-sample the labeled data, and inputting the obtained sampling results and the data labels corresponding to the sampling results into the large language model through the intelligent agent for labeling rule identification to obtain the updated labeling rules, after which the sampling results and the update prompt words containing the updated labeling rules are input into the large language model for data labeling processing to obtain the updated labeling results, and comparing the updated labeling results with the data labels. When the labeling evaluation index calculated based on the comparison result is greater than or equal to the preset evaluation index, it is determined that the reinforcement learning is completed, the updated labeling rule is used as the target labeling rule, and the target labeling prompt words are generated.

[0064] During the specific implementation process, in order to further improve the usability of the obtained target labeling rules and target labeling prompt words, the target labeling rules and target labeling prompt words can be further optimized for the data with incorrect labeling by the large language model. In an optional implementation provided in this embodiment, the target labeling rules are optimized in the following manner: inputting a labeling rule corresponding to a first data labeling result in the data labeling results into the large language model, so that the large language model optimizes the labeling rule corresponding to the first data labeling result; Optionally, the optimizing the labeling rules corresponding to the first data labeling result includes: adding sub-labeling rules to the labeling rules corresponding to the first data labeling result, and / or optimizing the rule details of the labeling rules corresponding to the first data labeling result.

[0065] Among them, the first data labeling result can be a result in the data labeling result that is inconsistent with the data label. For example, for the i input data and o output data of the first model, the data labeling result of the large language model under the r labeling rule is "relevant", while the data labels of the i input data and o output data are "irrelevant", that is, the data labeling result is inconsistent with the data label, so the data labeling result of the i input data and o output data is used as the first data labeling result.

[0066] For example, the labeling rules corresponding to the first data labeling result and the prompt text such as "Under this labeling rule, the first data labeling result is inconsistent with the data label. How can the labeling rule be adjusted?" are used to instruct the large language model to add sub-labeling rules to the labeling rules or modify the strictness of the labeling rules.

[0067] It should be noted that the above-mentioned step S206 can also be replaced by iteratively updating the labeling rules based on the data labeling results until the target labeling rules of the user-generated target labeling prompt words are obtained after reaching the preset number of iterations. In this embodiment, in the process of reinforcement learning, in addition to setting the end conditions of reinforcement learning (that is, whether the labeling evaluation index calculated based on the comparison result is greater than or equal to the preset evaluation index), a preset number of iterations can also be set, and the labeling rules are updated according to the preset number of iterations. The following specifically describes the iterative update process with the preset number of iterations as n: the labeled data and data labels are input into the large language model through the intelligent agent for labeling rule identification to obtain the labeling rules; the labeled data and the labeling prompt words containing the labeling rules are input into the large language model for data labeling processing to obtain the data labeling results; the labeling rules are iteratively updated according to the data labeling results to obtain the updated labeling rules. If the current number of iterations is greater than or equal to the preset number of iterations, the current updated labeling rules are used as the target labeling rules.

[0068] In specific implementation, after obtaining the target labeling rules and generating target prompt words based on the target labeling rules, the target prompt words can be used to label the data to be labeled, that is, the target prompt words and the data to be labeled are input into the large language model for labeling processing to obtain the target labeling results.

[0069] For example, the target prompt words and the data to be evaluated of the first model are input into the large language model so that the large language model can label the data to be evaluated. After that, the use effect of the first model can be judged according to the target labeling results obtained by the labeling process. If the first labeling results account for a large proportion of the target labeling results, it indicates that the use effect of the first model is better and can be applied online. If the second labeling results account for a large proportion of the target labeling results, it indicates that the use effect of the first model is not good and further iterative training is required.

[0070] In summary, the prompt word processing method provided in this embodiment is to input the labeled data and data labels into the large language model through the intelligent agent to perform labeling rule identification to obtain the labeling rules, input the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain the data labeling results, compare the data labeling results with the data labels, and update the labeling rules according to the comparison results. After the reinforcement learning is completed, the target labeling rules for generating target labeling prompt words are obtained, and the large language model is guided by the obtained target labeling prompt words, so that the large language model can label the data to be labeled under the guidance of the target labeling prompt words, thereby obtaining more accurate and usable target labeling results. Furthermore, in the process of comparing the data annotation result with the data label, the semantic similarity between the data annotation result and the data label can be calculated, and the comparison result of the first data annotation result whose semantic similarity is greater than the similarity threshold is determined as a passed comparison. The annotation evaluation index is calculated based on the first data annotation result and the data label. If the annotation evaluation index is less than the preset evaluation index, the annotation rule is updated, providing precise conditions for the completion of reinforcement learning, so that the obtained target annotation rule is more suitable for data annotation of large language models.

[0071] The following takes the application of a prompt word processing method provided by this embodiment in the scenario of generating a marked prompt word as an example, combined with Figure 3 and Figure 4 , further explanation of the prompt word processing method provided in this embodiment is given in Figure 4 The prompt word processing method applied to the annotation prompt word generation scenario specifically includes the following steps.

[0072] like Figure 3As shown, the environment module stores correlation labels and labeled data subsets corresponding to correlation labels as an example, wherein the correlation labels include "correlated" sub-data label 1 and "irrelevant" sub-data label 2. After that, the agent samples the labeled data corresponding to the "relevant" sub-data label 1 and the labeled data corresponding to the "irrelevant" sub-data label 2, and inputs the obtained sample data sets into the large language model, so that the large language model determines the reasons why the labeled data subsets are labeled with their respective sub-data labels, and generates labeling rules based on the labeling reasons. The ion module inputs the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain the data labeling results, and then uses the observation module to compare the data labeling results with the data labels in a regular matching manner to obtain a comparison result, and calculates the labeling evaluation index based on the comparison result. If the labeling evaluation index is less than the preset evaluation index, the labeling rules are updated. If the labeling evaluation index is greater than or equal to the preset evaluation index, the reinforcement learning is completed, and the target labeling rules for generating target labeling prompt words are obtained, and the evaluation data of the first model are labeled based on the target labeling prompt words.

[0073] Step S402 : determining the label type of the data label, classifying the labeled data according to the sub-data labels under the label type, and obtaining each labeled data subset.

[0074] Step S404: sampling each labeled data subset to obtain each sampled data subset.

[0075] In step S406 , the intelligent agent inputs each sampled data subset and the data label corresponding to each sampled data subset into the large language model for labeling rule recognition to obtain the labeling rule.

[0076] Step S408: determining the annotation prompt word template corresponding to each annotation data subset.

[0077] Step S410: writing the marking rules into the marking prompt word template to obtain the marking prompt word.

[0078] In step S412, the labeled data and the labeling prompt words are input into the large language model for data labeling processing to obtain data labeling results.

[0079] In step S414, the data annotation result is compared with the data label by using a regular matching method to obtain a comparison result.

[0080] Step S416: Calculate the annotation evaluation index based on the comparison result.

[0081] In step S418 , if the labeling evaluation index is less than the preset evaluation index, the labeling rules are updated to obtain target labeling rules for generating target labeling prompt words after the reinforcement learning is completed.

[0082] In addition, if the labeling evaluation index is greater than or equal to the preset evaluation index, the current labeling rule is used as the target labeling rule, and a target labeling prompt word is generated.

[0083] It should be noted that any one of steps S402 to S418 or any combination of multiple steps can be combined with any one of steps S202 to S206 to form a new implementation method according to the needs of implementation deployment; in addition, according to the needs of actual deployment, any one or multiple technical features can be selected from steps S402 to S418 and combined with any one or multiple technical features provided by steps S202 to S206 to form a new implementation method; or, any one or multiple technical features in steps S402 to S418 can also be replaced by any one or multiple technical features provided by steps S202 to S206 to form a new implementation method according to the needs of actual deployment, which will not be repeated here.

[0084] One or more embodiments of a data annotation method based on prompt words provided in this specification are as follows: Reference Figure 5 The data tagging method based on prompt words provided in this embodiment specifically includes steps S502 to S506.

[0085] Step S502: Acquire target labeling rules obtained through reinforcement learning.

[0086] Optionally, the reinforcement learning includes: inputting the labeled data and data labels into a large language model through an intelligent agent to identify labeling rules and obtain labeling rules, inputting the labeled data and labeling prompt words containing the labeling rules into the large language model to perform data labeling processing to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

[0087] The labeled data in this embodiment refers to data that has been labeled by machine or manually, and the labeled data includes evaluation data for evaluating the effectiveness of the use of the first model, wherein the evaluation data can be the input data and / or output data of the first model; in addition, the labeled data may also include the training data of the first model; the data label refers to the label of the labeled data, which is obtained by evaluating or judging the labeled data. Optionally, the data label includes at least one of the following types of labels: intention label, emotion label, relevance label, timeliness label, and quality label, wherein the emotion label can be positive, neutral, or negative, and the relevance label can be relevant or irrelevant. The first model can be a large financial insurance model. The large language model can be a Qwen2-14B model and / or a gpt4o model, or other models that can perform general tasks.

[0088] Here, the labeling process of the labeled data of the first model is explained by taking the correlation label as an example: after obtaining the input data and output data of the first model, the correlation of the input data and output data is judged to obtain the conclusion that the input data and output data of the first model are related, so the input data and output data of the first model are marked with a "correlated" label. In this example, the evaluation data composed of the input data and output data of the first model is used as the labeled data, and the obtained data label is the "correlated" sub-label in the correlation label.

[0089] The labeling rule refers to the rule for labeling the labeled data. The intelligent agent asks questions to the large language model so that the large language model can determine the labeling reasons of the labeled data being labeled with their respective data labels, and generates labeling rules based on the labeling reasons, that is, the labeling rule stipulates the reason why the data label of the labeled data is determined to be a specific label; using the above example, when the data label is a correlation label, the labeling rule can be a rule for the data label to be determined to be relevant, or a rule for the data label to be determined to be irrelevant, or a rule for the data label to be determined to be relevant and irrelevant, wherein the data label Rules that are considered relevant include: "1. The user input question is consistent with the topic of the question; 2. The model output directly answers the user input question; 3. The input question, question, and model output discuss the same topics, with a high degree of consistency. If any of these criteria are met, the data label is considered relevant." Rules that are considered irrelevant include: "1. The user input question is inconsistent with the topic of the question and model output; 2. The question and model output do not provide the specific information required by the user; 3. The question and model output lack specific information, making comparison and correlation impossible; 4. The input and output types do not match and fail to meet user needs. If any of these criteria are met, the data label is considered irrelevant." The agent is an entity that can operate autonomously in a specific environment and make decisions based on the information it perceives. The agent can be a software program, a robot, or other type of automated system that achieves a target task by interacting with the environment. Optionally, the agent includes a proxy agent, which is used to process tasks but does not directly participate in the final decision-making process. The agent in this embodiment can call upon a large language model, making the large language model the core component of the agent. The agent is located within a reinforcement learning framework. In addition, the reinforcement learning framework may also include one or more of an environment, an action, and an observation. The environment can be used to store labeled data and data labels. The agent is used to observe the labeled data and data labels, self-heuristically discover and summarize labeling rules. The action simulates the labeling process using labeling prompts. The observation uses the labeled data and data labels stored in the environment to observe the results of the labeling process.

[0090] In an optional implementation manner provided by this embodiment, before executing the step of obtaining the target labeling rule obtained after the reinforcement learning is completed, the method further includes: Obtain the data to be annotated submitted by the user through the annotation platform, and detect whether there is a preset annotation result corresponding to the data to be annotated; If not, executing the step of obtaining the target labeling rule obtained after the reinforcement learning is completed; If so, the preset annotation results are visualized on the interactive interface of the annotation platform.

[0091] The annotation platform may be a platform for users to provide data to be annotated, and the data to be annotated may include input data and / or output data of the first model.

[0092] Step S504: generating target labeling prompt words based on the target labeling rules.

[0093] During the specific execution process, the target marking rule can be written into the prompt word template corresponding to the target marking rule to obtain the target marking prompt word. In an optional implementation provided by this embodiment, the target marking prompt word is obtained in the following manner: Identifying a labeling prompt word template corresponding to the target labeling rule; The target labeling rule is written into the labeling prompt word template to obtain the labeling prompt word.

[0094] Step S506: input the target annotation prompt word and the data to be annotated into the large language model for annotation processing to obtain a target annotation result.

[0095] Optionally, the data to be labeled includes input data and / or output data of the first model. By inputting the target labeling prompt words and the data to be labeled into the large language model, the data to be labeled can be labeled by the large language model to obtain the target labeling result.

[0096] For example, the target labeling prompt is "Please determine whether the input data and output data of the first model are related. Your answer can only be related or not related." The target labeling result output by the large language model is "related", where "related" is a data label.

[0097] For another example, the target prompt words and the data to be evaluated of the first model are input into the large language model so that the large language model can label the data to be evaluated. After that, the use effect of the first model can be judged according to the target labeling results obtained by the labeling process. If the first labeling results account for a large proportion of the target labeling results, it indicates that the use effect of the first model is better and can be applied online. If the second labeling results account for a large proportion of the target labeling results, it indicates that the use effect of the first model is not good and further iterative training is required.

[0098] The following is an example of the application of a data annotation method based on prompt words in the evaluation data annotation scenario provided by this embodiment. Figure 6 , further explanation of the prompt word processing method provided in this embodiment is given in Figure 5 , a data annotation method based on prompt words applied to evaluation data annotation scenarios specifically includes the following steps.

[0099] Step S602: Obtain the data to be annotated submitted by the user through the annotation platform.

[0100] Prior to executing step S602, the agent may input the labeled data and data labels into the large language model to identify the labeling rules and obtain the labeling rules. The labeled data and the labeling prompt words containing the labeling rules may be input into the large language model for data labeling processing to obtain the data labeling results. The data labeling results may be compared with the data labels, and the labeling rules may be updated based on the comparison results to obtain the target labeling rules after the reinforcement learning is completed. Optionally, the data to be labeled includes evaluation data of the first model.

[0101] Step S604: Detect whether there is a preset annotation result corresponding to the data to be annotated.

[0102] Step S606: If not, obtain the target labeling rules obtained through reinforcement learning.

[0103] Optionally, if yes, a visual display of preset annotation results is performed on the interactive interface of the annotation platform.

[0104] Step S608: Generate target annotation prompt words based on the target annotation rules.

[0105] In step S610 , the target annotation prompt word and the data to be annotated are input into the large language model for annotation processing to obtain the target annotation result.

[0106] After that, the target annotation results can be visualized in the interactive interface of the annotation platform.

[0107] It should be noted that any one of steps S602 to S610 or any combination of multiple steps can be combined with any one of steps S502 to S506 to form a new implementation method according to the needs of implementation deployment; in addition, according to the needs of actual deployment, any one or multiple technical features can be selected from steps S602 to S610 and combined with any one or multiple technical features provided by steps S502 to S506 to form a new implementation method; or, any one or multiple technical features in steps S602 to S610 can be replaced with any one or multiple technical features provided by steps S502 to S506 to form a new implementation method according to the needs of actual deployment, which will not be repeated here.

[0108] An embodiment of a prompt word processing device provided in the specification is as follows: In the above embodiment, a prompt word processing method is provided, and correspondingly, a prompt word processing device is also provided, which will be described below with reference to the accompanying drawings.

[0109] Reference Figure 7 , which shows a schematic diagram of an embodiment of a prompt word processing device provided by this embodiment.

[0110] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiment described below is only illustrative.

[0111] This embodiment provides a prompt word processing device, the device comprising: The rule identification module 702 is configured to input the labeled data and data labels into the large language model through the intelligent agent to perform labeling rule identification and obtain the labeling rules; The annotation processing module 704 is configured to input the annotated data and the annotation prompt words containing the annotation rules into the large language model to perform data annotation processing and obtain data annotation results; The rule updating module 706 is configured to compare the data annotation result with the data label and update the annotation rule according to the comparison result to obtain the target annotation rule for generating the target annotation prompt word after the reinforcement learning is completed.

[0112] The specification provides an embodiment of a data annotation device based on prompt words as follows: In the above embodiment, a data labeling method based on prompt words is provided. Correspondingly, a data labeling device based on prompt words is also provided, which will be described below with reference to the accompanying drawings.

[0113] Reference Figure 7 , which shows a schematic diagram of an embodiment of a data tagging device based on prompt words provided in this embodiment.

[0114] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiment described below is only illustrative.

[0115] This embodiment provides a data tagging device based on prompt words, the device comprising: A rule acquisition module 802 is configured to acquire target labeling rules obtained through reinforcement learning; A prompt word generation module 804 is configured to generate a target labeling prompt word based on the target labeling rule; The annotation processing module 806 is configured to input the target annotation prompt word and the data to be annotated into the large language model for annotation processing to obtain a target annotation result; The reinforcement learning includes: obtaining labeling rules by calling a large language model through an intelligent agent to identify labeling rules, performing data labeling processing based on labeling prompt words containing the labeling rules to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

[0116] An embodiment of a prompt word processing device provided in this specification is as follows: Corresponding to the above-described prompt word processing method, based on the same technical concept, one or more embodiments of this specification further provide a prompt word processing device, which is used to execute the above-described prompt word processing method. Figure 8 A schematic diagram of the structure of a prompt word processing device provided in one or more embodiments of this specification.

[0117] This embodiment provides a prompt word processing device, including: like Figure 9As shown, the prompt word processing device can vary significantly due to different configurations or performance. It may include one or more processors 901 and memory 902. The memory 902 may store one or more applications or data. The memory 902 may be either transient or persistent. The applications stored in the memory 902 may include one or more modules (not shown), each of which may include a series of computer-executable instructions for the prompt word processing device. Furthermore, the processor 901 may be configured to communicate with the memory 902, allowing the prompt word processing device to execute the series of computer-executable instructions in the memory 902. The prompt word processing device may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input / output interfaces 905, one or more keyboards 906, and the like.

[0118] In a specific embodiment, the prompt word processing device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the prompt word processing device, and the one or more programs are configured to be executed by one or more processors, including computer-executable instructions for performing the following: The intelligent agent inputs the labeled data and data labels into the large language model to identify the labeling rules and obtain the labeling rules; Inputting the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain a data labeling result; The data annotation result is compared with the data label, and the annotation rule is updated according to the comparison result, so as to obtain the target annotation rule for generating the target annotation prompt word after the reinforcement learning is completed.

[0119] An embodiment of a data annotation device based on prompt words provided in this specification is as follows: Corresponding to the data tagging method based on prompt words described above, based on the same technical concept, one or more embodiments of this specification further provide a data tagging device based on prompt words, which is used to execute the data tagging method based on prompt words provided above. Figure 10 A schematic diagram of the structure of a prompt word processing device provided in one or more embodiments of this specification.

[0120] This embodiment provides a data annotation device based on prompt words, including: like Figure 10As shown, a data annotation device based on prompt words can vary significantly due to different configurations or performance. It may include one or more processors 1001 and memory 1002. Memory 1002 may store one or more applications or data. Memory 1002 may be either transient or persistent storage. The application stored in memory 1002 may include one or more modules (not shown), each of which may include a series of computer-executable instructions in the prompt word processing device. Furthermore, processor 1001 may be configured to communicate with memory 1002, executing the series of computer-executable instructions in memory 1002 on the prompt word processing device. The prompt word processing device may also include one or more power supplies 1003, one or more wired or wireless network interfaces 1004, one or more input / output interfaces 1005, one or more keyboards 1006, and the like.

[0121] In a specific embodiment, a data tagging device based on prompt words includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the prompt word processing device, and the one or more programs are configured to be executed by one or more processors, including computer-executable instructions for performing the following: Obtain target labeling rules obtained through reinforcement learning; generating a target annotation prompt word based on the target annotation rule; Inputting the target annotation prompt word and the data to be annotated into the large language model for annotation processing to obtain a target annotation result; The reinforcement learning includes: obtaining labeling rules by calling a large language model through an intelligent agent to identify labeling rules, performing data labeling processing based on labeling prompt words containing the labeling rules to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

[0122] An embodiment of a computer-readable storage medium provided in this specification is as follows: Corresponding to the prompt word processing method described above, based on the same technical concept, one or more embodiments of this specification further provide a computer-readable storage medium.

[0123] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions. When the computer-executable instructions are executed, the following process is implemented: The agent inputs the labeled data and data labels into the large language model to identify the labeling rules and obtain the labeling rules; Inputting the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain a data labeling result; The data annotation result is compared with the data label, and the annotation rule is updated according to the comparison result, so as to obtain the target annotation rule for generating the target annotation prompt word after the reinforcement learning is completed.

[0124] It should be noted that the embodiment of a computer-readable storage medium in this specification and the embodiment of a prompt word processing method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned corresponding method, and the repeated parts will not be repeated.

[0125] An embodiment of a computer-readable storage medium provided in this specification is as follows: Corresponding to the prompt word processing method described above, based on the same technical concept, one or more embodiments of this specification further provide a computer-readable storage medium.

[0126] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions. When the computer-executable instructions are executed, the following process is implemented: Obtain target labeling rules obtained through reinforcement learning; generating a target annotation prompt word based on the target annotation rule; Inputting the target annotation prompt word and the data to be annotated into the large language model for annotation processing to obtain a target annotation result; The reinforcement learning includes: obtaining labeling rules by calling a large language model through an intelligent agent to identify labeling rules, performing data labeling processing based on labeling prompt words containing the labeling rules to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

[0127] It should be noted that the embodiment of a computer-readable storage medium in this specification and the embodiment of a prompt word processing method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned corresponding method, and the repeated parts will not be repeated.

[0128] An embodiment of a computer program product provided in this specification is as follows: Corresponding to the prompt word processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.

[0129] A computer program product comprising a computer program / instructions, which, when executed by a processor, implements the following steps: The intelligent agent inputs the labeled data and data labels into the large language model to identify the labeling rules and obtain the labeling rules; Inputting the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain a data labeling result; The data annotation result is compared with the data label, and the annotation rule is updated according to the comparison result, so as to obtain the target annotation rule for generating the target annotation prompt word after the reinforcement learning is completed.

[0130] It should be noted that the embodiment of a computer program product in this specification and the embodiment of a prompt word processing method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned corresponding method, and the repeated parts will not be repeated.

[0131] An embodiment of a computer program product provided in this specification is as follows: Corresponding to the data tagging method based on prompt words described above, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.

[0132] A computer program product comprising a computer program / instructions, which, when executed by a processor, implements the following steps: Obtain target labeling rules obtained through reinforcement learning; generating a target annotation prompt word based on the target annotation rule; Inputting the target annotation prompt word and the data to be annotated into the large language model for annotation processing to obtain a target annotation result; The reinforcement learning includes: obtaining labeling rules by calling a large language model through an intelligent agent to identify labeling rules, performing data labeling processing based on labeling prompt words containing the labeling rules to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

[0133] It should be noted that the embodiment of a computer program product in this specification and the embodiment of a data labeling method based on prompt words in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned corresponding method, and the repeated parts will not be repeated.

[0134] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. For example, the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments are similar to the method embodiments, so the description is relatively simple. For relevant content in the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments, please refer to the partial description of the method embodiments.

[0135] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0136] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0137] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.

[0138] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0139] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0140] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0141] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0142] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0144] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0145] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0146] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0147] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising at least one ..." does not exclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.

[0148] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0149] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.

Claims

1. A prompt word processing method, comprising: The intelligent agent inputs the labeled data and data labels into the large language model to identify the labeling rules and obtain the labeling rules; Inputting the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain a data labeling result; The data annotation result is compared with the data label, and the annotation rule is updated according to the comparison result, so as to obtain the target annotation rule for generating the target annotation prompt word after the reinforcement learning is completed.

2. The prompt word processing method according to claim 1, wherein comparing the data annotation result with the data label and updating the annotation rule according to the comparison result comprises: Comparing the data annotation result with the data label using a regular matching method to obtain a comparison result; A labeling evaluation index is calculated based on the comparison result, and the labeling rule is updated according to the labeling evaluation index.

3. The prompt word processing method according to claim 2, wherein updating the annotation rule according to the annotation evaluation index comprises: If the labeling evaluation index is less than or equal to the preset evaluation index, the labeled data is sampled, and the sampling results and the data labels corresponding to the sampling results are input into the large language model by the intelligent agent for labeling rule recognition to obtain updated labeling rules.

4. The prompt word processing method according to claim 1, wherein the step of inputting the labeled data and data labels into the large language model to identify the labeling rules by the intelligent agent and obtaining the labeling rules further comprises: Determining a label type of the data label; Classify the labeled data according to the label type to obtain each labeled data set; Accordingly, the agent inputs the labeled data and data labels into the large language model for labeling rule recognition, including: The intelligent agent inputs the labeled data sets and the data labels corresponding to the labeled data sets into the large language model to perform labeling rule recognition to obtain the labeling rules.

5. The prompt word processing method according to claim 1, wherein the tagging rule identification comprises: Identify the marking reason based on the marked data and the data label to obtain the marking reason; The marking rule is generated according to the marking reason.

6. The prompt word processing method according to claim 1, wherein the step of inputting the labeled data and data labels into a large language model to identify labeling rules and obtain labeling rules by an intelligent agent comprises: Determining associated sub-data tags having an associated relationship in the data tags; The agent inputs the associated sub-data tag and the annotated data set corresponding to the associated sub-data tag into the large language model to perform annotating rule recognition, thereby obtaining an associated annotating rule; The associated labeling rules are spliced to obtain the labeling rules.

7. The prompt word processing method according to claim 1, wherein the marking prompt word is obtained by: Identifying a labeled data set included in the labeled data, and determining a labeling prompt word template corresponding to the labeled data set; The marking rule is written into the marking prompt word template to obtain the marking prompt word.

8. The prompt word processing method according to claim 1, further comprising: inputting a labeling rule corresponding to a first data labeling result in the data labeling results into the large language model, so that the large language model optimizes the labeling rule corresponding to the first data labeling result; The optimizing the annotation rule corresponding to the first data annotation result includes: adding a sub-annotation rule to the annotation rule corresponding to the first data annotation result, and / or optimizing the rule details of the annotation rule corresponding to the first data annotation result.

9. A data annotation method based on prompt words, comprising: Obtain target labeling rules obtained through reinforcement learning; generating a target annotation prompt word based on the target annotation rule; Inputting the target annotation prompt word and the data to be annotated into the large language model for annotation processing to obtain a target annotation result; The reinforcement learning includes: obtaining labeling rules by calling a large language model through an intelligent agent to identify labeling rules, performing data labeling processing based on labeling prompt words containing the labeling rules to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

10. The data labeling method based on prompt words according to claim 1, before executing the step of obtaining the target labeling rules obtained after the reinforcement learning is completed, further comprising: Obtain the data to be annotated submitted by the user through the annotation platform, and detect whether there is a preset annotation result corresponding to the data to be annotated; If not, executing the step of obtaining the target labeling rule obtained after the reinforcement learning is completed; If so, the preset annotation results are visualized on the interactive interface of the annotation platform.

11. A prompt word processing device, comprising: A rule recognition module is configured to input the labeled data and data labels into the large language model through the intelligent agent to perform labeling rule recognition and obtain the labeling rules; a labeling processing module configured to input the labeled data and the labeling prompt words containing the labeling rules into the large language model to perform data labeling processing and obtain a data labeling result; The rule updating module is configured to compare the data annotation result with the data label and update the annotation rule according to the comparison result to obtain the target annotation rule for generating the target annotation prompt word after the reinforcement learning is completed.

12. A data tagging device based on prompt words, comprising: A rule acquisition module is configured to acquire target labeling rules obtained through reinforcement learning; a prompt word generation module, configured to generate a target labeling prompt word based on the target labeling rule; a tagging processing module configured to input the target tagging prompt word and the data to be tagged into the large language model for tagging processing to obtain a target tagging result; The reinforcement learning includes: obtaining labeling rules by calling a large language model through an intelligent agent to identify labeling rules, performing data labeling processing based on labeling prompt words containing the labeling rules to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

13. A prompt word processing device, comprising: processor; and a memory configured to store computer-executable instructions that, when executed, cause the processor to: The intelligent agent inputs the labeled data and data labels into the large language model to identify the labeling rules and obtain the labeling rules; Inputting the labeled data and the labeling prompt words containing the labeling rules into the large language model for data labeling processing to obtain a data labeling result; The data annotation result is compared with the data label, and the annotation rule is updated according to the comparison result, so as to obtain the target annotation rule for generating the target annotation prompt word after the reinforcement learning is completed.

14. A data annotation device based on prompt words, comprising: processor; and a memory configured to store computer-executable instructions that, when executed, cause the processor to: Obtain target labeling rules obtained through reinforcement learning; generating a target annotation prompt word based on the target annotation rule; Inputting the target annotation prompt word and the data to be annotated into the large language model for annotation processing to obtain a target annotation result; The reinforcement learning includes: obtaining labeling rules by calling a large language model through an intelligent agent to identify labeling rules, performing data labeling processing based on labeling prompt words containing the labeling rules to obtain data labeling results, comparing the data labeling results with the data labels, and updating the labeling rules according to the comparison results.

15. A computer-readable storage medium for storing computer-executable instructions, wherein the computer-executable instructions implement the steps of the method according to claim 1 or 10 when executed.

Citation Information

Cited By

  • Method, device and equipment for generating data classification and grading rule

    CN121166626A

  • Semantic recognition method and device for customer complaint work order

    CN121902812A