Intelligent agent labeling method and device, equipment and storage medium
By obtaining agent configuration information and using pre-marking model generation and adjustment tool call planning, the problem that the existing technology cannot effectively support agent annotation is solved, and personalized and efficient agent annotation is achieved.
Patent Information
- Application Number
- CN202510076261.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art cannot effectively support the annotation process of the agent, especially in tool planning, and it is difficult to improve the tool planning capabilities of the agent.
By obtaining the agent configuration information, including the human prompt words, tool pools and pre-notation models, the pre-notation model is used to generate preliminary tool call plans, and the adjusted tool call plans are obtained through the human-computer interaction interface, and iterative training is carried out to build the agent.
Personalized agent labeling is realized, the agent labeling efficiency is improved, and the agent's tool planning and labeling process can more effectively support the agent's tool planning and labeling process.
Smart Images

Figure CN119940395A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent body labeling method, device, equipment and storage medium. Background Art
[0002] With the development of deep learning, artificial intelligence has been applied in many fields. The effectiveness of artificial intelligence technology is largely affected by the quality of training data, and annotation platforms for training data annotation have emerged.
[0003] Agent is an important concept in the field of artificial intelligence. It refers to an autonomous entity that can perceive the environment, make decisions, and perform actions to achieve specific goals. Agents are different from traditional models. For example, traditional dialogue models automatically reply to preset sentences through user-triggered keywords, while dialogue agents need to manage a tool library and plan how to effectively use these tools before finally generating answers. The process of calling agent tools is very complicated, and the existing annotation platform cannot support agent annotation. Therefore, how to improve the ability of tool planning in the annotation process of agents is an urgent problem to be solved. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide an intelligent agent tagging method, device, equipment and storage medium, which can support personalized intelligent agent tagging and improve the efficiency of intelligent agent tagging. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses an intelligent agent labeling method, comprising:
[0006] Obtaining agent configuration information and questions to be annotated; the agent configuration information includes a character prompt word, a tool pool and a pre-annotated model; the tool pool is used to record agent tools;
[0007] Based on the tool information of the intelligent agent tool, the character prompt words and the question to be labeled, using the pre-labeling model to generate a preliminary tool call plan corresponding to the question to be labeled, and obtaining the tool output result of the intelligent agent tool;
[0008] The adjusted tool call plan corresponding to the preliminary tool call plan is obtained through a human-computer interaction interface, and the pre-labeling model is iteratively trained based on the tool information of the intelligent agent tool, the character prompt words, the question to be labeled, the adjusted tool call plan and the tool output result, so as to construct an intelligent agent based on the trained labeling model.
[0009] Optionally, before obtaining the agent configuration information and the question to be labeled, the following is also included:
[0010] The corresponding intelligent agent tool is obtained by encapsulating the collected application programming interface; the tool information of the intelligent agent tool includes a name, a description field, an access method, a parameter field and an instance field.
[0011] Optionally, obtain agent configuration information and questions to be labeled, including:
[0012] The questions to be labeled uploaded by the administrator are obtained to obtain a question set, so as to perform preliminary tool call planning and model iteration for each question to be labeled in the question set in turn, until all the questions to be labeled are traversed to obtain a trained labeling model.
[0013] Optionally, the iterative training includes:
[0014] Based on the tool information of the agent tool, the character prompt word, the question to be labeled, the adjusted tool call plan and the tool output result, using the pre-labeling model to generate an Nth tool call plan corresponding to the question to be labeled;
[0015] Acquire the adjusted tool call plan corresponding to the Nth tool call plan through the human-computer interaction interface, and stop iterating until N reaches a preset target number of times;
[0016] The iteration process also includes:
[0017] An iteration stop instruction is received, and iteration is stopped according to the iteration stop instruction; the iteration stop instruction is selectively initiated by the user after analyzing the tool output result of the intelligent agent tool.
[0018] Optionally, based on the tool information of the agent tool, the character prompt words and the question to be labeled, using the pre-labeling model to generate a preliminary tool call plan corresponding to the question to be labeled, including:
[0019] The tool information of the agent tool and the question to be annotated are embedded into the character prompt word through a placeholder to obtain input data;
[0020] The input data is input into the pre-annotation model, and a preliminary tool call plan corresponding to the question to be annotated is obtained according to the output of the pre-annotation model; the pre-annotation model is a large language model.
[0021] Optionally, iteratively training the pre-labeled model to construct an agent based on the trained labeled model includes:
[0022] According to the preliminary tool call plan generated each time during the iteration process and the corresponding adjusted tool call plan, the pre-labeling model is fine-tuned to obtain a trained labeling model.
[0023] Optionally, iteratively training the pre-labeled model to construct an agent based on the trained labeled model includes:
[0024] Obtain adjusted tool output results corresponding to all tool output results in the iteration process through a human-computer interaction interface;
[0025] The preference learning of the pre-labeling model is performed by utilizing the initial tool call plan and the adjusted tool call plan corresponding to each iteration, as well as the tool output result and the corresponding adjusted tool output result.
[0026] In a second aspect, the present application discloses an intelligent agent labeling device, comprising:
[0027] An information acquisition module is used to acquire agent configuration information and questions to be annotated; the agent configuration information includes a character prompt word, a tool pool and a pre-annotated model; the tool pool is used to record agent tools;
[0028] A preliminary tool call plan generation module is used to generate a preliminary tool call plan corresponding to the question to be labeled using the pre-labeling model based on the tool information of the intelligent agent tool, the character prompt words and the question to be labeled, and obtain the tool output result of the intelligent agent tool;
[0029] An iterative module is used to obtain the adjusted tool call plan corresponding to the preliminary tool call plan through a human-computer interaction interface, and iteratively train the pre-labeling model based on the tool information of the intelligent agent tool, the character prompt words, the question to be labeled, the adjusted tool call plan and the tool output result, so as to construct an intelligent agent based on the trained labeled model.
[0030] In a third aspect, the present application discloses an electronic device, comprising:
[0031] Memory, used to store computer programs;
[0032] A processor is used to execute the computer program to implement the aforementioned intelligent agent labeling method.
[0033] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein the computer program implements the aforementioned intelligent agent labeling method when executed by a processor.
[0034] In the present application, agent configuration information and questions to be labeled are obtained; the agent configuration information includes character prompts, a tool pool and a pre-labeling model; the tool pool is used to record agent tools; based on the tool information of the agent tool, the character prompts and the questions to be labeled, a preliminary tool call plan corresponding to the questions to be labeled is generated using the pre-labeling model, and the tool output results of the agent tool are obtained; an adjusted tool call plan corresponding to the preliminary tool call plan is obtained through a human-computer interaction interface, and based on the tool information of the agent tool, the character prompts, the questions to be labeled, the adjusted tool call plan and the tool output results, the pre-labeling model is iteratively trained so as to construct an agent based on the trained labeled model. It can be seen that the agent configuration information supports users to configure character prompts and agent tools, thereby supporting personalized agent labeling; by first using the pre-labeling model to generate a preliminary tool call plan, and then obtaining manual corrections to the preliminary tool call plan, that is, the adjusted tool call plan, and then using the pre-labeling model based on the adjusted tool call plan to generate a tool call plan, the efficiency of generating the optimal tool call plan is improved, and the efficiency of agent labeling is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0036] Figure 1 A flow chart of an intelligent agent labeling method provided in this application;
[0037] Figure 2 A schematic diagram of a specific adjusted tool call planning input box provided for this application;
[0038] Figure 3 A schematic diagram of the structure of an intelligent body labeling device provided in this application;
[0039] Figure 4 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] In the prior art, the annotation platform cannot support intelligent body annotation. In order to overcome the above technical problems, the present application proposes an intelligent body annotation method that can support personalized intelligent body annotation and improve the efficiency of intelligent body annotation.
[0042] The present application embodiment discloses a method for labeling an intelligent agent. Figure 1 As shown, the method may include the following steps:
[0043] Step S11: Obtain agent configuration information and questions to be annotated; the agent configuration information includes character prompt words, a tool pool and a pre-annotated model; the tool pool is used to record agent tools.
[0044] In this embodiment, firstly, the configuration information of the agent is obtained, and the agent configuration information includes a person setting prompt, a tool pool and a pre-annotation model. The person setting prompt is a text used to guide the large language model to generate a specific type, theme or format output; in the process of agent annotation, the prompt plays a key role, which contains the person setting information of the agent, tool description and other contents, and guides the model to generate a reasonable tool call plan and answer; the person setting prompt also clarifies the principles that the agent needs to follow when calling tools and answering questions. For example, the prompt can clarify the professional principles that the agent should follow when answering stock-related questions, and how to build an effective query strategy based on the functions of different tools. Among them, the tool pool is used to record the agent tools, that is, to generate a tool call plan based on the tools in the tool pool, specifically including which tools in the tool pool are selected and the order of calling. The pre-annotation model is a model used to annotate the agent, that is, to determine the tool call plan, and the above-mentioned pre-annotation model can be a general large language model. At the same time, the question to be annotated is obtained so that the model can annotate the question.
[0045] Tool call planning is very important for intelligent agents. For example, if the question is "What are the reasons for the increase of the three stocks with the largest increase today?", the intelligent agent must first use the tool to obtain the latest stock increase information, and then search for the specific reasons for the increase of these three stocks in the second step. In this process, various unexpected situations may be encountered. For example, if the first step fails to successfully obtain the latest increase information, it is necessary to repeatedly iteratively adjust its strategy in subsequent steps, such as changing tools or adjusting search query statements, until the required information is obtained. This series of dynamic adjustments and attempts is called iterative planning. Existing annotation platforms are limited by the inability to configure prompt words and tools, so it is difficult to provide personalized intelligent agents.
[0046] Before obtaining the agent configuration information and the questions to be annotated, it also includes: obtaining the corresponding agent tool by encapsulating the collected application programming interface; the tool information of the agent tool includes a name, a description field, an access method, a parameter field and an instance field. That is, this application constructs an agent tool by encapsulating the application programming interface (API interface), that is, by collecting available API interfaces, including but not limited to: APIs provided by various service providers on the Internet, APIs customized by annotators, which can be encapsulated into tools for use by the agent. The tool consists of a name, access method, required parameters, optional parameters, and usage examples, see Table 1 below for details:
[0047] Table 1 Tool information of agent tools
[0048]
[0049] In other words, collect the interfaces provided by various service providers on the Internet or customized by the annotators, and encapsulate these interfaces into tools for the intelligent agent to use, so that the intelligent agent can obtain external data and services, such as obtaining stock price increase information, querying weather data, etc. Each API interface has its own specific specifications, including name, access method, parameter requirements, etc. The intelligent agent calls the corresponding API to obtain the required information by following these specifications. By defining the above tool information, a unified and standardized protocol is implemented to define the input and output format of the tool, and a large number of open source API interfaces are fully explored and utilized, so that the annotators can flexibly customize the intelligent agent tools and prompt words, thereby constructing a variety of intelligent agent annotation requirements. Accordingly, the intelligent agent tools in the tool pool in the intelligent agent configuration information can be selected from all constructed intelligent agent tools.
[0050] Step S12: Based on the tool information of the intelligent tool, the character prompt words and the question to be labeled, the pre-labeling model is used to generate a preliminary tool call plan corresponding to the question to be labeled, and the tool output result of the intelligent tool is obtained.
[0051] After determining the tool information, character prompt words and questions to be labeled of the intelligent agent tool, this information is input into the pre-labeling model, and the pre-labeling model is used to generate a preliminary tool call plan corresponding to the questions to be labeled, that is, an action list for the question, including the called tools and the tool input. The tool input may be the question to be labeled or other questions automatically derived from the question to be labeled; and the tool output results generated by each intelligent agent tool for the tool input are obtained.
[0052] In some embodiments, based on the tool information of the intelligent tool, the character prompt words and the question to be labeled, the pre-labeling model is used to generate a preliminary tool call plan corresponding to the question to be labeled, which may include: embedding the tool information of the intelligent tool and the question to be labeled into the character prompt words through placeholders to obtain input data; inputting the input data into the pre-labeling model, and obtaining a preliminary tool call plan corresponding to the question to be labeled according to the output of the pre-labeling model; the pre-labeling model is a large language model. That is, all the tool information The question q to be marked is embedded into the prompt word x through the placeholder method, and we get , generating a preliminary tool call plan, i.e. ; Where t represents the current step order, t=1 means the current step is the first step, It represents the action list generated by the current step, there are n actions in total, each action a consists of the tool name and the input of the tool.
[0053] Step S13: Obtain the adjusted tool call plan corresponding to the preliminary tool call plan through the human-computer interaction interface, and iteratively train the pre-labeling model based on the tool information of the intelligent agent tool, the character prompt words, the question to be labeled, the adjusted tool call plan and the tool output results, so as to construct an intelligent agent based on the trained labeling model.
[0054] After the model generates a preliminary tool call plan, the annotator will further confirm whether the plan is correct. The tool executor can help the annotator obtain the tool results. ; The annotator can judge whether the generated action list is correct based on the question and the tool results; if it is wrong, the tool can be selected by dragging and selecting, and the tool input can be modified by typing, such as Figure 2 As shown, Figure 2 The left input box is used to select tools (such as Figure 2 Select tools such as clarify, calculator, finquery, search, stocknews, finwiki, finkuaicha, etc. Figure 2 The tools and names are examples only). Figure 2The input box on the right (typing here) is used to modify the tool input; from this, you can get the action list after the modification is confirmed, that is, the adjusted tool call plan , and the corresponding tool output results . After the pre-annotation model is used to generate the preliminary tool call plan corresponding to the question to be annotated, the method further includes: obtaining an operation instruction for the preliminary tool call plan through a human-computer interaction interface; the operation instruction includes a continue instruction and an update instruction; the update instruction includes the adjusted tool call plan; if the operation instruction is the continue instruction, the pre-annotation model continues to iterate according to the continue instruction. That is, it should be noted that if the tool call plan output by the model is correct, no manual correction is required, and the next step can be continued.
[0055] In some embodiments, the iterative training may include: based on the tool information of the intelligent tool, the character prompt words, the question to be labeled, the adjusted tool call plan and the tool output results, using the pre-labeling model to generate the Nth tool call plan corresponding to the question to be labeled; obtaining the adjusted tool call plan corresponding to the Nth tool call plan through the human-computer interaction interface, and stopping the iteration until N reaches the preset target number; during the iteration process, it also includes: receiving an iteration stop instruction, and stopping the iteration according to the iteration stop instruction; the iteration stop instruction is selectively initiated by the user after analyzing the tool output results of the intelligent tool. In actual business applications, the latest version of the model will be selected as the pre-labeling model after each data iteration to implement an extremely efficient reinforcement fine-tuning strategy for the weak links of the current model. It can be understood that the adjusted tool call plan and the tool output results are concatenated after the previous prompt words to form a new input. , and then call the big model, and repeat this process. Specifically, you can set an upper limit T for the number of iterations, that is, the preset target number of times. If it exceeds this limit, it will be forced to stop. Of course, the model can decide whether to stop early. At this time, the big model will not output an action list, but a stop mark agreed in the prompt word: <finished>; You can also manually issue instructions to stop the iteration before reaching the upper limit of the number of iterations. In addition, the labeler has the final decision. If they feel that the collected information is not complete and cannot be stopped early, they can switch back to the iteration mode, manually add the action list, and increase the number of iterations. In addition, if the labeler is not satisfied with any of the above steps, they can go back to any step and execute it again.
[0056] Among them, obtaining the agent configuration information and the questions to be labeled includes: obtaining the questions to be labeled uploaded by the administrator to obtain a question set, so as to perform preliminary tool call planning and model iteration for each question to be labeled in the question set in turn, until all the questions to be labeled are traversed and a trained labeling model is obtained. That is, the labeling personnel can upload complex questions in batches through Excel at one time, and perform the following operations for each question: based on the tool information of the agent tool, the character prompt words and the questions to be labeled, the pre-labeling model is used to generate a preliminary tool call plan corresponding to the questions to be labeled, and the tool output results of the agent tool are obtained; the adjusted tool call plan corresponding to the preliminary tool call plan is obtained through the human-computer interaction interface, and the pre-labeling model is iteratively trained based on the tool information of the agent tool, the character prompt words, the questions to be labeled, the adjusted tool call plan and the tool output results. After the iteration is completed using all the questions to be labeled, the trained labeling model is obtained. After the tool call planning is completed, the additional information required to answer the user's questions is collected, that is, all the tool output results , at this time, based on this additional information, you can answer the user's questions in detail and fully.
[0057] In some embodiments, iteratively training the pre-labeled model to construct an agent based on the trained labeled model may include: fine-tuning the pre-labeled model according to the preliminary tool call plan generated each time during the iteration and the corresponding adjusted tool call plan to obtain the trained labeled model. Input the pre-annotated model together with the tool output to get the initial answer , , and then the annotator modifies and confirms it to get the final reply . At the same time, the tool call plan output by the model is also obtained And the output after manual confirmation ; and the input and output of the final answer, using this data, the pre-labeled model can be fine-tuned. For example, supervised fine-tuning (SFT) is used, that is, using data with labeled results (supervisory signals) to adjust and optimize the model; using the input and output data of the agent tool planning and answer collected during the labeling process as supervisory signals, the large model is fine-tuned, so that the model can better adapt to specific task requirements and improve accuracy and efficiency when dealing with similar problems.
[0058] In some embodiments, iteratively training the pre-labeled model to construct an intelligent agent based on the trained labeled model may include: obtaining adjusted tool output results corresponding to all tool output results in the iterative process through a human-computer interaction interface; using the corresponding preliminary tool call plan and the adjusted tool call plan each time in the iterative process, as well as the tool output results and their corresponding adjusted tool output results, to perform preference learning on the pre-labeled model. That is, collecting the pre-labeled and manually modified outputs at the same time, that is, ,as well as , these comparative data can be used for preference learning. Algorithms such as DPO (direct preference optimization) and PPO (proximal policy optimization) can be used to achieve this preference learning and improve the performance of the intelligent agent. By using pre-labeled and manually modified comparative data, the model can learn the tool selection and answering methods that humans prefer, so as to adjust its own decision-making and generation strategies, so that the model output is more in line with human expectations and actual needs.
[0059] This application can be applied to the annotation platform, which supports agent annotation through the above scheme, covering the complete process from tool packaging, agent configuration, question uploading, pre-annotation, manual correction to final answer generation and model fine-tuning, forming a closed-loop system, effectively solving the iterative agent annotation needs. And it has intelligent pre-annotation and flexible adjustment mechanism; that is, first use a large language model (such as GPT-4) for pre-annotation, generate the initial tool call plan and answer, provide a reference basis for manual annotation, significantly improve the annotation speed, and reduce manual workload. The pre-annotation of the agent not only accurately reproduces the planning and answering workflow of the agent, but also is equipped with auxiliary mechanisms such as tool executors, fallback mechanisms, and early termination. While providing great convenience for annotation personnel, it also effectively avoids the problem of mismatch between manual annotation style and large models. For example, when faced with a relatively clear question, the annotator only needs to continue to click "Continue" to quickly obtain the complete result; for questions that the current model cannot solve, the annotator can continue to advance the process after modifying the key nodes. At the same time, humans have the final decision-making power and can flexibly switch between iteration mode and stop mode according to actual conditions, or roll back to any step and re-execute, to ensure the controllability and optimizability of the annotation process. It also supports powerful tool customization and management, by collecting and encapsulating API interfaces from various sources into intelligent tools, and defining tool names, access methods, parameters and other information in detail, so that intelligent agents can flexibly call these tools to obtain the required information, enhancing the scope of intelligent agents' ability to handle problems; in addition, it supports customizing the tool pool according to needs during the annotation process, and annotators can select tools as needed to adapt to different annotation tasks and scenarios, improving the versatility and adaptability of the platform.
[0060] As can be seen from the above, in this embodiment, the agent configuration information and the questions to be labeled are obtained; the agent configuration information includes character prompts, a tool pool and a pre-labeling model; the tool pool is used to record the agent tool; based on the tool information of the agent tool, the character prompts and the questions to be labeled, the pre-labeling model is used to generate a preliminary tool call plan corresponding to the questions to be labeled, and the tool output results of the agent tool are obtained; the adjusted tool call plan corresponding to the preliminary tool call plan is obtained through the human-computer interaction interface, and based on the tool information of the agent tool, the character prompts, the questions to be labeled, the adjusted tool call plan and the tool output results, the pre-labeling model is iteratively trained so as to construct an agent based on the trained labeled model. It can be seen that the agent configuration information supports users to configure character prompts and agent tools, thereby supporting personalized agent labeling; by first using the pre-labeling model to generate a preliminary tool call plan, and then obtaining manual corrections to the preliminary tool call plan, that is, the adjusted tool call plan, and then using the pre-labeling model based on the adjusted tool call plan to generate a tool call plan, the efficiency of generating the optimal tool call plan is improved, and the efficiency of agent labeling is improved.
[0061] Correspondingly, the present application also discloses an intelligent body tagging device, see Figure 3 As shown, the device comprises:
[0062] The information acquisition module 11 is used to acquire agent configuration information and questions to be annotated; the agent configuration information includes a character prompt word, a tool pool and a pre-annotated model; the tool pool is used to record agent tools;
[0063] A preliminary tool call plan generating module 12 is used to generate a preliminary tool call plan corresponding to the question to be labeled using the pre-labeling model based on the tool information of the intelligent agent tool, the character prompt words and the question to be labeled, and obtain the tool output result of the intelligent agent tool;
[0064] The iteration module 13 is used to obtain the adjusted tool call plan corresponding to the preliminary tool call plan through the human-computer interaction interface, and iteratively train the pre-labeling model based on the tool information of the intelligent agent tool, the character prompt words, the question to be labeled, the adjusted tool call plan and the tool output result, so as to construct the intelligent agent based on the trained labeling model.
[0065] As can be seen from the above, in this embodiment, the agent configuration information and the questions to be labeled are obtained; the agent configuration information includes character prompts, a tool pool and a pre-labeling model; the tool pool is used to record the agent tool; based on the tool information of the agent tool, the character prompts and the questions to be labeled, the pre-labeling model is used to generate a preliminary tool call plan corresponding to the questions to be labeled, and the tool output results of the agent tool are obtained; the adjusted tool call plan corresponding to the preliminary tool call plan is obtained through the human-computer interaction interface, and based on the tool information of the agent tool, the character prompts, the questions to be labeled, the adjusted tool call plan and the tool output results, the pre-labeling model is iteratively trained so as to construct an agent based on the trained labeled model. It can be seen that the agent configuration information supports users to configure character prompts and agent tools, thereby supporting personalized agent labeling; by first using the pre-labeling model to generate a preliminary tool call plan, and then obtaining manual corrections to the preliminary tool call plan, that is, the adjusted tool call plan, and then using the pre-labeling model based on the adjusted tool call plan to generate a tool call plan, the efficiency of generating the optimal tool call plan is improved, and the efficiency of agent labeling is improved.
[0066] In some specific embodiments, the unit may include:
[0067] The intelligent agent tool generation unit is used to obtain the corresponding intelligent agent tool by encapsulating the collected application programming interface before obtaining the intelligent agent configuration information and the questions to be annotated; the tool information of the intelligent agent tool includes a name, a description field, an access method, a parameter field and an instance field.
[0068] In some specific embodiments, the information acquisition module 11 may specifically include:
[0069] The question-to-be-annotated acquisition unit is used to acquire the question-to-be-annotated uploaded by the administrator to obtain a question set, so as to perform preliminary tool call planning and model iteration for each question-to-be-annotated in the question set in turn, until all the question-to-be-annotated questions are traversed to obtain a trained annotation model.
[0070] In some specific embodiments, the iteration module 13 may specifically include:
[0071] an iterative unit, configured to generate an Nth tool call plan corresponding to the question to be labeled using the pre-labeling model based on the tool information of the agent tool, the character prompt word, the question to be labeled, the adjusted tool call plan, and the tool output result;
[0072] An iteration stopping unit, used for obtaining the adjusted tool calling plan corresponding to the Nth tool calling plan through a human-computer interaction interface, and stopping the iteration until N reaches a preset target number of times;
[0073] The iteration module 13 may further include:
[0074] An instruction acquisition unit is used to receive an iteration stop instruction and stop iteration according to the iteration stop instruction; the iteration stop instruction is selectively initiated by the user after analyzing the tool output result of the intelligent agent tool.
[0075] In some specific embodiments, the preliminary tool call plan generation module 12 may specifically include:
[0076] An input data generating unit, configured to embed the tool information of the agent tool and the question to be annotated into the character prompt word through a placeholder to obtain input data;
[0077] A preliminary tool call plan generating unit is used to input the input data into the pre-annotation model, and obtain a preliminary tool call plan corresponding to the question to be annotated according to the output of the pre-annotation model; the pre-annotation model is a large language model.
[0078] In some specific embodiments, the iteration module 13 may specifically include:
[0079] The model fine-tuning unit is used to fine-tune the pre-labeling model according to the preliminary tool call plan and the corresponding adjusted tool call plan generated each time during the iteration process to obtain a trained labeled model.
[0080] In some specific embodiments, the iteration module 13 may specifically include:
[0081] The adjusted tool output result acquisition unit is used to acquire the adjusted tool output results corresponding to all tool output results in the iteration process through the human-computer interaction interface;
[0082] The preference learning unit is used to perform preference learning on the pre-labeling model by using the corresponding preliminary tool call plan and the adjusted tool call plan each time in the iteration process, as well as the tool output result and its corresponding adjusted tool output result.
[0083] Furthermore, the present application also discloses an electronic device, see Figure 4 As shown, the contents in the figure cannot be considered as any limitation on the scope of use of the present application.
[0084] Figure 4 The present invention provides a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the intelligent body annotation method disclosed in any of the aforementioned embodiments.
[0085] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0086] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon include an operating system 221, a computer program 222, and data 223 including agent configuration information, etc. The storage method can be temporary storage or permanent storage.
[0087] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, so as to realize the operation and processing of the massive data 223 in the memory 22 by the processor 21, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the intelligent body labeling method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0088] Furthermore, an embodiment of the present application also discloses a computer storage medium, in which computer executable instructions are stored. When the computer executable instructions are loaded and executed by a processor, the steps of the intelligent agent labeling method disclosed in any of the aforementioned embodiments are implemented.
[0089] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0090] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0091] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0092] The above is a detailed introduction to the intelligent body labeling method, device, equipment and storage medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.< / finished>
Claims
1. An intelligent agent labeling method, characterized in that: include: Obtaining agent configuration information and questions to be annotated; the agent configuration information includes a character prompt word, a tool pool and a pre-annotated model; the tool pool is used to record agent tools; Based on the tool information of the intelligent agent tool, the character prompt words and the question to be labeled, using the pre-labeling model to generate a preliminary tool call plan corresponding to the question to be labeled, and obtaining the tool output result of the intelligent agent tool; The adjusted tool call plan corresponding to the preliminary tool call plan is obtained through a human-computer interaction interface, and the pre-labeling model is iteratively trained based on the tool information of the intelligent agent tool, the character prompt words, the question to be labeled, the adjusted tool call plan and the tool output result, so as to construct an intelligent agent based on the trained labeling model.
2. The intelligent agent labeling method according to claim 1, characterized in that: Before obtaining the agent configuration information and the question to be labeled, it also includes: The corresponding intelligent agent tool is obtained by encapsulating the collected application programming interface; the tool information of the intelligent agent tool includes a name, a description field, an access method, a parameter field and an instance field.
3. The intelligent agent labeling method according to claim 1, characterized in that: Get the agent configuration information and the questions to be labeled, including: The questions to be labeled uploaded by the administrator are obtained to obtain a question set, so as to perform preliminary tool call planning and model iteration for each question to be labeled in the question set in turn, until all the questions to be labeled are traversed to obtain a trained labeling model.
4. The intelligent agent labeling method according to claim 1, characterized in that: The iterative training comprises: Based on the tool information of the agent tool, the character prompt word, the question to be labeled, the adjusted tool call plan and the tool output result, using the pre-labeling model to generate an Nth tool call plan corresponding to the question to be labeled; Acquire the adjusted tool call plan corresponding to the Nth tool call plan through the human-computer interaction interface, and stop iterating until N reaches a preset target number of times; The iteration process also includes: An iteration stop instruction is received, and iteration is stopped according to the iteration stop instruction; the iteration stop instruction is selectively initiated by the user after analyzing the tool output result of the intelligent agent tool.
5. The intelligent agent labeling method according to claim 1, characterized in that: Based on the tool information of the agent tool, the character prompt words and the question to be labeled, using the pre-labeling model to generate a preliminary tool call plan corresponding to the question to be labeled, including: The tool information of the agent tool and the question to be annotated are embedded into the character prompt word through a placeholder to obtain input data; The input data is input into the pre-annotation model, and a preliminary tool call plan corresponding to the question to be annotated is obtained according to the output of the pre-annotation model; the pre-annotation model is a large language model.
6. The intelligent agent labeling method according to claim 1, characterized in that: Iteratively training the pre-labeled model to construct an agent based on the trained labeled model includes: According to the preliminary tool call plan generated each time during the iteration process and the corresponding adjusted tool call plan, the pre-labeling model is fine-tuned to obtain a trained labeling model.
7. The intelligent agent labeling method according to any one of claims 1 to 6, characterized in that: Iteratively training the pre-labeled model to construct an agent based on the trained labeled model includes: Obtain adjusted tool output results corresponding to all tool output results in the iteration process through a human-computer interaction interface; The preference learning of the pre-labeling model is performed by utilizing the initial tool call plan and the adjusted tool call plan corresponding to each iteration, as well as the tool output result and the corresponding adjusted tool output result.
8. An intelligent agent labeling device, characterized in that: include: An information acquisition module is used to acquire agent configuration information and questions to be annotated; the agent configuration information includes a character prompt word, a tool pool and a pre-annotated model; the tool pool is used to record agent tools; A preliminary tool call plan generation module is used to generate a preliminary tool call plan corresponding to the question to be labeled using the pre-labeling model based on the tool information of the intelligent agent tool, the character prompt words and the question to be labeled, and obtain the tool output result of the intelligent agent tool; An iterative module is used to obtain the adjusted tool call plan corresponding to the preliminary tool call plan through a human-computer interaction interface, and iteratively train the pre-labeling model based on the tool information of the intelligent agent tool, the character prompt words, the question to be labeled, the adjusted tool call plan and the tool output result, so as to construct an intelligent agent based on the trained labeled model.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the intelligent agent labeling method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein when the computer programs are executed by a processor, the intelligent agent labeling method as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method for enhancing personality and task linkage capability of role large model and related products
CN121743483A
Method for enhancing ability of character large model personality and task linkage and related product
CN121743483B