Method, device and equipment for generating model evaluation standard
By generating and optimizing model evaluation standards for agents, the problem of poor evaluation performance of model evaluation standards in the existing technology is solved, and a more efficient and objective model evaluation effect is achieved.
Patent Information
- Application Number
- CN202510080686.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The evaluation performance of existing model evaluation standards is poor, mainly due to the reliance on manual formulation, inefficient and susceptible to subjective factors, resulting in poor accuracy and consistency of evaluation results.
By obtaining standard-development requirements information and optimizing demand information, the agent automatically generates model evaluation standards, and optimizes the generated standards through the second agent to form a more suitable model evaluation standard.
It improves the efficiency and objectivity of the formulation of model evaluation standards, enhances the accuracy and consistency of the evaluation results, and thus improves the evaluation performance of the model evaluation standards.
Smart Images

Figure CN120012815A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model evaluation technology, and more particularly to a method for generating model evaluation standards. The present invention also relates to a device and equipment for generating model evaluation standards. Background Art
[0002] With the rapid development of big model technology, big models have been widely used in many fields of people's lives. In order to verify the service performance of big models in the application process, it is necessary to formulate model evaluation standards for evaluating the output of big models.
[0003] Traditional model evaluation standards are usually manually formulated. Based on manually formulated standards, their efficiency is relatively low, and the formulated model evaluation standards are easily affected by individual subjective factors. Using model evaluation standards formulated by different individuals to evaluate the output of the same model may result in inconsistent evaluation results. Therefore, the accuracy of the evaluation results of the model output evaluated based on the manually formulated model evaluation standards is relatively poor, that is, the evaluation performance of the manually formulated model evaluation standards is poor.
[0004] Based on this, how to improve the evaluation performance of model evaluation standards has become a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of this, the embodiments of the present application provide a method, apparatus and device for generating a model evaluation standard to solve the problem of poor evaluation performance of existing model evaluation standards.
[0006] According to a first aspect of an embodiment of the present application, a method for generating a model evaluation standard is provided, including: obtaining standard formulation requirement information on formulating a model evaluation standard; the standard formulation requirement information is used to indicate the generation of a model evaluation standard for evaluating model response information generated by a to-be-evaluated model in response to user question information in natural language form in at least a first dimension; inputting the standard formulation requirement information into a first intelligent agent used for standard formulation, and obtaining a first model evaluation standard output by the first intelligent agent; obtaining standard optimization requirement information on optimizing the first model evaluation standard; the standard optimization requirement information is used to indicate the optimization of the first model evaluation standard; inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent used for standard optimization, and obtaining a second model evaluation standard output by the second intelligent agent.
[0007] According to a second aspect of an embodiment of the present application, there is provided an apparatus for generating a model evaluation standard, comprising: a first requirement information acquisition module, for acquiring standard-setting requirement information for setting a model evaluation standard; the standard-setting requirement information is used to indicate the generation of a model evaluation standard for evaluating model response information generated by a model to be evaluated in response to user question information in natural language form in at least a first dimension; a standard generation module, for inputting the standard-setting requirement information into a first intelligent agent for performing standard setting, and obtaining a first model evaluation standard output by the first intelligent agent; a second requirement information acquisition module, for acquiring standard optimization requirement information for optimizing the first model evaluation standard; the standard optimization requirement information is used to indicate the optimization of the first model evaluation standard; a standard optimization module, for inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for performing standard optimization, and obtaining a second model evaluation standard output by the second intelligent agent.
[0008] According to a third aspect of an embodiment of the present application, there is provided a device for generating a model evaluation standard, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: obtain standard formulation requirement information on formulating a model evaluation standard; the standard formulation requirement information is used to indicate the generation of a model evaluation standard for evaluating model response information generated by a to-be-evaluated model in response to user question information in natural language form in at least a first dimension; input the standard formulation requirement information into a first intelligent agent used for standard formulation, to obtain a first model evaluation standard output by the first intelligent agent; obtain standard optimization requirement information on optimizing the first model evaluation standard; the standard optimization requirement information is used to indicate the optimization of the first model evaluation standard; input the standard optimization requirement information and the first model evaluation standard into a second intelligent agent used for standard optimization, to obtain a second model evaluation standard output by the second intelligent agent.
[0009] At least one embodiment in the present specification can achieve the following beneficial effects: inputting standard formulation requirement information for formulating model evaluation standards into an intelligent agent used for standard formulation, so as to automatically generate model evaluation standards based on the intelligent agent, without relying on manual standard formulation, thereby improving the efficiency of standard formulation, and also improving the objectivity of the formulated model evaluation standards. When evaluating the output of the model based on the model evaluation standards formulated by the intelligent agent, the accuracy of the evaluation results can also be improved, thereby also improving the evaluation performance of the model evaluation standards. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0011] Figure 1 It is a schematic diagram of an application scenario of a method for generating a model evaluation standard provided in an embodiment of this specification;
[0012] Figure 2 It is a flowchart of a method for generating a model evaluation standard provided in an embodiment of this specification;
[0013] Figure 3 This specification provides an embodiment corresponding to Figure 2 A swimlane flow diagram of the method for generating model evaluation criteria;
[0014] Figure 4 It is a flowchart of automatically generating a model evaluation standard, automatically evaluating a model to be evaluated based on the model evaluation standard, and automatically optimizing the model evaluation standard based on the evaluation result of the model to be evaluated, provided by an embodiment of this specification;
[0015] Figure 5 This specification provides an embodiment corresponding to Figure 2 A schematic diagram of a structure of a device for generating a model evaluation standard;
[0016] Figure 6 This specification provides an embodiment corresponding to Figure 2 A schematic diagram of the structure of a device for generating model evaluation criteria. DETAILED DESCRIPTION
[0017] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0018] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms of "a", "said" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.
[0019] It should be understood that, although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0020] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards in the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0021] Currently, the evaluation criteria for evaluating models are mainly formulated manually. Manual formulation of standards is not only inefficient, but also easily affected by subjective factors, which can easily lead to poor consistency and accuracy of evaluation results based on the evaluation criteria.
[0022] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0023] Figure 1 It is a schematic diagram of an application scenario of a method for generating model evaluation criteria provided in an embodiment of this specification.
[0024] like Figure 1 As shown, the application scenario diagram includes a user terminal 101 and a server 102.
[0025] In the embodiments of this specification, the user terminal may include but is not limited to smart phones, tablet computers, laptop computers, smart interactive devices, wearable devices, vehicle-mounted smart terminals, etc., wherein wearable devices may include but are not limited to: smart bracelets, smart watches, smart glasses, etc., and the user terminal is any one of the user terminals 101.
[0026] The server may include but is not limited to any device, equipment, platform, equipment cluster or cloud computing service center with computing and processing capabilities. The server 102 is any type of server.
[0027] A communication connection is established between the user terminal 101 and the server 102, wherein the communication connection mode may include but is not limited to a local area network connection, a wide area network connection, an Internet connection, a short-distance communication connection or other types of data network connections. The short-distance communication connection includes but is not limited to near field communication (NFC), local area network, Bluetooth, infrared and other connection modes.
[0028] The user using the user terminal 101 can send a request to the server 102 through the user terminal 101. In response to different requests sent by the user terminal 101, the server 102 can call the agent 1 to generate the model evaluation standard, call the agent 2 to optimize the generated model evaluation standard, or call the agent 3 to evaluate the model to be evaluated. If the memory resources of the user terminal 101 are sufficient to support the operation of the agent, the corresponding operations of each agent can also be run on the user terminal 101.
[0029] In the present application, a method for generating a model evaluation criterion is provided, which is described in detail one by one in the following embodiments.
[0030] Figure 2 It is a flowchart of a method for generating model evaluation criteria provided in an embodiment of this specification.
[0031] From the program perspective, the execution subject of the process can be a program installed on the device used to generate the model evaluation standard. It can be understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.
[0032] like Figure 2 As shown, the process may include the following steps.
[0033] Step 202: Obtain standard formulation requirement information on formulating model evaluation standards; the standard formulation requirement information is used to indicate the generation of model evaluation standards for evaluating model response information generated by the model to be evaluated in response to user question information in natural language form in at least a first dimension.
[0034] In the embodiments of this specification, the evaluation criteria refer to a series of indicators and criteria used to evaluate and measure the performance of a certain model, system or method. These standards are usually stored in text form and clearly describe the specific content, methods and requirements of the evaluation. The purpose of the evaluation criteria is to ensure the scientificity and objectivity of the evaluation process and provide users with reliable evaluation results. The model evaluation criteria can be a standard for evaluating the output results of the model to be evaluated, and the model to be evaluated can be a model or a group of models. In practical applications, from the perspective of business type, the model to be evaluated can include question dialogue models, sentiment analysis models, common sense reasoning models, inductive reasoning models, translation task models, text generation models, and mathematical and computational problem models, etc., wherein the question dialogue model can include factual question answering models, open-ended question answering models, explanatory question answering models, etc. From the structural form, the model to be evaluated can include a large language model, an intelligent agent, a multi-agent system, etc. Model evaluation criteria can be different for different application scenarios. For example, for scenarios where users purchase insurance, model evaluation criteria can include criteria for evaluating whether the model's response can encourage users to increase their credit limits, or whether the model's response can encourage insurance companies to increase sales. Another example: for scenarios where users ask questions, model evaluation criteria can include criteria for evaluating whether the model's response is professional, or whether the model's response is continuous, etc.
[0035] In practical applications, the model evaluation criteria can be in text form.
[0036] In the embodiments of this specification, the dimension can be understood as the angle from which the model evaluation standard evaluates the model to be evaluated. For example: the dimension of evaluating whether the model response information of the model to be evaluated is professional, the dimension of evaluating whether the model response information of the model to be evaluated is continuous, the dimension of evaluating whether the model response information of the model to be evaluated can promote the increase in credit limit, the dimension of evaluating whether the model response information of the model to be evaluated can promote sales, the dimension of evaluating whether the model response information of the model to be evaluated can reflect emotional care, the dimension of evaluating whether the model response information of the model to be evaluated is easy to understand, etc. The model evaluation standard can be a standard including one evaluation dimension, and the model evaluation standard can also be a standard including multiple evaluation dimensions, which is not limited. For example, for the scenario of users purchasing insurance, the model evaluation standard can include a standard for the increase in credit limit dimension of whether the response of the evaluation model can prompt users to increase their credit limit, and / or a standard for the sales dimension of whether the response of the evaluation model can prompt insurance companies to increase sales.
[0037] In the embodiments of this specification, the standard formulation requirement information may be the expected information extracted from the user's input information for indicating the generation of the evaluation standard, such as: please help formulate the standard for whether the response of the evaluation model in the insurance scenario can prompt the user to increase the credit limit, please help formulate the standard for whether the response of the evaluation model in the questioning scenario is professional, and please help formulate the standard for whether the response of the evaluation model in the questioning scenario is professional and continuous, etc. The standard formulation requirement information may be used to indicate the generation of a model evaluation standard for evaluating the model response information output by the evaluation model in at least the first dimension.
[0038] In practical applications, the user may be the first evaluation user who provides the standard formulation requirement information, and the first evaluation user may refer to a user who uses the method of the embodiment of this specification to obtain the model evaluation standard. Optionally, the user who uses the model evaluation standard obtained by the method based on the embodiment of this specification to evaluate the model to be evaluated may be the first evaluation user or other users except the first evaluation user, and this is not limited.
[0039] Step 204: Input the standard-setting requirement information into a first agent used for standard setting, and obtain a first model evaluation standard output by the first agent.
[0040] In the embodiments of this specification, the agent is the AI Agent model. Different from the traditional machine learning model, the agent model is a model with highly intelligent capabilities that can perceive the environment, make decisions and perform corresponding tasks. The core of the AIAgent model is its ability to think independently and call tools, so that it can gradually complete the given goals. During the use of the agent model, the specified tasks can be gradually completed according to the pre-defined roles and path planning.
[0041] In the embodiments of this specification, Chain of Thought (COT for short) is a method of generating more accurate and coherent output by guiding the model to reason step by step. The traditional prompt design usually directly gives a question or task, and the model generates output based on the input. However, this method often ignores the intermediate steps of the model in the reasoning process, resulting in the output may not be accurate or logical. The core idea of the thought chain COT is to guide the model to think according to a logical chain by introducing a series of step-by-step reasoning steps in the prompt. In this way, the model can not only generate the final answer, but also show the reasoning process, thereby improving the accuracy and interpretability of the output. In the embodiments of this specification, the prompt words used may contain information reflecting the thought chain, that is, the thought chain prompt information.
[0042] In the embodiment of this specification, the first agent may be an agent obtained by fine-tuning training using a sample containing thought chain prompt information based on a pre-trained model, and the thought chain prompt information may include step information for formulating a model evaluation standard based on standard formulation requirement information. The first agent may generate a first model evaluation standard corresponding to the standard formulation requirement information based on the step information.
[0043] Step 206: Obtain standard optimization requirement information about optimizing the first standard for model evaluation; the standard optimization requirement information is used to indicate optimization of the first standard for model evaluation.
[0044] In the embodiments of this specification, the standard optimization requirement information may be the optimization requirement information determined by the evaluation result of the model to be evaluated using the first model evaluation standard. Alternatively, the standard optimization requirement information may also be the optimization requirement information determined by the evaluator according to his / her own optimization expectations, which is not limited to this.
[0045] In actual applications, if the evaluation results reflect that the first model evaluation standard is a qualified standard, the standard optimization requirement information may be optimization requirement information used to reflect enhancement processing for the first model evaluation standard; if the evaluation results reflect that the first model evaluation standard is an unqualified standard, the standard optimization requirement information may be optimization requirement information used to reflect modification processing for the first model evaluation standard.
[0046] Optionally, the standard optimization requirement information can be used to indicate that the first standard of the model evaluation is optimized in at least a first dimension; or optionally, the standard optimization requirement information can be used to indicate that the first standard of the model evaluation is optimized in at least a second dimension, without limitation.
[0047] In the embodiments of this specification, the standard optimization requirement information may be requirement information indicating the optimization standard, such as: please optimize the increase standard of the model evaluation standard in the insurance scenario, please optimize the professional standard of the model evaluation standard in the questioning scenario, and please optimize the professional standard and continuity standard of the model evaluation standard in the questioning scenario, etc.
[0048] Optionally, the standard optimization requirement information may be provided by the first evaluation user; or optionally, the standard optimization requirement information may be provided by a second evaluation user other than the first evaluation user. Optionally, the user who uses the model evaluation standard obtained by the method according to the embodiment of this specification to evaluate the model to be evaluated may be the second evaluation user or other users other than the second evaluation user, and this is not limited.
[0049] Step 208: Input the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for standard optimization, and obtain a second model evaluation standard output by the second intelligent agent.
[0050] In the embodiment of this specification, the second agent may be an agent obtained by fine-tuning training using samples containing thought chain prompt information based on the pre-trained model, and the thought chain prompt information may include step information for optimizing the model evaluation standard based on the standard optimization requirement information. The second agent may perform standard optimization on the first model evaluation standard based on the step information to generate a second model evaluation standard corresponding to the standard optimization requirement information.
[0051] In the embodiment of this specification, the second agent can be used to automatically optimize the first model evaluation standard generated by the first agent, so that the model evaluation standard can adapt to changes in different tasks and environments, thereby improving the adaptability of the model evaluation standard.
[0052] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification can be interchanged according to actual needs, or some steps can be omitted or deleted.
[0053] Figure 2 The method in the embodiment of the invention can automatically generate model evaluation standards by using the first agent without relying on manual standard setting, thereby improving the efficiency of standard setting and the objectivity of the model evaluation standards. When the output of the model is evaluated based on the model evaluation standards formulated by the first agent, the accuracy of the evaluation results can be improved, thereby improving the evaluation performance of the model evaluation standards. The second agent is used to automatically optimize the first model evaluation standard generated for the first agent, so that the model evaluation standard can adapt to changes in different tasks and environments, thereby improving the adaptability of the model evaluation standard.
[0054] based on Figure 2 The present specification also provides some improved implementation methods of the method, which are described below.
[0055] In the embodiments of this specification, a specific embodiment is proposed for a method of evaluating a first standard based on a first agent generation model.
[0056] Optionally, the step of inputting the standard-setting requirement information into a first intelligent agent used for standard setting to obtain a first model evaluation standard output by the first intelligent agent may specifically include: inserting the standard-setting requirement information into a standard-setting prompt template containing standard-setting thought chain information to obtain standard-setting prompt information; the standard-setting thought chain information is used to represent a first logic for generating a standard based on the standard-setting requirement information; and inputting the standard-setting prompt information into a first intelligent agent used for standard setting to obtain a first model evaluation standard generated by the first intelligent agent according to the first logic.
[0057] In the embodiments of this specification, the standard formulation prompt template may include at least standard formulation thought chain information and a "user demand" section. The standard formulation requirement information may be inserted into the "user demand" section of the standard formulation prompt template. The expression "user demand" is only exemplary, and may be other expressions in actual application, such as "user requirements", etc.
[0058] In practical applications, the standard formulation thinking chain information may be information for guiding the model to think step by step according to the steps of the logic chain of the first logic, wherein the first logic may be logic information for generating the evaluation standard based on the standard formulation requirement information.
[0059] In the embodiment of this specification, the standard formulation requirement information is input into the "user demand" part of the standard formulation prompt template to obtain the standard formulation prompt information including the standard formulation thought chain information and the standard formulation requirement information. The standard formulation prompt information is input into the first intelligent agent so that the first intelligent agent responds to the standard formulation requirement information and generates the first model evaluation standard according to the logical steps indicated by the standard formulation thought chain information.
[0060] In the embodiment of the present specification, the standard formulation prompt information for generating the first standard for model evaluation includes thought chain information, so that the first intelligent agent can think according to the logical steps indicated by the thought chain information, thereby improving the accuracy of the first intelligent agent generating the first standard for model evaluation.
[0061] In order to further improve the accuracy of the first standard for evaluating the model generated by the first intelligent agent, a standard reference case may also be provided to the first intelligent agent.
[0062] Optionally, before inputting the standard-setting requirement information into the first intelligent agent for standard setting and obtaining the first model evaluation standard output by the first intelligent agent, it may also include: obtaining a first standard reference case; the first standard reference case includes user problem samples and case analysis information on at least the first dimension.
[0063] In the embodiments of this specification, the first standard reference case may be a case related to the current task extracted from historical data. In practical applications, the first standard reference case related to the current task may be extracted from historical data by experts.
[0064] In the embodiments of the present specification, the user question sample may be a question in natural language form raised by the user. For example, in the scenario where the user purchases insurance, the user question sample may be "Is it better to pay monthly or annually for this insurance?"
[0065] In practical applications, the case analysis information may include the case analysis conclusion and the case analysis reasons for reaching the case analysis conclusion.
[0066] Optionally, when the first standard reference case does not include a model response sample, the case analysis information specifically includes question analysis information for the user question sample in the at least first dimension.
[0067] Furthermore, the problem analysis information may include a problem analysis conclusion and a problem analysis reason for the problem analysis conclusion. In practical applications, for example, in a scenario where a user purchases insurance, for the user question example "Is it better to pay monthly or annually for this insurance?", the corresponding problem analysis information may be "This case is a case of increasing the insurance limit, and annual payment is better than monthly payment. Users should be encouraged to pay annually." The problem analysis conclusion is "Users should be encouraged to pay annually," and the problem analysis reason is "This case is a case of increasing the insurance limit, and annual payment is better than monthly payment."
[0068] Optionally, the step of inputting the standard-setting requirement information into a first intelligent agent for standard setting to obtain a first model evaluation standard output by the first intelligent agent may specifically include: inputting the standard-setting requirement information and the first standard reference case into a first intelligent agent for standard setting to obtain a first model evaluation standard output by the first intelligent agent.
[0069] In an embodiment of the present specification, the standard-setting requirement information and the first standard reference case can be inserted into a standard-setting prompt template containing standard-setting thought chain information to obtain standard-setting prompt information; then the standard-setting prompt information is input into a first intelligent agent used for standard setting to obtain a first standard for model evaluation generated by the first intelligent agent based on the first logic.
[0070] In actual application, the standard formulation prompt template may also include a "related case" section, and the first reference case may be inserted into the "related case" section of the standard formulation prompt template. The expression "related case" is only exemplary, and other expressions may be used in actual application, such as "reference case".
[0071] In the embodiment of the present specification, the standard formulation requirement information is input into the "user demand" part of the standard formulation prompt template, and the first reference case is inserted into the "related case" part of the standard formulation prompt template to obtain the standard formulation prompt information including the standard formulation thought chain information, the standard formulation requirement information and the first reference case. The standard formulation prompt information is input into the first intelligent agent so that the first intelligent agent responds to the standard formulation requirement information, refers to the first reference case, and generates the first standard for model evaluation according to the logical steps indicated by the standard formulation thought chain information.
[0072] In the embodiment of this specification, the first intelligent agent also refers to the first standard reference case in the process of generating the first standard for model evaluation, so that the first intelligent agent can generate the first standard for model evaluation that meets the requirements, thereby further improving the accuracy of the first standard for model evaluation.
[0073] Optionally, the first standard reference case may also include a model response sample for the user question sample; the case analysis information specifically includes response analysis information for the model response sample in at least the first dimension.
[0074] Furthermore, the reply analysis information may include the reply analysis conclusion and the reply analysis reason for the reply analysis conclusion. In practical applications, for example, in the scenario where users purchase insurance, for the user question example "What information is needed for insurance claims?", the corresponding model reply example may be "Claims usually require relevant documents, such as copies of ID cards, and may also require medical certificates or accident reports, which vary from insurance company to insurance company. Which company and which insurance did you buy? I will help you analyze it in detail." The corresponding reply analysis information may be "There are no factual errors or logical errors in the answer to this case, and it responds to the question positively, with good professionalism; the answer can connect the previous and the next, and guide the user to continue the conversation, with good continuity." The reply analysis conclusion is "good professionalism; good continuity", and the reply analysis reason is "There are no factual errors or logical errors in the answer to this case, and it responds to the question positively; the answer can connect the previous and the next, and guide the user to continue the conversation."
[0075] In the embodiments of the present specification, the first standard reference case may also include model response samples, thereby providing more reference information for the first intelligent agent, so as to be more conducive to the first intelligent agent generating a model evaluation first standard that meets the requirements, thereby further improving the accuracy of the model evaluation first standard.
[0076] In order to improve the diversity of the first standard reference case, the first standard reference case may also include positive reference cases and negative reference cases.
[0077] Optionally, the first standard reference case includes at least one of a positive reference case and a negative reference case; the model response sample in the positive reference case meets the target evaluation requirements in at least the first dimension; and the model response sample in the negative reference case does not meet the target evaluation requirements in at least the first dimension.
[0078] In practical applications, for example, in the scenario of purchasing insurance, the response of the model to be tested is evaluated to see whether it meets the standards of the professional dimension. If the model response meets the standards of the professional dimension, the model response is a positive reference case; if the model response does not meet the standards of the professional dimension, the model response is a negative reference case.
[0079] In the embodiments of this specification, the first standard reference case may include a positive reference case and a negative reference case. The first standard reference case may be a case formulated by an expert, or a case generated by other optimized models, and there is no limitation on this.
[0080] In practical applications, the first standard of model evaluation under ideal conditions can meet the target evaluation requirements, that is, the target evaluation requirements can be understood as evaluating the positive cases formulated by experts according to the first standard of model evaluation, and the evaluation results need to reflect that the positive cases formulated by experts are positive reference cases; and evaluating the negative cases formulated by experts according to the first standard of model evaluation, and the evaluation results need to reflect that the negative cases formulated by experts are negative reference cases. Among them, the target evaluation requirements are the requirements proposed by the evaluators for evaluating the output information of the model to be evaluated, and the first standard of model evaluation can be a standard document formulated by the first intelligent agent to characterize the target evaluation requirements.
[0081] In an embodiment of the present specification, the first standard reference case provided to the first intelligent agent may include both positive reference cases and negative reference cases, thereby improving the diversity of the first standard reference cases and further improving the accuracy of the first standard for model evaluation formulated by the first intelligent agent.
[0082] In order to improve the standardization of the first standard for model evaluation generated by the first intelligent agent, the historical standard for model evaluation may also be input into the first intelligent agent.
[0083] Optionally, before inputting the standard-setting requirement information into a first intelligent agent for standard setting and obtaining the first model evaluation standard output by the first intelligent agent, the process may further include: obtaining a historical model evaluation standard.
[0084] In the embodiments of this specification, the historical model evaluation standard can be a historical evaluation standard read from an external database. Specifically, the historical model evaluation standard can be read from the external database based on the retrieval enhancement technology (rag), so that the first agent can not only use the powerful generation ability of the model, but also supplement the insufficient information in the external database, so as to avoid the model relying only on the limited knowledge obtained during pre-training.
[0085] Optionally, the step of inputting the standard-setting requirement information into a first intelligent agent for standard setting to obtain a first model evaluation standard output by the first intelligent agent may specifically include: inputting the standard-setting requirement information, the first standard reference case and the model evaluation historical standard into a first intelligent agent for standard setting to obtain a first model evaluation standard output by the first intelligent agent.
[0086] In an embodiment of the present specification, the standard formulation requirement information, the first standard reference case and the model evaluation historical standard can be inserted into a standard formulation prompt template containing standard formulation thought chain information to obtain standard formulation prompt information; then the standard formulation prompt information is input into the first intelligent agent used for standard formulation to obtain the first model evaluation first standard generated by the first intelligent agent according to the first logic.
[0087] In actual application, the standard setting prompt template may also include a "historical evaluation standard" section, and the model evaluation historical standard may be inserted into the "historical evaluation standard" section of the standard setting prompt template. The expression "historical evaluation standard" is only exemplary, and other expressions may also be used in actual application, such as "reference evaluation standard".
[0088] In the embodiment of this specification, the standard formulation requirement information is input into the "user demand" part of the standard formulation prompt template, the first reference case is inserted into the "related case" part of the standard formulation prompt template, and the model evaluation history standard is inserted into the "historical evaluation standard" part of the standard formulation prompt template to obtain the standard formulation prompt information containing the standard formulation thought chain information, the standard formulation requirement information, the first reference case and the model evaluation history standard. The standard formulation prompt information is input into the first intelligent agent so that the first intelligent agent responds to the standard formulation requirement information, refers to the first reference case and the model evaluation history standard, and generates the first model evaluation standard according to the logical steps indicated by the standard formulation thought chain information.
[0089] In the embodiments of the present specification, the first intelligent agent also refers to the historical model evaluation standard during the process of generating the first model evaluation standard, thereby further prompting the first intelligent agent to generate the first model evaluation standard that meets the requirements (such as format requirements), thereby further improving the accuracy of the first model evaluation standard.
[0090] Optionally, the standard-setting prompt template may also include standard-setting personality information, which may be information used to reflect the knowledge and skills that the first intelligent agent needs to possess. For example, the standard-setting personality information may be "You are a senior model evaluation standard-setting expert. You are understanding user demands and developing evaluation standards that meet the requirements based on relevant cases provided by users and with reference to historical evaluation standards." Another example is that the standard-setting personality information may be "You are a senior evaluation standard-setting expert. You are understanding user demands and developing evaluation standards that meet the requirements based on relevant cases provided by users and with reference to historical evaluation standards."
[0091] In the embodiment of this specification, the standard formulation requirement information is input into the "user demand" part of the standard formulation prompt template, the first reference case is inserted into the "related case" part of the standard formulation prompt template, and the model evaluation history standard is inserted into the "historical evaluation standard" part of the standard formulation prompt template to obtain the standard formulation prompt information containing the standard formulation thinking chain information, the standard formulation requirement information, the first reference case, the model evaluation history standard and the standard formulation human setting information. The standard formulation prompt information is input into the first intelligent agent so that the first intelligent agent responds to the standard formulation requirement information, refers to the first reference case and the model evaluation history standard, combines the knowledge and skills reflected by the standard formulation human setting information, and generates the first model evaluation standard according to the logical steps indicated by the standard formulation thinking chain information.
[0092] In the embodiment of the present specification, the first intelligent agent, in the process of generating the first standard for model evaluation, possesses the knowledge and skills reflected by the standard setting personal information, which can further encourage the first intelligent agent to generate the first standard for model evaluation that meets the requirements, and further improve the accuracy of the first standard for model evaluation.
[0093] In order to facilitate technical personnel in this field to understand the present solution, a specific example is provided in the embodiments of this specification for a standard setting prompt template, and Example 1 is as follows.
[0094] ##Information on the person setting the standards:
[0095] You are a senior expert in setting evaluation standards. You are understanding user demands, combining the cases provided by users, and formulating evaluation standards that meet the requirements based on reference to historical evaluation standards.
[0096] ##User demands:
[0097] Please help formulate the standards for increasing the credit limit in insurance scenarios, and specify which users' problems should have their credit limit increased and which users' problems should not have their credit limit increased.
[0098] ##Related cases:
[0099] Refer to Case 1, the user asked: What information is needed for insurance claims? Analysis and conclusion: The user may have insurance claims needs, but the user did not provide specific insurance products. The user should be informed of the general materials required for insurance claims and asked about the specific products the user is consulting.
[0100] Refer to case 2. The user asked: What information is needed for insurance claims? Model response: Claims usually require relevant documents, such as a copy of your ID card. Medical certificates or accident reports may also be required, depending on the insurance company. Which company and which insurance did you buy? I will help you analyze it in detail. Analysis and conclusion: There are no factual or logical errors in the answer to this case, and the question is answered positively, so the answer is quite professional.
[0101] Refer to case 3, the user asked: What information is needed for insurance claims? Model response: Relevant documents are usually required for claims, you can consult the staff for details. Analysis and conclusion: The answer to this case did not directly address the question, and the answer was not professional.
[0102] ##Historical evaluation criteria: None.
[0103] ##Standard setting thinking chain information:
[0104] Step 1: Please read the "User Demands" and fully understand them. Step 2: Please carefully understand the "User Demands" and formulate new evaluation criteria based on the "Related Cases" given by users. Step 3: Please follow the target format for the output format, which is:
[0105] json{
[0106] \"Overview of the new standard\":\"\",
[0107] \"Specific content of the new standard\":\"\",
[0108] \"Case examples and explanations for each section of the new standard\":\"\"
[0109] }.
[0110] Optionally, the standard formulation thinking chain information in the standard formulation prompt template may include output format information. For example, the standard formulation thinking chain information in Example 1 includes output format information.
[0111] In the embodiments of this specification, a specific embodiment is also proposed for the method of generating the first intelligent agent.
[0112] Optionally, before inputting the standard-setting requirement information into the first intelligent agent used for standard setting and obtaining the first model evaluation standard output by the first intelligent agent, the process may also include: obtaining a first training sample set; a training sample in the first training sample set includes a standard-setting requirement information sample; the standard-setting requirement information sample is used to indicate the generation of a prediction standard for evaluating the model response information output by the model to be evaluated; inserting the standard-setting requirement information sample into a standard-setting prompt template containing standard-setting thought chain information to obtain a standard-setting prompt information sample; the standard-setting thought chain information is used to represent the first logic for generating a standard based on the standard-setting requirement information; inputting the standard-setting prompt information sample into a pre-trained large language model to obtain a first prediction standard generated by the large language model according to the first logic; and fine-tuning the large language model based on the first prediction standard to obtain the first intelligent agent.
[0113] In the embodiments of this specification, the standard-setting requirement information samples, standard-setting prompt templates, standard-setting thought chain information, and standard-setting prompt information samples involved in the fine-tuning training process of generating the first intelligent agent can be respectively consistent in form with the standard-setting requirement information, standard-setting prompt templates, standard-setting thought chain information, and standard-setting prompt information involved in the process of the first intelligent agent generating a model to evaluate the first standard. The explanation of the relevant content above can be referred to and will not be repeated here.
[0114] Optionally, the fine-tuning training of the large language model based on the first prediction standard to obtain the first intelligent agent may include: obtaining manual feedback information of the evaluation user on the first prediction standard; and fine-tuning the large language model according to the manual feedback information. The manual feedback information may specifically include information on whether the evaluation user believes that the first prediction standard meets expectations. In practical applications, the device for fine-tuning the large speech model may have a human-computer interaction interface, and the evaluation user inputs the manual feedback information on the first prediction standard into the human-computer interaction interface. If the manual feedback information indicates that the first prediction standard meets expectations, the large language model after the current fine-tuning training is determined to be the first intelligent agent; if the manual feedback information indicates that the first prediction standard does not meet expectations, the large language model after the current fine-tuning training is fine-tuned according to the manual feedback information until the preset training termination condition is met to obtain the first intelligent agent.
[0115] Optionally, the fine-tuning training of the large language model based on the first prediction standard to obtain the first intelligent agent may include: fine-tuning the large language model based on the difference between the first prediction standard and a preset first sample standard corresponding to the standard formulation requirement information sample to obtain the first intelligent agent. Specifically, a first loss function value between the first prediction standard and the first sample standard is calculated, and then the parameters of the large language model are adjusted according to the first loss function value until a preset training termination condition is met to obtain the first intelligent agent.
[0116] In actual application, the preset training termination condition may, optionally, include a loss function value being less than or equal to a preset loss threshold, or, optionally, may include a training round reaching a preset training number threshold.
[0117] In an embodiment of the present specification, a standard-setting prompt information sample is generated based on a standard-setting prompt template that includes standard-setting thought chain information, and the standard-setting prompt information sample is used to perform fine-tuning training on a pre-trained large language model, so that the large language model can think and learn according to the logical steps indicated by the thought chain information, so as to improve the performance of the first intelligent agent after fine-tuning training, and further improve the accuracy of the model evaluation standard generated based on the first intelligent agent.
[0118] In order to further improve the performance of the first intelligent agent, the standard setting prompt information sample can also include reference case information.
[0119] Optionally, before inserting the standard-setting requirement information sample into a standard-setting prompt template containing standard-setting thought chain information to obtain the standard-setting prompt information sample, it may also include: obtaining a first reference case; the first reference case includes a first question sample and first analysis information on at least the first dimension.
[0120] In the embodiments of this specification, the explanation of the first reference case can refer to the explanation of the first standard reference case, which will not be repeated here. In practical applications, the first reference case and the first standard reference case have different uses. The first reference case is used in the model training stage for fine-tuning the pre-trained large language model, and the first standard reference case is used in the model reasoning stage for the first agent to formulate the model evaluation standard.
[0121] Optionally, inserting the standard-setting requirement information sample into a standard-setting prompt template containing standard-setting thought chain information to obtain a standard-setting prompt information sample may specifically include: inserting the standard-setting requirement information sample and the first reference case into a standard-setting prompt template containing standard-setting thought chain information to obtain a standard-setting prompt information sample.
[0122] In an embodiment of the present specification, the standard-setting requirement information sample and the first reference case can be inserted into a standard-setting prompt template containing standard-setting thought chain information to obtain a standard-setting prompt information sample; then the standard-setting prompt information sample can be input into a pre-trained large language model to obtain a first prediction standard generated by the large language model based on the first logic; and then the large language model can be fine-tuned based on the first prediction standard to obtain a first intelligent agent.
[0123] In the embodiments of this specification, in the process of generating the first intelligent agent, a first reference case is also referred to so that the generated first intelligent agent can better meet the requirements, so as to improve the service performance of the first intelligent agent, and further improve the accuracy of the model evaluation standard generated by the first intelligent agent.
[0124] In the embodiments of this specification, a specific embodiment is also proposed for the method of optimizing the first standard of the model evaluation generated by the first intelligent agent using the second intelligent agent.
[0125] Optionally, the step of inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for performing standard optimization to obtain a second model evaluation standard output by the second intelligent agent may specifically include: inserting the standard optimization requirement information and the first model evaluation standard into a standard optimization prompt template containing standard optimization thinking chain information to obtain standard optimization prompt information; the standard optimization thinking chain information is used to represent the second logic for optimizing the standard based on the standard optimization requirement information; and inputting the standard optimization prompt information into a second intelligent agent for performing standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
[0126] In the embodiment of this specification, the standard optimization prompt template may include at least standard optimization thought chain information, a "user demand" part, and a "standard to be optimized" part. The standard optimization demand information may be inserted into the "user demand" part of the standard optimization prompt template, and the first model evaluation standard may be inserted into the "standard to be optimized" part of the standard optimization prompt template.
[0127] In practical applications, the standard optimization thinking chain information may be information for guiding the model to think step by step according to the steps of the logic chain of the second logic, wherein the second logic may be logic information for optimizing the standard based on the standard optimization requirement information.
[0128] In the embodiment of this specification, the standard optimization requirement information is input into the "user demand" part of the standard optimization prompt template, and the first model evaluation standard is inserted into the "standard to be optimized" part of the standard optimization prompt template to obtain standard optimization prompt information including standard optimization thinking chain information, standard optimization requirement information and the first model evaluation standard. The standard optimization prompt information is input into the second intelligent agent so that the second intelligent agent responds to the standard optimization requirement information and optimizes the first model evaluation standard according to the logical steps indicated by the standard optimization thinking chain information to obtain the second model evaluation standard.
[0129] In the embodiments of the present specification, the standard optimization prompt information used to optimize the first model evaluation standard includes thought chain information, so that the second intelligent agent can think according to the logical steps indicated by the thought chain information, thereby improving the accuracy of the second model evaluation standard obtained after the second intelligent agent optimizes the first model evaluation standard.
[0130] In the embodiment of this specification, the second agent can be used to automatically update and optimize the first model evaluation standard generated by the first agent, so that the model evaluation standard can adapt to changes in different tasks and environments, thereby improving the adaptability of the model evaluation standard.
[0131] In order to further improve the accuracy of the second model evaluation standard obtained after the second intelligent agent optimizes the first model evaluation standard, a standard reference case can also be provided to the second intelligent agent.
[0132] Optionally, the process of inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for standard optimization, before obtaining the second model evaluation standard output by the second intelligent agent, may also include: obtaining a second standard reference case; the second standard reference case includes user problem samples and first case analysis information on at least the first dimension, or the second standard reference case includes user problem samples and second case analysis information on at least the second dimension.
[0133] In the embodiments of this specification, the at least first dimension may be represented as the dimension information corresponding to the first standard of model evaluation, and the at least second dimension may be represented as the dimension information other than the dimension information corresponding to the first standard of model evaluation. For example, if the first standard of model evaluation is a standard for evaluating whether the output of a model is professional, then the dimension information corresponding to the first standard of model evaluation is the professional dimension, then the first case analysis information may be the analysis information in terms of the professional dimension, and the second case analysis information may be the analysis information in terms of other dimensions except the professional dimension.
[0134] In the embodiment of this specification, the user question example may be a question in natural language form raised by the user. For example, in the scenario where the user purchases insurance, the user question example may be "What information is required for insurance claims?"
[0135] In practical applications, the first case analysis information and the second case analysis information may both include a case analysis conclusion and case analysis reasons for reaching the case analysis conclusion.
[0136] Optionally, when the second standard reference case does not include a model response sample, the case analysis first information specifically includes question analysis information for the user question sample in the at least first dimension. Alternatively, the case analysis second information specifically includes question analysis information for the user question sample in the at least second dimension.
[0137] Furthermore, the problem analysis information may include a problem analysis conclusion and the problem analysis reasons for reaching the problem analysis conclusion.
[0138] Optionally, the second standard reference case may also include model response samples for the user question samples; the case analysis first information specifically includes response analysis information for the model response samples on at least the first dimension; or, the case analysis second information specifically includes response analysis information for the model response samples on at least the second dimension.
[0139] Optionally, the second standard reference case may also include evaluation result information after evaluating the model response sample output by the model in response to the user question sample based on the first model evaluation standard, and the first case analysis information may specifically include evaluation and analysis information for the evaluation result information in the at least first dimension; or the second case analysis information may specifically include evaluation and analysis information for the evaluation result information in the at least second dimension. The evaluation and analysis information may be analysis information determined by the evaluator based on the evaluation result.
[0140] Furthermore, the evaluation and analysis information may include the evaluation and analysis conclusions and the evaluation and analysis reasons for the evaluation and analysis conclusions. For example, in the scenario of users purchasing insurance, a user question sample may be "What information is required for insurance claims?" A model response sample may be "Claims usually require relevant documents, such as a copy of your ID card, and may also require a medical certificate or accident report, which varies from insurance company to insurance company. Which company and which insurance did you buy? Let me analyze it in detail for you?" The evaluation result information for the model response sample based on the first model evaluation standard is "The model response is quite professional."
[0141] The evaluation and analysis information for the evaluation result information may be "There are no factual errors or logical errors in the model's answer, and it responds to the question positively, which indicates that the model's answer is highly professional, and the result of evaluating the model's answer based on the first model evaluation standard is also that the model's answer is highly professional, so the first model evaluation standard performs well in evaluating whether the model's answer is professional." Among them, the evaluation and analysis conclusion is "The first model evaluation standard performs well in evaluating whether the model's answer is professional." The evaluation and analysis reason is "There are no factual errors or logical errors in the model's answer, and it responds to the question positively, which indicates that the model's answer is highly professional, and the result of evaluating the model's answer based on the first model evaluation standard is also that the model's answer is highly professional."
[0142] Optionally, the step of inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for performing standard optimization to obtain the second model evaluation standard output by the second intelligent agent may specifically include: inputting the standard optimization requirement information, the first model evaluation standard and the second standard reference case into a second intelligent agent for performing standard optimization to obtain the second model evaluation standard output by the second intelligent agent.
[0143] Specifically, the standard optimization requirement information, the first model evaluation standard and the second standard reference case can be inserted into a standard optimization prompt template containing standard optimization thinking chain information to obtain standard optimization prompt information; then the standard optimization prompt information can be input into a second intelligent agent used for standard optimization to obtain the second model evaluation standard output by the second intelligent agent.
[0144] In the embodiment of this specification, the standard optimization prompt template may further include a "related case" section. In practical applications, the second standard reference case may be inserted into the "related case" section of the standard optimization prompt template.
[0145] In the embodiment of this specification, the standard optimization requirement information is input into the "user demand" part of the standard optimization prompt template, the first model evaluation standard is inserted into the "standard to be optimized" part of the standard optimization prompt template, and the second standard reference case is inserted into the "related case" part of the standard optimization prompt template to obtain standard optimization prompt information including standard optimization thinking chain information, standard optimization requirement information, the first model evaluation standard and the second standard reference case. The standard optimization prompt information is input into the second intelligent agent so that the second intelligent agent responds to the standard optimization requirement information, refers to the second standard reference case, and optimizes the first model evaluation standard according to the logical steps indicated by the standard optimization thinking chain information to obtain the second model evaluation standard.
[0146] In an embodiment of the present specification, the standard optimization prompt information used to optimize the first standard of model evaluation also includes a second standard reference case, so that the second intelligent agent can refer to the second standard reference case during the process of optimizing the first standard of model evaluation, thereby further improving the accuracy of the second standard of model evaluation obtained after the second intelligent agent optimizes the first standard of model evaluation.
[0147] Optionally, the standard optimization prompt template may also include standard optimization character information, which may be information used to reflect the knowledge and skills that the second agent needs to possess. For example, the standard optimization character information may indicate that you are a senior expert in setting evaluation standards, and are currently understanding user demands, combining the cases provided by users, and optimizing the current evaluation standards based on the original evaluation standards.
[0148] In the embodiment of this specification, the standard optimization demand information is input into the "user demand" part of the standard optimization prompt template, the first model evaluation standard is inserted into the "standard to be optimized" part of the standard optimization prompt template, and the second standard reference case is inserted into the "related case" part of the standard optimization prompt template to obtain standard optimization prompt information containing standard optimization thinking chain information, standard optimization demand information, first model evaluation standard, second standard reference case and standard optimization personality information. The standard optimization prompt information is input into the second intelligent agent so that the second intelligent agent responds to the standard optimization demand information, refers to the second standard reference case, combines the knowledge and skills reflected by the standard optimization personality information, and optimizes the first model evaluation standard according to the logical steps indicated by the standard optimization thinking chain information to obtain the second model evaluation standard.
[0149] In the embodiment of this specification, the standard optimization prompt information used to optimize the first model evaluation standard also includes standard optimization personality information to further improve the accuracy of the second model evaluation standard obtained after the second intelligent agent optimizes the first model evaluation standard.
[0150] In the embodiments of this specification, the relevant examples of the standard optimization prompt template can refer to Example 1 corresponding to the standard formulation prompt template. Based on the standard optimization personality information, standard optimization requirement information, second standard reference case, and standard optimization thinking chain information in the standard optimization prompt template, the standard formulation personality information, standard formulation requirement information, first standard reference case, and standard formulation thinking chain information in Example 1 are replaced to obtain Example 2 corresponding to the standard optimization prompt template, wherein Example 2 also needs to add the first standard for model evaluation.
[0151] In the embodiments of this specification, a specific embodiment is also proposed for the method of generating a second intelligent agent.
[0152] Optionally, the step of inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for standard optimization and obtaining the second model evaluation standard output by the second intelligent agent may also include: obtaining a second training sample set; a training sample in the second training sample set includes a standard optimization requirement information sample and a second sample standard to be optimized; the standard optimization requirement information sample is used to indicate optimization of the second sample standard; inserting the standard optimization requirement information sample and the second sample standard into a standard optimization prompt template containing standard optimization thinking chain information to obtain a standard optimization prompt information sample; the standard optimization thinking chain information is used to represent the second logic for optimizing the standard based on the standard optimization requirement information; inputting the standard optimization prompt information sample into a pre-trained large language model to obtain a second prediction standard obtained after the large language model optimizes the second sample standard according to the second logic; and fine-tuning the large language model based on the second prediction standard to obtain a second intelligent agent.
[0153] In the embodiments of this specification, for the explanation of the standard optimization requirement information samples, standard optimization prompt templates, standard optimization thinking chain information, standard optimization prompt information samples and other contents involved in the fine-tuning training process of generating the second intelligent agent, reference can be made to the explanation of the standard optimization requirement information, standard optimization prompt templates, standard optimization thinking chain information, standard optimization prompt information and other contents involved in the process of evaluating the first standard of the above-mentioned second intelligent agent optimization model, and no further details will be given here.
[0154] Optionally, the fine-tuning training of the large language model based on the second prediction standard to obtain the second intelligent agent may include: obtaining manual feedback information of the evaluation user on the second prediction standard; and fine-tuning the large language model according to the manual feedback information. The manual feedback information may specifically include information on whether the evaluation user believes that the second prediction standard meets expectations. In practical applications, the device for fine-tuning the large speech model may have a human-computer interaction interface, and the evaluation user inputs the manual feedback information on the second prediction standard into the human-computer interaction interface. If the manual feedback information indicates that the second prediction standard meets expectations, the large language model after the current fine-tuning training is determined to be the second intelligent agent; if the manual feedback information indicates that the second prediction standard does not meet expectations, the large language model after the current fine-tuning training is fine-tuned according to the manual feedback information until the preset training termination condition is met to obtain the second intelligent agent.
[0155] Optionally, the fine-tuning training of the large language model based on the second prediction standard to obtain the second agent may include: fine-tuning the large language model based on the difference between the second prediction standard and a preset third sample standard corresponding to the standard optimization requirement information sample. Specifically, a second loss function value between the second prediction standard and the third sample standard is calculated, and then the parameters of the large language model are adjusted according to the second loss function value until a preset training termination condition is met to obtain the second agent.
[0156] In practical applications, the large language model used when training the second agent and the large language model used when training the first agent can be structurally the same or different base large models. The base large model can be a pre-trained large language model, such as a model of the GPT series or a model of the Qianwen series, etc., without limitation.
[0157] In an embodiment of the present specification, a standard optimization prompt information sample is generated based on a standard optimization prompt template containing standard optimization thought chain information, and the standard optimization prompt information sample is used to perform fine-tuning training on a pre-trained large language model, so that the large language model can think and learn according to the logical steps indicated by the thought chain information, so as to improve the performance of the second intelligent agent after fine-tuning training, and further improve the accuracy of the model evaluation standard after optimization based on the second intelligent agent.
[0158] In order to further improve the performance of the second agent, the standard optimization prompt information sample can also include reference case information.
[0159] Optionally, before inserting the standard optimization requirement information sample and the second sample standard into a standard optimization prompt template containing standard optimization thinking chain information to obtain the standard optimization prompt information sample, it may also include: obtaining a second reference case; the second reference case includes a second problem sample and second analysis information on at least the first dimension, or the second reference case includes a second problem sample and third analysis information on at least the second dimension.
[0160] In the embodiments of this specification, the explanation of the second reference case can refer to the explanation of the second standard reference case, which will not be repeated here.
[0161] In actual applications, the second reference case and the second standard reference case have different uses. The second reference case is used in the model training stage for fine-tuning the pre-trained large language model, and the second standard reference case is used in the model reasoning stage of the second agent optimization model evaluation standard.
[0162] Optionally, the step of inserting the standard optimization requirement information sample and the second sample standard into a standard optimization prompt template containing standard optimization thinking chain information to obtain a standard optimization prompt information sample may specifically include: inserting the standard optimization requirement information sample, the second sample standard and the second reference case into a standard optimization prompt template containing standard optimization thinking chain information to obtain a standard optimization prompt information sample.
[0163] In an embodiment of the present specification, the standard optimization requirement information sample, the second sample standard and the second reference case can be inserted into a standard optimization prompt template containing standard optimization thinking chain information to obtain a standard optimization prompt information sample; then the standard optimization prompt information sample can be input into a pre-trained large language model to obtain a second prediction standard generated by the large language model according to the second logic; and then the large language model can be fine-tuned based on the second prediction standard to obtain a first intelligent agent.
[0164] In the embodiments of this specification, in the process of generating the second intelligent agent, a second reference case is also referred to so that the generated second intelligent agent can better meet the requirements, so as to improve the service performance of the second intelligent agent, and further improve the accuracy of the model evaluation standard after the second intelligent agent is optimized.
[0165] In the embodiments of this specification, a specific embodiment is also proposed for a method of automatically evaluating the model output based on the first model evaluation standard.
[0166] Optionally, after inputting the standard-setting requirement information into the first intelligent agent used for standard setting and obtaining the first model evaluation standard output by the first intelligent agent, the method may also include: obtaining a test case generated by the model to be evaluated; the test case includes actual problem information input by a user into the model to be evaluated and actual response information generated by the model to be evaluated in response to the actual problem information; inputting the test case and the first model evaluation standard into a third intelligent agent used for model evaluation, and obtaining an evaluation result obtained by the third intelligent agent through evaluation of the test case according to the first model evaluation standard; the evaluation result includes conclusion information and reason information; the conclusion information includes first conclusion information or second conclusion information, the first conclusion information is used to indicate that the actual response information meets the first model evaluation standard in at least the first dimension, and the second conclusion information is used to indicate that the actual response information does not meet the first model evaluation standard in at least the first dimension; the reason information is used to indicate the reason why the third intelligent agent draws the conclusion information.
[0167] In the embodiment of the present specification, the method of obtaining the test case generated by the model to be evaluated may include: extracting the test case from the operation log data and task record data of the model to be evaluated.
[0168] In the embodiment of this specification, the third agent can be a model generated by fine-tuning training based on the pre-trained large language model. The case to be evaluated, the first standard for model evaluation, and the model evaluation requirement information can be input into the third agent for model evaluation. The third agent responds to the model evaluation requirement information and evaluates the case to be evaluated according to the first standard for model evaluation to obtain the evaluation result. Among them, the model evaluation requirement information can be information for requesting the third agent to evaluate the case to be evaluated according to the first standard for model evaluation.
[0169] In the embodiments of this specification, in order to facilitate those skilled in the art to understand this solution, this specification uses a specific example to illustrate the evaluation results. Example 2 is as follows.
[0170] For the dialogue scenario between the user and the dialogue model, the model to be evaluated is AI Friend, and the case to be evaluated is a dialogue case between the user and AI Friend. The user's actual question is "I have a cold, what should I do?", AI Friend's actual reply 1 "You should take some cold medicine", AI Friend's actual reply 2 "You should go to the hospital to see a doctor as soon as possible, and receive treatment according to the doctor's advice. In the next few days, you should drink more water, eat a lighter diet, exercise appropriately, and keep warm." The first standard for model evaluation is used to evaluate whether the AI Friend's reply has emotional care, and the standard requires that reasonable opinions should be given to the user in all aspects before the reply information can be judged to have reached the standard of emotional care. The model evaluation requirement information is to evaluate whether the AI Friend's reply information has reached the standard of emotional care according to the first standard of model evaluation.
[0171] The evaluation result may include conclusion information and reason information. The conclusion information may include information used to indicate that the actual reply information meets the first standard of the model evaluation in the at least first dimension. For example, the conclusion information given for the actual reply 2 of the above-mentioned AI friend is "the actual reply 2 of the AI friend meets the first standard of the model evaluation in the dimension of emotional care", and the corresponding reason information is "because the actual reply 2 of the AI friend expresses the meaning of care to the user from multiple aspects." The conclusion information may also include information used to indicate that the actual reply information does not meet the first standard of the model evaluation in the at least first dimension. For example, the conclusion information given for the actual reply 1 of the above-mentioned AI friend is "the actual reply 1 of the AI friend does not meet the first standard of the model evaluation in the dimension of emotional care", and the corresponding reason information is "because the actual reply 1 of the AI friend does not express the meaning of care to the user from multiple aspects."
[0172] Optionally, the step of inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for performing standard optimization to obtain the second model evaluation standard output by the second intelligent agent may specifically include: inputting the standard optimization requirement information, the first model evaluation standard and label information determined based on the evaluation result into a second intelligent agent for performing standard optimization to obtain the second model evaluation standard output by the second intelligent agent; the label information includes a first sample label or a second sample label; the first sample label is used to indicate that the evaluation result of the case to be evaluated does not meet the target evaluation requirement; the second sample label is used to indicate that the evaluation result of the case to be evaluated meets the target evaluation requirement.
[0173] In an embodiment of the present specification, the standard optimization requirement information, the first model evaluation standard, and the label information determined based on the evaluation results can be inserted into a standard optimization prompt template containing standard optimization thinking chain information to obtain standard optimization prompt information; then the standard optimization prompt information is input into a second intelligent agent used for standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
[0174] In the embodiment of this specification, the label information of the evaluation result may be information determined by the evaluator. If the evaluation result generated by the third agent for the case to be evaluated does not meet the target evaluation requirement, a first sample label is generated for the evaluation result. If the evaluation result generated by the third agent for the case to be evaluated meets the target evaluation requirement, a second sample label is generated for the evaluation result.
[0175] In the embodiments of the present specification, the first intelligent agent automatically generates a first model evaluation standard; the third intelligent agent can automatically evaluate the case to be evaluated based on the first model evaluation standard and generate the evaluation results; the evaluator signs the evaluation results to obtain label information of the evaluation results; the third intelligent agent can automatically optimize the first model evaluation standard based on the first model evaluation standard and the label information of the evaluation results, thereby improving the degree of automation of the generation process of the first model evaluation standard, the evaluation process of the case to be evaluated, and the optimization process of the first model evaluation standard.
[0176] In the embodiments of this specification, the first model evaluation standard is optimized using the evaluation cases extracted from the actual operation data of the model to be evaluated, thereby improving the ability of the first model evaluation standard to adapt to the actual environment and improving the adaptability of the first model evaluation standard.
[0177] In the embodiments of this specification, a specific embodiment is also proposed for optimizing the first criterion of model evaluation by using label information determined by evaluation results.
[0178] Optionally, the step of inputting the standard optimization requirement information, the first model evaluation standard, and the label information determined based on the evaluation result into a second intelligent agent for standard optimization to obtain the second model evaluation standard output by the second intelligent agent may specifically include: sending the case to be evaluated and the evaluation result corresponding to the case to be evaluated to an evaluation user; obtaining the label information marked by the evaluation user on the case to be evaluated based on the evaluation result; and inputting the case to be evaluated carrying the label information, the standard optimization requirement information, and the first model evaluation standard into a second intelligent agent for standard optimization to obtain the second model evaluation standard output by the second intelligent agent.
[0179] In an embodiment of the present specification, the standard optimization requirement information, the first model evaluation standard, and the case to be evaluated carrying the label information can be inserted into a standard optimization prompt template containing standard optimization thinking chain information to obtain standard optimization prompt information; then the standard optimization prompt information is input into a second intelligent agent used for standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
[0180] In the embodiment of this specification, the evaluation user may be a user who has the ability to judge whether the evaluation result is accurate. If the evaluation result generated for the case to be evaluated does not meet the target evaluation requirements, a first sample label (bad sample label) is set for the case to be evaluated; if the evaluation result generated for the case to be evaluated meets the target evaluation requirements, a second sample label (good sample label) is set for the case to be evaluated.
[0181] In actual applications, the evaluation result meets the target evaluation requirements, which may include: the third agent gives the first conclusion information (indicating that the actual response information meets the first standard of the model evaluation in at least the first dimension), and the evaluation user believes that the actual response information meets the target evaluation requirements in at least the first dimension. Alternatively, the third agent gives the second conclusion information (indicating that the actual response information does not meet the first standard of the model evaluation in at least the first dimension), and the evaluation user believes that the actual response information does not meet the target evaluation requirements in at least the first dimension.
[0182] In actual applications, the evaluation result does not meet the target evaluation requirements, which may include: the third agent gives the first conclusion information (indicating that the actual response information meets the first standard of the model evaluation in at least the first dimension), but the evaluation user believes that the actual response information does not meet the target evaluation requirements in at least the first dimension. Alternatively, the third agent gives the second conclusion information (indicating that the actual response information does not meet the first standard of the model evaluation in at least the first dimension), but the evaluation user believes that the actual response information meets the target evaluation requirements in at least the first dimension.
[0183] Optionally, the inputting of the standard optimization requirement information and the first model evaluation standard into the second intelligent agent for standard optimization, before obtaining the second model evaluation standard output by the second intelligent agent, may also include: obtaining a first evaluated case carrying a first sample label; the first evaluated case is obtained by labeling at least part of the cases to be evaluated based on the evaluation result; the first sample label is used to indicate that the evaluation result does not meet the target evaluation requirement.
[0184] Optionally, the step of inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent that performs standard optimization to obtain a second model evaluation standard output by the second intelligent agent may specifically include: inputting the first evaluated case carrying the first sample label, the standard optimization requirement information and the first model evaluation standard into a second intelligent agent that performs standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
[0185] Optionally, the process of inputting the standard optimization requirement information and the first model evaluation standard into the second intelligent agent for standard optimization, before obtaining the second model evaluation standard output by the second intelligent agent, may also include: obtaining a second evaluated case carrying a second sample label; the second evaluated case is obtained by labeling at least part of the cases to be evaluated based on the evaluation result; the second sample label is used to indicate that the evaluation result meets the target evaluation requirement.
[0186] Optionally, the step of inputting the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for performing standard optimization to obtain a second model evaluation standard output by the second intelligent agent may specifically include: inputting the first evaluated case carrying the first sample label, the second evaluated case carrying the second sample label, the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for performing standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
[0187] In the embodiments of the present specification, the evaluation results obtained based on the first model evaluation standard can be manually labeled to select good sample label cases and bad sample label cases, and the good sample label cases and / or bad sample label cases can be used as sample cases to optimize the first model evaluation standard, so that the second intelligent agent can optimize the first model evaluation standard from the perspective of good sample label cases and / or bad sample label cases, so as to more accurately generate a model evaluation standard that meets the evaluation requirements.
[0188] In the embodiments of this specification, a specific embodiment is also proposed for a method of generating evaluation results based on a third intelligent agent.
[0189] Optionally, the step of inputting the case to be evaluated and the first model evaluation standard into a third intelligent agent for model evaluation to obtain an evaluation result obtained by the third intelligent agent evaluating the case to be evaluated according to the first model evaluation standard, may specifically include: inserting the case to be evaluated and the first model evaluation standard into an evaluation prompt template containing evaluation thinking chain information to obtain evaluation prompt information; the evaluation thinking chain information is used to represent the third logic for evaluating the case to be evaluated; inputting the evaluation prompt information into a third intelligent agent for model evaluation to obtain an evaluation result obtained by the third intelligent agent evaluating the case to be evaluated according to the third logic.
[0190] In the embodiment of this specification, the evaluation prompt template may at least include evaluation thought chain information, a "user demand" part, a "model evaluation standard" part, a "case to be evaluated" part, and evaluation person setting information.
[0191] In practical applications, the evaluation thinking chain information may be information for guiding the model to think step by step according to the steps of the logic chain of the third logic, wherein the third logic may be logic information for generating evaluation results based on the evaluation requirement information.
[0192] In actual applications, user demands can be evaluation demand information. For example, in the scenario of evaluating whether the reply information of AI friends meets the standards of the emotional care dimension, the evaluation demand information can be "Please tell me which replies have emotional care and which replies do not have emotional care."
[0193] In actual application, the case to be evaluated, the first standard for model evaluation and the evaluation requirement information can be inserted into an evaluation prompt template containing evaluation thinking chain information to obtain evaluation prompt information containing evaluation thinking chain information, the case to be evaluated, the first standard for model evaluation, evaluation requirement information and preset personality information; and then the evaluation prompt information is input into a third intelligent agent for model evaluation to obtain an evaluation result obtained by the third intelligent agent evaluating the case to be evaluated according to the first standard for model evaluation.
[0194] In the embodiment of the present specification, the evaluation prompt information used to generate the evaluation results includes thought chain information, so that the third intelligent agent can think according to the logical steps indicated by the thought chain information, thereby improving the accuracy of the evaluation results generated by the third intelligent agent for the case to be evaluated.
[0195] In the embodiments of this specification, a specific embodiment is also proposed for the method of generating a third intelligent agent.
[0196] Optionally, before the method of inputting the case to be evaluated and the first model evaluation standard into a third intelligent agent for model evaluation and obtaining the evaluation result obtained by the third intelligent agent through evaluation of the case to be evaluated according to the first model evaluation standard, it may also include: obtaining a third training sample set and an evaluation reference standard; a training sample in the third training sample set includes a sample to be evaluated and an evaluation result label corresponding to the sample to be evaluated; inserting the sample to be evaluated and the evaluation reference standard into an evaluation prompt template containing evaluation thinking chain information to obtain third sample prompt information; the evaluation thinking chain information is used to represent the third logic for evaluating the sample to be evaluated; inputting the third sample prompt information into a pre-trained large language model to obtain a prediction result generated by the large language model according to the third logic; and fine-tuning the large language model based on the difference between the prediction result and the evaluation result label to obtain the third intelligent agent.
[0197] In an embodiment of the present specification, the sample to be evaluated, the evaluation reference standard and the training requirement information can be inserted into an evaluation prompt template containing evaluation thinking chain information to obtain third sample prompt information, and then the third sample prompt information can be input into a pre-trained large language model, so that the large language model responds to the training requirement information, refers to the evaluation reference standard, and predicts the sample to be evaluated according to the logical steps of the evaluation thinking chain information to obtain a prediction result.
[0198] In the embodiments of this specification, for the explanation of the training requirement information, evaluation prompt template, evaluation thinking chain information, third sample prompt information and other contents involved in the fine-tuning training process of generating the third intelligent agent, reference can be made to the explanation of the evaluation requirement information, evaluation prompt template, evaluation thinking chain information, evaluation prompt information and other contents involved in the process of the third intelligent agent evaluating the case to be evaluated, and no further details will be given here.
[0199] In the embodiments of this specification, the evaluation reference standard may be a standard set by an evaluator or a standard set by an intelligent agent, and this is not limited.
[0200] In the embodiment of this specification, the evaluation result label may be a label generated by an evaluator.
[0201] Optionally, fine-tuning the large language model based on the difference between the prediction result and the evaluation result label to obtain the third agent may include: fine-tuning the large language model based on the difference between the prediction result and the evaluation result label to obtain the third agent. Specifically, a third loss function value between the prediction result and the evaluation result label is calculated, and then the parameters of the large language model are adjusted according to the third loss function value until a preset training termination condition is met to obtain the third agent.
[0202] Optionally, based on the prediction results, the large language model is fine-tuned to obtain a third agent.
[0203] Optionally, fine-tuning the large language model based on the prediction result to obtain a third agent may include: obtaining manual feedback information on the prediction result from the evaluation user; and fine-tuning the large language model according to the manual feedback information. The manual feedback information may specifically include information on whether the evaluation user believes that the prediction result meets expectations. In actual applications, the device for fine-tuning the large speech model may have a human-computer interaction interface, and the evaluation user inputs the manual feedback information on the prediction result into the human-computer interaction interface. If the manual feedback information indicates that the prediction result meets expectations, the large language model after the current fine-tuning training is used to determine the third agent; if the manual feedback information indicates that the prediction result does not meet expectations, the large language model after the current fine-tuning training is fine-tuned according to the manual feedback information until the preset training termination condition is met to obtain the third agent.
[0204] Optionally, fine-tuning the large language model based on the first prediction standard to obtain the first agent may include: fine-tuning the large language model based on the difference between the first prediction standard and a preset first sample standard corresponding to the standard formulation requirement information sample to obtain the first agent. Specifically, a first loss function value between the first prediction standard and the first sample standard is calculated, and then the parameters of the large language model are adjusted according to the first loss function value until a preset training termination condition is met to obtain the first agent.
[0205] In an embodiment of the present specification, a third sample prompt information is generated based on an evaluation prompt template including evaluation thought chain information, and the third sample prompt information is used to perform fine-tuning training on a pre-trained large language model, so that the large language model can think and learn according to the logical steps indicated by the thought chain information, so as to improve the performance of the third intelligent agent after fine-tuning training, and further improve the accuracy of the evaluation results after the third intelligent agent evaluates the model to be evaluated.
[0206] In order to have a more intuitive understanding of the performance of the model to be evaluated, a model evaluation report can also be generated for the model to be evaluated.
[0207] Optionally, after the standard optimization requirement information and the first model evaluation standard are input into a second intelligent agent for standard optimization and the second model evaluation standard is obtained by the second intelligent agent, it may also include: obtaining multiple cases to be evaluated generated by the model to be evaluated; the cases to be evaluated include actual problem information input by the user into the model to be evaluated and actual reply information generated by the model to be evaluated in response to the actual problem information; the cases to be evaluated and the second model evaluation standard are input into a third intelligent agent for model evaluation, and the evaluation results obtained by the third intelligent agent on the cases to be evaluated according to the second model evaluation standard; statistical analysis is performed on the evaluation results obtained by the third intelligent agent for the multiple cases to be evaluated, and a model evaluation report of the model to be evaluated on at least the first dimension.
[0208] In the embodiments of this specification, the second model evaluation standard may be an optimized evaluation standard that meets the target evaluation requirements.
[0209] In practical applications, the evaluation results may include a positive evaluation result reflecting that the conclusion of the actual response information of the model to be evaluated on at least the first dimension is qualified information, and a negative evaluation result reflecting that the conclusion of the actual response information of the model to be evaluated on at least the first dimension is unqualified information. The first number of positive evaluation results and the second number of negative evaluation results are counted, and the first number of positive evaluation results, and / or the second number of negative evaluation results, and / or the ratio between the first number and the second number, and / or the ratio between the first number and the total number, and / or the ratio between the second number and the total number are used to determine the model evaluation report included in the at least the first dimension of the model to be evaluated. The total number is the sum of the first number and the second number.
[0210] In practical applications, if the at least first dimension involves multiple dimensions, statistics may be performed for each dimension separately.
[0211] Optionally, the model evaluation report is output to the evaluation user and / or other users.
[0212] In the embodiments of this specification, the model evaluation report can provide the evaluation user with intuitive and easy-to-understand model evaluation results to help the evaluation user better select and optimize the model.
[0213] Figure 3 This specification provides an embodiment corresponding to Figure 2 A swimlane flow diagram of a method for generating a model evaluation standard. The business execution process of the method for generating a model evaluation standard may involve execution subjects such as a first agent, a second agent, and a third agent.
[0214] like Figure 3 As shown, the process may include the following steps.
[0215] In the embodiments of the present specification, the process may include a model evaluation standard generation phase and a model evaluation standard use and optimization phase, wherein the model evaluation standard generation phase may include steps 302 to 310, and the model evaluation standard use and optimization phase may include steps 312 to 332.
[0216] Step 302: Obtain standard formulation requirement information for formulating model evaluation standards.
[0217] Step 304: Optionally, obtain a first standard reference case; the first standard reference case includes user question samples and case analysis information on at least the first dimension.
[0218] Step 306: Optionally, obtain historical standards for model evaluation.
[0219] Step 308: If standard-setting requirement information is obtained, the standard-setting requirement information is inserted into a standard-setting prompt template containing standard-setting thinking chain information to obtain standard-setting prompt information; if standard-setting requirement information and a first standard reference case are obtained, the standard-setting requirement information and the first standard reference case are inserted into a standard-setting prompt template containing standard-setting thinking chain information to obtain standard-setting prompt information; if standard-setting requirement information, a first standard reference case and a model evaluation history standard are obtained, the standard-setting requirement information, the first standard reference case and the model evaluation history standard are inserted into a standard-setting prompt template containing standard-setting thinking chain information to obtain standard-setting prompt information.
[0220] Step 310: Input the standard-setting prompt information into the first agent used for standard setting, and obtain the first standard for model evaluation generated by the first agent according to the first logic.
[0221] Step 312: Obtain the test case generated by the model to be tested.
[0222] Step 314: Input the case to be evaluated and the first model evaluation standard into a third agent for model evaluation, and obtain an evaluation result obtained by the third agent by evaluating the case to be evaluated according to the first model evaluation standard.
[0223] Step 316: Send the case to be evaluated and the evaluation result corresponding to the case to be evaluated to the evaluation user.
[0224] Step 318: Obtain label information that the evaluation user marks on the case to be evaluated based on the evaluation result.
[0225] Step 320: Obtain standard optimization requirement information about optimizing the first criterion for evaluating the model.
[0226] Step 322: Acquire a second standard reference case, where the second standard reference case may include the case to be evaluated that carries the label information, or the second standard reference case may not include the case to be evaluated that carries the label information.
[0227] Step 324: If standard optimization requirement information is obtained, the standard optimization requirement information and the first model evaluation standard are inserted into the standard optimization prompt template containing standard optimization thinking chain information to obtain standard optimization prompt information; the standard optimization requirement information, the first model evaluation standard and the second standard reference case are inserted into the standard optimization prompt template containing standard optimization thinking chain information to obtain standard optimization prompt information.
[0228] Step 326: Input the standard optimization prompt information into the second agent used to perform standard optimization, and obtain the second model evaluation standard output by the second agent.
[0229] Step 328: Acquire multiple cases to be evaluated generated by the model to be evaluated.
[0230] Step 330: Input the case to be evaluated and the second model evaluation standard into a third agent for model evaluation, and obtain an evaluation result obtained by the third agent by evaluating the case to be evaluated according to the second model evaluation standard.
[0231] Step 332: Perform statistical analysis on the evaluation results obtained by the third agent for the multiple cases to be evaluated, and obtain a model evaluation report of the model to be evaluated in at least the first dimension.
[0232] Figure 4 It is a flowchart of automatically generating a model evaluation standard, automatically evaluating a model to be evaluated based on the model evaluation standard, and automatically optimizing the model evaluation standard based on the evaluation result of the model to be evaluated, provided by an embodiment of this specification.
[0233] In the embodiments of this specification, Figure 4 As shown, the historical standards for model evaluation, standard formulation requirements information, and standard reference cases are input into the first intelligent agent. The first intelligent agent can automatically generate model evaluation standards based on the historical standards for model evaluation, standard formulation requirements information, and standard reference cases, according to the standard formulation thinking chain information. The model evaluation standards and the cases to be evaluated generated by the first intelligent agent are input into the third intelligent agent. The third intelligent agent can use the model evaluation standards and the evaluation thinking chain information to automatically evaluate the cases to be evaluated generated by the model to be evaluated, and obtain the evaluation results. The evaluators mark the cases in the evaluation results, and extract good sample cases with good sample labels and bad sample cases with bad sample labels. The good sample cases with good sample labels and the bad sample cases with bad sample labels are input into the second intelligent agent. The second intelligent agent refers to the good sample cases with good sample labels and the bad sample cases with bad sample labels, and automatically optimizes the model evaluation standards used by the third intelligent agent when evaluating the cases to be evaluated according to the standard optimization thinking chain information, and obtains the optimized model evaluation standards. Based on the following Figure 4 The process of the embodiment of this specification shown in the figure can realize a closed-loop process of automatically generating model evaluation standards, automatically evaluating the model to be evaluated based on the model evaluation standards, and automatically optimizing the model evaluation standards based on the evaluation results. It can also realize continuous optimization of the model evaluation standards to improve the generation efficiency of the model evaluation standards and the applicability and accuracy of the model evaluation standards.
[0234] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.
[0235] Figure 5 This specification provides an embodiment corresponding to Figure 2 A schematic diagram of the structure of a device for generating model evaluation criteria.
[0236] like Figure 5 As shown, the device may include:
[0237] A first requirement information acquisition module 502 is used to acquire standard formulation requirement information for formulating a model evaluation standard; the standard formulation requirement information is used to indicate the generation of a model evaluation standard for evaluating model response information generated by the model to be evaluated in response to user question information in natural language form in at least a first dimension;
[0238] A standard generation module 504 is used to input the standard formulation requirement information into a first agent for standard formulation, and obtain a first model evaluation standard output by the first agent;
[0239] The second requirement information acquisition module 506 is used to acquire standard optimization requirement information about optimizing the first standard of the model evaluation; the standard optimization requirement information is used to indicate to optimize the first standard of the model evaluation;
[0240] The standard optimization module 508 is used to input the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for standard optimization, and obtain the second model evaluation standard output by the second intelligent agent.
[0241] based on Figure 5 The present specification also provides some improved implementation schemes of the device, which are described below.
[0242] Optionally, the standard generating module 504 may specifically include:
[0243] A standard formulation requirement information insertion unit is used to insert the standard formulation requirement information into a standard formulation prompt template containing standard formulation thought chain information to obtain standard formulation prompt information; the standard formulation thought chain information is used to represent the first logic of generating a standard based on the standard formulation requirement information.
[0244] The first input unit is used to input the standard-setting prompt information into a first intelligent agent used for standard setting, so as to obtain a first standard for model evaluation generated by the first intelligent agent according to the first logic.
[0245] Optionally, the device may further include:
[0246] The first standard reference case acquisition module is used to acquire a first standard reference case; the first standard reference case includes a user question sample and case analysis information on at least the first dimension.
[0247] Optionally, the standard generating module 504 may specifically include:
[0248] The second input unit is used to input the standard formulation requirement information and the first standard reference case into the first intelligent agent used for standard formulation, and obtain the first model evaluation standard output by the first intelligent agent.
[0249] Optionally, the first standard reference case also includes model response samples for the user question samples; the case analysis information specifically includes response analysis information for the model response samples in at least the first dimension.
[0250] Optionally, the first standard reference case includes at least one of a positive reference case and a negative reference case; the model response sample in the positive reference case meets the target evaluation requirements in at least the first dimension; and the model response sample in the negative reference case does not meet the target evaluation requirements in at least the first dimension.
[0251] Optionally, the device may further include:
[0252] The model evaluation historical standard acquisition module is used to obtain the model evaluation historical standard.
[0253] Optionally, the standard generating module 504 may specifically include:
[0254] The third input unit is used to input the standard formulation requirement information, the first standard reference case and the model evaluation historical standard into the first intelligent agent used for standard formulation, so as to obtain the first model evaluation standard output by the first intelligent agent.
[0255] Optionally, the device may further include:
[0256] The first training sample set acquisition module is used to acquire a first training sample set; a training sample in the first training sample set includes a standard setting requirement information sample; the standard setting requirement information sample is used to indicate the generation of a prediction standard for evaluating the model response information output by the model to be evaluated.
[0257] A standard formulation requirement information sample insertion module is used to insert the standard formulation requirement information sample into a standard formulation prompt template containing standard formulation thought chain information to obtain a standard formulation prompt information sample; the standard formulation thought chain information is used to represent the first logic of generating a standard based on the standard formulation requirement information.
[0258] The standard setting prompt information sample input module is used to input the standard setting prompt information sample into a pre-trained large language model to obtain a first prediction standard generated by the large language model according to the first logic.
[0259] The first fine-tuning training module is used to perform fine-tuning training on the large language model based on the first prediction standard to obtain a first intelligent agent.
[0260] Optionally, the device may further include:
[0261] The first reference case acquisition module is used to acquire a first reference case; the first reference case includes a first problem sample and first analysis information on at least the first dimension.
[0262] The standard formulation requirement information sample insertion module may specifically include:
[0263] The first insertion unit is used to insert the standard formulation requirement information sample and the first reference case into a standard formulation prompt template containing standard formulation thought chain information to obtain a standard formulation prompt information sample.
[0264] Optionally, the standard optimization module 508 may specifically include:
[0265] The second insertion unit is used to insert the standard optimization requirement information and the first model evaluation standard into a standard optimization prompt template containing standard optimization thinking chain information to obtain standard optimization prompt information; the standard optimization thinking chain information is used to represent the second logic of optimizing the standard based on the standard optimization requirement information.
[0266] The fourth input unit is used to input the standard optimization prompt information into the second intelligent agent used for standard optimization to obtain the second model evaluation standard output by the second intelligent agent.
[0267] Optionally, the device may further include:
[0268] The second standard reference case acquisition module is used to acquire a second standard reference case; the second standard reference case includes a user problem sample and first case analysis information on at least the first dimension, or the second standard reference case includes a user problem sample and second case analysis information on at least the second dimension.
[0269] Optionally, the standard optimization module 508 may specifically include:
[0270] The fifth input unit is used to input the standard optimization requirement information, the first model evaluation standard and the second standard reference case into the second intelligent agent used for standard optimization, so as to obtain the second model evaluation standard output by the second intelligent agent.
[0271] Optionally, the device may further include:
[0272] The second training sample set acquisition module is used to acquire the second training sample set; a training sample in the second training sample set includes a standard optimization requirement information sample and a second sample standard to be optimized; the standard optimization requirement information sample is used to indicate the optimization of the second sample standard.
[0273] The standard optimization requirement information sample and the second sample standard insertion module are used to insert the standard optimization requirement information sample and the second sample standard into a standard optimization prompt template containing standard optimization thinking chain information to obtain a standard optimization prompt information sample; the standard optimization thinking chain information is used to represent the second logic of optimizing the standard based on the standard optimization requirement information.
[0274] The standard optimization prompt information sample input module is used to input the standard optimization prompt information sample into a pre-trained large language model to obtain a second prediction standard obtained by optimizing the second sample standard by the large language model according to the second logic.
[0275] The second fine-tuning training module is used to perform fine-tuning training on the large language model based on the second prediction standard to obtain a second intelligent agent.
[0276] Optionally, the device may further include:
[0277] The second reference case acquisition module is used to acquire a second reference case; the second reference case includes a second problem sample and second analysis information on at least the first dimension, or the second reference case includes a second problem sample and third analysis information on at least the second dimension.
[0278] The standard optimization requirement information sample and the second sample standard insertion module may specifically include:
[0279] The third insertion unit is used to insert the standard optimization requirement information sample, the second sample standard and the second reference case into a standard optimization prompt template containing standard optimization thinking chain information to obtain a standard optimization prompt information sample.
[0280] Optionally, the device may further include:
[0281] The module for obtaining cases to be evaluated is used to obtain cases to be evaluated generated by the model to be evaluated; the cases to be evaluated include actual question information input by the user into the model to be evaluated and actual reply information generated by the model to be evaluated in response to the actual question information.
[0282] The module for inputting cases to be evaluated and the first standard for model evaluation is used to input the cases to be evaluated and the first standard for model evaluation into a third intelligent agent used for model evaluation, and obtain an evaluation result obtained by the third intelligent agent evaluating the cases to be evaluated according to the first standard for model evaluation; the evaluation result includes conclusion information and reason information; the conclusion information includes first conclusion information or second conclusion information, the first conclusion information is used to indicate that the actual response information meets the first standard for model evaluation in at least the first dimension, and the second conclusion information is used to indicate that the actual response information does not meet the first standard for model evaluation in at least the first dimension; the reason information is used to indicate the reason why the third intelligent agent draws the conclusion information.
[0283] Optionally, the standard optimization module 508 may specifically include:
[0284] The sixth input unit is used to input the standard optimization requirement information, the first model evaluation standard and the label information determined based on the evaluation result into the second intelligent agent used for standard optimization, so as to obtain the second model evaluation standard output by the second intelligent agent; the label information includes a first sample label or a second sample label; the first sample label is used to indicate that the evaluation result of the case to be evaluated does not meet the target evaluation requirement; the second sample label is used to indicate that the evaluation result of the case to be evaluated meets the target evaluation requirement.
[0285] Optionally, the sixth input unit may specifically include:
[0286] The sending subunit is used to send the case to be evaluated and the evaluation result corresponding to the case to be evaluated to the evaluation user.
[0287] The acquisition subunit is used to obtain label information marked by the evaluation user on the case to be evaluated based on the evaluation result.
[0288] The input subunit is used to input the case to be evaluated carrying the label information, the standard optimization requirement information and the first model evaluation standard into the second intelligent agent for standard optimization, and obtain the second model evaluation standard output by the second intelligent agent.
[0289] Optionally, the case to be evaluated and the first standard input module for model evaluation may specifically include:
[0290] A fourth inserting unit is used to insert the case to be evaluated and the first standard of model evaluation into an evaluation prompt template containing evaluation thought chain information to obtain evaluation prompt information; the evaluation thought chain information is used to represent the third logic for evaluating the case to be evaluated;
[0291] The seventh input unit is used to input the evaluation prompt information into the third agent used for model evaluation, and obtain the evaluation result obtained by the third agent through evaluating the case to be evaluated according to the third logic.
[0292] Optionally, the device may further include:
[0293] The third training sample set and evaluation reference standard acquisition module is used to acquire the third training sample set and the evaluation reference standard; a training sample in the third training sample set includes a sample to be evaluated and an evaluation result label corresponding to the sample to be evaluated.
[0294] The module for inserting samples to be evaluated and evaluation reference standards is used to insert the samples to be evaluated and the evaluation reference standards into an evaluation prompt template containing evaluation thinking chain information to obtain third sample prompt information; the evaluation thinking chain information is used to represent the third logic for evaluating the samples to be evaluated.
[0295] The third sample prompt information input module is used to input the third sample prompt information into a pre-trained large language model to obtain a prediction result generated by the large language model according to the third logic.
[0296] The third fine-tuning training module is used to fine-tune the large language model based on the difference between the prediction result and the evaluation result label to obtain a third intelligent agent.
[0297] Optionally, the device may further include:
[0298] A module for acquiring multiple cases to be evaluated is used to acquire multiple cases to be evaluated generated by the model to be evaluated; the cases to be evaluated include actual question information input by the user into the model to be evaluated and actual response information generated by the model to be evaluated in response to the actual question information.
[0299] The module for inputting the case to be evaluated and the second standard for model evaluation is used to input the case to be evaluated and the second standard for model evaluation into a third intelligent agent for model evaluation, and obtain the evaluation result obtained by the third intelligent agent when evaluating the case to be evaluated according to the second standard for model evaluation.
[0300] The statistical analysis module is used to perform statistical analysis on the evaluation results obtained by the third agent for the multiple cases to be evaluated, so as to obtain a model evaluation report of the model to be evaluated in at least the first dimension.
[0301] Figure 5 In the device, the model evaluation standard can be automatically generated by using the first intelligent agent without relying on manual standard setting, thereby improving the efficiency of standard setting and the objectivity of the formulated model evaluation standard. When the output of the model is evaluated based on the model evaluation standard formulated by the first intelligent agent, the accuracy of the evaluation result can be improved, and thus the evaluation performance of the model evaluation standard can be improved. The first model evaluation standard generated for the first intelligent agent is automatically optimized by using the second intelligent agent, so that the model evaluation standard can adapt to changes in different tasks and environments, thereby improving the adaptability of the model evaluation standard.
[0302] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method.
[0303] Figure 6This specification provides an embodiment corresponding to Figure 2 A schematic diagram of a device for generating model evaluation criteria. Figure 6 As shown, the computing device 600 may include:
[0304] at least one processor 620; and,
[0305] A memory 610 is communicatively connected to the at least one processor; wherein,
[0306] The memory 610 stores instructions that can be executed by the at least one processor 620, and the instructions are executed by the at least one processor 620 to enable the at least one processor 620 to: obtain standard formulation requirement information on formulating a model evaluation standard; the standard formulation requirement information is used to indicate the generation of a model evaluation standard for evaluating model response information generated by a to-be-evaluated model in response to user question information in natural language form in at least a first dimension; input the standard formulation requirement information into a first intelligent agent used for standard formulation to obtain a first model evaluation standard output by the first intelligent agent; obtain standard optimization requirement information on optimizing the first model evaluation standard; the standard optimization requirement information is used to indicate the optimization of the first model evaluation standard; input the standard optimization requirement information and the first model evaluation standard into a second intelligent agent used for standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
[0307] Figure 6 The structure block diagram of a computing device 600 for generating a model evaluation standard according to an embodiment of the present application is shown. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and the database 650 is used to store data.
[0308] Computing device 600 also includes access device 640 , which enables computing device 600 to communicate via one or more networks 660 .
[0309] In one embodiment of the present application, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 6 The device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0310] Among them, when the processor 620 executes the computer instructions, it implements the steps of the method for generating a model evaluation standard.
[0311] The above is a schematic scheme of a device for generating a model evaluation standard in this embodiment. It should be noted that the technical scheme of the device and the technical scheme of the method for generating a model evaluation standard are of the same concept, and the details not described in detail in the technical scheme of the device can be referred to the description of the technical scheme of the method for generating a model evaluation standard.
[0312] Figure 6 Examples of the network 660 in the example include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1440 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0313] The computing device 600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 600 may also be a mobile or stationary server.
[0314] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The devices, equipment and methods provided in the embodiments of this specification correspond to each other, so the devices and equipment also have beneficial technical effects similar to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding devices and equipment will not be repeated here.
[0315] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0316] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0317] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0318] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0319] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0320] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0321] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0322] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0323] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0324] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0325] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0326] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0327] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0328] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0329] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for generating a model evaluation criterion, comprising: Obtain information on the standard setting requirements for developing model evaluation standards; The standard formulation requirement information is used to instruct the generation of a model evaluation standard for evaluating model response information generated by the model to be evaluated in response to user question information in natural language form in at least a first dimension; Inputting the standard formulation requirement information into a first agent used for standard formulation, and obtaining a first model evaluation standard output by the first agent; Obtaining standard optimization requirement information about optimizing the first standard of the model evaluation; The standard optimization requirement information is used to indicate the optimization of the first standard of the model evaluation; The standard optimization requirement information and the first model evaluation standard are input into a second intelligent agent for standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
2. The method according to claim 1, wherein the inputting the standard formulation requirement information into a first agent for standard formulation to obtain a first model evaluation standard output by the first agent specifically comprises: Inserting the standard formulation requirement information into a standard formulation prompt template containing standard formulation thought chain information to obtain standard formulation prompt information; The standard formulation thought chain information is used to represent a first logic for generating a standard based on the standard formulation requirement information; The standard-setting prompt information is input into a first agent used for standard setting, and a first standard for model evaluation generated by the first agent according to the first logic is obtained.
3. The method according to claim 1, before inputting the standard formulation requirement information into a first agent for standard formulation and obtaining the first model evaluation standard output by the first agent, further comprising: Obtain the first standard reference case; The first standard reference case includes a user problem sample and case analysis information on at least the first dimension; The step of inputting the standard formulation requirement information into a first agent for standard formulation to obtain a first model evaluation standard output by the first agent specifically includes: The standard formulation requirement information and the first standard reference case are input into a first intelligent agent for standard formulation to obtain a first model evaluation standard output by the first intelligent agent.
4. The method of claim 3, wherein: The first standard reference case also includes a model response sample for the user question sample; the case analysis information specifically includes response analysis information for the model response sample in at least the first dimension.
5. The method of claim 4, wherein: The first standard reference case includes at least one of a positive reference case and a negative reference case; the model response sample in the positive reference case meets the target evaluation requirement in the at least first dimension; The model response examples in the negative reference cases do not meet the target evaluation requirements in at least the first dimension.
6. The method according to claim 3, before inputting the standard formulation requirement information into a first agent for standard formulation and obtaining the first model evaluation standard output by the first agent, further comprising: Obtain historical standards for model evaluation; The step of inputting the standard formulation requirement information into a first agent for standard formulation to obtain a first model evaluation standard output by the first agent specifically includes: The standard formulation requirement information, the first standard reference case and the model evaluation historical standard are input into a first intelligent agent for standard formulation to obtain a first model evaluation standard output by the first intelligent agent.
7. The method according to claim 1, before inputting the standard formulation requirement information into a first agent for standard formulation and obtaining the first model evaluation standard output by the first agent, further comprising: Obtaining a first training sample set; A training sample in the first training sample set includes a standard formulation requirement information sample; The standard formulation requirement information sample is used to indicate the generation of a prediction standard for evaluating the model response information output by the model to be evaluated; Inserting the standard formulation requirement information sample into a standard formulation prompt template containing standard formulation thought chain information to obtain a standard formulation prompt information sample; the standard formulation thought chain information is used to represent the first logic of generating a standard based on the standard formulation requirement information; Inputting the standard setting prompt information sample into a pre-trained large language model to obtain a first prediction standard generated by the large language model according to the first logic; The large language model is fine-tuned and trained based on the first prediction criterion to obtain a first intelligent agent.
8. The method according to claim 7, before inserting the standard formulation requirement information sample into the standard formulation prompt template containing the standard formulation thought chain information to obtain the standard formulation prompt information sample, further comprising: Get the first reference case; The first reference case includes a first problem sample and first analysis information on the at least first dimension; The step of inserting the standard formulation requirement information sample into a standard formulation prompt template containing standard formulation thought chain information to obtain the standard formulation prompt information sample specifically includes: The standard formulation requirement information sample and the first reference case are inserted into a standard formulation prompt template containing standard formulation thought chain information to obtain a standard formulation prompt information sample.
9. The method of claim 1, wherein the inputting the standard optimization requirement information and the first model evaluation standard into a second agent for standard optimization to obtain a second model evaluation standard output by the second agent specifically comprises: Inserting the standard optimization requirement information and the first model evaluation standard into a standard optimization prompt template containing standard optimization thought chain information to obtain standard optimization prompt information; The standard optimization thinking chain information is used to represent the second logic of optimizing the standard based on the standard optimization requirement information; The standard optimization prompt information is input into a second agent for standard optimization to obtain a second model evaluation standard output by the second agent.
10. The method of claim 1, wherein before the standard optimization requirement information and the first model evaluation standard are input to a second agent for standard optimization and the second model evaluation standard output by the second agent is obtained, the method further comprises: Obtain the second standard reference case; The second standard reference case includes a user problem sample and first case analysis information on at least the first dimension, or the second standard reference case includes a user problem sample and second case analysis information on at least the second dimension; The step of inputting the standard optimization requirement information and the first model evaluation standard into a second agent for standard optimization to obtain a second model evaluation standard output by the second agent specifically includes: The standard optimization requirement information, the first model evaluation standard and the second standard reference case are input into a second intelligent agent for standard optimization to obtain the second model evaluation standard output by the second intelligent agent.
11. The method of claim 1, wherein before the standard optimization requirement information and the first model evaluation standard are input to a second agent for standard optimization and the second model evaluation standard output by the second agent is obtained, the method further comprises: Obtaining a second training sample set; A training sample in the second training sample set includes a standard optimization requirement information sample and a second sample standard to be optimized; The standard optimization requirement information sample is used to indicate to optimize the second sample standard; Inserting the standard optimization requirement information sample and the second sample standard into a standard optimization prompt template containing standard optimization thought chain information to obtain a standard optimization prompt information sample; the standard optimization thought chain information is used to represent the second logic of optimizing the standard based on the standard optimization requirement information; Inputting the standard optimization prompt information sample into a pre-trained large language model to obtain a second prediction standard obtained by optimizing the second sample standard by the large language model according to the second logic; The large language model is fine-tuned and trained based on the second prediction criterion to obtain a second intelligent agent.
12. The method according to claim 11, before inserting the standard optimization requirement information sample and the second sample standard into a standard optimization prompt template containing standard optimization thought chain information to obtain the standard optimization prompt information sample, further comprising: Obtain a second reference case; The second reference case includes a second problem example and second analysis information on the at least first dimension, or the second reference case includes a second problem example and third analysis information on the at least second dimension; The step of inserting the standard optimization requirement information sample and the second sample standard into a standard optimization prompt template containing standard optimization thought chain information to obtain a standard optimization prompt information sample specifically includes: The standard optimization requirement information sample, the second sample standard and the second reference case are inserted into a standard optimization prompt template containing standard optimization thinking chain information to obtain a standard optimization prompt information sample.
13. The method according to claim 1, wherein after inputting the standard formulation requirement information into a first agent for standard formulation and obtaining a first model evaluation standard output by the first agent, the method further comprises: Obtaining a test case generated by the test model; The case to be evaluated includes actual question information input by a user into the model to be evaluated and actual response information generated by the model to be evaluated in response to the actual question information; The case to be evaluated and the first model evaluation standard are input into a third agent for model evaluation, and an evaluation result obtained by the third agent evaluating the case to be evaluated according to the first model evaluation standard is obtained; the evaluation result includes conclusion information and reason information; the conclusion information includes first conclusion information or second conclusion information, the first conclusion information is used to indicate that the actual response information meets the first model evaluation standard in at least the first dimension, and the second conclusion information is used to indicate that the actual response information does not meet the first model evaluation standard in at least the first dimension; the reason information is used to indicate the reason why the third agent draws the conclusion information; The step of inputting the standard optimization requirement information and the first model evaluation standard into a second agent for standard optimization to obtain a second model evaluation standard output by the second agent specifically includes: Inputting the standard optimization requirement information, the first model evaluation standard, and label information determined based on the evaluation result into a second agent for standard optimization, to obtain a second model evaluation standard output by the second agent; the label information includes a first sample label or a second sample label; The first sample label is used to indicate that the evaluation result of the case to be evaluated does not meet the target evaluation requirement; the second sample label is used to indicate that the evaluation result of the case to be evaluated meets the target evaluation requirement.
14. The method of claim 13, wherein the step of inputting the standard optimization requirement information, the first model evaluation standard, and the label information determined based on the evaluation result into a second agent for standard optimization to obtain the second model evaluation standard output by the second agent comprises: Sending the case to be evaluated and the evaluation result corresponding to the case to be evaluated to the evaluation user; Obtaining label information marked by the evaluation user on the case to be evaluated based on the evaluation result; The case to be evaluated carrying the label information, the standard optimization requirement information and the first model evaluation standard are input into a second intelligent agent for standard optimization to obtain a second model evaluation standard output by the second intelligent agent.
15. The method according to claim 13, wherein the step of inputting the case to be evaluated and the first model evaluation standard into a third agent for model evaluation to obtain an evaluation result obtained by the third agent evaluating the case to be evaluated according to the first model evaluation standard, specifically comprises: Inserting the case to be evaluated and the first standard of model evaluation into an evaluation prompt template containing evaluation thought chain information to obtain evaluation prompt information; The evaluation thinking chain information is used to represent the third logic for evaluating the case to be evaluated; The evaluation prompt information is input into a third agent used for model evaluation, and an evaluation result obtained by the third agent through evaluation of the case to be evaluated according to the third logic is obtained.
16. The method of claim 13, wherein before inputting the case to be evaluated and the first model evaluation standard into a third agent for model evaluation and obtaining an evaluation result obtained by the third agent evaluating the case to be evaluated according to the first model evaluation standard, the method further comprises: Obtaining a third training sample set and an evaluation reference standard; A training sample in the third training sample set includes a sample to be evaluated and an evaluation result label corresponding to the sample to be evaluated; Inserting the sample to be evaluated and the evaluation reference standard into an evaluation prompt template containing evaluation thought chain information to obtain third sample prompt information; the evaluation thought chain information is used to represent the third logic for evaluating the sample to be evaluated; Inputting the third sample prompt information into a pre-trained large language model to obtain a prediction result generated by the large language model according to the third logic; Based on the difference between the prediction result and the evaluation result label, the large language model is fine-tuned to obtain a third agent.
17. The method of claim 1, wherein after inputting the standard optimization requirement information and the first model evaluation standard into a second agent for standard optimization and obtaining the second model evaluation standard output by the second agent, the method further comprises: Acquire multiple test cases generated by the test model; The case to be evaluated includes actual question information input by a user into the model to be evaluated and actual response information generated by the model to be evaluated in response to the actual question information; Inputting the case to be evaluated and the second model evaluation standard into a third agent for model evaluation, and obtaining an evaluation result obtained by the third agent evaluating the case to be evaluated according to the second model evaluation standard; Perform statistical analysis on the evaluation results obtained by the third agent for the multiple cases to be evaluated, and obtain a model evaluation report of the model to be evaluated in at least the first dimension.
18. A device for generating a model evaluation standard, comprising: A first requirement information acquisition module is used to obtain standard formulation requirement information on formulating model evaluation standards; The standard formulation requirement information is used to instruct the generation of a model evaluation standard for evaluating model response information generated by the model to be evaluated in response to user question information in natural language form in at least a first dimension; A standard generation module, used for inputting the standard formulation requirement information into a first agent for standard formulation, and obtaining a first model evaluation standard output by the first agent; A second requirement information acquisition module is used to obtain standard optimization requirement information about optimizing the first standard of the model evaluation; The standard optimization requirement information is used to indicate the optimization of the first standard of the model evaluation; The standard optimization module is used to input the standard optimization requirement information and the first model evaluation standard into a second intelligent agent for standard optimization, and obtain the second model evaluation standard output by the second intelligent agent.
19. A device for generating a model evaluation standard, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquiring standard formulation requirement information on formulating a model evaluation standard; the standard formulation requirement information is used to indicate the generation of a model evaluation standard for evaluating model response information generated by the model to be evaluated in response to user question information in natural language form in at least a first dimension; Inputting the standard formulation requirement information into a first agent used for standard formulation, and obtaining a first model evaluation standard output by the first agent; Acquire standard optimization requirement information about optimizing the first standard of the model evaluation; the standard optimization requirement information is used to indicate to optimize the first standard of the model evaluation; The standard optimization requirement information and the first model evaluation standard are input into a second intelligent agent for standard optimization to obtain a second model evaluation standard output by the second intelligent agent.