Evaluation methods and apparatus for large models, electronic devices and computer-readable storage media
By using preset evaluation rules and multi-dimensional evaluation methods to evaluate the response information of large language models, the problem of inaccurate evaluation in existing technologies is solved, and accurate evaluation and personalized feedback of the response capabilities of large language models are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2024-09-18
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to accurately assess the responsiveness of large language models when user input conflicts with system settings, fail to effectively identify critical issues such as system-customized jailbreaks and leaks, and cannot comprehensively uncover obvious problems in model responses and lack of fine-grained capability assessment.
The response information of M large language models is initially evaluated using preset evaluation rules. Combining coarse-grained and fine-grained evaluation dimensions, the response capability of the large language models is determined by comparing the degree of consistency between the response information and the prompt information and the semantic matching of multiple evaluation dimensions.
It improves the accuracy and comprehensiveness of large language model response capability assessment, can identify obvious problems and determine the degree to which response information meets the assessment dimensions, and ensures the accuracy and flexibility of assessment results.
Smart Images

Figure CN119272857B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of large models and deep learning, and more specifically, to a method and apparatus for evaluating large models, electronic devices, and computer-readable storage media. Background Technology
[0002] With the development of computer and network technologies, Large Language Models (LLMs) have emerged. LLMs are deep learning-based artificial intelligence models primarily used for processing and generating natural language. These models are trained on large amounts of data and are capable of understanding, generating, and translating text. Summary of the Invention
[0003] This disclosure provides a method and apparatus for evaluating large models, an electronic device, and a computer-readable storage medium.
[0004] According to one aspect of this disclosure, a method for evaluating large language models is provided, comprising: evaluating each of the response information of M large language models to an input instruction based on a preset evaluation rule to obtain first evaluation information for each of the response information, where M is a positive integer greater than 1; in response to the consistency between the first evaluation information of the M large language models, evaluating each of the response information based on multiple evaluation dimensions to obtain second evaluation information for each of the response information; and determining an evaluation result based on the second evaluation information for each of the response information, wherein the evaluation result characterizes the response capability of each of the M large language models.
[0005] According to another aspect of this disclosure, an evaluation apparatus for large language models is provided, comprising: a first evaluation module, configured to evaluate each of the response information of M large language models to an input instruction based on a preset evaluation rule, to obtain first evaluation information for each of the response information, where M is a positive integer greater than 1; a second evaluation module, configured to evaluate each of the response information based on multiple evaluation dimensions in response to the consistency between the first evaluation information of the M large language models, to obtain second evaluation information for each of the response information; and a determination module, configured to determine an evaluation result based on the second evaluation information for each of the response information, wherein the evaluation result characterizes the response capability of each of the M large language models.
[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0007] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the steps of the above-described method.
[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0011] Figure 1 The illustration schematically shows a system architecture for applying large-scale model evaluation methods according to embodiments of the present disclosure;
[0012] Figure 2 A flowchart illustrating an evaluation method for a large model according to an embodiment of this disclosure is shown schematically;
[0013] Figure 3 An example schematic diagram illustrating a preset evaluation rule according to an embodiment of the present disclosure is shown.
[0014] Figure 4 The illustration shows an example diagram of the evaluation process for M large models according to an embodiment of the present disclosure;
[0015] Figure 5 The illustration shows an example diagram of the evaluation process for M large models according to another embodiment of the present disclosure;
[0016] Figure 6 The illustration shows an example schematic diagram of the evaluation process of a large model according to an embodiment of the present disclosure;
[0017] Figure 7 A block diagram schematically illustrates an evaluation apparatus for a large model according to an embodiment of the present disclosure; and
[0018] Figure 8A block diagram of an electronic device suitable for implementing an evaluation method for large models according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0019] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0022] When using expressions such as "at least one of A, B, and C", the expression should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).
[0023] Large language models possess system setting information, which provides initial configuration or context for the large language model and influences its behavior and output. This system setting information needs to be globally effective and has strong control over the large language model's responses.
[0024] In one example, methods for measuring the quality of responses generated by a large language model under the guidance of system-defined information include at least one of the following: automated evaluation methods and human evaluation methods. Automated evaluation methods refer to measuring the quality of model-generated text through algorithms or pre-defined metrics. Human evaluation methods refer to measuring the quality of model-generated text by relying on subjective scoring of the generated responses by human evaluators.
[0025] However, the methods described above struggle to make accurate assessments when user input conflicts with system settings, failing to effectively identify critical issues such as system-customized jailbreaks and data leaks within large models. Furthermore, these methods cannot comprehensively and effectively uncover obvious problems in model responses, and they also lack sufficient ability to assess the model's fine-grained capabilities.
[0026] To address this, embodiments of this disclosure propose an evaluation scheme for large language models. For example, for the response information of M large language models to input instructions, each response information is evaluated separately based on preset evaluation rules to obtain first evaluation information for each response information, where M is a positive integer greater than 1; since the first evaluation information of the M large language models is consistent with each other, each response information is evaluated separately based on multiple evaluation dimensions to obtain second evaluation information for each response information; and, based on the second evaluation information of each response information, an evaluation result is determined, and the evaluation result characterizes the response capability of each of the M large language models.
[0027] According to embodiments of this disclosure, preset evaluation rules can be used to perform preliminary evaluation of the response information of each large language model. When the first evaluation information of each large language model is consistent with each other, multiple evaluation dimensions can be used to further evaluate the response information of each large language model. By combining coarse-grained preset evaluation rules and fine-grained evaluation dimensions to evaluate the response information of each large model, not only can obvious problems in the response information be found, but also the degree to which the response information meets the evaluation dimensions can be determined, thereby improving the accuracy of the evaluation results and realizing an accurate evaluation of the response capability of the large language model.
[0028] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0029] In the technical solution of the present invention, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0030] Figure 1 The illustration schematically depicts a system architecture for applying large-scale model evaluation methods according to embodiments of this disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0031] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0032] Users can interact with server 105 via network 104 using at least one of the first terminal device 101, second terminal device 102, and third terminal device 103 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, second terminal device 102, and third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0035] It should be noted that the large model evaluation method provided in this embodiment can generally be executed by server 105. Correspondingly, the large model evaluation apparatus provided in this embodiment can generally be located in server 105. The large model evaluation method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the large model evaluation apparatus provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0036] Alternatively, the large model evaluation method provided in this embodiment of the present disclosure can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the large model evaluation apparatus provided in this embodiment of the present disclosure can also be located in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0037] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0038] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.
[0039] The above describes the system architecture of the evaluation method for large models provided in this disclosure. The following will use... Figure 2 As an example, the evaluation process of the large model disclosed herein will be further explained.
[0040] Figure 2 A flowchart illustrating an evaluation method for a large model according to an embodiment of this disclosure is shown schematically.
[0041] like Figure 2 As shown, the evaluation method 200 of this large model includes operations S210 to S230.
[0042] In operation S210, for the response information of each of the M large language models to the input command, each response information is evaluated based on the preset evaluation rules to obtain the first evaluation information of each response information, where M is a positive integer greater than 1.
[0043] In operation S220, in response to the consistency of the first evaluation information of each of the M large language models, each response information is evaluated separately based on multiple evaluation dimensions to obtain the second evaluation information of each response information.
[0044] In operation S230, the evaluation result is determined based on the second evaluation information of each response information. The evaluation result characterizes the response capability of each of the M large language models.
[0045] Large language models can be considered large-scale artificial intelligence models, which are machine learning models with extremely large parameters and complex computational structures. Large language models can process massive amounts of data and complete various complex text processing tasks. Examples include natural language processing tasks, dialogue tasks, text generation tasks, sentiment analysis tasks, translation tasks, summarization tasks, intelligent search tasks, speech recognition and synthesis tasks, etc.
[0046] Large language models can include prompts and perform text processing tasks based on these prompts (System Settings). Prompts refer to the settings or parameters of a large language model, used to control its behavior and performance during runtime, guiding its response to input commands. By setting different prompts, the output performance of the large language model can be optimized according to specific needs.
[0047] Input instructions (User Messages) are questions, requests, or instructions posed by the user to guide the large language model in its response. The same input instruction can be input into multiple large language models, yielding response information from each model. Response information (Model Output) refers to the output text generated by the large language model based on the input instruction, answering user questions, providing information, or performing corresponding actions.
[0048] After obtaining the response information of multiple large language models to input instructions, each response can be evaluated separately based on preset evaluation rules to obtain the initial evaluation information for each response. The evaluation process based on preset evaluation rules can be understood as a coarse-grained evaluation, used to assess the responsiveness of the large language models. The initial evaluation information can be used to characterize the responsiveness of the large language models.
[0049] The preset evaluation rules can be configured according to actual business needs and are not limited here. For example, a preset evaluation rule could be to determine that the first evaluation information passes the verification if the degree of consistency between the response information and the prompt information is greater than a first preset threshold. Alternatively, a preset evaluation rule could be to determine that the first evaluation information passes the verification if the degree of consistency between the response information and the input command is greater than a second preset threshold. Alternatively, a preset evaluation rule could be to prioritize the prompt information over the input command.
[0050] After obtaining the initial evaluation information for each large language model, it can be determined whether multiple initial evaluation information are consistent with each other. If the initial evaluation information for multiple large language models is inconsistent, the responsiveness of the large language model can be determined directly based on these multiple initial evaluation information. Inconsistency among multiple initial evaluation information can be understood as follows: a first number of initial evaluation information points indicate that the large language model passes the verification of the preset evaluation rules, a second number of initial evaluation information points indicate that the large language model fails the verification of the preset evaluation rules, and the sum of the first and second numbers is M. In this case, the responsiveness of the large language model corresponding to the initial evaluation information indicating that the large language model passes the preset evaluation rules can be determined as the first level. Alternatively, the responsiveness of the large language model corresponding to the second evaluation information indicating that the large language model fails the preset evaluation rules can be determined as the second level.
[0051] When the initial evaluation information of multiple large language models is consistent with each other, the models can be further evaluated based on multiple evaluation dimensions to obtain the second evaluation information for each response. Consistency of multiple initial evaluation information can be understood as: all M initial evaluation information points indicate that the large language model passes the verification of the preset evaluation rules, or all M initial evaluation information points indicate that the large language model fails the verification of the preset evaluation rules. The evaluation process based on evaluation dimensions can be understood as fine-grained evaluation, i.e., used to assess the quality of the large language model's response. The second evaluation information can be used to characterize the quality of the large language model's response.
[0052] After obtaining the second evaluation information for each response, the evaluation results characterizing the responsiveness of each large language model can be determined. The responsiveness of a large language model refers to its ability to respond promptly, accurately, and appropriately to input content. Responsiveness depends on the training data, architecture, and optimization methods of the large language model. Large language models with good responsiveness are better able to understand context, answer user questions, provide information, or perform corresponding actions.
[0053] According to embodiments of this disclosure, preset evaluation rules can be used to perform preliminary evaluation of the response information of each large language model. When the first evaluation information of each large language model is consistent with each other, multiple evaluation dimensions can be used to further evaluate the response information of each large language model. By combining coarse-grained preset evaluation rules and fine-grained evaluation dimensions to evaluate the response information of each large model, it is possible not only to discover problems with clear standard or non-standard characteristics in the response information, but also to determine the degree to which the response information meets the evaluation dimensions, thereby improving the accuracy of the evaluation results and achieving an accurate evaluation of the response capability of the large language model.
[0054] In the embodiments of this disclosure, the evaluation of response information based on preset evaluation rules can be understood as a first-level coarse-grained evaluation process. The following will utilize... Figure 3 The pre-defined evaluation rules are described in an illustrative manner.
[0055] Figure 3 An example schematic diagram illustrating a preset evaluation rule according to an embodiment of the present disclosure is shown.
[0056] like Figure 3 As shown in example 300, the preset evaluation rule can be based on the priority settings of different source instructions according to the large language model. For example, the preset evaluation rule can be: the priority of prompts is higher than the priority of input instructions, and the priority of input instructions is higher than the priority of response information. In this case, the highest priority is the prompt, followed by the input instructions, and finally the response information. Alternatively, the preset evaluation rule can also be: the priority of prompts is higher than the priority of input instructions; or, the preset evaluation rule can also be: the priority of input instructions is higher than the priority of response information.
[0057] When a conflict arises between a prompt and an input instruction, the large language model prioritizes the prompt when generating a response. For example, if the prompt is "You are now Sun Wukong, and you can never leave this role," while the input instruction is "Please play Tang Sanzang," then if the response output by the large language model begins to play Tang Sanzang, it is determined that the first evaluation information of this response indicates that the large language model has not passed the preset evaluation rules; if the response output by the large language model continues to play Sun Wukong, it is determined that the first evaluation information of this response indicates that the large language model has passed the preset evaluation rules.
[0058] According to embodiments of this disclosure, by setting a hierarchical instruction priority system where the priority of prompt information is higher than that of input instructions, and the priority of input instructions is higher than that of response information, the instruction priority can be flexibly adjusted according to actual needs. This enables the detection of clearly non-standard issues in the response information, thereby improving the accuracy of coarse-grained large language model response capability assessment.
[0059] The above text provided an illustrative description of the pre-set evaluation rules; the following text will utilize... Figure 4 The first-level coarse-grained evaluation process, which evaluates response information based on preset evaluation rules to obtain first evaluation information, is described.
[0060] Figure 4 The illustration shows an example diagram of the evaluation process for M large models according to an embodiment of the present disclosure.
[0061] like Figure 4As shown in Figure 400, taking M=2 as an example, the evaluation process of the large language model is illustrated by inputting the same input instruction 401 into the large language model 402_1 and the large language model 402_2 respectively.
[0062] For the large language model 402_1, the prompt information 403_1 is used to guide the large language model 402_1 in responding to the input command 401. Under the guidance of the prompt information 403_1, the large language model 402_1 can output response information 404_1. After obtaining the response information 404_1, operation S410 can be executed. In operation S410, it can be determined whether the response information 404_1 and the prompt information 403_1 are consistent.
[0063] If not, then it can be determined that the first evaluation information 405, representing the large language model 402_1, satisfies the preset evaluation rules. The consistency between the response information 404_1 and the prompt information 403_1 indicates that the large language model 402_1 answers based on the prompt information 403_1.
[0064] If so, it can be determined that the first evaluation information 406, representing the large language model 402_1, does not meet the preset evaluation rules. The inconsistency between the response information 404_1 and the prompt information 403_1 indicates that the large language model 402_1 did not respond based on information 403_1.
[0065] For the large language model 402_2, the prompt information 403_2 is used to guide the large language model 402_2 in responding to the input command 401. Under the guidance of the prompt information 403_2, the large language model 402_2 can output response information 404_2. After obtaining the response information 404_2, operation S420 can be executed. In operation S420, it can be determined whether the response information 404_2 and the prompt information 403_2 are consistent.
[0066] If not, then it can be determined that the first evaluation information 407, representing the large language model 402_2, satisfies the preset evaluation rules. The consistency between the response information 404_2 and the prompt information 403_2 indicates that the large language model 402_2 responded based on the prompt information 403_2.
[0067] If so, it can be determined that the first evaluation information 408, representing the large language model 402_2, does not meet the preset evaluation rules. The inconsistency between the response information 404_2 and the prompt information 403_2 indicates that the large language model 402_2 did not respond based on the prompt information 403_2.
[0068] According to embodiments of this disclosure, by comparing the relationship between response information and prompt information, it is determined whether the performance of the large language model conforms to the preset evaluation rules, and different first evaluation information is given when the response information and prompt information are consistent and inconsistent, thereby ensuring the automation of the coarse-grained evaluation process and improving evaluation efficiency.
[0069] After obtaining the initial evaluation information for each large language model, it can be determined whether these multiple initial evaluation information are consistent. If the multiple initial evaluation information are inconsistent, the response capability of each large language model can be directly determined without performing further evaluation operations based on evaluation dimensions.
[0070] For example, when the first evaluation information of the large language model 402_1 and the first evaluation information of the large language model 402_2 are inconsistent, taking the first evaluation information 406 representing that the large language model 402_1 does not meet the preset evaluation rules, and the first evaluation information 407 representing that the large language model 402_2 meets the preset evaluation rules, as an example, the response capability of the large language model representing that it does not meet the preset evaluation rules can be determined as the second level 409, and the response capability of the large language model representing that it meets the preset evaluation rules can be determined as the first level 410. That is, the response capability of the large language model 402_1 is determined as the second level 409, and the response capability of the large language model 402_2 is determined as the first level 410. For example, the first level 410 is high-level, and the second level 409 is low-level.
[0071] According to embodiments of this disclosure, by comparing the first evaluation information of multiple large language models based on preset evaluation rules, it is possible to focus on identifying clearly defined non-standard issues in the participating models during the coarse-grained evaluation stage, and directly classify their response capabilities when the first evaluation information of each large language model is inconsistent, thereby improving the evaluation efficiency of the performance of large language models.
[0072] The above section provided an illustrative description of the first-level coarse-grained evaluation process. The following section will utilize... Figure 5 The process of evaluating response information based on multiple evaluation dimensions to obtain second evaluation information at the second level is described.
[0073] Figure 5 The illustration shows an example diagram of the evaluation process for M large models according to another embodiment of the present disclosure.
[0074] like Figure 5 As shown in Figure 500, taking M=2 as an example, the evaluation process of the large language model is illustrated by inputting the same input instruction 501 into the large language model 502_1 and the large language model 502_2 respectively.
[0075] For the large language model 502_1, response information 504_1 can be output under the guidance of prompt information 503_1. After obtaining response information 504_1, operation S510 can be executed. In operation S510, it can be determined whether response information 504_1 and prompt information 503_1 are consistent.
[0076] If not, then the first evaluation information 505 representing the large language model 502_1 satisfying the preset evaluation rules can be determined. If yes, then the first evaluation information 506 representing the large language model 502_1 not satisfying the preset evaluation rules can be determined.
[0077] For the large language model 502_2, response information 504_2 can be output under the guidance of prompt information 503_2. After obtaining response information 504_2, operation S520 can be executed. In operation S520, it can be determined whether response information 504_2 and prompt information 503_2 are consistent.
[0078] If not, then the first evaluation information 507 representing the large language model 502_2 satisfying the preset evaluation rules can be determined. If yes, then the first evaluation information 508 representing the large language model 502_2 not satisfying the preset evaluation rules can be determined.
[0079] After obtaining the initial evaluation information for each large language model, it can be determined whether these multiple initial evaluation information are consistent. If the multiple initial evaluation information are consistent, then an evaluation operation based on evaluation dimensions is required.
[0080] For example, when the first evaluation information of the large language model 402_1 and the first evaluation information of the large language model 402_2 are consistent, the first evaluation information 505 representing that the large language model 502_1 meets the preset evaluation rules and the first evaluation information 507 representing that the large language model 502_2 meets the preset evaluation rules can be used as examples.
[0081] For the large language model 502_1, semantic matching needs to be performed on the response information 504_1 based on the prompt information 503_1 for each evaluation dimension, resulting in matching information 509 for each evaluation dimension. Semantic matching refers to evaluating the semantic similarity between two texts in natural language processing. The method of semantic matching can be configured according to actual business needs and is not limited here. For example, semantic matching methods can include at least one of the following: word-level semantic matching, sentence-level semantic matching, and semantic matching based on deep learning models.
[0082] According to embodiments of this disclosure, semantic matching is performed based on prompts from multiple evaluation dimensions, resulting in a more detailed and comprehensive evaluation of the response information of the large language model. By generating matching information from each dimension and integrating it into a second evaluation, the comprehensiveness, accuracy, and flexibility of the evaluation are ensured, which helps to improve the practical application effect of the large language model.
[0083] In one example, the prompts for each evaluation dimension may include at least one of the following: character customization information, role customization information, ability customization information, and style customization information. The prompts for each evaluation dimension can better meet the personalized needs of different users.
[0084] Character customization information specifies the persona the large language model will portray, including personalized settings such as appearance, personality, and behavioral characteristics. For example, character customization information could be "mobile phone assistant." Role customization information specifies the specific role or entity the large language model will portray, including personalized settings such as appearance, skills, and personality. For example, role customization information could be an existing role or a fictional role. Ability customization information specifies the large language model's abilities, skills, or attributes. For example, ability customization information could be "providing purchasing services." Style customization information specifies the language style the large language model will use in response. For example, style customization information could be "tsundere" or "aloof."
[0085] According to embodiments of this disclosure, by employing an evaluation mechanism based on different dimensions of prompts, such as character customization information, role customization information, ability customization information, and style customization information, the evaluation results can provide developers with more detailed feedback, clarifying which specific areas or abilities the model is lacking in. This enhances the personalization and targeting of the evaluation process, ensuring that the evaluation process can more precisely detect performance differences in the model and clarify which specific areas or abilities the model is lacking in, thereby improving the adaptability and optimization potential of the large language model. This mechanism also provides more flexible and diverse options for the evaluation process, making the evaluation process more comprehensive and adaptable to actual needs.
[0086] After obtaining the matching information 509 for each evaluation dimension, the second evaluation information of the response information can be determined based on the matching information 509 for each evaluation dimension. On this basis, the evaluation information can be weighted according to the preset weights of each evaluation dimension to obtain the weighted evaluation information 510 of the large language model 502_1. The preset weights refer to fixed weight values set for different evaluation dimensions to measure their contribution to the overall evaluation. The sum of the preset weights for each evaluation dimension is 1. By multiplying the information obtained from different evaluation dimensions by their corresponding weights according to these preset weights, they are incorporated into a comprehensive consideration, thereby reflecting the importance of each evaluation dimension to the final weighted evaluation information 510.
[0087] Similarly, for the large language model 502_2, semantic matching is performed on the response information 504_2 based on the prompt information 503_2 for each evaluation dimension, resulting in matching information 511 for each evaluation dimension. Based on the matching information 511 for each evaluation dimension, the second evaluation information of the response information is determined. On this basis, the evaluation information can be weighted according to the preset weights of each evaluation dimension to obtain the weighted evaluation information 512 of the large language model 502_2.
[0088] After obtaining the weighted evaluation information 510 of the large language model 502_1 and the weighted evaluation information 512 of the large language model 502_2, the large language models 502_1 and 502_2 can be ranked according to the weighted evaluation information 510 and 512. For example, if the weighted evaluation information 510 is greater than the weighted evaluation information 512, the large language model 502_1 can be ranked first and the large language model 502_2 last. Thus, the large language model 502_1 can be determined as the first level and the large language model 502_2 as the second level, resulting in the evaluation result 513.
[0089] It should be noted that when evaluating multiple large language models, predetermined rankings can be set in advance. The response capabilities of the top-ranked large language models can be designated as the first level, while the response capabilities of the remaining models can be designated as the second level, thus obtaining the evaluation results. The predetermined rankings can be set based on historical experience and the number of large language models participating in the evaluation; for example, the predetermined ranking could be 3.
[0090] According to embodiments of this disclosure, a flexible and accurate multi-dimensional large language model evaluation method is provided through weighted processing, sorting, and ranking mechanisms. This method not only allows for flexible adjustment of evaluation weights for different application scenarios but also improves evaluation efficiency through sorting and ranking, thereby ensuring the rapid and accurate selection of the best-performing large language model. At the same time, it provides clear feedback directions for the optimization and improvement of the large language model.
[0091] The evaluation process for multiple large language models has been illustrated above. The following section will utilize... Figure 6 The evaluation process for a single large language model is explained.
[0092] Figure 6 The illustration shows an example schematic diagram of the evaluation process of a large model according to an embodiment of the present disclosure.
[0093] like Figure 6 As shown, in 600, the large language model 602 has a prompt message 603 to guide the large language model 602 in responding to the input command 601. Under the guidance of the prompt message 603, the large language model 602 can output response information 604.
[0094] After obtaining response information 604, it can be evaluated based on preset evaluation rules to obtain first evaluation information. For example, operation S610 can be executed. In operation S610, it can be determined whether response information 604 and prompt information 603 are consistent.
[0095] If not, then the first evaluation information 605 representing that the large language model 602 meets the preset evaluation rules can be determined, and the response capability of the large language model 602 can be determined to be at the first level 606. If yes, then the first evaluation information 406 representing that the large language model 602 does not meet the preset evaluation rules can be determined, and the response capability of the large language model 602 can be determined to be at the second level 608.
[0096] According to embodiments of this disclosure, for the evaluation of a single large language model, the performance of the large language model can be evaluated directly using an evaluation method based on preset evaluation rules. By automating the evaluation and grading of the response information of the large language model, the evaluation efficiency and accuracy are improved.
[0097] The above are merely exemplary embodiments, but are not limited thereto. Other evaluation methods for large models known in the art may also be included, as long as they can achieve an accurate evaluation of the responsiveness of large language models.
[0098] Based on the evaluation method for large models provided in this disclosure, this disclosure also provides an evaluation apparatus for large models. The following will utilize... Figure 7 The device is described in detail.
[0099] Figure 7 A block diagram of an evaluation apparatus for a large model according to an embodiment of the present disclosure is shown schematically.
[0100] like Figure 7 As shown, the large model evaluation device 700 may include a first evaluation module 710, a second evaluation module 720, and a determination module 730.
[0101] The first evaluation module 710 is used to evaluate each response information of the M large language models in response to the input command based on the preset evaluation rules, and obtain the first evaluation information of each response information, where M is a positive integer greater than 1.
[0102] The second evaluation module 720 is used to evaluate each response information based on multiple evaluation dimensions in response to the consistency between the first evaluation information of each of the M large language models, thereby obtaining the second evaluation information of each response information.
[0103] The determination module 730 is used to determine the evaluation result based on the second evaluation information of each response information, wherein the evaluation result characterizes the response capability of each of the M large language models.
[0104] According to embodiments of this disclosure, each large language model has its own prompt information, which guides the large language model to respond to input commands. Preset evaluation rules include at least one of the following: the prompt information has a higher priority than the input command; or, the input command has a higher priority than the response information.
[0105] According to embodiments of this disclosure, the first evaluation module 710 may include a first determining unit and a second determining unit.
[0106] The first determining unit is used to determine, for each large language model, the first evaluation information characterizing that the large language model meets the preset evaluation rules, provided that the response information and the prompt information are consistent.
[0107] The second determining unit is used to determine, when the response information and the prompt information are inconsistent, that the first evaluation information represents that the large language model does not meet the preset evaluation rules.
[0108] According to embodiments of this disclosure, the second evaluation module 720 may include a matching unit and a third determination unit.
[0109] The matching unit is used to perform semantic matching on each response information based on the prompt information of each evaluation dimension, so as to obtain the matching information of each evaluation dimension.
[0110] The third determining unit is used to determine the second evaluation information of the response information based on the matching information of each evaluation dimension.
[0111] According to embodiments of this disclosure, the prompts for each evaluation dimension include at least one of the following: character customization information, role customization information, ability customization information, and style customization information.
[0112] According to embodiments of this disclosure, the determining module 730 may include a weighted processing unit, a sorting unit, and a fourth determining unit.
[0113] The weighted processing unit is used to perform weighted processing on the second evaluation information according to the preset weights of each evaluation dimension for each response information, so as to obtain the weighted evaluation information of each response information.
[0114] The sorting unit is used to sort the M large language models according to the weighted evaluation information of each response information, so as to obtain the sorted M large language models.
[0115] The fourth determining unit is used to determine the response capability of the second-to-first-ranked large language model among the sorted M large language models as the first level, and the response capability of the large language models other than the second-to-first-ranked large language models as the second level, so as to obtain the evaluation result.
[0116] According to embodiments of this disclosure, the large model evaluation apparatus 700 may further include a first determining module and a second determining module.
[0117] The first determining module is used to determine the response capability of the large language model whose first evaluation information satisfies the preset evaluation rules among the M large language models in response to the inconsistency of their respective first evaluation information.
[0118] The second determining module is used to determine the response capability of the large language model whose first evaluation information representation does not meet the preset evaluation rules as the second level among the M large language models.
[0119] According to embodiments of this disclosure, the large model evaluation device 700 may further include a third evaluation module, a fourth determination module, and a fifth determination module.
[0120] The third evaluation module is used to evaluate the response information of the large language model to the input command based on preset evaluation rules to obtain the first evaluation information.
[0121] The fourth determination module is used to determine the response capability of the large language model as the first level, provided that the first evaluation information characterization of the large language model meets the preset evaluation rules.
[0122] The fifth determination module is used to determine the response capability of the large language model as the second level when the first evaluation information characterization does not meet the preset evaluation rules.
[0123] Figure 8 The diagram schematically illustrates an electronic device suitable for implementing an evaluation method for large models according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0124] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0125] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0126] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as large model evaluation methods. For example, in some embodiments, the large model evaluation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the large model evaluation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform large model evaluation methods by any other suitable means (e.g., by means of firmware).
[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0128] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0129] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0131] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0132] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0133] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0134] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for evaluating large models, comprising: For each of the M large language models outputting text to an input command, each output text is evaluated based on a preset evaluation rule to obtain first evaluation information for each output text. The large language models are used to perform text processing tasks. Each large language model has globally effective prompt information, which provides initial configuration or context to control the behavior and performance of the large language model during runtime and guides the large language model to respond to the input command. The preset evaluation rule evaluates the priority of compliance when at least two of the prompt information, the input command, and the output text conflict. The priority of compliance includes at least one of the following: the priority of the prompt information is higher than the priority of the input command; or the priority of the input command is higher than the priority of the output text, where M is a positive integer greater than 1. In response to the consistency of the first evaluation information of each of the M large language models, each output text is evaluated based on multiple evaluation dimensions to determine the semantic similarity between the prompt information of each evaluation dimension and the output text, thereby obtaining the second evaluation information of each output text. as well as An evaluation result is determined based on the second evaluation information of each of the output texts, wherein the evaluation result characterizes the responsiveness of each of the M large language models in generating the output texts under the guidance of the prompt information.
2. The method according to claim 1, wherein, The evaluation of each output text based on preset evaluation rules, to obtain the first evaluation information for each output text, includes: For each of the aforementioned large language models, If the output text matches the prompt information, it is determined that the first evaluation information indicates that the large language model satisfies the preset evaluation rule; and If the output text is inconsistent with the prompt information, it is determined that the first evaluation information indicates that the large language model does not meet the preset evaluation rules.
3. The method according to claim 1 or 2, wherein, The second evaluation information for each output text, based on multiple evaluation dimensions, includes: For each of the output texts, Based on the prompt information for each of the evaluation dimensions, semantic matching is performed on the output text to obtain matching information for each of the evaluation dimensions; and Based on the matching information for each of the evaluation dimensions, the second evaluation information of the output text is determined.
4. The method according to claim 3, wherein, Each of the assessment dimensions includes at least one of the following prompts: character customization information, role customization information, ability customization information, and style customization information.
5. The method according to claim 1 or 2, wherein, The step of determining the evaluation result based on the second evaluation information of each of the output texts includes: For each output text, the second evaluation information is weighted according to the preset weight of each evaluation dimension to obtain the weighted evaluation information of each output text. Based on the weighted evaluation information of each output text, the M large language models are sorted to obtain the sorted M large language models; and The response capability of the top two largest language models in the sorted M large language models is determined as the first level, and the response capability of the remaining large language models is determined as the second level, thus obtaining the evaluation result.
6. The method according to claim 1, further comprising: In response to the inconsistency of the first evaluation information of the M large language models, the response capability of the large language model whose first evaluation information represents the satisfaction of the preset evaluation rules is determined as the first level among the M large language models; as well as Among the M large language models, the response capability of the large language model whose first evaluation information does not meet the preset evaluation rules is determined as the second level.
7. The method according to claim 1, further comprising: Based on the preset evaluation rules, the output text of the large language model in response to the input instruction is evaluated to obtain the first evaluation information. If the first evaluation information indicates that the large language model meets the preset evaluation rules, the response capability of the large language model is determined to be at the first level. as well as If the first evaluation information indicates that the large language model does not meet the preset evaluation rules, the response capability of the large language model is determined to be at the second level.
8. An evaluation device for a large model, comprising: The first evaluation module is used to evaluate the output text of each of the M large language models in response to the input command, based on a preset evaluation rule, to obtain first evaluation information for each output text. Each large language model has globally effective prompt information, which is used to provide initial configuration or context to control the behavior and performance of the large language model during runtime and to guide the large language model to respond to the input command. The preset evaluation rule is used to evaluate the priority of compliance when at least two of the prompt information, the input command, and the output text conflict. The priority of compliance includes at least one of the following: the priority of the prompt information is higher than the priority of the input command; or the priority of the input command is higher than the priority of the output text, where M is a positive integer greater than 1. The second evaluation module is used to evaluate each of the output texts based on multiple evaluation dimensions in response to the consistency of the first evaluation information of each of the M large language models, so as to determine the semantic similarity between the prompt information of each evaluation dimension and the output text, and obtain the second evaluation information of each output text. as well as A determination module is used to determine an evaluation result based on the second evaluation information of each of the output texts, wherein the evaluation result characterizes the responsiveness of each of the M large language models in generating the output texts under the guidance of the prompt information.
9. The apparatus according to claim 8, wherein, The first evaluation module includes: The first determining unit is configured to, for each of the large language models, determine, when the output text matches the prompt information, that the first evaluation information indicates that the large language model satisfies the preset evaluation rule; and The second determining unit is used to determine, when the output text is inconsistent with the prompt information, that the first evaluation information indicates that the large language model does not meet the preset evaluation rules.
10. The apparatus according to claim 8 or 9, wherein, The second evaluation module includes: A matching unit is configured to perform semantic matching on each output text based on the prompt information for each evaluation dimension, thereby obtaining matching information for each evaluation dimension; and The third determining unit is used to determine the second evaluation information of the output text based on the matching information of each of the evaluation dimensions.
11. The apparatus according to claim 10, wherein, Each of the assessment dimensions includes at least one of the following prompts: character customization information, role customization information, ability customization information, and style customization information.
12. The apparatus according to claim 8 or 9, wherein, The determining module includes: The weighted processing unit is used to perform weighted processing on the evaluation information for each output text according to the preset weight of each evaluation dimension, so as to obtain the weighted evaluation information of each output text. A sorting unit is configured to sort the M large language models according to the weighted evaluation information of each output text, thereby obtaining M sorted large language models; and The fourth determining unit is used to determine the response capability of the large language model located in the first predetermined position among the sorted M large language models as the first level, and the response capability of the large language models other than the first predetermined position as the second level, so as to obtain the evaluation result.
13. The apparatus according to claim 8, further comprising: The first determining module is used to determine the response capability of the large language model whose first evaluation information satisfies the preset evaluation rules as the first level in response to the inconsistency of the first evaluation information of the M large language models. as well as The second determining module is used to determine the response capability of the large language model among the M large language models whose first evaluation information characterization does not meet the preset evaluation rules as the second level.
14. The apparatus of claim 8, further comprising: The third evaluation module is used to evaluate the output text of the large language model in response to the input instruction based on the preset evaluation rules, and obtain the first evaluation information. The fourth determining module is used to determine the response capability of the large language model as the first level when the first evaluation information indicates that the large language model meets the preset evaluation rules; as well as The fifth determining module is used to determine the response capability of the large language model as the second level when the first evaluation information indicates that the large language model does not meet the preset evaluation rules.
15. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
16. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Model evaluation method and device, electronic equipment and storage medium
CN117272011A
Quality evaluation method, system and equipment of multi-modal large model and storage medium
CN117909702A
Model determination method and device, equipment, storage medium and program product
CN118349812A