Large language model evaluation method, server and computer readable storage medium

Through the automated evaluation method, the output results of the large language model are compared with the test set and the annotation set, the problem of low evaluation efficiency of large language models in the existing technology is solved, and a more efficient and accurate evaluation process is achieved.

CN120196520AActive Publication Date: 2025-06-24HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311729417.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-24
Estimated Expiration
2043-12-14

AI Technical Summary

Technical Problem

The evaluation efficiency of large language models in the prior art is low, and developers need to manually evaluate large language models at each stage, resulting in low efficiency.

Method used

Provide a large language model evaluation method, which automates the evaluation process and improves the evaluation efficiency by obtaining the test set, generating the result set and comparing it with the annotation set.

Benefits of technology

Through the automated evaluation process, the efficiency of large language model evaluation is significantly improved, manual intervention is reduced, and the accuracy and efficiency of evaluation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196520A_ABST
    Figure CN120196520A_ABST
Patent Text Reader

Abstract

The invention provides a large language model evaluation method, a server and a computer readable storage medium, and relates to the technical field of computer processing, and the method comprises the steps: obtaining a first test set; the first test set comprises a plurality of first test data, the first test data comprises first format information, and the first format information is used for specifying an output format of the language model as a first format; obtaining a first result set output by the first large language model based on the first test set; the first result set comprises multiple pieces of first result data, the multiple pieces of first result data are in one-to-one correspondence with the multiple pieces of first test data, and the format of the first result data is a first format; obtaining a first evaluation result based on the first result set and the label set; the first evaluation result is used for evaluating the capability of the first large language model; the annotation set comprises a plurality of annotation data, the plurality of annotation data correspond to the plurality of first test data, and the format of the annotation data is a first format. By means of the method, the efficiency of large language model evaluation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer processing technologies, and in particular, to a method for evaluating large language models, a server, and a computer-readable storage medium. Background Art

[0002] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, which can generate corresponding natural language text according to the input natural language text, or understand the meaning of the input natural language text. Currently, large language models can be used to process tasks such as text classification, question answering, and dialogue.

[0003] The process from the development to the completion of a large language model includes a pre-training stage, a fine-tuning stage, and a stage of going live for use. In each stage, it is necessary to evaluate the capabilities of the large language model. Since there are differences in large language models at different stages, developers need to manually evaluate the large language models at each stage. Thus, there is a problem of low efficiency in evaluating large language models. Summary of the Invention

[0004] In view of this, this application provides a method for evaluating large language models, a server, and a computer-readable storage medium, which improves the efficiency of evaluating large language models.

[0005] In a first aspect, this application provides a method for evaluating a large language model. The method includes: obtaining a first test set; the first test set includes multiple first test data, and the first test data includes first format information, where the first format information is used to specify that the output format of the language model is the first format; obtaining a first result set output by the first large language model based on the first test set; the first result set includes multiple first result data, and the multiple first result data correspond one-to-one to the multiple first test data, and the format of the first result data is the first format; obtaining a first evaluation result based on the first result set and an annotation set; the first evaluation result is used to evaluate the capabilities of the first large language model; the annotation set includes multiple annotation data, and the multiple annotation data correspond one-to-one to the multiple first test data, and the format of the annotation data is the first format.

[0006] Among them, the first test data can be of a type that the text large language model can process, and the specific type of the first test data is not limited here.

[0007] The capabilities of large language models include one or more of factual Q&A, reading comprehension, framework generation, passage rewriting, summary extraction, math problem-solving, reasoning, poetry generation, or programming fields, etc. Understandably, the capabilities of large language models are not substantially limited to the above-listed content, and there may be similarities between two capabilities of large language models. For example, the reading comprehension ability and the summary extraction ability of large language models may both require the large language model to perform information extraction. In addition, the capabilities of large language models can also be described in other ways, such as information extraction ability, sentiment analysis ability, etc.

[0008] In the above method, the first result data is output by the first large language model processing the first test data, so the first result data and the first test data are in one-to-one correspondence. Also, since the first test data and the labeled data are in one-to-one correspondence, therefore, the first result data and the labeled data are in one-to-one correspondence. Moreover, the formats of the first result data and the labeled data are both the first format, so the server can efficiently process the result data and the labeled data with the same format, thereby efficiently obtaining the first evaluation result. Thus, through the above method, the efficiency of large language model evaluation can be effectively improved.

[0009] In a possible implementation manner of the first aspect, obtaining the first test set includes: obtaining a first data set, a first prompt template set, and a first format; the first data set includes multiple data, and the first prompt template set includes at least one first prompt template; based on the first data set, the first prompt template set, and the first format, generate the first test set.

[0010] In the above implementation manner, the server can first merge the data in the first data set, the first prompt templates in the first prompt template set, and the first format to generate multiple first test data. In this way, a first test set including multiple first test data can be generated. By combining a small amount of data and a small number of first prompt templates, a large amount of first test data can be obtained, and the large language model can be more effectively evaluated using the obtained multiple first test data.

[0011] In a possible implementation manner of the first aspect, the first format is a pre-configured format; or, the first format is a custom format.

[0012] In a possible implementation manner of the first aspect, the method further includes: receiving a first instruction from an electronic device; the first instruction is used to specify that the output format of the language model is the first format.

[0013] In a possible implementation of the first aspect, obtaining the first evaluation result based on the first result set and the annotation set includes: using test code to compare and analyze each first result data in the first result set with the corresponding annotation data in the annotation set to obtain multiple comparison results; obtaining the first evaluation result based on the multiple comparison results and the evaluation metrics.

[0014] Since the format of the first result data in the first result set is the first format and the format of the annotation data in the annotation set is the first format, test code can be used to compare and analyze the first result data and the annotation data with the same format, thereby effectively improving the efficiency of evaluating the large language model.

[0015] In a possible implementation of the first aspect, the first large language model can understand the first format information;

[0016] Before obtaining the first test set, the method further includes: obtaining a second test set for testing the first large language model's ability to understand the first format information; testing the first large language model's ability to understand the first format information based on the second test set.

[0017] In the above implementation process, when the server tests the first large language model's ability to understand the first format information based on the second test set, it can first obtain the test result output by the first large language model based on the second test set, and determine the first large language model's ability to understand the format information based on the test result; where the test data included in the second test set contains format information, and the test result includes the test result data output by the first large language model based on the test data in the input second test set; if the format of the test result data is the first format, it indicates that the first large language model can understand the format information, and if the format of the test result data is not the first format, it indicates that the first large language model cannot understand the format information.

[0018] In the case where the first large language model can understand the format information, it can ensure that the format of the first result data output by the first large language model is the first format, thereby efficiently obtaining the first evaluation result for the first result set and the annotation set with the same first format.

[0019] In a possible implementation of the first aspect, the method further includes: in response to the first large language model being unable to understand the first format information, obtaining a third test set, where the third test set includes a plurality of second test data, and the third test set is generated based on the first data set and the first prompt template set; obtaining a second result set output by the first large language model based on the third test set; the second result set includes a plurality of second result data, and the plurality of second result data corresponds one-to-one to the plurality of second test data; obtaining a fourth test set based on the third test set, the second result set, and the second format information; the fourth test set is used to instruct the second large language model to analyze the capabilities of the first large language model, the fourth test set includes a plurality of third test data, and one third test data includes a second test data, the second result data corresponding to the second test data, and the second format information; the second format information is used to specify that the output format of the language model is the second format; obtaining a third result set output by the second large language model based on the fourth test set; the third result set includes a plurality of third result data, and the plurality of third result data corresponds one-to-one to the plurality of third test data, and the format of the third result data is the second format; obtaining a second evaluation result based on the third result set. Among them, the second format information and the first format information may be the same or different. Specifically, the first format information and the second format information can be determined according to the actual needs of the user.

[0020] In a possible implementation of the first aspect, obtaining the second evaluation result based on the third result set includes: obtaining the second evaluation result based on the third result set and the evaluation metric.

[0021] In a possible implementation of the first aspect, the method further includes: receiving an evaluation metric from an electronic device. The evaluation metric may include at least one of accuracy, precision, and recall.

[0022] In the above implementation, when determining based on the third result set and the evaluation metric, the server may determine the number of correct results and the number of incorrect results in the second result set representing the output of the first large language model based on the third test set in the third result set, and then, based on the evaluation metric, determine the evaluation result. For example, if the evaluation metric is accuracy, the server may divide the number of correct results in the second result set representing the output of the first large language model based on the third test set in the third result set by the total number of results in the third result set, so as to obtain the evaluation result represented as accuracy.

[0023] In a possible implementation of the first aspect, the method further includes: obtaining a second set of prompt templates; the second set of prompt templates includes at least one second prompt template; obtaining a fifth test set based on the first data set, the second set of prompt templates, and the first format; obtaining a fourth result set output by the first large language model based on the fifth test set; the format of the result data included in the fourth result set is the first format; obtaining a third evaluation result based on the fourth result set and the annotation set; the third evaluation result is used to evaluate the capabilities of the first large language model; obtaining a prompt evaluation result based on the first evaluation result and the third evaluation result, and the prompt evaluation result is used to evaluate the capabilities of the first prompt template and the second prompt template.

[0024] In a second aspect, the present application provides a server, which includes a communication module, a memory, and one or more processors; the communication module, the memory, and the processor are coupled; the communication module is used to establish a communication connection and send and receive data through the communication connection, the memory is used to store computer program code, and the computer program code includes computer instructions; when the processor executes the computer instructions, the server is caused to execute the method described in the first aspect and any of its possible design manners.

[0025] In a third aspect, the present application provides a computer-readable storage medium, including computer instructions, which when running on a server, cause the server to execute the method described in the first aspect and any of its possible design manners.

[0026] In a fourth aspect, the present application provides a computer program product, which when running on a server, causes the server to execute the method described in the first aspect and any of its possible design manners above.

[0027] In a fifth aspect, the present application provides a device, which is included in the server and has the function of implementing the server behavior in any of the methods in the above aspects and possible implementation manners. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, an allocation module or unit, a scanning module or unit, a recycling module or unit, a moving module or unit, and a storage module or unit, etc.

[0028] In a sixth aspect, the embodiments of the present application provide a chip system, which includes a processor and may further include a memory for implementing any of the methods provided in the first aspect and any of its possible design manners. The chip system may be composed of chips or may include chips and other discrete devices.

[0029] Understandably, the server described in the second aspect and any possible design method thereof provided above, the computer-readable storage medium described in the third aspect, and the computer program product described in the fourth aspect are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here. Description of the Drawings

[0030] Figure 1 Schematic diagram of the development process of a large language model provided by an embodiment of the present application;

[0031] Figure 2 Schematic diagram of the evaluation of a large language model provided by an embodiment of the present application;

[0032] Figure 3 Schematic diagram of a method for evaluating a large language model provided by an embodiment of the present application;

[0033] Figure 4 Schematic diagram of a model evaluation system provided by an embodiment of the present application;

[0034] Figure 5 Schematic diagram of the hardware structure of a server provided by an embodiment of the present application;

[0035] Figure 6 Flow chart of a method for evaluating a large language model provided by an embodiment of the present application Figure 1 ;

[0036] Figure 7 Schematic diagram of the operation display interface for evaluating a large language model provided by an embodiment of the present application;

[0037] Figure 8 Schematic diagram of the evaluation of a large language model provided by an embodiment of the present application Figure 1 ;

[0038] Figure 9 Flow chart of a method for evaluating a large language model provided by an embodiment of the present application Figure 2 ;

[0039] Figure 10 Schematic diagram of the evaluation of a large language model provided by an embodiment of the present application Figure 2 ;

[0040] Figure 11 Flow chart of a method for evaluating a large language model provided by an embodiment of the present application Figure 3 ;

[0041] Figure 12 Schematic diagram of the evaluation of a large language model provided by an embodiment of the present application Figure 3 。 DETAILED DESCRIPTION

[0042] In the following, the terms "first" and "second" are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0043] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0044] Before introducing the embodiments of the present application, the technologies involved in the embodiments of the present application are first introduced in detail.

[0045] 1. Large Language Model

[0046] A large language model is a deep learning model that can process and generate complex natural language. The training process of the large language model uses a large amount of text data, so the large language model can generate accurate response text for the input natural language.

[0047] 2. Prompt

[0048] The input text of the large language model is also called a prompt. Prompt can be understood as a way for users to interact with the large language model based on natural language. In related technologies, prompt can be summarized as input text that contains any one of instructions, context, input data, and output indicators.

[0049] In order to better evaluate large language models using a small amount of data during the evaluation process, the prompt in the embodiments of this application can be understood as an input text including two parts: data (such as the above-mentioned input data) and a prompt template (such as the above-mentioned context). Among them, the data can be understood as the questions asked to the large language model. The prompt template can be understood as supplementing some examples or background knowledge before and after the question when asking questions to the large language model. The main purpose of the prompt template is to stimulate the large language model's capabilities in context few-shot learning, zero-shot learning, and chain of thought, and guide the large language model to generate content satisfactory to users for the questions raised. In this way, a large amount of evaluation data can be obtained by combining a small amount of data and a small amount of prompt templates, and the obtained evaluation data can be used to more effectively evaluate the large language model.

[0050] Example 1:

[0051] Data: What is a cat?

[0052] Prompt template: Can you explain () to me?

[0053] The above data and prompt template can form prompt1: Can you explain to me what a cat is?

[0054] Taking this prompt1 as the input text of the large speech model, the large language model can output the result text. The result text output by the large speech model can be called result text1, specifically: A cat belongs to the cat family. A cat is small in size, with body color ranging from blue-gray to brown-yellow, with a slender build, a body length of 0.3 - 0.5 meters, its whole body covered with dense and soft hair, a small collarbone, a short snout, round eyes, a thick neck, relatively short limbs, and several spherical pads under the feet; the tongue surface is covered with filamentous hooked papillae with a horny layer. Male and female cats are similar to each other, only the male's head is thick and round. A cat's reproductive and breeding period lasts throughout its life, and its lifespan is 12 - 17 years.

[0055] Example 2:

[0056] Data: What is a cat?

[0057] Prompt template: If I were a primary school student and you were a teacher, can you explain () to me?

[0058] The above data and prompt template can form prompt2: If I were a primary school student and you were a teacher, can you explain to me what a cat is?

[0059] Use this prompt2 as the input text for the large language model, and the large language model can output the result text. The result text output by the large language model can be called result text2, specifically: Cats are cute animals. They have pointed ears, wet noses, small mouths, and a pair of big bright eyes that shine in the dark.

[0060] In the prompt1 of Example 1, the questioner only provided the question to the large language model and did not inform the large language model of the questioner's identity. Therefore, after processing prompt1, the large language model's returned answer did not consider the questioner's identity and output the relevant information about cats as the result text. It can be seen that the result text1 output by the large language model in Example 1 above describes cats from a biological aspect.

[0061] In the prompt2 of Example 2, the questioner provided the same question to the large language model and also informed the large language model that the questioner's identity is a primary school student. That is to say, prompt1 and prompt2 are composed of different prompt templates. Therefore, after processing prompt2, in the result text2 output by the large language model, the language that is easier for primary school students to understand is used to introduce cats.

[0062] According to the above two examples, when the questioner wants to know about "cats", since the text input by the questioner to the large language model includes different knowledge backgrounds, that is, the prompt template changes, the answer content of the large language model changes.

[0063] The following introduces the large language model evaluation scheme provided by the embodiments of the present application in conjunction with the accompanying drawings.

[0064] The development process of the large language model includes a pre-training stage and a fine-tuning stage. In the pre-training stage, developers need to train the model using a large amount of unlabeled natural language text so that the large language model can learn the general language patterns of natural language, thereby improving the expression ability and generalization ability of the large language model. In the fine-tuning stage, developers need to train the model with carefully designed question-and-answer dialogues (users imitate the dialogue between humans and the large language model) so that the large language model can process the input text more accurately and output higher-quality text. As Figure 1 shown, before the pre-training stage, developers need to evaluate the large language model A so as to formulate corresponding pre-training strategies according to the evaluation results to pre-train the large language model to obtain the large language model B. Before the fine-tuning stage, developers need to evaluate the large language model B so as to formulate corresponding fine-tuning strategies to continue training the model to obtain the large language model C.

[0065] During the pre-training stage or the fine-tuning stage, since the large language model has been trained and its capabilities have changed, the formats of the outputs made by the large language model for the same input are different. Thus, during the pre-training stage or the fine-tuning stage, developers need to customize corresponding test codes for the result texts in different formats output by the large language model to analyze the result texts in different formats output by the large language model at different stages, and then conduct evaluations based on the analysis results.

[0066] In addition, before the large language model goes live, developers may also need to evaluate the large language model to make a final confirmation of the capabilities of the large language model before it goes live. As Figure 1 shown, developers need to evaluate the large language model C before it goes live in order to formulate corresponding pre-live adjustment strategies based on the evaluation results to adjust the large language model C to obtain the large language model D. After obtaining the large language model D, developers may also need to evaluate the large language model D after it goes live to facilitate the timely maintenance or modification of the large language model D by the developers. In the above process, there may be differences between the large speech models A, B, C, and D. Therefore, developers also need to re-customize the test codes to analyze the result texts in different formats output by the large language model at different stages, so as to obtain the evaluation results of the large language model at different stages based on the analysis results.

[0067] Exemplarily, before evaluating the large language model, developers need to first determine the test input set and the annotation set for testing the large language model. Among them, the test input set includes multiple input texts, and these input texts can be used as the input of the large language model. The annotation set includes the annotation texts corresponding to each input text in the test input set. After inputting the input texts in the test input set into the large language model, the large language model can output corresponding response texts, which can be called result texts. By comparing the annotation texts corresponding to each input text in the test input set with the result texts, developers can determine the differences between the annotation texts and the result texts, and thus can evaluate the capabilities of the large language model based on these differences.

[0068] Example 3:

[0069] Input text: Please determine the time and location in the following given text. Text: "Please go to the park this evening."

[0070] Result text: This evening; Park.

[0071] Annotation text: Time: This evening; Location: Park.

[0072] Example 4:

[0073] Input text: Please identify the time and place in the text given below. Text: "Please go to the park this evening.".

[0074] Result text: Evening; park.

[0075] Annotated text: Time: this evening; Location: park.

[0076] In the above example 3, by comparing the result text and the annotated text, it can be found that the large language model can accurately obtain the result text that the user wants based on the input text. Therefore, it can be considered that the ability of the large language model can meet the user's usage needs. In the above example 4, by comparing the result text and the annotated text, it can be found that in the result text obtained by the large language model based on the input text, the time item only has evening, and it does not accurately indicate that it is tonight. Therefore, it can be considered that the ability of the large language model is poor, and developers still need to continue to train the large language model to make it meet the user's usage needs.

[0077] The above example uses an input text to evaluate the ability of a large language model. It is understandable that in order to more accurately evaluate the ability of a large language model, the number of input texts included in the test input set is relatively large, and accordingly, the number of annotated texts in the annotation set is also relatively large. Figure 2 As shown, the developer inputs multiple input texts in the test input set into the large language model so that the large language model can output the corresponding result texts. The multiple result texts outputted can form a result set. Due to the large number of texts, manually comparing each result text in the result set with the corresponding annotated text in the annotation set requires a lot of manpower. Therefore, the developer needs to write appropriate test code for the result set including multiple result texts and the corresponding annotated texts, so as to use the test code to compare the annotated text in the annotation set with the result text in the result set, analyze the difference between the two, and thus evaluate the ability of the large language model based on the comparison results.

[0078] However, in the case of large language model changes, Figure 2 The input text in the test input set shown is input into the changed large language model. Due to the change in the language processing capability of the changed large language model, the changed large language model outputs a new result set based on the test input set. Figure 2 That is, the format of the result text in the new result set is different from the format of the annotation text, which causes the developer to Figure 2 The test code written in cannot compare the result text in the new result set with the annotated text in the annotation set, so it is impossible to use this test code to evaluate the capabilities of the changed large language model.

[0079] To be able to evaluate the changed large language model, developers need to rewrite the test code according to the new result set and annotation set. Moreover, every time the large language model changes, developers need to rewrite the test code. Thus, the efficiency of evaluating the large language model is reduced.

[0080] Therefore, an embodiment of the present application provides a method for evaluating a large language model. As Figure 3 shown, developers can specify the output format of the large language model. In this way, when inputting the input text in the test input set into the large language model, the large language model can output the result text in the specified format. Since the format of the result text is specified, the server can execute the test code written according to the specified format to perform a differential comparison between the result text output by the large language model and the annotation text, so as to evaluate the ability of the large language model. Moreover, even if the large language model changes and the actual content of the changed result text also changes, since the format of the result text is specified, the server can still use the test code written according to the specified format to perform a differential comparison between the result text output by the changed large language model and the annotation text, so as to realize the evaluation of the ability of the changed large language model. Thus, it is not necessary for developers to write different test codes for different large language models for evaluation, improving the efficiency of large language model evaluation.

[0081] The above description is given by taking the input format of the large language model being specified by the developer as an example. In some other examples, the output format of the large language model can also be a pre-specified default format.

[0082] The method for evaluating a large language model provided by the embodiment of the present application can be applied to a model evaluation system.

[0083] In some embodiments, as Figure 4 shown, the model evaluation system can include at least one electronic device 100 and a server 200. The large language model can be deployed in the server 200. Users can send a specified instruction for the output format to the large language model in the server 200 through the electronic device 100, so that the large language model can output the result text in the specified output format. In addition, users can also input the input text to the large language model in the server 200 through the electronic device 100, so that the large language model outputs the corresponding result text according to the input text. Of course, the format of the result text is the specified output format. Subsequently, the evaluation result can be obtained by analyzing the difference between the result text in the specified output format and the corresponding annotation text. Then, the server 200 sends the evaluation result to the electronic device 100, so that the electronic device 100 can display the evaluation result to the user.

[0084] In some examples, the user can send an evaluation instruction to the large language model in the server 200 through the electronic device 100. The evaluation instruction includes the above instruction for specifying the output format of the large language model, multiple input texts, and multiple corresponding annotation texts. The server 200 can obtain a test input text based on each input text in the evaluation instruction and the instruction for specifying the output format, so that the large language model in the server 200 can process the test input text to obtain a result text in the specified output format. After that, the server 200 can use a general test code to analyze the difference between the result text in the specified output format and the annotation text, and obtain an evaluation result based on the analysis result. The server 200 can send the evaluation result to the electronic device 100 for the electronic device 100 to display to the user for viewing.

[0085] In some examples, the electronic device in the model evaluation system can specifically be an electronic device such as a mobile phone, a tablet computer, a smart screen, a laptop computer, a vehicle-mounted device, a wearable device (such as a smart watch), an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an artificial intelligence device, etc. The embodiments of the present application do not limit the specific type of the electronic device and the installed operating system.

[0086] In some examples, the server in the model evaluation system can be a device with natural language processing functions such as a cloud server or a network server. The above server can be a single server, a server cluster composed of multiple servers, or a cloud computing service center.

[0087] Next, the hardware structure of the server will be introduced.

[0088] As Figure 5 shown, the server 200 can include a processor 210, a memory 220, and a communication module 230.

[0089] The processor 210 can be used to read and execute computer-readable instructions. Specifically, the processor 210 may include a controller, an arithmetic unit, and registers. Among them, the controller is mainly responsible for instruction decoding and sending control signals for the operations corresponding to the instructions. The arithmetic unit is mainly responsible for storing register operands and intermediate operation results temporarily stored during the instruction execution process, etc. In a specific implementation, the hardware architecture of the processor 210 can be an application specific integrated circuit (ASIC) architecture, a MIPS (microprocessor without interlocked piped stages) architecture, an ARM (advanced risc machines) architecture, or a network processor (NP) architecture, etc.

[0090] The memory 220 is coupled to the processor 210 and is used to store various software programs and / or multiple sets of instructions. In a specific implementation, the memory 220 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 220 can store an operating system, such as embedded operating systems like uCOS, VxWorks, RTLinux, etc.

[0091] The communication module 230 can be used to establish a communication connection between the cloud server 100 and other communication terminals (such as Figure 4 multiple electronic devices 100 therein) through a network, and is used to send and receive data through the network.

[0092] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the server 200. In other embodiments, the server 200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0093] Figure 6 This is a flowchart illustration of a large language model evaluation method provided by an embodiment of the present application. Figure 1 Before the large language model is trained and not yet launched, or after the large language model is launched, the server can implement the evaluation of the large language model through the steps as Figure 6 shown. The large language model evaluation method includes the following steps:

[0094] S601: The server obtains instruction 1 for specifying the output format of the large language model.

[0095] Instruction 1 (i.e., the aforementioned "first instruction") is used to specify the output format of the large language model (i.e., the aforementioned "first format"). After the server obtains Instruction 1, it can specify the output format of the large language model through Instruction 1. This Instruction 1 can also be referred to as a general instruction, etc.

[0096] In some embodiments, the specified output format can be any one of the JavaScript Object Notation (JSON) format, the Extensible Markup Language (XML) format, and the YAML Ain't Markup Language (YAML) format. Among them, the JSON format is a sequence of tokens, which can include six construction characters and values. Among them, the six construction characters include the left square bracket ([), the left curly bracket ({), the right square bracket (]), the right curly bracket (}), the colon (:), and the comma (,), and the value can be an object, an array, a number, a string, or one of the three literals (false, null, true).

[0097] Specifically, an object is enclosed by curly brackets and consists of members separated by commas. The members are key-value pairs composed of strings and the values described above, separated by commas. An array consists of a group of values enclosed by square brackets. A string is a collection of any number of Unicode characters enclosed by double quotes ("").

[0098] For the explanations of the above different formats, reference can be made to the related technologies, and they will not be elaborated in this embodiment.

[0099] In some embodiments, the output format specified by Instruction 1 for the large speech model can be the default format built into the server (the "built-in default format" can also be described as "pre-configured format"), or it can be a user-defined output format (the "user-defined output format" can also be described as "customized format").

[0100] Exemplarily, Figure 7 A schematic diagram of an operation display interface for evaluating a large language model is shown. The operation display interface displayed on the electronic device is used for the user to configure the model evaluation parameters. For example, the operation display interface includes multiple configuration items, including a model output format configuration item for the user to configure the output format of the large language model. For example Figure 7In the operation display interface shown, there are two options corresponding to the model output format configuration item. One option is the default format, and the other option is the custom format. When the user selects the model output format configuration item as the default format, the electronic device can send the default format to the server as Instruction 1. The server can obtain Instruction 1. At this time, Instruction 1 is used to specify that the large language model uses the default format as the output format. When the user selects the model output format configuration item as the custom format, the user can also input a custom output format in the electronic device. At this time, the electronic device sends the output format input by the user to the server as Instruction 1. The server can obtain Instruction 1, and this Instruction 1 is used to specify that the large language model uses the user-defined output format as the output format. For example, in response to the user selecting Figure 7 the control 701 corresponding to the model output format configuration item in Figure 7 , and the format input in the input box 702, such as the JSON format, the electronic device can send the JSON format to the server as Instruction 1. The server can obtain that the output format of the specified large language model is the user-defined JSON format.

[0101] S602: The server obtains multiple test input texts based on Instruction 1 and multiple input texts.

[0102] It should be noted that the test input text is the aforementioned "first test data", and the form of the first test data can be a type that the text large language model can process. For the sake of easy understanding, the embodiments of the present application will hereinafter take the first test data as the text type as an example to introduce the large language model evaluation scheme.

[0103] In some embodiments, the server can obtain the output format of the large language model based on Instruction 1. After that, the server can merge the input text and the output format into a test input text. It can be understood that the server can merge Instruction 1 and the input text in any way, and the embodiments of the present application do not limit this. Among them, as described in the foregoing embodiments, in some examples, the input text can include two parts: a prompt template and data.

[0104] In some examples, the server can obtain a format specification text for specifying the output format of the large language model based on the output format specified by Instruction 1. The form of the format specification text can be a natural language text or a format example. After that, the server can merge the obtained format specification text with the input text to obtain a test input text.

[0105] For example, when Instruction 1 is used to specify that the output format of the large language model is the JSON format, the server can obtain that the output format of the specified large language model is the JSON format based on Instruction 1. After that, the server can obtain the corresponding natural language text based on this, such as: Please output in the JSON format, which is the above format specification text (i.e., the aforementioned "first format information") for specifying the output format of the large language model.

[0106] Taking the input text: Please determine the time and location in the following given text. Text: "Please go to the park this evening." as an example. It can be understood that in the input text: Please determine the time and location in the following given text. Text: "" is the prompt template; Please go to the park this evening is the data.

[0107] Then, the test input text obtained by the server can be: Please determine the time and location in the following given text. Text: "Please go to the park this evening." Please output in the JSON format.

[0108] Or, the test input text obtained by the server can be: Please output in the JSON format. Please determine the time and location in the following given text. Text: "Please go to the park this evening.".

[0109] Another example is that when Instruction 1 is used to specify that the output format of the large language model is the JSON format, the server can obtain that the output format of the specified large language model is the JSON format based on Instruction 1. After that, the server can obtain the corresponding format example based on this, such as: Please output in the JSON format, and the output format example is {"time": "tomorrow morning", "location": "school"}, which is the above format specification text for specifying the output format of the large language model.

[0110] Taking the input text: Please determine the time and location in the following given text. Text: "Please go to the park this evening." as an example.

[0111] Then, the test input text obtained by the server can be: Please determine the time and location in the following given text. Text: "Please go to the park this evening." Please output in the JSON format, and the output format example is {"time": "tomorrow morning", "location": "school"}.

[0112] Or, the test input text obtained by the server can be: Please output in the JSON format, and the output format example is {"time": "tomorrow morning", "location": "school"}. Please determine the time and location in the following given text. Text: "Please go to the park this evening.".

[0113] In some embodiments, to evaluate the capabilities of a large language model, the server needs to input multiple input texts into the large language model separately and obtain corresponding result texts. After that, the server analyzes the differences between each result text and the corresponding annotated text of the input text, and then can obtain the analysis results of multiple result texts, so as to calculate metrics that can be used to evaluate the capabilities of the large language model based on the analysis results. As mentioned above, the server can obtain a large number of input texts based on a small amount of data and a small number of prompt templates combined.

[0114] Exemplarily, the data (i.e., the aforementioned "first data") included in the data set (i.e., the aforementioned "first data set") are as follows:

[0115] Please go to the park tonight, please go to the park tomorrow morning, please go to the school tonight, please go to the school tomorrow morning.

[0116] The prompt templates (i.e., the aforementioned "first prompt templates") included in the prompt template set (i.e., the aforementioned "first prompt template set") are as follows:

[0117] Please identify the time and location in the following given text, please point out the time and location in the following text.

[0118] Then, the server can obtain the following eight input texts:

[0119] Input text 1: Please identify the time and location in the following given text. Text: "Please go to the park tonight.".

[0120] Input text 2: Please identify the time and location in the following given text. Text: "Please go to the park tomorrow morning.".

[0121] Input text 3: Please identify the time and location in the following given text. Text: "Please go to the school tonight.".

[0122] Input text 4: Please identify the time and location in the following given text. Text: "Please go to the school tomorrow morning.".

[0123] Input text 5: Please point out the time and location in the following text. Text: "Please go to the park tonight.".

[0124] Input text 6: Please point out the time and location in the following text. Text: "Please go to the park tomorrow morning.".

[0125] Input text 7: Please point out the time and location in the following text. Text: "Please go to the school tonight.".

[0126] Input Text 8: Please indicate the time and location in the following text. Text: "Please go to school tomorrow morning."

[0127] In some embodiments, as Figure 7 shown in the operation display interface, it may further include: a dataset configuration item and a prompt template configuration item. Among them, the dataset configuration item is used for the user to configure the data in the input text of the large language model, and the prompt template configuration item is used for the user to configure the prompt template in the input text of the large language model. As Figure 7 shown in the operation display interface, when the user selects the dataset configuration item as Dataset A and the prompt template configuration item as Prompt Template Set A, the electronic device can send Dataset A and Prompt Template Set A to the server. Therefore, the server can obtain Dataset A and Prompt Template Set A. After that, the server can obtain multiple input texts based on the data in Dataset A and the prompt templates in Prompt Template Set A.

[0128] For example, as shown in Figure 8 after the server obtains Instruction 1, it can obtain a test input text based on an input text and Instruction 1. When there are multiple input texts, the server can obtain a test input set A including multiple test input texts, that is, the aforementioned "first test set".

[0129] S603: The server inputs multiple test input texts into the large language model to be evaluated respectively, and obtains multiple result texts output by the large language model to be evaluated in a specified format.

[0130] As in the example in S602, for the scenario where there are multiple input texts, the server can obtain a test input set A including multiple test input texts. After the server obtains the test input set A, it can input the multiple test input texts included in the test input set A into the large language model to be evaluated respectively. After that, the large language model to be evaluated can output the result texts corresponding to each test input text in the test input set A in a specified format. The server can then obtain the result texts corresponding to each test input text in the test input set A output in a specified format (i.e., the aforementioned "first result data") to obtain a result set (i.e., the aforementioned "first result set").

[0131] Continuing with the example in S602 above, the test input text is: Please determine the time and location in the following given text. Text: "Please go to the park tonight." Please output in JSON format, and the output format example is {"time": "tonight", "location": "park"}.

[0132] Then, after the server inputs the above test input text into the evaluation large language model, the result text in the specified format obtained can be: {"time": "this evening", "location": "park"}.

[0133] For another example, the test input text is: Please determine the time and location in the following given text. Text: "Please go to school tomorrow morning." Please output in JSON format.

[0134] Then, after the server inputs the above test input text into the evaluation large language model, the result text in the specified format obtained can be: {"time": "tomorrow morning", "location": "school"}.

[0135] It can be understood that after the test input text is input into the large language model, the large language model can obtain the result text in the corresponding specified format based on the input test input text. Then, as Figure 8 shown, by inputting multiple test input texts in the test input set A into the large language model respectively, the result text in the specified format corresponding to each test input text can be obtained, that is, the server can obtain a result set including multiple result texts in the specified format.

[0136] S604: The server analyzes each result text in the specified format and the annotation text corresponding to each result text to obtain multiple analysis results.

[0137] As Figure 8 shown, the server analyzes each result text in the specified format in the result set and the annotation text corresponding to each result text in the specified format, and multiple analysis results can be obtained.

[0138] The annotation text is the correct result that the user expects the large language model to output by taking the test input text as the input of the large language model. Generally, the annotation text can be manually annotated by the annotator after analyzing the characteristics of the input text, and the format of the annotation text is the same as the output format specified in the above instruction 1. In this way, after the server inputs the test input text into the large language model and the large language model outputs the result text in the specified format, the server can compare the result text in the specified format and the corresponding annotation text to obtain the analysis result. Then, the server determines the language analysis ability of the large language model based on the analysis result, that is, completes the evaluation of the large language model.

[0139] In some embodiments, since the format of the result text in the specified format is the same as the format of the annotated text, the server can execute the test code written by the developer to analyze the differences between multiple result texts in the specified format and multiple annotated texts. It should be noted that since the format of the result text in the specified format is the same as the format of the annotated text, and the test code written by the developer is used to compare the result text and the annotated text, the test code written by the developer is a general-purpose code and can be used to analyze the differences between result texts in other formats and the corresponding annotated texts. For example, the server can analyze the result text in JSON format and the annotated text in JSON format through the test code, and can also analyze the result text in XML format and the annotated text in XML format.

[0140] For example, the result text is: {"time": "This evening", "location": "Park"}.

[0141] The annotated text is: {"time": "This evening", "location": "Park"}.

[0142] Then, the server can execute the pre-written test code to compare the differences between the two and obtain the comparison result. For example, during the execution of the test code, the server can extract the value of "time" as "This evening" and the value of "location" as "Park" from the result text. Also, the server can extract the value of "time" as "This evening" and the value of "location" as "Park" from the annotated text. After comparison, the server can obtain a comparison result indicating that the result text is the same as the annotated text, indicating that the result text is correct.

[0143] Another example, the result text is: {"time": "Evening", "location": "Park"}.

[0144] The annotated text is: {"time": "This evening", "location": "Park"}.

[0145] Then, during the execution of the test code, the server can extract the value of "time" as "Evening" and the value of "location" as "Park" from the result text. Also, the server can extract the value of "time" as "This evening" and the value of "location" as "Park" from the annotated text. After comparison, the server can obtain a comparison result indicating that the result text is different from the annotated text, indicating that the result text is incorrect. It can be understood that the comparison result in the above example is the analysis result.

[0146] S605: The server obtains the evaluation result (i.e., the aforementioned "first evaluation result") based on multiple analysis results.

[0147] In some embodiments, as Figure 8 shown, the test input set A includes multiple test input texts, and then the large language model correspondingly outputs multiple result texts in a specified format. The server can analyze each result text in the specified format in the result set to obtain an analysis result that includes the comparison results between multiple result texts in the specified format and the corresponding labeled texts. After that, the server can calculate an index for evaluating the large language model based on the multiple comparison results in the analysis result, that is, the evaluation result.

[0148] In some examples, the server can send the evaluation result to the electronic device, so that the electronic device can display the evaluation result to the user in any one or more forms such as graphics, data, text, etc.

[0149] In some examples, the indexes for evaluating the large language model can include one or more of accuracy, precision, recall, F1 parameter (F1 parameter = 2 * (precision * recall) / (precision + recall), and the F1 parameter is the weighted harmonic mean of precision and recall), etc. For the above evaluation indexes, specific information can be referred to in the relevant technologies and will not be elaborated here.

[0150] Optionally, the specific evaluation index can be set according to the actual needs of the user. As Figure 7 shown, the operation display interface can also include an evaluation index configuration item. When the user selects the evaluation index configuration item as index A, the electronic device can send index A to the server to instruct the server to calculate index A based on the analysis result.

[0151] Exemplarily, when the evaluation index is accuracy, if the test input set A includes N test input texts, then the large language model correspondingly outputs N result texts. If the server finds that M of the N result texts are incorrect when comparing the N result texts with the corresponding labeled texts, where M ≤ N, then the server can calculate the evaluation index accuracy = M / N.

[0152] In some embodiments, after the server analyzes the result text in the specified format and the labeled text to obtain the analysis result, it can compare the obtained evaluation result this time with the preset evaluation result or the historical evaluation result to obtain the change in the ability of the current large language model, so as to determine whether to continue training the model based on the change in ability. Or, the server can compare the evaluation result obtained by evaluating the large language model this time with the evaluation results of other large language models to obtain the large language model with more outstanding ability among multiple large language models.

[0153] In some examples, after the pre-training stage, the large language model version V1 is obtained. To test the information extraction ability of the large language model version V1, the server can obtain multiple test input texts for information extraction, and the test input texts include information specifying the output format of the large language model. Then, by inputting the obtained multiple test input texts into the large language model version V1, so that it outputs a result text in the specified format, that is, extracting the information in the test input text and outputting a result text including the extraction result in the specified format. In this way, the server can obtain a result set A including multiple result texts in the specified format. The server can use general test code to analyze each result text in the result set A and the annotation text corresponding to each result text in the annotation set to obtain an analysis result. Then, the server can obtain corresponding metrics based on the analysis result. Based on the obtained metrics, the server can analyze whether the information extraction ability of the large language model version V1 meets the standard. For example, the server inputs 10 test input texts into the large language model version V1 to make the large language model version V1 output 10 result texts in the specified format. Then, after the server uses the test code to compare the 10 result texts and the 10 annotation texts, it is determined that 5 of the 10 result texts included in the result set A are correct result texts. Then, it can be determined that the information extraction accuracy rate of the large language model version V1 is 5 / 10 = 50%. If the information extraction accuracy rate of the large language model above 70% is regarded as the information extraction ability meeting the standard, then, in the above example, the information extraction accuracy rate of the large language model version V1 does not reach above 70%, and the server can consider that the information extraction ability of the large language model version V1 does not meet the standard.

[0154] Correspondingly, in the case where the server considers that the information extraction ability of the large language model version V1 does not meet the standard, the server can obtain more training texts that can enhance the information extraction ability of the large language model and use these training texts to continue training the large language model version V1 to improve the information extraction ability of the large language model version V1.

[0155] Continuing with the above example, after the developer fine-tunes the large language model version V1, the large language model version V2 is obtained. Similarly, after the server inputs the above multiple test input texts for information extraction into the large language model version V2, the result set B can be obtained. If the server analyzes and finds that there are 10 result texts in the result set B, and 8 of them are correct result texts, then it can be determined that the information extraction accuracy rate of the large language model version V2 is 80%. If the information extraction accuracy rate of the large language model above 70% is regarded as the information extraction ability meeting the standard, then, in the above example, the information extraction accuracy rate of the large language model version V2 reaches above 70%, and the server can consider that the information extraction ability of the large language model version V2 meets the standard.

[0156] In some other examples, the server evaluates different versions of the large language model and obtains the corresponding evaluation metrics. If the server determines that the evaluation metrics of the previous version of the large language model are better than those of the later version of the large language model, it indicates that the ability of the previous version of the large language model is stronger than that of the later version of the large language model. In this case, it shows that the ability of the previous version of the large language model has decreased instead after training, which may be caused by problems in the training process. For example, problems in the training text used for training lead to a decrease in the ability of the large language model. That is to say, according to the evaluation results, it can also be determined whether there are problems with the training text for training the large language model and whether adjustment is needed.

[0157] It should be noted that when evaluating different versions of the large language model, the same test code is used, or in other words, a general test code is used. This is because although the capabilities of different versions of the large language model have changed and the content of the output result text may also have changed, they all output the result text in a specified output format. Therefore, the same test code can be used for evaluation.

[0158] In some other examples, after the pre-training stage, the large language model version V1 is obtained. To test the sentiment analysis ability of the large language model version V1, the server can obtain multiple test input texts for sentiment analysis, and the test input texts include information specifying the output format of the large language model. Then, by inputting the obtained multiple test input texts into the large language model version V1, it is enabled to output result texts in the specified format, that is, perform sentiment analysis on the information in the test input texts and output result texts including the sentiment analysis results in the specified format. In this way, the server can obtain result texts including multiple specified formats, that is, obtain result set C. The server can use the test code used to test the information extraction ability of the large language model in the above examples, that is, the general test code to analyze each result text in result set C and the annotation text corresponding to each result text in the annotation set to obtain the analysis result. Then, the server can obtain the corresponding metrics based on the analysis result. Based on the obtained metrics, the server can analyze whether the sentiment analysis ability of the large language model version V1 meets the standard. For example, the server inputs 10 test input texts into the large language model version V1 to enable the large language model version V1 to output 10 result texts in the specified format. Then, after the server uses the test code to compare the 10 result texts and the 10 annotation texts, it is obtained that 8 of the 10 result texts included in result set C are correct result texts. Then, it can be determined that the sentiment analysis accuracy rate of the large language model version V1 is 8 / 10 = 80%. If the sentiment analysis accuracy rate of the large language model above 70% is regarded as the sentiment analysis ability meeting the standard, then, in the above example, the sentiment analysis accuracy rate of the large language model version V1 reaches above 70%, and the server can consider that the sentiment analysis ability of the large language model version V1 meets the standard.

[0159] It can be understood that for different abilities of the same large language model, or the same ability of different large language models (such as different versions of the large language model, or large language models with different functions), due to specifying the output format of the large language model, therefore, the server can use the same test code to analyze the result texts and annotation texts with the same format to obtain the analysis result, so that the server can calculate the corresponding evaluation metrics based on the analysis result. In this way, the evaluation efficiency can be effectively improved.

[0160] In some embodiments, the server can use different types of test input texts to achieve the purpose of evaluating large language models in different fields. The server obtains test input texts based on the input text and instruction 1, and the input text includes data and a prompt template. When the type of the data changes, the type of the test input text also changes accordingly. For example, Figure 7In the operation display interface shown, the drop-down options of the data set configuration item may include information extraction capabilities and sentiment analysis capabilities corresponding to the above-mentioned embodiments. In addition, the capabilities of the large language model can be described in other descriptions. For example, the capabilities of the large language model can be described as factual question-answering capabilities, reading comprehension capabilities, framework generation capabilities, paragraph rewriting capabilities, summary extraction capabilities, mathematical problem-solving capabilities, reasoning capabilities, poetry generation capabilities, or programming field capabilities, etc., then correspondingly, the drop-down options of the data set configuration item may include different types of data sets in the fields of factual question-answering, reading comprehension, framework generation, paragraph rewriting, summary, mathematical problem-solving, reasoning, poetry generation, and programming, then the user can select the corresponding data set at the data set configuration item in the operation display interface according to the actual evaluation requirements, so as to realize the evaluation of the specified capabilities of the large language model. For example, if the data set configuration item is a data set of the reading comprehension type, then the type of the test input text obtained by the server is reading comprehension, then the reading comprehension ability of the large language model can be evaluated using this type of test input text.

[0161] It should be noted that developers can set the type of data set according to the actual application scenario of the large language model to evaluate the capabilities that the large language model focuses on in the application scenario. For example, the large language model used in the psychological counseling scenario focuses more on sentiment analysis capabilities. In this case, when evaluating the large language model, a sentiment analysis type data set can be selected. The large language model used in the data statistics and collation scenario focuses more on information extraction capabilities. In this case, when evaluating the large language model, an information extraction type data set can be selected.

[0162] Figure 9 A schematic diagram of a large language model evaluation method provided in an embodiment of the present application Figure 2 .

[0163] Before the large language model is trained and launched, or after the large speech model is launched, the server can Figure 9 The steps shown implement the evaluation of the large language model. The large language model evaluation method includes the following steps:

[0164] S901: The server inputs the input text into the large language model Mt to obtain the result text output by the large language model Mt.

[0165] The input text may be obtained based on a data set and a prompt template. For details, please refer to the description of the input text in the above embodiment, which will not be repeated here.

[0166] like Figure 10As shown, the server inputs each input text (i.e., the aforementioned "second test data") in the test input set (i.e., the aforementioned "third test set") into the large language model Mt, and can obtain multiple result texts (i.e., the aforementioned "second result data") output by the large language model Mt, that is, obtain the result set (i.e., the aforementioned "second result set").

[0167] After the input text is input into the large language model Mt, the large language model Mt can process the input text and obtain the corresponding result text. Since the output format of the large language model Mt is not specified, therefore, the large language model Mt actually generates the result text according to its logic of processing natural language, and the generated result text is not in a fixed format.

[0168] For example, input text 1:

[0169] Please identify the time and location in the following given text. Text: "Please go to the park this evening.".

[0170] Result text 1 corresponding to input text 1:

[0171] The time is this evening and the location is the park.

[0172] Input text 2:

[0173] Please point out the time and location in the following given text. Text: "Please go to school tomorrow morning.".

[0174] Result text 2 corresponding to input text 2:

[0175] {"time": "tomorrow morning", "location": "school"}

[0176] It can be understood that in the above example, two different input texts 1 and input text 2 are respectively input for the same large language model. Then, when the large language model analyzes different input texts, due to the difference between the action of "identifying" in input text 1 and the action of "pointing out" in input text 2, the large language model may think that different outputs need to be made for these two different actions. As in the above example, the large language model may think that the result text corresponding to the action of "identifying" in input text 1 needs to be expressed in natural language text. Therefore, the large language model outputs natural language as shown in result text 1. In addition, the large language model may think that the result text corresponding to the action of "pointing out" in input text 2 needs to be expressed in JSON format. Therefore, the large language model outputs the content in JSON format as shown in result text 2.

[0177] It should be noted that for the two relatively similar input texts in the above example, when the information extraction ability of the large language model is strong, the probability of content and / or format differences between the result text 1 and the result text 2 output by the large language model is relatively low. However, when the information extraction ability of the large language model is weak, the probability of format differences between the result text 1 and the result text 2 output by the large language model is relatively high. Thus, the server can obtain multiple result texts by inputting multiple input texts into the large language model, and then analyze the information extraction ability of the large language model based on the multiple result texts.

[0178] The formats of the above two result texts are different. Among them, the content of result text 1 is text content, and the server cannot use code to parse the keywords in result text 1, that is, the server cannot analyze result text 1. The format of the above result text 2 is JSON format, so the server can analyze result text 2.

[0179] Since the server can only analyze some of the result texts output by the large language model Mt and cannot analyze the other part of the result texts, the server cannot use the result texts to evaluate the ability of the large language model Mt. For this reason, the server can perform the following steps (S902 - S905):

[0180] S902: The server obtains instruction 2 for instructing the analysis of the large language model Ma to perform an analysis operation and specifying the output format of the large language model Ma.

[0181] Instruction 2 is used to instruct the analysis of the large language model Ma to perform an analysis operation and to instruct the large language model Ma to output the analysis result in the specified output format.

[0182] After the server obtains instruction 2, it can make the analysis large language model Ma perform an analysis operation through the following steps.

[0183] S903: The server combines instruction 2, the result text, and the input text to obtain an analysis text.

[0184] It should be noted that the specific implementation of the server combining instruction 2, the result text, and the input text can refer to Figure 6 the specific implementation of the corresponding content in the illustrated embodiment, which will not be elaborated here in detail.

[0185] Such as Figure 10As shown, based on the result set (i.e., the aforementioned "second result set") including multiple result texts (i.e., the aforementioned "second result data"), the test input set (i.e., the aforementioned "third test set") including multiple input texts (i.e., the aforementioned "second test data"), and instruction 2 (i.e., the aforementioned "second format information"), the server can obtain the analysis input set (i.e., the aforementioned "fourth test set") including multiple analysis texts (i.e., the aforementioned "third test data").

[0186] S904: The server inputs the analysis text into the analysis large language model Ma to obtain the analysis result output by the analysis large language model Ma.

[0187] In some embodiments, the server can obtain the analysis instruction for instructing the analysis large language model Ma to perform the analysis operation and the output format for instructing the output of the analysis large language model Ma based on instruction 2. The server can merge the analysis instruction, the output format, the input text of the input large language model Mt, and the result text output by the large language model Mt into the analysis text. It can be understood that the server can merge the analysis instruction, the output format, the input text of the input large language model Mt, and the result text output by the large language model Mt in any way, and the embodiments of the present application do not limit this.

[0188] As Figure 10 shown, the server inputs the merged analysis text (i.e., the aforementioned "third test data") into the analysis large language model Ma (i.e., the aforementioned "second large language model"), and the large language model Ma can output the analysis result (i.e., the aforementioned "third result data"), and the server obtains this analysis result.

[0189] For example, the input text of the input large language model Mt can be: Please extract the time and location in the text. Text: I want to eat hot pot at Haidilao today.

[0190] The result text output by the large language model Mt can be: The time is today, and the location is Haidilao.

[0191] Instruction 2 can be: The following is the answer of a certain model for the input. Please analyze the correctness of the result text. Correct is 1, and wrong is 0. Return in JSON format. For example: {"response": "1"}.

[0192] In the example of Instruction 2 above, it can be seen that in Instruction 2, the statement "The following is the response of a certain model to the input. Please analyze the correctness of the result text" indicates that the large language model Ma needs to analyze the correctness of the input text and the corresponding output text of a certain model (i.e., the large language model Mt to be evaluated), and the statement in Instruction 2 "1 for correct and 0 for incorrect. Return in JSON format. For example: {"response": "1"}" indicates that the large language model Ma needs to give the analysis result in the specified format after analysis. In Instruction 2, the specified format is JSON format.

[0193] The analysis text can be: The following is the response of a certain model to the input. Please analyze the correctness of the result. 1 for correct and 0 for incorrect. {Model input: Please extract the time and location from the text. Text: I want to have hot pot at Haidilao today. Model output: The time is today and the location is Haidilao}. Please return in JSON format, for example: {"response": "1"}.

[0194] After the above analysis text is input into the large language model Ma, the analysis result output by the large language model Ma is: {"response": "1"}.

[0195] The above analysis result indicates that the large language model Ma has analyzed and found that the result text returned by the large language model Mt for the input text is correct.

[0196] For another example, the input text of the large language model Mt can be: Please give the characteristics of a cat.

[0197] The result text output by the large language model Mt can be: Cats are cute animals. They have pointed ears, wet noses, small mouths, and a pair of big bright eyes that shine in the dark.

[0198] Instruction 2 can be: The following is the response of a certain model to the input. Please evaluate the result text according to the marked text and give a score. Return in JSON format, where the evaluation result indicates whether it is easy to understand, and score represents the scoring result, and the score range is 0 - 100. For example: {"Evaluation result": "Not easy to understand", "score": "60"}.

[0199] Based on the input text of the large language model Mt, the result text output by the large language model Mt, and Instruction 2, the server can obtain the analysis text as follows: The following is the answer of a certain model to the input. Please evaluate the result text and give a score. Return it in JSON format, where the evaluation result indicates whether it is easy to understand, and "score" represents the scoring result, and the score range is 0 - 100. For example: {"Evaluation result": "Easy to understand", "score": "80"}. {Model input: Please give the characteristics of a cat. Model output: A cat is a cute animal. They have pointed ears, a wet nose, a small mouth, and a pair of big bright eyes that shine in the dark}.

[0200] After the above analysis text is input into the analysis large language model Ma, the analysis result output by the analysis large language model Ma is: {"Evaluation result": "Easy to understand", "score": "80"}.

[0201] The above analysis result indicates that the analysis large language model Ma, through analysis, finds that the result text returned by the large language model Mt for the input text is easy to understand, and the degree of ease of understanding of the result text returned by the large language model Mt for the input text is scored 80 points.

[0202] For example, as combined with Figure 10 shown, after the server obtains Instruction 2, it can obtain the analysis text based on an input text input to the large language model Mt, the result text output by the large language model Mt based on this input text, and Instruction 2. When there are multiple input texts to the large language model Mt, there are also multiple result texts output by the large language model Mt. The server can obtain an analysis input set including multiple analysis texts. The server inputs each analysis text in the analysis input set into the large language model Ma respectively, and can obtain multiple analysis results output by the large language model Ma.

[0203] S905: The server obtains the evaluation result based on the analysis result.

[0204] As in the example in S904, for the scenario where the number of input texts for testing the large language model Mt is multiple, the number of result texts output by the large language model Mt is also multiple. Therefore, the number of analysis texts obtained by the server by combining Instruction 2 with the result texts and the input texts is also multiple. Then, for the multiple analysis texts input to the analysis large language model Ma, multiple analysis results can be obtained respectively. Thus, as Figure 10 shown, the server can obtain the evaluation result (i.e., the aforementioned "second evaluation result") based on multiple analysis results (i.e., the aforementioned "third result data", and multiple analysis results can form the aforementioned "third result set").

[0205] In the above embodiments, when the analysis result is used to represent the correctness of the result text output by the large language model Mt for the input text, the server can determine, from multiple analysis results, the number of correct result texts output by the large language model Mt for the input text and the number of incorrect result texts. In this way, the server can calculate the accuracy rate of the output of the large language model Mt. It can be understood that the evaluation result can also be other metrics, such as recall rate, F1 parameter, etc. In addition, the evaluation result can also be other metrics for representing the ability of the large language model Mt, such as the easy-to-understand rate representing the difficulty of understanding the output result of the large language model Mt.

[0206] Exemplarily, when the number of input texts for testing the large language model Mt is multiple, if 10 analysis results are obtained by analyzing multiple analysis texts input to the large language model Ma. And among the 10 analysis results, 1 analysis result indicates that the result text is easy to understand and the score is 90 or above, 6 analysis results indicate that the result text is easy to understand and the score is below 90 and 80 or above, 1 analysis result indicates that the result text is easy to understand and the score is below 80 and 70 or above, and 2 analysis results indicate that the result text is not easy to understand and the score is below 70. Then, based on the above analysis results, the server can obtain the following evaluation results. For example, the easy-to-understand rate of the result text output by the large language model Mt is (1 + 6 + 1) / 10 = 80%. Another example is that among the result texts that are easy to understand output by the large language model Mt, the probability of the result text with a higher score (score of 90 or above) appearing is 1 / (1 + 6 + 1) = 12.5%.

[0207] In the above embodiments, when the analysis result is used to represent the easy-to-understand degree of the result text output by the large language model Mt for the input text, the server can determine, from multiple analysis results, the number of result texts that are easy to understand output by the large language model Mt for the input text and the number of result texts that are not easy to understand. In this way, the server can calculate the probability of the result text that is easy to understand output by the large language model Mt as the evaluation result.

[0208] It should be noted that for different analysis results, the server can correspondingly obtain different evaluation results based on the analysis results. The specific type of the evaluation result is not limited in this embodiment.

[0209] It can be understood that for the convenience of calculation, the number of texts given in the above examples is small. In actual application scenarios, the number of input texts is not limited by the above examples.

[0210] Figure 11 The flow of a large language model evaluation method provided by an embodiment of the present applicationFigure 3 。

[0211] After the user sets the output format of the large language model Mt to the specified format, the user can click on the test control as shown in Figure 7 to test whether the large language model Mt can understand the specified format set by the user.

[0212] In response to the user clicking on the test control, the server can receive a test instruction and start executing the following steps (S1101 - S1103 as shown in Figure 11 ).

[0213] S1101: The server obtains the specified format and the format test input text.

[0214] Among them, the specified format (i.e., the aforementioned "first format") can be user - defined or the default format built into the server. For example, as shown in Figure 7 , when the default format corresponding to the model output format configuration item is in the selected state, it indicates that the user expects the large language model to use the default format as the output format. Then, the electronic device can send the default format to the server as the specified format. Correspondingly, the server can obtain the default format as the specified format. In the case where the custom format corresponding to the model output format configuration item shown in Figure 7 is in the selected state, it indicates that the user expects the large language model to use the custom format as the output format. For example, when the format entered in the input box 702 is the JSON format, the electronic device can send the JSON format to the server as the specified format, and the server can obtain the user - defined JSON format as the specified format.

[0215] The format test input text can be pre - set in the server.

[0216] S1102: The server combines the specified format and the format test input text into a combined input text and inputs it into the large language model Mt, and obtains the format test result text output by the large language model Mt.

[0217] In the above steps, the test data included in the second test set, that is, the combined input text obtained by the server combining the specified format and the format test input text. Among them, the specified format is the aforementioned "format information", the large language model Mt is the aforementioned "first large language model", and the format test result text output by the large language model Mt is the aforementioned "test result".

[0218] S1103: The server determines whether the large language model Mt can understand the specified format according to the format of the format test result text.

[0219] When the large language model Mt understands the specified format, the server can execute Figure 6 the large language model evaluation steps (S601 - S605) shown.

[0220] When the large language model Mt does not understand the specified format, execute Figure 9 the large language model evaluation steps (S901 - S905) shown.

[0221] It should be noted that the specific implementation of the server for merging the specified format and the format test input text can refer to Figure 6 the corresponding content implementation in the embodiments shown, which will not be elaborated here in detail.

[0222] Exemplarily, the format test input text built into the server is: Please indicate the time and location of the given text. Text: "I want to go home for hot pot tonight.".

[0223] The specified format obtained by the server is in JSON format.

[0224] The server merges the above - mentioned specified format and the format test input text into a merged input text: Please indicate the time and location of the given text. Text: "I want to go home for hot pot tonight.". The output format is in JSON format.

[0225] After that, when the server inputs the above - mentioned merged input text into the large language model Mt, it can obtain the format test result text output by the large language model Mt. For example, the obtained format test result text is: {"time": "today", "location": "home"}.

[0226] Then, the server determines whether the large language model Mt can understand the specified format according to the format of the above - mentioned format test result text.

[0227] In some examples, combining the format test result text in the above example, the server obtains that the keyword corresponding to time in the format test result text is "today", and the keyword corresponding to location is "home". Then, the server can determine that the format of the above - mentioned format test result text is in JSON format, and the specified format obtained by the server is also in JSON format. It can also be described as the format of the test result data is the first format. Then, the server can determine that the large language model Mt understands the specified format, that is, it is determined that the large language model Mt can output the result text in the specified format.

[0228] In some other examples, if the server inputs the merged input text in the above example into the large language model Mt and obtains the format test result text output by the large language model Mt as: The time is today and the location is home. The server cannot obtain the name and corresponding value in the JSON format from this format test result text. In other words, the server cannot analyze this test result text in the JSON format. Therefore, the server can determine that the format of the above format test result text is not the JSON format, and then can judge that the large language model Mt does not understand the specified format, that is, it is determined that the large language model Mt cannot output the result text in the specified format.

[0229] In some embodiments, when it is determined that the large language model Mt can output the result text in the specified format, the server can send a message confirming that the large language model Mt can output the specified format to the electronic device, so that the electronic device displays a prompt interface for prompting the user that the large language model Mt can output the specified format. Correspondingly, when it is determined that the large language model Mt cannot output the result text in the specified format, the server can send a message confirming that the large language model Mt cannot output the specified format to the electronic device, so that the electronic device displays a prompt interface for prompting the user that the large language model Mt cannot output the specified format.

[0230] In some embodiments, when the user needs to evaluate the large language model Mt, the user can directly click on the evaluation control in the operation display interface displayed on the electronic device. In response to the user's operation of clicking on the evaluation control, the electronic device can send an evaluation instruction to the server. Then, in response to receiving this evaluation instruction, the server can first execute S1101 - S1103 to determine whether the large language model Mt can output the result text in the specified format. When it is determined that the large language model Mt can output the result text in the specified format, that is, after S1103, the server can continue to execute S601 - S605 as Figure 6 shown. When it is determined that the large language model Mt cannot output the result text in the specified format, that is, after S1103, the server can continue to execute S901 - S905 as Figure 9 shown.

[0231] Using the above evaluation method, the prompt template can also be evaluated. For example, while keeping the dataset and the instruction of the specified output format unchanged, different test input texts are merged by using different prompt templates. Different analysis results can be obtained by using these different test input texts. By analyzing the obtained analysis results, the evaluation of the prompt template can be realized.

[0232] In some embodiments, such as Figure 12As shown, the server obtains multiple input texts A obtained by combining the data in the dataset (i.e., the aforementioned "first dataset") and at least one prompt template A (i.e., the aforementioned "first prompt template") in the prompt template set A (i.e., the aforementioned "first set of prompt templates"). After the server obtains the output format instruction 1 for specifying the large language model, the server can merge each input text A and the instruction 1 respectively to obtain a test input set A including multiple test input texts. The server inputs each test input text in the test input set A into the large language model respectively, and can obtain multiple result texts in the specified format output by the large language model. In this way, the server can use the test code to analyze the differences between each result text in the result set A and the corresponding annotated text to obtain the analysis result A. Based on the analysis result A, the server can obtain an evaluation result (i.e., the aforementioned "first evaluation result"). Similarly, the server can combine the data in the same dataset (i.e., the aforementioned "first dataset") as in the above example and at least one prompt template B (i.e., the aforementioned "second prompt template") in another prompt template set B (i.e., the aforementioned "second set of prompt templates") to obtain multiple input texts B. Based on each input text B and the instruction 1 (the instruction 1 is used to specify the output format of the large language model, that is, to specify the aforementioned "first format"), the test input text can be merged. After that, the server inputs each test input text in the test input set B (i.e., the aforementioned "fifth test set") into the large language model to obtain a result set B including multiple result texts (i.e., the aforementioned "fourth result set"). The server uses the test code to analyze the differences between each result text in the result set B and the corresponding annotated text to obtain the analysis result B. Based on the analysis result B, the server can obtain an evaluation result (i.e., the aforementioned "third evaluation result"). Then the server can analyze the differences between the two evaluation results to obtain the differences in the two large language model evaluations. Since this difference is caused by the change of the prompt template, based on the two evaluation results, the evaluation result of the prompt template can be obtained.

[0233] Optionally, by analyzing the analysis result A and the analysis result B, the server can obtain the differences in the two large language model evaluations. Since this difference is caused by the change of the prompt template, the evaluation result obtained by the server by analyzing the analysis result A and the analysis result B can actually be regarded as the evaluation result of the prompt template.

[0234] Exemplarily, if the data in the dataset is combined with the prompt template A to form the input text A, then analyze the resulting text A output by the large language model at this time. After analyzing each resulting text A in the result set A based on the annotated text, the server can obtain the accuracy rate of the large language model in this evaluation as 80% based on the analysis result. If the data in the dataset is combined with the prompt template B to form the input text B, then analyze the resulting text B output by the large language model at this time. After analyzing each resulting text B in the result set B based on the annotated text, the server can obtain the accuracy rate of the large language model in this evaluation as 90% based on the analysis result. Thus, the server can conclude that the accuracy rate of the large language model in the first evaluation is higher than that in the second evaluation. Since the large language model did not change during the two evaluation processes, but the prompt template changed, the evaluation result finally obtained by the server can be: The prompt template B used in the second evaluation of the large language model can obtain more accurate results.

[0235] The embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions run on the above-mentioned server, the server is enabled to execute each function or step in the above-mentioned method embodiment.

[0236] The embodiment of the present application also provides a computer program product, including a computer program. When the computer program runs on the server, the server is enabled to execute each function or step in the above-mentioned method embodiment.

[0237] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-mentioned division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0238] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0239] The unit described as the separation component may or may not be physically separated. The component displayed as a unit may be a single physical unit or multiple physical units, that is, it may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0240] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0241] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0242] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any change or replacement within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.

Claims

1. A method for evaluating large language models, characterized in that, The method includes: Obtaining a first test set; the first test set includes a plurality of first test data, and the first test data includes first format information for specifying that the output format of the language model is a first format; Obtaining a first result set output by the first large language model based on the first test set; the first result set includes a plurality of first result data, and the plurality of first result data corresponds one-to-one with the plurality of first test data, and the format of the first result data is the first format; Obtaining a first evaluation result based on the first result set and an annotation set; the first evaluation result is used to evaluate the ability of the first large language model; the annotation set includes a plurality of annotation data, and the plurality of annotation data corresponds one-to-one with the plurality of first test data, and the format of the annotation data is the first format.

2. The method according to claim 1, characterized in that The obtaining of the first test set includes: Obtaining a first data set, a first prompt template set, and the first format; the first data set includes a plurality of data, and the first prompt template set includes at least one first prompt template; Generating the first test set based on the first data set, the first prompt template set, and the first format.

3. The method according to claim 2, characterized in that, The first format is a pre-configured format; or, the first format is a custom format.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Receiving a first instruction from an electronic device; the first instruction is used to specify that the output format of the language model is the first format.

5. The method according to any one of claims 1-4, characterized in that, The obtaining of the first evaluation result based on the first result set and the annotation set includes: Using test code to compare and analyze each first result data in the first result set with the corresponding annotation data in the annotation set to obtain a plurality of comparison results; Obtaining the first evaluation result based on the plurality of comparison results and evaluation metrics.

6. The method according to any one of claims 1-5, characterized in that, The first large language model can understand the first format information; Before the obtaining of the first test set, the method further includes: Obtaining a second test set for testing the understanding ability of the first large language model for the first format information; Testing the understanding ability of the first large language model for the first format information based on the second test set.

7. The method according to claim 6, characterized in that, The method further includes: In response to the first large language model not being able to understand the first format information, obtaining a third test set, the third test set includes a plurality of second test data, and the third test set is generated based on a first data set and a first prompt template set; Obtaining a second result set output by the first large language model based on the third test set; the second result set includes a plurality of second result data, and the plurality of second result data corresponds one-to-one with the plurality of second test data; Obtain a fourth test set based on the third test set, the second result set, and the second format information; the fourth test set is used to instruct the second large language model to analyze the capabilities of the first large language model. The fourth test set includes multiple third test data. One third test data includes a second test data, the second result data corresponding to the second test data, and the second format information; the second format information is used to specify that the output format of the language model is the second format. Obtain a third result set output by the second large language model based on the fourth test set; the third result set includes multiple third result data, and the multiple third result data correspond one-to-one to the multiple third test data. The format of the third result data is the second format. Obtain a second evaluation result based on the third result set.

8. The method according to claim 7, wherein The obtaining the second evaluation result based on the third result set includes: Obtain the second evaluation result based on the third result set and the evaluation metrics.

9. The method according to claim 5 or 8, characterized in that The method further includes: Receive the evaluation metrics from the electronic device.

10. The method according to claim 2, wherein The method further includes: Obtain a second set of prompt templates; the second set of prompt templates includes at least one second prompt template. Obtain a fifth test set based on the first data set, the second set of prompt templates, and the first format. Obtain a fourth result set output by the first large language model based on the fifth test set; the format of the result data included in the fourth result set is the first format. Obtain a third evaluation result based on the fourth result set and the annotation set; the third evaluation result is used to evaluate the capabilities of the first large language model. Obtain a prompt evaluation result based on the first evaluation result and the third evaluation result. The prompt evaluation result is used to evaluate the capabilities of the first prompt template and the second prompt template.

11. A server, characterized in that, The server includes a communication module, a memory, and one or more processors; the communication module, the memory, and the processors are coupled. The communication module is used to establish a communication connection and send and receive data through the communication connection. The memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the server is caused to execute the method according to any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, Includes computer instructions that, when running on a server, cause the server to execute the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Language model evaluation method and device, electronic equipment and storage medium

    CN117093459A

  • Data processing instruction generation method and device, large model training method and device and electronic equipment

    CN117112572A

  • Text recognition model training method and related device

    CN117216257A

  • System and method for automatically generating executable tests in gherkin format

    US20220283929A1

  • Auditing artificial intelligence (AI) systems through common sense reasoning tasks

    US20230359835A1