Large language model evaluation method and device, equipment and storage medium

By evaluating large language models through multiple evaluation methods, we can address their lack of adaptability and response accuracy in complex situations, achieve effective evaluation in the absence of standard answers, and ensure the reliability and fairness of the model in practical applications.

CN120632387APending Publication Date: 2025-09-12ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510656425.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing large language models lack adaptability and response accuracy when faced with complex or specific questions, and it is difficult to effectively evaluate their effectiveness, especially in the absence of standard answers.

Method used

The large language model is evaluated using a variety of evaluation methods, including obtaining samples, annotating evaluation conclusions, and summarizing model evaluation conclusions. The execution results of the first and second language models are compared through a third language model to generate an overall evaluation conclusion to ensure the correctness and reliability of the evaluation.

Benefits of technology

It achieves a reliable measurement of the effectiveness of large language models, ensuring their effectiveness and fairness in practical applications, especially in the absence of standard answers, and improves the accuracy and reliability of evaluation conclusions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632387A_ABST
    Figure CN120632387A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a large language model evaluation method and device, equipment and a storage medium. The method comprises the steps of obtaining a sample containing a prompt corresponding to a target task, a first execution result obtained by executing the target task based on the prompt by a first large language model, and a second execution result obtained by executing the target task based on the prompt by a second large language model; obtaining an annotation evaluation conclusion annotated by the annotation party for the sample according to a good and bad comparison result of the first execution result and the second execution result; inputting the sample into at least one third big language model, so that each third big language model performs quality comparison on the first execution result and the second execution result based on the sample, and generating a model evaluation conclusion corresponding to the sample according to a quality comparison result; summarizing the annotation evaluation conclusion and the model evaluation conclusion into a total evaluation conclusion corresponding to the sample; wherein the evaluation conclusion is used for indicating a good and bad comparison result of the first execution result and the second execution result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and storage medium for evaluating a large language model. Background Art

[0002] Today, intelligent dialogue systems (IDSs) are widely used in a wide range of fields, including customer service, legal aid, online education, and medical consultation. Taking the rapidly developing field of medical consultation as an example, IDSs have become a key tool for improving the quality and efficiency of medical services. By simulating human communication, IDSs provide users with medical-related conversations, Q&A sessions, and query services, such as disease diagnosis, treatment recommendations, and medication instructions.

[0003] An intelligent dialogue system is a specific user-facing application of a large language model. It aims to understand and answer questions posed by users in natural language and generate clear and concise answers. Specifically, the intelligent dialogue system is based on a large language model, which understands and answers user questions and generates corresponding answers. Therefore, the performance of the large language model directly determines the accuracy of the content in intelligent dialogues, including factuality, relevance, and regional and temporal accuracy. Evaluation of large language models involves measuring their performance through specific model testing methods and performance metrics. This has become a critical step in ensuring the reliability, effectiveness, and fairness of large language models in practical applications. Summary of the Invention

[0004] One or more embodiments of the present application provide the following technical solutions:

[0005] This application provides a large language model evaluation method, the method comprising:

[0006] Obtaining a sample for evaluating the first large language model and the second large language model; wherein the sample includes a prompt corresponding to a target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt;

[0007] Obtaining a labeling evaluation conclusion for the sample labeling by the labeling party based on a comparison result of the first execution result and the second execution result;

[0008] Inputting the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generates a model evaluation conclusion corresponding to the sample according to the comparison results;

[0009] The annotation evaluation conclusion and the model evaluation conclusion are summarized into a total evaluation conclusion corresponding to the sample; wherein the evaluation conclusion is used to indicate a comparison result of the advantages and disadvantages of the first execution result and the second execution result.

[0010] The present application also provides a large language model evaluation device, the device comprising:

[0011] a sample acquisition module for acquiring samples for evaluating the first large language model and the second large language model; wherein the samples include a prompt corresponding to a target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt;

[0012] a labeling acquisition module, which acquires a labeling evaluation conclusion of the sample labeling by the labeling party based on a comparison result of the first execution result and the second execution result;

[0013] a model evaluation module, inputting the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generates a model evaluation conclusion corresponding to the sample based on the comparison results;

[0014] A conclusion generation module summarizes the annotation evaluation conclusion and the model evaluation conclusion into an overall evaluation conclusion corresponding to the sample; wherein the evaluation conclusion is used to indicate a comparison result of the advantages and disadvantages of the first execution result and the second execution result.

[0015] The present application also provides an electronic device, comprising:

[0016] processor;

[0017] a memory for storing processor-executable instructions;

[0018] The processor implements the steps of any of the above methods by running the executable instructions.

[0019] The present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of any of the methods described above.

[0020] In the above technical solution, first, a sample for evaluating the first large language model and the second large language model can be obtained. The sample may include a prompt corresponding to the target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt. Then, on the one hand, a labeling evaluation conclusion for the sample can be obtained based on the comparison result of the first execution result and the second execution result in the sample by the labeling party. On the other hand, the sample can be input into at least one third large language model, so that each third large language model can compare the first execution result and the second execution result in the sample based on the sample, and generate a model evaluation conclusion corresponding to the sample based on the comparison result. Finally, the labeling evaluation conclusion and the model evaluation conclusion can be summarized into a total evaluation conclusion corresponding to the sample. Among them, the evaluation conclusion can be used to indicate the comparison result of the first execution result and the second execution result in the sample.

[0021] The above approach, on the one hand, enables the evaluation of large language models, making it possible to measure the model effectiveness of large language models, thereby ensuring the reliability, effectiveness, and fairness of large language models in practical applications. Moreover, it is particularly suitable for the evaluation of large language models when there are no standard answers corresponding to questions. On the other hand, multiple evaluation methods are used to evaluate large language models, and the evaluation conclusions obtained through these methods are summarized, which can improve the correctness and reliability of the final overall evaluation conclusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The following is a description of the accompanying drawings required for describing the exemplary embodiments, in which:

[0023] Figure 1 It is a schematic diagram of an intelligent dialogue system shown in an exemplary embodiment of the present application.

[0024] Figure 2 This is a flowchart of a large language model evaluation method shown in an exemplary embodiment of the present application.

[0025] Figure 3 It is a schematic diagram of a process for generating a model evaluation conclusion shown in an exemplary embodiment of the present application.

[0026] Figure 4 This is a schematic diagram of a process for generating a general evaluation conclusion shown in an exemplary embodiment of the present application.

[0027] Figure 5 It is a structural diagram of a device shown in an exemplary embodiment of the present application.

[0028] Figure 6It is a block diagram of an evaluation device for a large language model shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0029] The exemplary embodiments will be described in detail herein, with examples thereof being illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of the present application. Instead, they are merely examples consistent with some aspects of one or more embodiments of the present application.

[0030] It should be noted that, in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this application. In some other embodiments, the method may include more or fewer steps than those described in this application. In addition, a single step described in this application may be broken down into multiple steps for description in other embodiments; and multiple steps described in this application may be combined into a single step for description in other embodiments.

[0031] With the development of financial technology, money laundering activities have become increasingly covert and complex, involving faster and more diverse flows of funds. In the field of anti-money laundering, faced with ever-changing tools and complex money laundering methods, traditional rule-based monitoring methods and simple data analysis technologies have significant limitations in addressing these emerging threats. For example, they are unable to effectively process unstructured data, such as key information from sources such as news reports and social media posts; at the same time, they often perform poorly for complex money laundering patterns that require a deep understanding of the context to identify. Therefore, current money laundering risk identification mainly relies on manual processing of anti-money laundering public opinion information or reviewing recent suspicious transaction reports. This method is not only time-consuming and labor-intensive, but also inefficient, making it difficult to automatically identify complex money laundering patterns and effectively predict and detect potential money laundering activities.

[0032] Given these challenges, there's a general desire for intelligent, efficient solutions to improve money laundering risk identification in anti-money laundering efforts. Large Language Models (LLMs) have garnered widespread attention for their exceptional natural language understanding and generation capabilities. Trained on large-scale text datasets, these models are capable of understanding and processing the nuances of human language, offering new insights into addressing these challenges.

[0033] Specifically, large language models can be used to automatically parse and classify massive amounts of unstructured data from various channels, quickly extracting characteristic information related to money laundering behavior; by learning complex patterns from historical cases, the accuracy of identifying new money laundering techniques can be improved; and they can be combined with real-time data for dynamic risk assessments, providing early warnings of potential money laundering activities. Furthermore, leveraging the power of deep learning algorithms, large language models can continuously optimize their performance without extensive human intervention, adapting to the ever-changing money laundering risk environment. Therefore, using large language models to improve the intelligence level of existing money laundering risk identification systems will not only help reduce manual workloads and improve work efficiency, but will also significantly enhance financial institutions' ability to prevent financial crime and ensure the safe and stable operation of the economic system.

[0034] In this application, the money laundering risk identification system can be implemented using an intelligent dialogue system (or intelligent question-and-answer system). Taking the intelligent dialogue system as an example, it simulates human communication methods to provide users with medical-related dialogue, question-and-answer, and query services. The intelligent dialogue system is a specific user-facing application of the large language model. It is designed to understand and answer questions posed by users in natural language and generate concise and clear answers. Specifically, the intelligent dialogue system is based on the large language model, which understands and answers user questions and generates corresponding answers.

[0035] In actual applications, using an intelligent dialogue system, users can include text materials such as anti-money laundering public opinion information or recent suspicious transaction reports in questions related to money laundering risk identification and provide them to the large language model. The large language model will understand and answer the questions raised by the users, that is, identify money laundering risks based on these text materials and generate corresponding answers, which describe whether money laundering risks are identified from these text materials.

[0036] Large language models are deep learning models trained using large amounts of text data. They can be used to generate natural language text or understand its meaning. Large language models can handle a variety of natural language tasks, such as text classification, named entity recognition (NER), question answering, and conversation, and are a key path to artificial intelligence.

[0037] In the field of natural language processing (NLP), large-scale text datasets are often referred to as corpuses. Corpuses can contain a variety of text data, such as literary works, academic papers, legal documents, news reports, everyday conversations, emails, and online forum posts. By learning from the text data in a corpus, large language models can acquire and understand the patterns and regularities of natural language, enabling effective processing and generation of human language.

[0038] Large language models typically use the Transformer architecture, meaning they are deep learning models based on the Transformer architecture. Transformer-based deep learning models are a type of neural network model that excels in fields like natural language processing.

[0039] Transformer is a neural network model used for sequence-to-sequence modeling. Transformer does not rely on recursive structures and can parallelize training and inference, speeding up model processing. In deep learning models based on the Transformer architecture, a multi-layer Transformer encoder is typically used to extract features from the input sequence, and a Transformer decoder is used to convert the extracted features into an output sequence. At the same time, such models typically also use a self-attention mechanism to capture long-distance dependencies in the input sequence, as well as residual connections and regularization methods to accelerate training and improve model performance.

[0040] A pretrained model is a large language model pretrained on large amounts of unlabeled text data. Pretrained models are general-purpose models; they are not designed or optimized for specific tasks. To adapt pretrained models to specific application scenarios and task requirements, they require fine-tuning to improve their performance on specific tasks. The large language model that is ultimately put into use is typically a pretrained model that has been further fine-tuned, performing supervised learning on labeled text data. Pretraining and fine-tuning are complementary processes: pretraining enables the model to acquire broad language understanding capabilities, while fine-tuning makes the model more specialized and accurate for specific tasks.

[0041] In other words, the training process of a large language model can be divided into two stages: pre-training and fine-tuning. During the pre-training stage, unsupervised learning (e.g., self-supervised learning) can be used on large-scale, unlabeled text datasets (e.g., online encyclopedias, online articles, books, etc.). Specifically, the model can predict missing parts or the next word based on the context, learn statistical laws such as semantics and syntax, and language structure. Backpropagation and optimization algorithms (e.g., gradient descent) are used to minimize prediction losses, iteratively update model parameters, and gradually improve the model's understanding of language. During the fine-tuning phase, you can select corresponding supervised learning tasks (for example, text classification, named entity recognition, question-answering systems, dialogue systems, etc.) based on the specific application scenarios and task requirements, and prepare task-specific text datasets. You can then use the pre-trained model as the starting point for fine-tuning, and use supervised learning to fine-tune on the task-specific text dataset. Specifically, you can perform the task based on the text dataset, and minimize the loss used to measure the performance of the model in processing specific tasks through backpropagation and optimization algorithms (for example, gradient descent). The model parameters are iteratively updated to gradually improve the model's performance on specific tasks. In practical applications, fine-tuning can flexibly choose supervised learning, unsupervised learning, or semi-supervised learning methods based on the specific application scenarios and the type of available data.

[0042] The language comprehension ability learned by the large language model during the pre-training and fine-tuning stages enables the large language model to perform logical inference, knowledge reasoning, or problem-solving by understanding, analyzing, and integrating text information when faced with complex problems or tasks. This ability is usually referred to as the reasoning ability of the large language model.

[0043] In practical applications, the pre-trained large language model is usually called the base model of the large language model, and the fine-tuned large language model is called the serving model of the large language.

[0044] Large language models are typically guided or stimulated by prompts (also known as prompts) to perform specific tasks. A prompt can be an initial text or text fragment provided to the large language model, such as a sentence, a question, or a conversation, intended to guide or stimulate the model to produce the corresponding output. Prompts are a key tool for guiding model output and can be very simple or quite complex, including instructions, examples, and descriptions of the desired output format. Prompts explicitly tell the large language model what task it is expected to perform, such as answering a question, simulating a conversation, writing an article, or translating text. Prompts also provide the large language model with necessary background information and context, enabling it to understand the logic, style, theme, or stance it should follow when generating content. Prompts can also inspire the large language model to demonstrate its inherent knowledge or specific language abilities, such as explaining complex concepts, citing regulations, or imitating the writing style of a specific author.

[0045] Since large language models are primarily used to understand and generate human language based on text processing, prompts usually appear in the form of text. However, in practical applications, large language models can also accept other forms of input as prompts, such as images, audio, and even video, provided that the large language model is designed or trained to process multimodal data.

[0046] Intelligent dialogue systems typically rely primarily on the knowledge acquired by their large language models during training by studying static corpora. Due to the limitations of this knowledge, the system may experience hallucinations when answering complex or specific questions. Hallucinations occur when the content generated by the large language model appears plausible and coherent, sometimes even mimicking human emotions and thought patterns, creating the illusion of "understanding" the input, even though in reality, the content is inaccurate or misleading. This reliance on static corpora limits the system's adaptability and response accuracy.

[0047] To improve the adaptability and response accuracy of intelligent dialogue systems, a method called Retrieval-Augmented Generation (RAG) can be used to combine information retrieval and model generation. This allows the intelligent dialogue system to answer user questions without relying solely on the knowledge acquired by the large language model during training by learning static corpus. Instead, the system can first perform information retrieval from a large document collection based on the question, then understand and address the question based on the retrieved relevant documents, and generate the corresponding answer. In other words, the document collection can be combined with the large language model, and relevant information can be retrieved from the document collection in real time during the model generation process to assist the model in making more accurate and comprehensive responses or decisions. Because the retrieved information and the context of the question are taken into account during the model generation process, the generated content can be guaranteed to meet actual needs while being accurate, reliable, coherent, and natural.

[0048] This shows that the performance of large language models directly determines the accuracy of intelligent conversation content, including factuality, relevance, and regional and temporal accuracy. Evaluation of large language models typically involves measuring their effectiveness through specific model testing methods and performance metrics. This has become a crucial step in ensuring the reliability, effectiveness, and fairness of large language models in practical applications.

[0049] One or more embodiments of the present application provide a technical solution for implementing the evaluation of a large language model. In this technical solution, a sample for evaluating a first large language model and a second large language model can first be obtained. The sample can include a prompt corresponding to a target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt. Then, on the one hand, an annotation evaluation conclusion for the sample can be obtained based on a comparison result of the first execution result and the second execution result in the sample by the annotation party. On the other hand, the sample can be input into at least one third large language model, so that each third large language model can compare the first execution result and the second execution result in the sample based on the sample, and generate a model evaluation conclusion corresponding to the sample based on the comparison result. Finally, the annotation evaluation conclusion and the model evaluation conclusion can be summarized into a total evaluation conclusion corresponding to the sample. The evaluation conclusion can be used to indicate a comparison result of the first execution result and the second execution result in the sample.

[0050] The above approach, on the one hand, enables the evaluation of large language models, making it possible to measure the model effectiveness of large language models, thereby ensuring the reliability, effectiveness, and fairness of large language models in practical applications. Moreover, it is particularly suitable for the evaluation of large language models when there are no standard answers corresponding to questions. On the other hand, multiple evaluation methods are used to evaluate large language models, and the evaluation conclusions obtained through these methods are summarized, which can improve the correctness and reliability of the final overall evaluation conclusion.

[0051] Please refer to Figure 1 , Figure 1 It is a schematic diagram of an intelligent dialogue system shown in an exemplary embodiment of the present application.

[0052] like Figure 1 As shown, the intelligent dialogue system may include a server and at least one client accessing the server via any type of wired or wireless network.

[0053] The above-mentioned server may correspond to a server comprising an independent physical host, or may be a server cluster consisting of multiple independent physical hosts; or, may correspond to a virtual server, cloud server, etc. hosted by a host cluster.

[0054] The above-mentioned client can correspond to terminal devices such as smart phones, tablet computers, laptops, desktop computers, PCs (Personal Computers), PDAs (Personal Digital Assistants), wearable devices (such as smart glasses, smart watches, etc.), smart car devices or game consoles.

[0055] Users can use the intelligent dialogue service provided by the intelligent dialogue system through the client; the client and the server can implement user-oriented intelligent dialogue services through data exchange between each other. The intelligent dialogue service can be an intelligent question-and-answer service that implements intelligent dialogue.

[0056] Specifically, the server can be equipped with a large language model, and the intelligent dialogue system can be based on the large language model, and the large language model can understand and answer the question text and generate an answer text corresponding to the question text.

[0057] For example, the client can output a corresponding user interface to the user, allowing the user to perform operations such as inputting question text (which may be called a query or question) and uploading documents or pictures used to assist in asking questions in the user interface, so as to ask questions to the intelligent dialogue system and use the intelligent dialogue service provided by the intelligent dialogue system. The client can send the question text input by the user to the server, which will generate a corresponding answer text (which may be called an answer) for the question text and output the answer text to the user, that is, return the answer text to the client, and the client will display the answer text to the user through the user interface so that the user can view the corresponding answer generated by the intelligent dialogue system, thereby realizing a user-oriented intelligent dialogue service.

[0058] It should be noted that question text can be regarded as a special kind of prompt. Question text refers to the specific problem that the user wants to solve, and this problem is usually expressed through a carefully designed prompt.

[0059] The server can also be equipped with a knowledge base and an information retrieval component. This knowledge base is external to the large language model installed on the server. That is, the data in this knowledge base is not the knowledge acquired by the large language model through learning during training. Instead, it serves as auxiliary data in the large language model's reasoning process, assisting the large language model in generating an answer text corresponding to the question text. During the reasoning process of the large language model, the information retrieval component can perform information retrieval in the knowledge base based on the question text, thereby assisting the large language model in generating an answer text corresponding to the question text through the retrieved relevant information.

[0060] It should be noted that the aforementioned server can also be equipped with other functional components or subsystems such as a prompt generation component. These components or subsystems can work in conjunction with the large language model installed on the server to jointly generate the answer text corresponding to the question text.

[0061] Furthermore, the intelligent dialogue system may have only one large language model, which can be used to perform dialogue tasks and, based on actual needs, can also be used to perform tasks that can assist in generating content for intelligent dialogue, such as language detection tasks, question rewriting tasks, question splitting tasks, intent recognition tasks, and answer aggregation tasks. Alternatively, the intelligent dialogue system may have multiple large language models, which, in addition to the large language model used to perform dialogue tasks, may also include large language models for performing language detection, large language models for performing question rewriting tasks, large language models for performing question splitting tasks, large language models for performing intent recognition tasks, and large language models for performing answer aggregation tasks, which are used to perform tasks that can assist in generating content for intelligent dialogue.

[0062] In this application, the Figure 1 The large language model used to perform the dialogue task in the intelligent dialogue system shown is used as the large language model to be evaluated (which can be called the first large language model); alternatively, all the large language models in the intelligent dialogue system can be collectively referred to as the first large language model. In order to evaluate the first large language model, several question texts and generated answer texts corresponding to each question text generated by the first large language model can be obtained. For a question text, the generated answer text corresponding to the question text is the answer text generated by the first large language model based on the question text through reasoning. In actual applications, a question text can be one or more words or sentences used to ask a question, and an answer text can be one or more words or sentences used to answer the question.

[0063] If there is a standard answer corresponding to each question (that is, the answer that is considered to be the truth), another large language model (which can be called the third language model) can be used. The third language model will compare the generated answer text corresponding to each question text with the standard answer text to achieve the correctness evaluation of the generated answer text corresponding to each question text. That is, the third language model will evaluate whether the generated answer text corresponding to each question text is correct, and generate corresponding model evaluation results. Based on these model evaluation results, specific model effect indicators can be calculated for the first language model to represent the model effect of the first language model through these model effect indicators.

[0064] Alternatively, multiple large language models (all referred to as third-largest language models) can be utilized, with each third-largest language model comparing the generated answer text corresponding to each question text with the standard answer text to evaluate the correctness of the generated answer text corresponding to each question text. That is, each third-largest language model evaluates whether the generated answer text corresponding to each question text is correct and generates a corresponding model evaluation result. In this case, taking one of the question texts and its corresponding generated answer text as an example, the corresponding model evaluation results generated by each third-largest language model can be summarized; further, based on the summarized model evaluation results, specific model effect indicators can be calculated for the first-largest language model, so that the model effect of the first-largest language model can be represented by these model effect indicators.

[0065] However, if there are no standard answers corresponding to each question, a large language model used to perform the dialogue task in another intelligent dialogue system can be used as the large language model to be evaluated (this can be called the second large language model); alternatively, all large language models in the intelligent dialogue system can be collectively referred to as the second large language model. To evaluate the first and second large language models, several question texts, generated answer texts corresponding to each question text generated by the first large language model, and generated answer texts corresponding to each question text generated by the second large language model can be obtained.

[0066] In this case, a third language model can be used to compare the generated answer texts corresponding to each question text generated by the first language model with the generated answer texts corresponding to each question text generated by the second language model. The third language model can compare the advantages and disadvantages of the generated answer texts corresponding to each question text generated by the first language model and the second language model respectively, and generate corresponding model evaluation conclusions based on these comparison results. The model evaluation conclusion reflects whether the answer generated by the first language model is better or the answer generated by the second language model is better.

[0067] Alternatively, multiple large language models (all of which can be referred to as third-largest language models) can be utilized, with each third-largest language model comparing the generated answer text corresponding to each question text generated by the first-largest language model with the generated answer text corresponding to each question text generated by the second-largest language model. Each third-largest language model can compare the quality of the generated answer text corresponding to each question text generated by the first-largest language model and the second-largest language model, and generate corresponding model evaluation conclusions based on these quality comparison results. In this case, taking one question text and the corresponding generated answer text generated by the first-largest language model and the second-largest language model as an example, the corresponding model evaluation conclusions generated by each third-largest language model can be summarized. The summarized model evaluation conclusion reflects whether the answer corresponding to the question generated by the first-largest language model is better or the answer corresponding to the question generated by the second-largest language model is better. Furthermore, based on the summarized model evaluation conclusions corresponding to each question text, an overall model evaluation conclusion can be generated. The model evaluation conclusion reflects whether the answer generated by the first-largest language model is better or the answer generated by the second-largest language model is better.

[0068] The evaluation process of the first language model and the second language model is described in detail below.

[0069] refer to Figure 2 , Figure 2This is a flowchart of a large language model evaluation method shown in an exemplary embodiment of the present application.

[0070] In this embodiment, the large language model evaluation method described above can be applied to an intelligent evaluation platform that interfaces with various third-largest language models. The intelligent evaluation platform and the third-largest language model can interact via various types of wired or wireless networks to establish a communication connection, thereby enabling the intelligent evaluation platform to interface with the third-largest language model. Alternatively, the third-largest language model can be part of the intelligent evaluation platform.

[0071] like Figure 2 As shown, the evaluation method of the large language model may include the following steps:

[0072] Step 202: Obtain samples for evaluating the first large language model and the second large language model; wherein the samples include a prompt corresponding to a target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt.

[0073] In this embodiment, in order to evaluate the first and second language models, several prompts corresponding to the target task, execution results corresponding to each prompt generated by the first language model (which may be referred to as first execution results), and execution results corresponding to each prompt generated by the second language model (which may be referred to as second execution results) may be obtained. For a prompt corresponding to the target task, the first execution result corresponding to the prompt is the first execution result obtained by the first language model executing the target task based on the prompt, and the second execution result corresponding to the prompt is the second execution result obtained by the second language model executing the target task based on the prompt.

[0074] It should be noted that in some embodiments, the target task may be an intelligent question-answering task. Accordingly, the prompt may be a question text, the first execution result may be a first answer text, and the second execution result may be a second answer text. This allows for a comprehensive evaluation of the large language model's performance in intelligent question-answering tasks.

[0075] Alternatively, in some embodiments, the intelligent question-answering task can be split to separate the link nodes from the task chain of the intelligent question-answering task to obtain the subtasks included in the intelligent question-answering task. Among them, the subtasks included in the intelligent question-answering task may be language detection tasks, question rewriting tasks, question splitting tasks, intent recognition tasks, and / or answer aggregation tasks. This allows for staged evaluation of the subtasks of the large language model in the process of performing the intelligent question-answering task, rather than just the results of the overall evaluation, thereby increasing the interpretability of the evaluation.

[0076] In this case, the first or second language model can refer to the serving model of the large language model. In practical applications, the constructed large language model can be pre-trained using unsupervised learning on a large, unlabeled text dataset to obtain a base model of the large language model. Furthermore, the dialogue task can be used as a supervised learning task for fine-tuning, and a text dataset specific to the dialogue task can be prepared. The base model of the large language model can then be used as the starting point for fine-tuning, using supervised learning to fine-tune the data on the text dataset specific to the dialogue task, resulting in a serving model of the large language model.

[0077] In practical applications, a prompt corresponding to the target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt can constitute a sample for evaluating the first large language model and the second large language model. Accordingly, a number of prompts corresponding to the target task, a first execution result obtained by the first large language model performing the target task based on each of the prompts, and a second execution result obtained by the second large language model performing the target task based on each of the prompts can constitute a sample set for evaluating the first large language model and the second large language model.

[0078] In some embodiments, the question text may be a question text related to anti-money laundering.

[0079] Step 204: Obtaining a labeling evaluation conclusion for the sample labeling by the labeling party based on a comparison result of the first execution result and the second execution result.

[0080] In this embodiment, on the one hand, for a sample consisting of a prompt corresponding to the above-mentioned target task, a first execution result corresponding to the prompt, and a second execution result corresponding to the prompt, the annotator can compare the pros and cons of the first execution result and the second execution result in the sample, and then annotate the sample with a corresponding evaluation conclusion (which can be called an annotation evaluation conclusion) based on the comparison result. The annotation evaluation conclusion can be used to indicate the annotation policy's comparison result of the pros and cons of the first execution result and the second execution result in the sample. For example, if the annotator determines through comparison that the first execution result in the sample is better than the second execution result, the sample can be annotated with "Result 1 wins"; if the annotator determines through comparison that the second execution result in the sample is better than the first execution result, the sample can be annotated with "Result 2 wins".

[0081] In actual applications, the above-mentioned annotation party may refer to the annotation personnel or the electronic equipment used by the annotation personnel, or may refer to other intelligent annotation systems, and this application does not impose any special restrictions on this.

[0082] Step 206: Input the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generates a model evaluation conclusion corresponding to the sample according to the comparison result.

[0083] In this embodiment, on the other hand, for the above-mentioned sample, the above-mentioned third largest language model can also compare the pros and cons of the first execution result and the second execution result in the sample, so as to mark the sample with a corresponding evaluation conclusion (which can be called a model evaluation conclusion) based on the pros and cons comparison result. Among them, the model evaluation conclusion can be used to indicate the pros and cons comparison result of the third largest language model for the first execution result and the second execution result in the sample. For example, if the third largest language model determines through comparison that the first execution result in the sample is better than the second execution result, then "Result 1 wins" can be output as the model evaluation conclusion corresponding to the sample; if the third largest language model determines through comparison that the second execution result in the sample is better than the first execution result, then "Result 2 wins" can be output as the model evaluation conclusion corresponding to the sample.

[0084] Specifically, the above-mentioned samples can be input into each third-largest language model, so that each third-largest language model can compare the first execution result and the second execution result based on the sample, and generate a model evaluation conclusion corresponding to the sample according to the comparison result.

[0085] In practical applications, a prompt can be constructed based on the above sample, and the prompt can be input into each third-largest language model. Under the guidance of the prompt, each third-largest language model can compare the first execution result and the second execution result based on the sample, and generate a model evaluation conclusion corresponding to the sample based on the comparison result.

[0086] The third language model mentioned above can refer to the serving model of the large language model. In practical applications, the constructed large language model can be pre-trained on a large-scale, unlabeled text dataset using unsupervised learning to obtain the base model of the large language model. Furthermore, the task of comparing the first and second execution results in a sample can be used as a supervised learning task during fine-tuning, and a text dataset specific to this task can be prepared. The base model of the large language model can then be used as the starting point for fine-tuning, and fine-tuned on the task-specific text dataset using supervised learning to obtain the serving model of the large language model.

[0087] In some embodiments, taking any third-largest language model as an example, the above-mentioned sample can be input into the third-largest language model, and the third-largest language model calculates at least one task evaluation indicator corresponding to the first execution result and the second execution result in the sample based on the sample, and compares the at least one task evaluation indicator corresponding to the first execution result and the second execution result in the sample, so as to generate a model evaluation conclusion corresponding to the sample based on the size comparison result, thereby realizing that the third-largest language model compares the first execution result and the second execution result in the sample based on the sample, and generates a model evaluation conclusion corresponding to the sample based on the comparison result.

[0088] For example, assume that 4 task evaluation metrics are set, namely metric M1, metric M2, metric M3, and metric M4, and corresponding weights are set for these 4 task evaluation metrics. Among them, the weight of M1 is w1, the weight of M2 is w2, the weight of M3 is w3, and the weight of M4 is w4. Further assume that the 4 task evaluation metrics corresponding to the first answer text in the above sample calculated by a certain third large language model are M11, M21, M31, and M41, and the 4 task evaluation metrics corresponding to the second answer text in the sample are M12, M22, M32, and M42. Then, the total metric corresponding to the first answer text in the sample can be calculated as M10 = (w1 * M11 + w2 * M21 + w3 * M31 + w4 * M41), and the total metric corresponding to the second answer text in the sample can be calculated as M20 = (w1 * M12 + w2 * M22 + w3 * M32 + w4 * M42). Furthermore, the magnitudes of M10 and M20 can be compared. If M10 > M20, the model evaluation conclusion generated by the third large language model corresponding to the sample can be "Answer 1 wins". If M10 < M20, the model evaluation conclusion generated by the third large language model corresponding to the sample can be "Answer 2 wins".

[0089] Alternatively, the magnitudes of M11 and M12, M21 and M22, M31 and M32, M41 and M42 can be compared. Assume that M11 > M12, M21 > M22, M31 < M32, M41 < M42. Then, the magnitudes of the weights (w1 + w2) by which the first answer text is superior and the weights (w3 + w4) by which the second answer text is superior can be further compared. If (w1 + w2) > (w3 + w4), the model evaluation conclusion generated by the third large language model corresponding to the sample can be "Answer 1 wins". If (w1 + w2) < (w3 + w4), the model evaluation conclusion generated by the third large language model corresponding to the sample can be "Answer 2 wins".

[0090] In some embodiments, if the above target task is an intelligent question - answering task, the above at least one task evaluation metric may include one or more of the following metrics: relevance for indicating whether the answer is directly related to the question; accuracy for indicating whether there are factual errors in the answer; fluency for indicating whether the language expression of the answer is fluent; logic for indicating whether the logic of the answer is clear; professionalism for indicating whether the use of professional terms in the answer is accurate; completeness for indicating whether the answer completely answers the question; depth for indicating whether the answer deeply explores the question; practicality for indicating whether the answer has practical application value.

[0091] Specifically, relevance reflects whether the answer is directly related to the question and whether it includes irrelevant content. Accuracy reflects whether the answer contains factual errors. Fluency reflects whether the answer's language is fluent and smooth, and whether there are grammatical errors or inappropriate wording. Logic reflects whether the answer's structure and logic are clear, and whether there are any inconsistencies or logical jumps. Professionalism reflects whether the answer uses professional terminology accurately. Completeness reflects whether the answer fully answers the question and whether it omits important details or information. Depth reflects whether the answer provides sufficient depth and insight, and whether it merely provides a superficial response or delves into the issue. Practicality reflects whether the answer has practical application value and provides readers with specific guidance or suggestions.

[0092] In actual applications, different task evaluation indicators can be set for different tasks (intelligent question and answer tasks or subtasks split from intelligent question and answer tasks) according to actual situations and needs. This application does not impose any special restrictions on this.

[0093] In some embodiments, when there are multiple third-largest language models, taking any one of the third-largest language models as an example, the model evaluation conclusions corresponding to the sample generated by each third-largest language model can be summarized into the total model evaluation conclusion corresponding to the sample.

[0094] like Figure 3 As shown, assuming that there are four third-largest language models, namely large language model 1, large language model 2, large language model 3, and large language model 4 (for example: GPT-4, KIMI, Tongyi Qianwen, Wenxin Yiyan), the model evaluation conclusions corresponding to the above sample generated by large language model 1, large language model 2, and large language model 4 are all "answer 1 wins", and the model evaluation conclusion corresponding to the sample generated by large language model 3 is "answer 2 wins". It can be regarded as that in the process of comparing the pros and cons of the answers, large language model 1, large language model 2, and large language model 4 voted for answer 1, and large language model 3 voted for answer 2. Since the number of third-largest language models that voted for answer 1 (the number is 3) is greater than the number of third-largest language models that voted for answer 2 (the number is 1), it can be determined that the total model evaluation conclusion corresponding to the sample is "answer 1 wins".

[0095] Alternatively, corresponding weights can be set for these 4 third largest language models respectively. The weight of language model 1 is w1, the weight of language model 2 is w2, the weight of language model 3 is w3, and the weight of language model 4 is w4. Then, the total weight of the third largest language models voting for answer 1 can be calculated as W1 = (w1 + w2 + w4), and the total weight of the third largest language models voting for answer 2 is W2 = w3. If W1 > W2, it can be determined that the overall evaluation conclusion of the model corresponding to this sample is "answer 1 wins". If W1 < W2, it can be determined that the overall evaluation conclusion of the model corresponding to this sample is "answer 2 wins".

[0096] Step 208: Aggregate the labeled evaluation conclusion and the model evaluation conclusion into the overall evaluation conclusion corresponding to the sample; where the evaluation conclusion is used to indicate the comparison result of the advantages and disadvantages between the first execution result and the second execution result.

[0097] In this embodiment, for the above sample, when the labeled evaluation conclusion corresponding to this sample and the model evaluation conclusion corresponding to this sample are obtained, the labeled evaluation conclusion and the model evaluation conclusion can be aggregated into the overall evaluation conclusion corresponding to this sample. Since the labeled evaluation conclusion can be used to indicate the comparison result of the advantages and disadvantages between the first execution result and the second execution result in the above labeling policy for this sample, and the model evaluation conclusion can be used to indicate the comparison result of the advantages and disadvantages between the first execution result and the second execution result in the sample by the third largest language model, the overall evaluation conclusion can also be used to indicate the comparison result of the advantages and disadvantages between the first execution result and the second execution result in this sample overall.

[0098] It should be noted that if the majority of the above labeled evaluation conclusion and the above model evaluation conclusion indicate that the first execution result in the sample is better than the second execution result, it can be determined that the overall evaluation conclusion corresponding to this sample is that the first execution result is better than the second execution result; if the majority of the labeled evaluation conclusion and the model evaluation conclusion indicate that the second execution result in the sample is better than the first execution result, it can be determined that the overall evaluation conclusion corresponding to this sample is that the second execution result is better than the first execution result; if half of the labeled evaluation conclusion and the model evaluation conclusion indicate that the first execution result in the sample is better than the second execution result, and the other half indicates that the second execution result in the sample is better than the first execution result, it can be determined that the overall evaluation conclusion corresponding to this sample is that the first execution result and the second execution result have the same effect.

[0099] Alternatively, corresponding weights may be set for the aforementioned annotation evaluation conclusions and the aforementioned model evaluation conclusions, respectively, and the total weight corresponding to the evaluation conclusion indicating that the first execution result in the aforementioned sample is better than the second execution result, and the total weight corresponding to the evaluation conclusion indicating that the second execution result in the aforementioned sample is better than the first execution result, may be calculated. If the former is greater than the latter, then the overall evaluation conclusion corresponding to the aforementioned sample may be determined to be that the first execution result is better than the second execution result; if the latter is greater than the former, then the overall evaluation conclusion corresponding to the aforementioned sample may be determined to be that the second execution result is better than the first execution result; if the two are the same, then the overall evaluation conclusion corresponding to the aforementioned sample may be determined to be that the first execution result and the second execution result have the same effect.

[0100] In some embodiments, the above-mentioned labeled evaluation conclusions and the model evaluation conclusions corresponding to the above-mentioned sample generated by each third largest language model can be used as parallel evaluation conclusions, and the total evaluation conclusion corresponding to the sample can be determined based on the comparison result of the number of evaluation conclusions indicating that the first execution result in the above-mentioned sample is better than the second execution result in these evaluation conclusions and the number of evaluation conclusions indicating that the second execution result in the sample is better than the first execution result; or, the total evaluation conclusion corresponding to the sample can be determined based on the comparison result of the total weight of the evaluation conclusions indicating that the first execution result in the above-mentioned sample is better than the second execution result in these evaluation conclusions and the total weight of the evaluation conclusions indicating that the second execution result in the sample is better than the first execution result.

[0101] In some embodiments, the model evaluation conclusions corresponding to the above-mentioned samples generated by each third language model can be first summarized into a total model evaluation conclusion corresponding to the sample, and then the above-mentioned labeled evaluation conclusions and the total model evaluation conclusion can be summarized into a total evaluation conclusion corresponding to the sample. Specifically, if the labeled evaluation conclusion and the total model evaluation conclusion both indicate that the first execution result in the above-mentioned sample is better than the second execution result, then it can be determined that the total evaluation conclusion corresponding to the sample is that the first execution result is better than the second execution result; if the labeled evaluation conclusion and the total model evaluation conclusion both indicate that the second execution result in the sample is better than the first execution result, then it can be determined that the total evaluation conclusion corresponding to the sample is that the second execution result is better than the first execution result; if one of the labeled evaluation conclusion and the total model evaluation conclusion indicates that the first execution result in the sample is better than the second execution result, and the other indicates that the second execution result in the sample is better than the first execution result, then it can be determined that the total evaluation conclusion corresponding to the sample is that the first execution result and the second execution result have the same effect. Alternatively, assuming that the labeled evaluation conclusion indicates that the first execution result in the above-mentioned sample is better than the second execution result, and the overall evaluation conclusion of the model indicates that the second execution result in the sample is better than the first execution result, then when the weight of the labeled evaluation conclusion is greater than the weight of the overall evaluation conclusion of the model, it can be determined that the overall evaluation conclusion corresponding to the sample is that the first execution result is better than the second execution result; when the weight of the overall evaluation conclusion of the model is greater than the weight of the labeled evaluation conclusion, it can be determined that the overall evaluation conclusion corresponding to the sample is that the second execution result is better than the first execution result; when the weight of the labeled evaluation conclusion is equal to the weight of the overall evaluation conclusion of the model, it can be determined that the overall evaluation conclusion corresponding to the sample is that the first execution result and the second execution result have the same effect.

[0102] In some embodiments, if the target task is an intelligent question-answering task, the total evaluation conclusion corresponding to the sample can be used as the evaluation conclusion corresponding to the intelligent question-answering task. At this time, the evaluation conclusion can reflect the model effects of the first language model and the second language model when performing the intelligent question-answering task. For example, assuming that the evaluation conclusion is "Answer 1 (i.e., the answer generated by the first language model) wins", it can be determined that the model effect of the first language model when performing the intelligent question-answering task is better than the model effect of the second language model when performing the intelligent question-answering task; and so on.

[0103] In order to avoid the randomness of the evaluation, multiple samples for evaluating the above-mentioned first language model and the above-mentioned second language model can be obtained, and the total evaluation conclusion corresponding to each of these samples can be obtained through the above-mentioned steps 202 to 208. If the total evaluation conclusions corresponding to the majority of these samples indicate that the above-mentioned first execution result is better than the above-mentioned second execution result, it can be determined that the model effect of the first language model when performing the above-mentioned intelligent question and answer task is better than the model effect of the second language model when performing the intelligent question and answer task; if the total evaluation conclusions corresponding to the majority of these samples indicate that the second execution result is better than the first execution result, it can be determined that the model effect of the second language model when performing the intelligent question and answer task is better than the model effect of the first language model when performing the intelligent question and answer task; if the total evaluation conclusions corresponding to half of these samples indicate that the first execution result is better than the second execution result, and the total evaluation conclusions corresponding to the other half indicate that the second execution result is better than the first execution result, it can be determined that the model effect of the second language model when performing the intelligent question and answer task is the same as the model effect of the first language model when performing the intelligent question and answer task.

[0104] In some embodiments, if the target task is a subtask split from an intelligent question-and-answer task, the overall evaluation conclusions corresponding to the samples of each subtask can be aggregated into the evaluation conclusion corresponding to the intelligent question-and-answer task. Specifically, the overall evaluation conclusion corresponding to the samples of the majority of subtasks can be determined as the evaluation conclusion corresponding to the intelligent question-and-answer task; alternatively, the overall evaluation conclusion corresponding to the samples of multiple subtasks with a greater total weight can be determined as the evaluation conclusion corresponding to the intelligent question-and-answer task.

[0105] Similar to the above content, when an evaluation conclusion corresponding to the above-mentioned intelligent question-answering task is obtained, the evaluation conclusion can be used to reflect the model effects of the above-mentioned first language model and the above-mentioned second language model when performing the intelligent question-answering task.

[0106] In some embodiments, if the above-mentioned first largest language model and the above-mentioned second largest language model adopt the RAG method, combined with information retrieval and model generation, to generate an answer text corresponding to the question text, then for a sample containing a question text, a first answer text generated by the first largest language model, and a second answer text generated by the second largest language model, the first answer text is specifically an answer text generated by the first largest language model based on the question text and the retrieved information (which may be called the first retrieval context) by reasoning, and the second answer text is specifically an answer text generated by the second largest language model based on the question text and the retrieved information (which may be called the second retrieval context) by reasoning.

[0107] In the above case, in addition to obtaining the annotation evaluation conclusion and model evaluation conclusion corresponding to the above sample through the above steps 202 to 208, it is also possible to calculate at least one RAG evaluation indicator corresponding to the first answer text and the second answer text in the sample based on the sample, the above first retrieval context and the above second retrieval context, and compare the at least one RAG evaluation indicator corresponding to the first answer text and the second answer text in the sample to generate a RAG evaluation conclusion corresponding to the sample based on the size comparison result.

[0108] It should be noted that the method of generating the RAG evaluation conclusion corresponding to the above-mentioned sample based on the size comparison result of the above-mentioned at least one RAG evaluation indicator can refer to the method of generating the model evaluation conclusion corresponding to the sample based on the size comparison result of the above-mentioned at least one task evaluation indicator, and this application will not go into details here.

[0109] For the above-mentioned samples, once the labeled evaluation conclusion corresponding to the sample, the model evaluation conclusion corresponding to the sample, and the RAG evaluation conclusion corresponding to the sample are obtained, the labeled evaluation conclusion, the model evaluation conclusion, and the RAG evaluation conclusion can be summarized into the overall evaluation conclusion corresponding to the sample.

[0110] It should be noted that if the majority of the above-mentioned annotation evaluation conclusions, the above-mentioned model evaluation conclusions (which can be the model evaluation conclusions corresponding to the sample generated by each third language model, or the total model evaluation conclusion corresponding to the sample obtained by summarizing) and the above-mentioned RAG evaluation conclusions indicate that the first answer text in the above-mentioned sample is better than the second answer text, then it can be determined that the total evaluation conclusion corresponding to the sample is that the first answer text is better than the second answer text; if the majority of the annotation evaluation conclusions, the model evaluation conclusions and the RAG evaluation conclusions indicate that the second answer text in the sample is better than the first answer text, then it can be determined that the total evaluation conclusion corresponding to the sample is that the second answer text is better than the answer text; if half of the annotation evaluation conclusions, the model evaluation conclusions and the RAG evaluation conclusions indicate that the first answer text in the sample is better than the second answer text, and the other half indicate that the second answer text in the sample is better than the first answer text, then it can be determined that the total evaluation conclusion corresponding to the sample is that the first answer text and the second answer text have the same effect.

[0111] Alternatively, corresponding weights may be set for the above-mentioned annotation evaluation conclusions, the above-mentioned model evaluation conclusions, and the above-mentioned RAG evaluation conclusions, respectively, and the total weight corresponding to the evaluation conclusion indicating that the first answer text in the above-mentioned sample is better than the second answer text, and the total weight corresponding to the evaluation conclusion indicating that the second answer text in the sample is better than the first answer text, may be calculated. If the former is greater than the latter, it may be determined that the overall evaluation conclusion corresponding to the sample is that the first answer text is better than the second answer text; if the latter is greater than the former, it may be determined that the overall evaluation conclusion corresponding to the sample is that the second answer text is better than the first answer text; if the two are the same, it may be determined that the overall evaluation conclusion corresponding to the sample is that the first answer text and the second answer text have the same effect.

[0112] In some embodiments, the at least one RAG evaluation indicator mentioned above may include one or more indicators shown below: fidelity for indicating the degree of factual consistency between the answer and the retrieval context; answer relevance for indicating the relevance between the answer and the question; context relevance for indicating the relevance between the retrieval context and the question; and context recall for indicating the degree of consistency between the retrieval context and the standard answer.

[0113] Specifically, fidelity can be used to measure the factual consistency between the answer and the retrieval context. Answer relevance aims to assess the relevance between the answer and the question. For example, the vector similarity between the embedding vectors corresponding to the answer and the question can be used as the relevance between the answer and the question. If the answer is incomplete or contains redundant information, it indicates that the relevance to the question is low. Contextual relevance can be used to measure whether real information that is more relevant to the question in the retrieval context is ranked higher.

[0114] It should be noted that, when there is no standard answer corresponding to the question, the at least one RAG evaluation metric may also include contextual recall, which can be used to measure the degree of consistency between the retrieval context and the standard answer that is considered to be the truth.

[0115] like Figure 4As shown, for the above-mentioned sample, the above-mentioned annotator can annotate the sample with the corresponding annotation evaluation conclusion based on the comparison results of the first answer text and the second answer text in the sample. The above-mentioned third language model can generate the model evaluation conclusion corresponding to the sample by respectively evaluating the relevance, accuracy, fluency, logic, professionalism, completeness, depth, practicality and other indicators of the first answer text and the second answer text in the sample. The RAG evaluation conclusion corresponding to the sample can also be generated by respectively evaluating the loyalty, answer relevance, context relevance and other indicators of the first answer text and the second answer text in the sample. In this case, assuming that the annotation evaluation conclusion corresponding to the sample is "Answer 2 wins", and the model evaluation conclusion and RAG evaluation conclusion corresponding to the sample are both "Answer 1 wins", then it can be determined that the overall model evaluation conclusion corresponding to the sample is "Answer 1 wins".

[0116] It should be noted that, to avoid model evaluation homogeneity, the aforementioned first language model, the aforementioned second language model, and each third language model are typically different large language models. However, in some cases, these large language models can also refer to the same large language model, that is, these large language models can also refer to different services provided by the same large language model. The specific configuration can be based on actual needs, and this application does not impose any special restrictions on this.

[0117] In the above technical solution, first, a sample for evaluating the first large language model and the second large language model can be obtained. The sample may include a prompt corresponding to the target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt. Then, on the one hand, a labeling evaluation conclusion for the sample can be obtained based on the comparison result of the first execution result and the second execution result in the sample by the labeling party. On the other hand, the sample can be input into at least one third large language model, so that each third large language model can compare the first execution result and the second execution result in the sample based on the sample, and generate a model evaluation conclusion corresponding to the sample based on the comparison result. Finally, the labeling evaluation conclusion and the model evaluation conclusion can be summarized into a total evaluation conclusion corresponding to the sample. Among them, the evaluation conclusion can be used to indicate the comparison result of the first execution result and the second execution result in the sample.

[0118] The above approach, on the one hand, enables the evaluation of large language models, making it possible to measure the model effectiveness of large language models, thereby ensuring the reliability, effectiveness, and fairness of large language models in practical applications. Moreover, it is particularly suitable for the evaluation of large language models when there are no standard answers corresponding to questions. On the other hand, multiple evaluation methods are used to evaluate large language models, and the evaluation conclusions obtained through these methods are summarized, which can improve the correctness and reliability of the final overall evaluation conclusion.

[0119] Corresponding to the aforementioned method embodiments, the present application also provides device embodiments.

[0120] Please refer to Figure 5 , Figure 5 : This is a structural diagram of a device shown in an exemplary embodiment of the present application. At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510, and of course may also include other required hardware. One or more embodiments of the present application can be implemented based on software, such as the processor 502 reading the corresponding computer program from the non-volatile memory 510 into the memory 508 and then running it. Of course, in addition to software implementation, one or more embodiments of the present application do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic module, and can also be hardware or logic devices.

[0121] Please refer to Figure 6 , Figure 6 It is a block diagram of an evaluation device for a large language model shown in an exemplary embodiment of the present application.

[0122] The above-mentioned large language model evaluation device can be applied to Figure 4 The device shown in the figure is used to implement the technical solution of this application. The device includes:

[0123] A sample acquisition module 602 acquires samples for evaluating the first large language model and the second large language model; wherein the samples include a prompt corresponding to a target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt;

[0124] The annotation acquisition module 604 acquires an annotation evaluation conclusion of the sample annotation by the annotator based on a comparison result of the first execution result and the second execution result;

[0125] A model evaluation module 606 inputs the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generates a model evaluation conclusion corresponding to the sample based on the comparison results.

[0126] The conclusion generation module 608 summarizes the annotation evaluation conclusion and the model evaluation conclusion into a total evaluation conclusion corresponding to the sample; wherein the evaluation conclusion is used to indicate a comparison result of the first execution result and the second execution result.

[0127] In some embodiments, the target task is an intelligent question-answering task; the prompt includes a question text; the first execution result includes a first answer text; and the second execution result includes a second answer text.

[0128] In some embodiments, the first answer text is generated by the first large language model through reasoning based on the question text and the first search context; the second answer text is generated by the second large language model through reasoning based on the question text and the second search context;

[0129] The device further comprises:

[0130] A RAG evaluation module, based on the sample, the first retrieval context, and the second retrieval context, calculates at least one RAG evaluation indicator corresponding to the first answer text and the second answer text, respectively, and compares the at least one RAG evaluation indicator corresponding to the first answer text and the second answer text, respectively, to generate a RAG evaluation conclusion corresponding to the sample according to the comparison result;

[0131] The summarizing of the annotation evaluation conclusion and the model evaluation conclusion into a total evaluation conclusion corresponding to the sample includes:

[0132] The annotation evaluation conclusion, the model evaluation conclusion and the RAG evaluation conclusion are summarized into an overall evaluation conclusion corresponding to the sample.

[0133] In some embodiments, the at least one RAG evaluation indicator includes one or more indicators shown below: fidelity for indicating the degree of factual consistency between the answer and the retrieval context; answer relevance for indicating the relevance between the answer and the question; context relevance for indicating the relevance between the retrieval context and the question; context recall for indicating the degree of consistency between the retrieval context and the standard answer.

[0134] In some embodiments, inputting the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generating a model evaluation conclusion corresponding to the sample based on the comparison results, includes:

[0135] The sample is input into at least one third-largest language model, so that each third-largest language model calculates at least one task evaluation indicator corresponding to the first execution result and the second execution result based on the sample, and compares the at least one task evaluation indicator corresponding to the first execution result and the second execution result, so as to generate a model evaluation conclusion corresponding to the sample based on the size comparison result.

[0136] In some embodiments, the target task is an intelligent question-answering task;

[0137] The at least one task evaluation indicator includes one or more indicators shown below: relevance, used to indicate whether the answer is directly related to the question; accuracy, used to indicate whether the answer contains factual errors; fluency, used to indicate whether the language expression of the answer is fluent; logic, used to indicate whether the logic of the answer is clear; professionalism, used to indicate whether the use of professional terms in the answer is accurate; completeness, used to indicate whether the answer completely answers the question; depth, used to indicate whether the answer explores the question in depth; practicality, used to indicate whether the answer has actual application value.

[0138] In some embodiments, the target task is a subtask split from the intelligent question-answering task.

[0139] In some embodiments, the apparatus further comprises:

[0140] The conclusion summary module summarizes the total evaluation conclusions corresponding to the samples of each subtask split from the intelligent question-answering task into an evaluation conclusion corresponding to the intelligent question-answering task.

[0141] In some embodiments, aggregating the annotation evaluation conclusion and the model evaluation conclusion into an overall evaluation conclusion corresponding to the sample includes:

[0142] Summarizing the model evaluation conclusions generated by each third language model into an overall model evaluation conclusion corresponding to the sample;

[0143] The labeled evaluation conclusion and the model overall evaluation conclusion are summarized into an overall evaluation conclusion corresponding to the sample.

[0144] For the device embodiments, they basically correspond to the method embodiments, so for relevant details, please refer to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the technical solution of this application.

[0145] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or physical devices, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0146] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0147] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0148] Computer-readable media include permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0149] It should be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0150] The above description is of specific embodiments of the present application. Other embodiments are within the scope of this application. In some cases, the actions or steps described in this application can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0151] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. The term "and / or" refers to and includes any or all possible combinations of one or more of the associated listed items.

[0152] The terms "one embodiment," "some embodiments," "example," "specific example," or "one implementation" used in one or more embodiments of the present application mean that the specific features or characteristics described in conjunction with the embodiment are included in at least one embodiment of the present application. The schematic descriptions of these terms do not necessarily refer to the same embodiment. Moreover, the specific features or characteristics described can be combined in an appropriate manner in one or more embodiments of the present application. In addition, different embodiments and specific features or characteristics in different embodiments can be combined without conflict.

[0153] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0154] The above description is merely a preferred embodiment of one or more embodiments of the present application and is not intended to limit one or more embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application shall be included in the scope of protection of one or more embodiments of the present application.

[0155] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

Claims

1. A method for evaluating a large language model, the method comprising: Obtaining a sample for evaluating the first large language model and the second large language model; wherein the sample includes a prompt corresponding to a target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt; Obtaining a labeling evaluation conclusion for the sample labeling by the labeling party based on a comparison result of the first execution result and the second execution result; Inputting the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generates a model evaluation conclusion corresponding to the sample according to the comparison results; The annotation evaluation conclusion and the model evaluation conclusion are summarized into a total evaluation conclusion corresponding to the sample; wherein the evaluation conclusion is used to indicate a comparison result of the advantages and disadvantages of the first execution result and the second execution result.

2. According to the method of claim 1, the target task is an intelligent question-answering task; the prompt includes a question text; the first execution result includes a first answer text; and the second execution result includes a second answer text.

3. The method according to claim 2, wherein the first answer text is generated by the first large language model through reasoning based on the question text and the first search context; The second answer text is generated by the second language model through reasoning based on the question text and the second retrieval context; The method further comprises: Based on the sample, the first retrieval context, and the second retrieval context, calculating at least one RAG evaluation indicator corresponding to the first answer text and the second answer text, respectively, and performing a size comparison on the at least one RAG evaluation indicator corresponding to the first answer text and the second answer text, respectively, to generate a RAG evaluation conclusion corresponding to the sample according to the size comparison result; The summarizing of the annotation evaluation conclusion and the model evaluation conclusion into a total evaluation conclusion corresponding to the sample includes: The annotation evaluation conclusion, the model evaluation conclusion and the RAG evaluation conclusion are summarized into an overall evaluation conclusion corresponding to the sample.

4. The method according to claim 3, wherein the at least one RAG evaluation indicator comprises one or more indicators shown below: fidelity for indicating the degree of factual consistency between the answer and the retrieval context; Answer relevance, which indicates the relevance of the answer to the question; Context relevance, which indicates the relevance between the retrieval context and the question; Context recall is used to indicate the degree of consistency between the retrieval context and the standard answer.

5. The method according to claim 1, wherein inputting the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generating a model evaluation conclusion corresponding to the sample based on the comparison results, comprises: The sample is input into at least one third-largest language model, so that each third-largest language model calculates at least one task evaluation indicator corresponding to the first execution result and the second execution result based on the sample, and compares the at least one task evaluation indicator corresponding to the first execution result and the second execution result, so as to generate a model evaluation conclusion corresponding to the sample based on the size comparison result.

6. The method according to claim 5, wherein the target task is an intelligent question-answering task; The at least one task evaluation metric includes one or more metrics shown below: relevance for indicating whether the answer is directly related to the question; Accuracy, which indicates whether the answer contains factual errors; Fluency, which indicates whether the answer is fluent in language expression; Logicality, used to indicate whether the logic of the answer is clear; Professionalism, used to indicate whether the use of professional terms in the answer is accurate; Completeness, used to indicate whether the answer fully answers the question; Used to indicate whether the answer explores the depth of the question in depth; used to indicate whether the answer has practical application value.

7. The method according to claim 1, wherein the target task is a subtask split from the intelligent question-answering task.

8. The method according to claim 7, further comprising: The total evaluation conclusions corresponding to the samples of each subtask split from the intelligent question-answering task are summarized as the evaluation conclusion corresponding to the intelligent question-answering task.

9. The method according to claim 1, wherein aggregating the annotation evaluation conclusion and the model evaluation conclusion into an overall evaluation conclusion corresponding to the sample comprises: Summarizing the model evaluation conclusions generated by each third language model into an overall model evaluation conclusion corresponding to the sample; The annotation evaluation conclusion and the model overall evaluation conclusion are summarized into an overall evaluation conclusion corresponding to the sample.

10. A large language model evaluation device, comprising: a sample acquisition module for acquiring samples for evaluating the first large language model and the second large language model; wherein the samples include a prompt corresponding to a target task, a first execution result obtained by the first large language model performing the target task based on the prompt, and a second execution result obtained by the second large language model performing the target task based on the prompt; a labeling acquisition module, which acquires a labeling evaluation conclusion of the sample labeling by the labeling party based on a comparison result of the first execution result and the second execution result; a model evaluation module, inputting the sample into at least one third language model, so that each third language model compares the first execution result and the second execution result based on the sample, and generates a model evaluation conclusion corresponding to the sample based on the comparison results; A conclusion generation module summarizes the annotation evaluation conclusion and the model evaluation conclusion into an overall evaluation conclusion corresponding to the sample; wherein the evaluation conclusion is used to indicate a comparison result of the advantages and disadvantages of the first execution result and the second execution result.

11. An electronic device comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 9 by running the executable instructions.

12. A computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.