Evaluation method, device and equipment, computer readable storage medium and program product
By using automated evaluation methods and pre-set prompts and similarity calculations, the problem of inconsistent evaluation results from large language models was solved, achieving efficient, accurate, and consistent evaluation results and improving evaluation efficiency.
Patent Information
- Application Number
- CN202410592773.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies for evaluating large language models lack uniformity and standardization, resulting in low accuracy and efficiency of evaluation results, inconsistent user script quality, and difficulty in achieving consistent and efficient evaluation.
By acquiring a test dataset, using preset prompts to instruct the large language model to respond, determining the dataset to be evaluated, and calculating the score through similarity, the large language model is automated for evaluation, ensuring consistency in evaluation results for the same model among different users.
The system enables automated evaluation of large language models, ensuring the consistency and accuracy of evaluation results, improving evaluation efficiency, and reducing resource waste and evaluation delays.
Smart Images

Figure CN120973656A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computers, and particularly, the present disclosure relates to a measurement and evaluation method, device, equipment, computer readable storage medium and program product. BACKGROUND
[0002] In the prior art, a large language model is an artificial intelligence model aiming to understand and generate human language. The large language model can be trained on a large amount of text data. The large language model can perform a wide range of tasks, such as text summarization, translation, sentiment analysis, etc. In the prior art, the performance of the large language model is measured and evaluated in an artificial manner. For example, different users evaluate the same large language model by using scripts written by themselves. Since the evaluation results of the same large language model by different users are quite different, the accuracy of the evaluation results of the large language model is also low, thereby resulting in low evaluation efficiency of the large language model. SUMMARY
[0003] The present disclosure proposes a measurement and evaluation method, device, equipment, computer readable storage medium and computer program product to solve the problem of how to improve the evaluation efficiency of the large language model, aiming at the shortcomings of the prior art.
[0004] In a first aspect, the present disclosure provides a measurement and evaluation method, comprising: obtaining a test data set, the test data set comprising a plurality of first test data and second test data corresponding to each first test data, the first test data being used to represent a question, and the second test data being used to represent an actual answer corresponding to the question represented by the first test data; determining a to-be-evaluated data set corresponding to the test data set by having the large language model to be evaluated reply to the question represented by any first test data in the test data set based on the first test data and a preset first prompt word, the to-be-evaluated data set comprising a plurality of to-be-evaluated data corresponding to the plurality of first test data in the test data set, the to-be-evaluated data corresponding to any first test data being used to represent a predicted answer corresponding to the question represented by any first test data, and the preset first prompt word comprising a first task instruction for the large language model to be evaluated, the first task instruction being used to instruct the large language model to be evaluated to output the to-be-evaluated data corresponding to any first test data; determining a similarity between any to-be-evaluated data and any second test data in the test data set by a target large language model based on the to-be-evaluated data, the second test data and a preset second prompt word, the preset second prompt word comprising a second task instruction for the target large language model, and the second task instruction being used to instruct the target large language model to output the similarity between any to-be-evaluated data and any second test data. determine a score for the large language model to be evaluated based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set.
[0005] In one embodiment, the test data set is obtained, including: obtaining evaluation indication information for the large language model to be evaluated, the evaluation indication information including a model identifier of the large language model to be evaluated and a storage address of the test data set in the cloud; obtaining the test data set from the cloud based on the storage address in the cloud.
[0006] In one embodiment, the evaluation indication information for the large language model to be evaluated is obtained, including: in response to an evaluation start operation on the evaluation option in the evaluation interface, initiating an evaluation task for the large language model to be evaluated based on a preset query rate per second, and generating evaluation indication information for the large language model to be evaluated; The evaluation option in the evaluation interface includes a model identifier of the large language model to be evaluated and a storage address in the cloud.
[0007] In one embodiment, based on any first test data in the test data set and a preset first prompt word, the question represented by any first test data is replied to by the large language model to be evaluated to determine the to-be-evaluated data set corresponding to the test data set, including: based on any first test data in the test data set and the model identifier of the large language model to be evaluated, determining a test task corresponding to any first test data in the original test data set; based on the test task and the preset first prompt word, the question represented by any first test data is replied to by the large language model to be evaluated to determine any to-be-evaluated data in the to-be-evaluated data set.
[0008] In one embodiment, based on the test task and the preset first prompt word, the question represented by any first test data is replied to by the large language model to be evaluated to determine any to-be-evaluated data in the to-be-evaluated data set, including: based on the model identifier of the large language model to be evaluated included in the test task, the large language model to be evaluated is called from the preset large language model set, and based on the preset first prompt word and any first test data included in the test task, the question represented by any first test data is replied to by the large language model to be evaluated to determine the to-be-evaluated data corresponding to any first test data, and the to-be-evaluated data set includes the to-be-evaluated data corresponding to any test data.
[0009] In one embodiment, based on the preset first prompt word and any first test data included in the test task, the question represented by any first test data is answered by the large language model to be evaluated to determine the to-be-evaluated data corresponding to any first test data, comprising: Write any first test data included in the test task into the preset first prompt word to obtain an updated first prompt word; The updated first prompt word is input into the large language model to be evaluated, and the question represented by any first test data is answered to determine the to-be-evaluated data corresponding to any first test data.
[0010] In one embodiment, based on any to-be-evaluated data in the to-be-evaluated data set, any second test data in the test data set and the preset second prompt word, the similarity between any to-be-evaluated data and any second test data is determined by the target large language model, comprising: Write any to-be-evaluated data in the to-be-evaluated data set and any second test data in the test data set into the preset second prompt word to obtain an updated second prompt word; The updated second prompt word is input into the target large language model, and the similarity between any to-be-evaluated data and any second test data is determined by similarity calculation processing.
[0011] In one embodiment, based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set, the score for the large language model to be evaluated is determined, comprising: Based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set, the conversion processing is performed by the target large language model to determine the score corresponding to each similarity; Based on the score corresponding to each similarity, the score for the large language model to be evaluated is determined.
[0012] In one embodiment, after determining the score for the large language model to be evaluated based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set, further comprising: Through the evaluation interface, at least one of the model identifier of the large language model to be evaluated, any to-be-evaluated data, the score for the large language model to be evaluated, the evaluation progress for the large language model to be evaluated, the evaluation state for the large language model to be evaluated, and the collection identifier of the test data set is displayed.
[0013] In a second aspect, the present disclosure provides an evaluation device, comprising: The first processing module is configured to obtain a test data set, the test data set including a plurality of first test data and second test data corresponding to each of the first test data, the first test data being used to represent a question, and the second test data being used to represent an actual answer corresponding to the question represented by the first test data; The second processing module is configured to determine a to-be-evaluated data set corresponding to the test data set by having a large language model to be evaluated reply to a question represented by any first test data in the test data set based on the first test data and a preset first prompt word, the to-be-evaluated data set including a plurality of to-be-evaluated data corresponding to the plurality of first test data in the test data set, the to-be-evaluated data corresponding to the any first test data being used to represent a predicted answer corresponding to the question represented by the any first test data, and the preset first prompt word including a first task instruction for the large language model to be evaluated, the first task instruction being used to instruct the large language model to be evaluated to output the to-be-evaluated data corresponding to the any first test data. The third processing module is configured to determine a similarity between any to-be-evaluated data in the to-be-evaluated data set and any second test data in the test data set based on the to-be-evaluated data and the second test data and a preset second prompt word by a target large language model, the preset second prompt word including a second task instruction for the target large language model, the second task instruction being used to instruct the target large language model to output the similarity between the any to-be-evaluated data and the any second test data. The fourth processing module is configured to determine a score for the large language model to be evaluated based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set.
[0014] In a third aspect, the present disclosure provides an electronic device, including a processor, a memory and a bus; The bus is configured to connect the processor and the memory. The memory is configured to store operation instructions. The processor is configured to execute the evaluation method of the first aspect of the present disclosure by invoking the operation instructions.
[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium storing a computer program, the computer program being used to execute the evaluation method of the first aspect of the present disclosure.
[0016] In a fifth aspect, the present disclosure provides a computer program product including a computer program, the computer program being executed by a processor to implement the steps of the evaluation method in the first aspect of the present disclosure.
[0017] The technical scheme provided by the embodiments of the present disclosure has at least the following beneficial effects: The technical scheme provided by the embodiments of the present disclosure has at least the following beneficial effects: BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required to be used in the description of the embodiments of the present disclosure will be briefly introduced.
[0019] Figure 1 The architecture schematic diagram of the evaluation system provided by the embodiments of the present disclosure is shown in the figure. Figure 2A flowchart of a measurement and evaluation method provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 3 A flowchart of a measurement and evaluation method provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 4 A flowchart of a measurement and evaluation method provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 5 A flowchart of a measurement and evaluation method provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 6 A flowchart of a measurement and evaluation method provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 7 A flowchart of a measurement and evaluation method provided by an embodiment of the present disclosure is shown in FIG. 1. Figure 8 A flowchart of a measurement and evaluation method provided by an embodiment of the present disclosure is shown in FIG. 1. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions of the technical solutions of the embodiments of the present disclosure, and do not limit the technical solutions of the embodiments of the present disclosure.
[0021] Those skilled in the art can understand that the singular forms "a", "an" and "the" used herein include plural forms unless specifically stated otherwise. It should be further understood that the terms "include" and "contain" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude other features, information, data, steps, operations, elements, components and / or their combinations supported by the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element establish a connection relationship through an intermediate element. In addition, "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates that at least one of the items defined by the term, for example, "A and / or B" indicates implementation as "A", or implementation as "B", or implementation as "A and B".
[0022] It can be understood that in the specific embodiments of the present disclosure, data related to measurement and evaluation is involved, and when the above embodiments of the present disclosure are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0023] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be described in further detail below with reference to the accompanying drawings.
[0024] The embodiment of the present disclosure is a kind of evaluation method provided by evaluation system, which relates to artificial intelligence, map and the like field.
[0025] Artificial intelligence (AI) is to use digital computer or digital computer controlled machine simulation, extension and extension of human intelligence, perception environment, knowledge acquisition and use knowledge to obtain the best results of theory, method, technology and application system.In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence.Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0026] Artificial intelligence technology is a comprehensive discipline, which involves a wide range of fields, both hardware and software level technology.Artificial intelligence basic technology generally includes such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies.Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, automatic driving, intelligent transportation and other several major directions.
[0027] Intelligent transportation system (ITS) is also called intelligent transportation system (Intelligent Transportation System), which is to effectively integrate advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) in transportation, service control and vehicle manufacturing, to strengthen the relationship between vehicles, roads and users, so as to form a comprehensive transportation system to ensure safety, improve efficiency, improve the environment and save energy.
[0028] In order to better understand and illustrate the scheme of the embodiment of the present disclosure, some technical terms involved in the embodiment of the present disclosure are simply explained as follows.
[0029] LLM: LLM (Large Language Model) is also known as a large language model, which is an artificial intelligence model designed to understand and generate human language; LLM can be trained on a large amount of text data, and LLM can perform a wide range of tasks such as text summarization, translation, sentiment analysis, etc.; LLM has changed the learning paradigm in various research fields and shown great potential in bridging the gap between classic recommenders and open-world knowledge; LLM has shown excellent capabilities such as problem solving, logical reasoning, creative writing, etc. with its huge model size and corpus size; LLM learns from a large amount of Internet text and encodes a large amount of open-world knowledge, i.e. from basic factual information to complex social norms and logical structures, so LLM can perform basic logical reasoning consistent with known facts and relationships.
[0030] Prompt: Prompt is a piece of text describing the task input by the user when interacting with the large language model.
[0031] GPT-4: GPT-4 is a large language model released for the chat robot ChatGPT.
[0032] QPS: QPS (Queries Per Second) represents the number of evaluation tasks that the evaluation system can handle per second; for example, if the evaluation system can handle 5 evaluation tasks per second, then QPS is 5.
[0033] The existing technology has problems such as: (1) Since each user needs to maintain their own scripts independently, this not only increases the user's workload, but also leads to inconsistent script quality; the reliability and evaluation efficiency of the script vary from user to user, which may affect the accuracy of the evaluation results of the large language model.
[0034] (2) Lack of unified evaluation visualization tools, making it difficult to intuitively present the evaluation results of the large language model to users, and unable to fully determine the subtle differences in the performance of the large language model.
[0035] (3) Lack of standardized evaluation system.
[0036] It should be noted that a standardized evaluation system can provide all users with a unified evaluation tool and evaluation process, thereby improving the evaluation efficiency of the large language model, ensuring the consistency of the evaluation results, and promoting the sharing of knowledge and best practices.
[0037] (4) Due to a lack of coordination, different users may conduct repeated evaluations of the same large language model, which not only wastes evaluation resources but may also delay the evaluation work progress, resulting in low evaluation efficiency.
[0038] Based on this, the present disclosure provides an evaluation method, apparatus, device, computer-readable storage medium, and program product, and the specific technical solutions will be described below.
[0039] The solutions provided in this disclosure relate to artificial intelligence technology. The technical solutions of this disclosure will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.
[0040] To better understand the solution provided in this disclosure, the solution will be described below in conjunction with a specific application scenario.
[0041] In one embodiment, Figure 1 The diagram shows an architecture schematic of an evaluation system applicable to embodiments of this disclosure. It is understood that the evaluation methods provided in these embodiments can be applied to, but are not limited to, applications such as... Figure 1 In the application scenarios shown.
[0042] In this example, as Figure 1 As shown, the architecture of the evaluation system in this example may include, but is not limited to, server 10, server 20, terminal 30, and database 40. Server 10, server 20, terminal 30, and database 40 can interact with each other via network 50. For example, server 10 serves as the evaluation backend service, while server 20 and terminal 30 constitute the frontend platform.
[0043] The server 10 acquires a test data set, the test data set including a plurality of first test data and second test data corresponding to each first test data, the first test data being used to represent a question, and the second test data being used to represent an actual answer corresponding to the question represented by the first test data; the server 10 determines a to-be-evaluated data set corresponding to the test data set by causing the large language model to be evaluated to reply to the question represented by any first test data in the test data set based on the first test data and a preset first prompt word, the to-be-evaluated data set including a plurality of to-be-evaluated data corresponding to the plurality of first test data in the test data set, the to-be-evaluated data corresponding to any first test data being used to represent an estimated answer corresponding to the question represented by the first test data, and the preset first prompt word including a first task instruction for the large language model to be evaluated, the first task instruction being used to instruct the large language model to be evaluated to output the to-be-evaluated data corresponding to any first test data; the server 10 determines a similarity between any to-be-evaluated data and any second test data in the test data set based on the to-be-evaluated data, the second test data, and a preset second prompt word by using a target large language model, the preset second prompt word including a second task instruction for the target large language model, the second task instruction being used to instruct the target large language model to output the similarity between any to-be-evaluated data and any second test data; the server 10 determines a score for the large language model to be evaluated based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set; the server 10 sends the score for the large language model to be evaluated to the server 20, and the server 20 displays, through a test interface in the terminal 30, a model identifier of the large language model to be evaluated, any to-be-evaluated data, the score for the large language model to be evaluated, a test progress of the large language model to be evaluated, a test state of the large language model to be evaluated, a set identifier of the test data set, and the like; and the server 10 stores the score for the large language model to be evaluated in the database 40.
[0044] It can be understood that the above is only an example, and the present embodiment is not limited thereto.
[0045] The terminal includes, but is not limited to, a smart phone (such as an Android phone, an iOS phone, etc.), a mobile phone simulator, a tablet computer, a notebook computer, a digital broadcast receiver, a MID (Mobile Internet Device), a PDA (Personal Digital Assistant), a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc.
[0046] The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server or a server cluster providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.
[0047] Cloud computing is a computing mode that distributes computing tasks on a resource pool composed of a large number of computing devices, so that various application systems can obtain computing power, storage space and information services according to needs. The network providing resources is called "cloud". The resources in the "cloud" are infinitely expandable to users and can be obtained at any time, used on demand, expanded at any time, and paid according to use.
[0048] As a basic capability provider of cloud computing, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established, and various types of virtual resources are deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, network devices.
[0049] According to logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer is deployed on the PaaS layer, or the SaaS is directly deployed on the IaaS. PaaS is a platform for software running, such as databases, web containers, etc. SaaS is various business software, such as web portal websites, SMS mass senders, etc. Generally, SaaS and PaaS are upper layers relative to IaaS.
[0050] The so-called artificial intelligence cloud service is generally also referred to as AIaaS (AI as a Service, which is “AI as a Service” in Chinese). This is a current mainstream service method for artificial intelligence platforms. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI-themed mall: all developers can access and use one or more artificial intelligence services provided by the platform through the API interface. Some senior developers can also use the AI frameworks and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services.
[0051] The above network can include but is not limited to: wired networks, wireless networks. Among them, the wired network includes: local area networks, metropolitan area networks, and wide area networks, and the wireless network includes: Bluetooth, Wi-Fi, and other networks that implement wireless communication. Specifically, it can also be determined based on the actual application scenario requirements and is not limited here.
[0052] See Figure 2 , Figure 2 shows a schematic flowchart of an evaluation method provided by an embodiment of the present disclosure. Among them, this method can be executed by any electronic device, such as an evaluation system, etc.; as an optional implementation, this method can be executed by an evaluation system. For the convenience of description, in the description of some optional embodiments below, the evaluation system will be used as an example of the execution subject of this method for illustration. As Figure 2 shown, the evaluation method provided by an embodiment of the present disclosure includes the following steps: S201, obtain a test data set, the test data set includes multiple first test data and second test data corresponding to each first test data. The first test data is used to represent a problem, and the second test data is used to represent the actual answer corresponding to the problem represented by the first test data.
[0053] Specifically, the test data set includes multiple first test data and multiple second test data. The first test data is used to represent a problem, and the second test data is used to represent the actual answer. One first test data among the multiple first test data and one second test data among the multiple second test data correspond to each other, forming a (problem, actual answer) test data pair. For example, the first test data is “Which city is the capital of Hubei?”; the second test data is “Wuhan”; among them, “Which city is the capital of Hubei?” is the problem, and the actual answer is “Wuhan”. Another example of the first test data is “Translate ‘how are you?’ into Chinese”; the second test data is “你好吗?”; among them, “Translate ‘how are you?’ into Chinese” is the problem, and the actual answer is “你好吗?”.
[0054] S202, based on any first test data in the test data set and the preset first prompt word, reply to the question represented by any first test data through the large language model to be evaluated, determine the data set to be evaluated corresponding to the test data set, the data set to be evaluated includes a plurality of data to be evaluated corresponding to a plurality of first test data in the test data set, and the data to be evaluated corresponding to any first test data is used to represent the estimated answer corresponding to the question represented by any first test data. The preset first prompt word includes a first task instruction for the large language model to be evaluated, and the first task instruction is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data.
[0055] Specifically, the large language model to be evaluated is LLM, the preset first prompt word includes a first task instruction for the large language model to be evaluated, and the first task instruction includes a task description, which is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data. For example, the task description is: please provide the corresponding data to be evaluated (estimated answer) for the first test data (question). The first test data is used to represent the question, and the data to be evaluated is used to represent the estimated answer corresponding to the question; the first test data is, for example, "Which city is the capital of Hubei?"; the data to be evaluated is, for example, "Wuhan"; wherein "Which city is the capital of Hubei?" is a question, and the estimated answer is "Wuhan". For example, write the first test data "Which city is the capital of Hubei?" into the preset first prompt word, that is, write "Which city is the capital of Hubei?" into the task description of the preset first prompt word, to obtain an updated first prompt word, and the task description of the updated first prompt word is: please provide the corresponding data to be evaluated (estimated answer) for "Which city is the capital of Hubei?"; input the updated first prompt word into the large language model to be evaluated, and the large language model to be evaluated replies to "Which city is the capital of Hubei?" and outputs the data to be evaluated (estimated answer) "Wuhan".
[0056] S203, based on any first test data in the test data set and the preset first prompt word, reply to the question represented by any first test data through the large language model to be evaluated, determine the data set to be evaluated corresponding to the test data set, the data set to be evaluated includes a plurality of data to be evaluated corresponding to a plurality of first test data in the test data set, and the data to be evaluated corresponding to any first test data is used to represent the estimated answer corresponding to the question represented by any first test data. The preset first prompt word includes a first task instruction for the large language model to be evaluated, and the first task instruction is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data.
[0057] Specifically, the preset second prompt word includes a second task instruction for the target large language model, the second task instruction includes a task description, the task description is used to instruct the target large language model to output a similarity between any to-be-evaluated data and any second test data, and a score for any to-be-evaluated data, i.e., a score for the large language model to be evaluated. For example, the task description is: please determine the similarity between the to-be-evaluated data (estimated answer) and the second test data (actual answer), and determine the score for the to-be-evaluated data based on the similarity. The to-be-evaluated data is used to represent the estimated answer corresponding to the question, and the second test data is used to represent the actual answer. For example, the to-be-evaluated data is “Ezhou”, and the second test data is “Wuhan”, wherein the estimated answer is “Ezhou”, and the actual answer is “Wuhan”; the similarity between the to-be-evaluated data “Ezhou” and the second test data “Wuhan” is 0. For example, the to-be-evaluated data is “Wuhan”, and the second test data is “Wuhan”, wherein the estimated answer is “Wuhan”, and the actual answer is “Wuhan”; the similarity between the to-be-evaluated data “Wuhan” and the second test data “Wuhan” is 1. For example, the to-be-evaluated data “Ezhou” and the second test data “Wuhan” are written into the preset second prompt word, i.e., the to-be-evaluated data “Ezhou” and the second test data “Wuhan” are written into the task description of the preset second prompt word, to obtain an updated second prompt word, and the task description of the updated second prompt word is: please determine the similarity between “Ezhou” and “Wuhan”, and determine the score for the to-be-evaluated data based on the similarity; the updated second prompt word is input into the target large language model, the target large language model performs similarity calculation processing, determines that the similarity between “Ezhou” and “Wuhan” is 0, and based on the similarity, determines that the score for the to-be-evaluated data is 0, i.e., the score for the large language model to be evaluated is 0.
[0058] In S204, based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set, a score for the large language model to be evaluated is determined.
[0059] Specifically, the to-be-evaluated data set includes a plurality of to-be-evaluated data, each to-be-evaluated data in the plurality of to-be-evaluated data corresponds to a similarity, and each similarity corresponds to a score for the to-be-tested large language model. For example, the to-be-evaluated data set includes 10 to-be-evaluated data, and the similarities corresponding to the 10 to-be-evaluated data are 0.1, 0.3, 0.2, 0.9, 0.8, 0.5, 0.3, 0.6, 0.2, and 1 respectively. The score value range is [0, 100], and the target large language model is converted and processed to determine that the scores corresponding to the similarities 0.1, 0.3, 0.2, 0.9, 0.8, 0.5, 0.3, 0.6, 0.2, and 1 are 10, 30, 20, 90, 80, 50, 30, 60, 20, and 10 respectively, that is, the scores for the to-be-tested large language model are 10, 30, 20, 90, 80, 50, 30, 60, 20, and 10, and the scores 10, 30, 20, 90, 80, 50, 30, 60, 20, and 10 constitute a score set for the to-be-tested large language model.
[0060] In the embodiments of the present disclosure, a test data set is obtained, the test data set includes a plurality of first test data and second test data corresponding to each first test data, the first test data is used to represent a question, and the second test data is used to represent an actual answer corresponding to the question represented by the first test data; based on any first test data in the test data set and a preset first prompt word, the question represented by any first test data is replied by a large language model to be evaluated, and a data set to be evaluated corresponding to the test data set is determined, the data set to be evaluated includes a plurality of data to be evaluated corresponding to the plurality of first test data in the test data set, and the data to be evaluated corresponding to any first test data is used to represent an estimated answer corresponding to the question represented by any first test data, the preset first prompt word includes a first task instruction for the large language model to be evaluated, and the first task instruction is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data; based on any data to be evaluated in the data set to be evaluated, any second test data in the test data set and a preset second prompt word, a similarity between any data to be evaluated and any second test data is determined by a target large language model, the preset second prompt word includes a second task instruction for the target large language model, and the second task instruction is used to instruct the target large language model to output the similarity between any data to be evaluated and any second test data; based on the similarity corresponding to each data to be evaluated in the data set to be evaluated, a score for the large language model to be evaluated is determined; in this way, the large language model to be evaluated is instructed by the preset first prompt word to reply to the first test data (question) and obtain the data to be evaluated (estimated answer), the target large language model is instructed by the preset second prompt word to determine the similarity between the data to be evaluated (estimated answer) and the second test data (actual answer), and the data to be evaluated (estimated answer) is scored based on the similarity, that is, the large language model to be evaluated is scored, and the evaluation result for the large language model to be evaluated is obtained, thereby realizing the automatic evaluation of the large language model to be evaluated; through the automatic evaluation, different users evaluate the same large language model to be evaluated, and the same evaluation result will be obtained, ensuring the consistency of the evaluation result and improving the accuracy of the evaluation result, thereby improving the evaluation efficiency of the large language model to be evaluated.
[0061] In one embodiment, a test data set is obtained, including steps A1-A2: Step A1, obtaining evaluation instruction information for the large language model to be evaluated, the evaluation instruction information including a model identifier of the large language model to be evaluated and a storage address of a cloud end of the test data set.
[0062] Specifically, for example, as Figure 3As shown, the front-end platform in the evaluation system generates evaluation instruction information for the large language model to be evaluated, that is, the evaluation system obtains the evaluation instruction information for the large language model to be evaluated; the front-end platform sends the evaluation instruction information for the large language model to be evaluated to the evaluation background service in the evaluation system, and the evaluation background service receives the evaluation instruction information for the large language model to be evaluated.
[0063] Step A2, based on the cloud storage address, obtain the test data set from the cloud.
[0064] Specifically, for example, as shown in Figure 3 As shown, the evaluation background service in the evaluation system analyzes and processes the evaluation instruction information for the large language model to be evaluated, and obtains the model identifier of the large language model to be evaluated, the cloud storage address of the test data set, etc.; the evaluation background service obtains the test data set from the cloud based on the cloud storage address of the test data set, and the cloud is, for example Figure 3 cloud storage as shown in
[0065] In one embodiment, the evaluation instruction information for the large language model to be evaluated is obtained, including: In response to the evaluation start operation of the evaluation option in the evaluation interface, initiate the evaluation task for the large language model to be evaluated based on the preset query rate per second, and generate the evaluation instruction information for the large language model to be evaluated; The evaluation option in the evaluation interface includes the model identifier of the large language model to be evaluated and the cloud storage address.
[0066] Specifically, the evaluation interface of the front-end platform in the evaluation system, for example Figure 4As shown, the evaluation interface includes a plurality of evaluation options, such as uploading a test data set, a test data set address, a Venus service ID (IdentityDocument), a base model selection, a QPS option, a Venus appGroupID, a GPT4 option, etc. The evaluation option "upload test data set" allows the user to upload a custom test data set to the cloud, i.e., the user uploads the test data set to the cloud for storage through the front-end platform. The test data set name is the collection identifier of the test data set. The user inputs a certain test data set address in the evaluation option "test data set address". The test data set address is the storage address of the test data set in the cloud, and the evaluation background service in the evaluation system can obtain the test data set from the cloud. The Venus service ID is the ID of the user-defined large language model. The user inputs the ID of a certain custom large language model in the evaluation option "Venus service ID", and then uses the custom large language model as the large language model to be evaluated. The server ̠ id represents the ID of the user-defined large language model. The base model is a general large language model. Under the evaluation option "base model selection", a plurality of large language model options are set, such as large language model A, large language model B, large language model C, large language model D, large language model E, and large language model F. The user selects a certain large language model from the plurality of large language models, and then uses the large language model as the large language model to be evaluated. The QPS option is, for example, "the default value of QPS is 2. Please decide whether to modify according to the situation (it is recommended not to modify)". If the user does not modify the QPS option, the QPS is 2. If the user modifies the QPS option, i.e., the user customizes the QPS, the QPS in the QPS option is the user-modified QPS, e.g., the user modifies the QPS to 5, i.e., the QPS in the QPS option is 5. The Venus appGroupID and the Venus service ID are related and correspond to each other. The Venus appGroupID represents the ID of a legitimate user who can call the custom large language model. The ID of the custom large language model is the Venus service ID. The Venus appGroupID is, for example, 876. The GPT4 option is, for example, "whether to use GPT4 to evaluate the result (if not needed, please do not select)". If the user clicks the button to use GPT4 to evaluate the result, the user can input a custom GPT4 prompt in the option box corresponding to the GPT4 option. The custom GPT4 prompt is a preset second prompt. The result refers to any data to be evaluated. The user clicks the "start test" button to initiate an evaluation task for the large language model to be evaluated based on a preset query rate (QPS), and generates evaluation instruction information for the large language model to be evaluated.
[0067] It should be noted that the evaluation system employs an efficient scheduling mechanism to ensure that each evaluation task is processed promptly and accurately. For example, the system initiates each of the multiple evaluation tasks sequentially according to a pre-set QPS (Queries Per Second). This scheduling mechanism is not only efficient but also ensures the stability of the evaluation system, avoiding potential problems caused by overload. For instance, if the QPS is 2, the system initiates 2 evaluation tasks per second, each including one test task and one evaluation task, thus ensuring that both the test and evaluation tasks within each evaluation task are processed efficiently. Timely and accurate processing is ensured. The testing task includes determining any data to be evaluated in the dataset using the large language model to be evaluated. The evaluation task includes determining the score for the large language model to be evaluated using the target large language model. The target large language model is, for example, GPT-4. There are four evaluation tasks, including four testing tasks and four evaluation tasks. These four evaluation tasks constitute the GPT-4 evaluation task set. The GPT-4 evaluation task set is initiated, and based on an efficient scheduling mechanism, it is ensured that each of the four evaluation tasks can be evaluated in a timely and accurate manner.
[0068] For example, such as Figure 4 As shown, after the user clicks, selects, or inputs multiple assessment options on the assessment interface, the user clicks the "Start Test" button, which initiates the assessment for the assessment option on the assessment interface. In response to this assessment initiation, the front-end platform in the assessment system launches an assessment task for the large language model to be assessed based on a preset QPS, and generates assessment instruction information for the large language model to be assessed. This assessment instruction information includes the model identifier of the large language model to be assessed and the cloud storage address of the test dataset. The front-end platform then sends the assessment instruction information to the assessment back-end service in the assessment system.
[0069] In one embodiment, based on any first test data in the test dataset and a preset first prompt word, the question represented by any first test data is answered using the large language model to be evaluated, thereby determining the dataset to be evaluated corresponding to the test dataset, including steps B1-B2: Step B1: Based on any first test data in the test dataset and the model identifier of the large language model to be evaluated, determine the test task corresponding to any first test data in the original test dataset.
[0070] Specifically, a test task is constructed based on a first test data in the test dataset and the model identifier of the large language model to be evaluated. The test task includes the first test data and the model identifier of the large language model to be evaluated.
[0071] The application scenarios of the large language model to be evaluated are such as text generation, question answering system, translation, abstract generation, image description, code generation, sentiment analysis, dialogue system, etc.
[0072] The application scenario is text generation. The first test data is, for example, "Please write a story about the school library", and the data to be evaluated output by the large language model to be evaluated is, for example, a content description of a story about the school library.
[0073] The application scenario is the question answering system. The first test data is, for example, "Which city is the capital of France?", and the data to be evaluated output by the large language model to be evaluated is, for example, France.
[0074] The application scenario is translation. The first test data is, for example, "Translate 'how are you?' into Chinese", and the data to be evaluated output by the large language model to be evaluated is, for example, "你好吗?".
[0075] The application scenario is abstract generation. The first test data is, for example, "Please summarize the main points of the debate", and the data to be evaluated output by the large language model to be evaluated is, for example, a content description of the main points of the debate.
[0076] The application scenario is image description. The first test data is, for example, "Please describe this picture", and the data to be evaluated output by the large language model to be evaluated is, for example, a content description of the picture.
[0077] The application scenario is code generation. The first test data is, for example, "Please write a function to calculate the factorial of a number", and the data to be evaluated output by the large language model to be evaluated is, for example, a certain function.
[0078] The application scenario is sentiment analysis. The first test data is, for example, "Please provide the sentiment analysis of the product review", and the data to be evaluated output by the large language model to be evaluated is, for example, the result of the sentiment analysis of the product review.
[0079] The application scenario is the dialogue system. The first test data is, for example, "You are a robot that recommends movies. Please recommend a comedy movie", and the data to be evaluated output by the large language model to be evaluated is, for example, Comedy Movie B.
[0080] Step B2, based on the test task and the preset first prompt, use the large language model to be evaluated to reply to the question represented by any first test data, and determine any data to be evaluated in the set of data to be evaluated.
[0081] Specifically, write any first test data included in the test task into a preset first prompt to obtain an updated first prompt; input the updated first prompt into the large language model to be evaluated to answer the question represented by any first test data, and determine the data to be evaluated corresponding to any first test data.
[0082] The application scenarios of the large language model to be evaluated include, for example, text generation, question answering systems, translation, abstract generation, image description, code generation, sentiment analysis, dialogue systems, etc.
[0083] The application scenario is text generation. For example, write the first test data "Please write a story about the school library" into a preset first prompt to obtain an updated first prompt; input the updated first prompt into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "Content description of a story about the school library".
[0084] The application scenario is a question answering system. For example, write the first test data "Which city is the capital of France?" into a preset first prompt to obtain an updated first prompt; input the updated first prompt into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "Paris".
[0085] The application scenario is translation. For example, write the first test data "Translate 'how are you?' into Chinese" into a preset first prompt to obtain an updated first prompt; input the updated first prompt into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "你好吗?".
[0086] The application scenario is abstract generation. For example, write the first test data "Please summarize the main points of the debate" into a preset first prompt to obtain an updated first prompt; input the updated first prompt into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "Content description of the main points of the debate".
[0087] The application scenario is image description. For example, write the first test data "Please describe this picture" into a preset first prompt to obtain an updated first prompt; input the updated first prompt into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "Content description of the picture".
[0088] The application scenario is code generation. For example, the first test data "please write a function to calculate the factorial of a number" is written into the preset first prompt word to obtain an updated first prompt word. The updated first prompt word is input into the large language model to be evaluated, and the problem represented by the first test data is replied to. The large language model to be evaluated outputs the evaluation data "a certain function".
[0089] The application scenario is sentiment analysis. For example, the first test data "please provide sentiment analysis of product reviews" is written into the preset first prompt word to obtain an updated first prompt word. The updated first prompt word is input into the large language model to be evaluated, and the problem represented by the first test data is replied to. The large language model to be evaluated outputs the evaluation data "sentiment analysis result of product reviews".
[0090] The application scenario is a dialogue system. For example, the first test data "you are a movie recommendation robot, please recommend a comedy movie" is written into the preset first prompt word to obtain an updated first prompt word. The updated first prompt word is input into the large language model to be evaluated, and the problem represented by the first test data is replied to. The large language model to be evaluated outputs the evaluation data "comedy movie B".
[0091] In one embodiment, based on the test task and the preset first prompt word, the problem represented by any first test data is replied to by the large language model to be evaluated, and any evaluation data in the evaluation data set is determined, including: Based on the model identifier of the large language model to be evaluated included in the test task, the large language model to be evaluated is called from the preset large language model set, and based on the preset first prompt word and any first test data included in the test task, the problem represented by any first test data is replied to by the large language model to be evaluated. Determine the evaluation data corresponding to any first test data, and the evaluation data set includes the evaluation data corresponding to any test data.
[0092] Specifically, the preset large language model set includes a plurality of large language models and a model identifier of each large language model in the plurality of large language models. For example, if the model identifier A of the large language model to be evaluated is included in the test task, the large language model corresponding to the model identifier A is queried from the preset large language model set, and the large language model corresponding to the model identifier A is determined as the large language model to be evaluated.
[0093] For example, if the evaluation system initiates multiple evaluation tasks, each evaluation task in the multiple evaluation tasks includes a test task and an evaluation task, that is, the multiple evaluation tasks correspond to multiple test tasks; write the first test data included in a certain test task among the multiple test tasks into a preset first prompt word to obtain an updated first prompt word; input the updated first prompt word into the large language model to be evaluated to answer the question represented by the first test data, determine the data to be evaluated corresponding to the first test data, that is, the multiple test tasks correspond to multiple data to be evaluated, and construct the multiple data to be evaluated into a data set to be evaluated.
[0094] In one embodiment, based on a preset first prompt word and any first test data included in a test task, using the large language model to be evaluated to answer the question represented by any first test data, and determining the data to be evaluated corresponding to any first test data includes: Write any first test data included in the test task into a preset first prompt word to obtain an updated first prompt word; Input the updated first prompt word into the large language model to be evaluated, and by answering the question represented by any first test data, determine the data to be evaluated corresponding to any first test data.
[0095] Specifically, the application scenario of the large language model to be evaluated is text generation. For example, write the first test data "Please write a story about the school library" into a preset first prompt word to obtain an updated first prompt word; input the updated first prompt word into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "A content description of a story about the school library".
[0096] The application scenario of the large language model to be evaluated is a question-and-answer system. For example, write the first test data "What is the capital city of France?" into a preset first prompt word to obtain an updated first prompt word; input the updated first prompt word into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "Paris".
[0097] The application scenario of the large language model to be evaluated is translation. For example, write the first test data "Translate 'how are you?' into Chinese" into a preset first prompt word to obtain an updated first prompt word; input the updated first prompt word into the large language model to be evaluated to answer the question represented by the first test data, and the large language model to be evaluated outputs the data to be evaluated "你好吗?".
[0098] The application scenario of the to-be-evaluated large language model is image description, for example, the first test data "please describe this picture" is written into the preset first prompt word to obtain an updated first prompt word; the updated first prompt word is input into the to-be-evaluated large language model, and a question represented by the first test data is replied, and the to-be-evaluated large language model outputs the to-be-evaluated data "content description of the picture".
[0099] The application scenario of the to-be-evaluated large language model is image description, for example, the first test data "please describe this picture" is written into the preset first prompt word to obtain an updated first prompt word; the updated first prompt word is input into the to-be-evaluated large language model, and a question represented by the first test data is replied, and the to-be-evaluated large language model outputs the to-be-evaluated data "content description of the picture".
[0100] The application scenario of the to-be-evaluated large language model is code generation, for example, the first test data "please write a function to calculate the factorial of a number" is written into the preset first prompt word to obtain an updated first prompt word; the updated first prompt word is input into the to-be-evaluated large language model, and a question represented by the first test data is replied, and the to-be-evaluated large language model outputs the to-be-evaluated data "a certain function".
[0101] The application scenario of the to-be-evaluated large language model is sentiment analysis, for example, the first test data "please provide a sentiment analysis of product reviews" is written into the preset first prompt word to obtain an updated first prompt word; the updated first prompt word is input into the to-be-evaluated large language model, and a question represented by the first test data is replied, and the to-be-evaluated large language model outputs the to-be-evaluated data "sentiment analysis result of product reviews".
[0102] The application scenario of the to-be-evaluated large language model is a dialogue system, for example, the first test data "you are a movie recommendation robot, please recommend a comedy movie" is written into the preset first prompt word to obtain an updated first prompt word; the updated first prompt word is input into the to-be-evaluated large language model, and a question represented by the first test data is replied, and the to-be-evaluated large language model outputs the to-be-evaluated data "comedy movie B".
[0103] It should be noted that the prompt splicing process of the to-be-evaluated large language model includes: writing the first test data into the preset first prompt to obtain an updated first prompt. The preset first prompt is a pre-designed and optimized prompt template, so as to ensure that the first test data is correctly integrated into the prompt required by the to-be-evaluated large language model, that is, the first test data is written into the preset first prompt to obtain the updated first prompt. The prompt splicing process of the to-be-evaluated large language model ensures seamless connection between the first test data and the to-be-evaluated large language model, and creates conditions for obtaining accurate test results.
[0104] In one embodiment, based on any to-be-evaluated data in the to-be-evaluated data set, any second test data in the test data set, and a preset second prompt, the similarity between any to-be-evaluated data and any second test data is determined by a target large language model, including: writing any to-be-evaluated data in the to-be-evaluated data set and any second test data in the test data set into the preset second prompt to obtain an updated second prompt; inputting the updated second prompt into the target large language model to determine the similarity between any to-be-evaluated data and any second test data through similarity calculation processing.
[0105] Specifically, the test data set includes a plurality of first test data and a plurality of second test data, the first test data is used to represent a question, the second test data is used to represent an actual answer, one first test data in the plurality of first test data and one second test data in the plurality of second test data correspond to each other to form a (question, actual answer) test data pair. For example, the first test data is “Which city is the capital of France?”, the second test data corresponding to the first test data is “Paris”, the to-be-evaluated data and the second test data “Paris” are written into the preset second prompt to obtain an updated second prompt; the updated second prompt is input into the target large language model, and the target large language model determines the similarity between the to-be-evaluated data and the second test data “Paris” through similarity calculation processing; if the to-be-evaluated data is “Paris”, the similarity between the to-be-evaluated data “Paris” and the second test data “Paris” is 1; the target large language model converts the similarity into a corresponding score.
[0106] The target large language model is, for example, GPT4, and the preset second prompt is, for example: "{\"appGroupId\":1362,\"messages\":[{\"role\":\"system\",\"content\":[{\"type\":\"text\",\"text\":\"You are a very smart and wise chatbot that is good at assessing whether a paragraph of machine output is consistent with a paragraph of actual output. Your task is to compare the text description of the machine output with the actual text description and determine whether they are consistent. You can complete this task in the following way:\\n1. Pay attention to whether the content described by the machine output and the actual output is the same.\\n2. Synonyms or similar expressions should be considered as consistent descriptions. \"}]},{\"role\":\"user\",\"content\":[{\"type\":\"text\",\"text\":\"Please evaluate whether the machine output description is consistent with the actual description. The machine output content will be given with the keyword \\\"Predicted Description\\\". The actual description will be given with \\\"Correct Description\\\". As follows:\\nPredicted Description:%s\\nCorrect Description:%s\\n\\nPlease evaluate the consistency of Predicted Description and Correct Description. Your evaluation result should be an integer with a value of 0 to 100. 0 means that the machine output description and the actual description content are not related at all, and 100 means that the machine output content and the actual description content are completely consistent in meaning. Note that you need to give it in the format of a Python dictionary, with the key \\\"score\\\" and the value of your evaluation result. You are prohibited from outputting anything other than this. \"}]}]}".
[0107] It should be noted that the chat robot refers to the role played by the target large language model, and the "machine" in "the content output by the machine will be given with the keyword "Predicted Description"" refers to the large language model to be evaluated, Predicted Description represents the data to be evaluated, Correct Description represents the second test data, "Please evaluate the consistency of Predicted Description and Correct Description" refers to determining the similarity between the data to be evaluated and the second test data, "key is "score"" indicates that the keyword is the score, "value is your evaluation result" The evaluation result is the value of the score, and "your evaluation result should be an integer with a value of 0 to 100" indicates that the score value ranges from 0 to 100, and the similarity can be converted into the corresponding score.
[0108] In one embodiment, based on the similarity corresponding to each piece of data to be evaluated in the set of data to be evaluated, the score for the large language model to be evaluated is determined, comprising: Based on the similarity corresponding to each piece of data to be evaluated in the set of data to be evaluated, the similarity corresponding to each piece of data to be evaluated is converted by the target large language model to determine the score corresponding to each similarity. Based on the score corresponding to each similarity, the score for the large language model to be evaluated is determined.
[0109] Specifically, the set of data to be evaluated includes a plurality of data to be evaluated, each piece of data to be evaluated in the plurality of data to be evaluated corresponds to a similarity, and each similarity corresponds to a score for the large language model to be evaluated. For example, the set of data to be evaluated includes 10 pieces of data to be evaluated, and the similarities corresponding to the 10 pieces of data to be evaluated are 0.1, 0.3, 0.2, 0.9, 0.8, 0.5, 0.3, 0.6, 0.2 and 1. The score value range is [0, 100], and the similarities 0.1, 0.3, 0.2, 0.9, 0.8, 0.5, 0.3, 0.6, 0.2 and 1 correspond to the scores 10, 30, 20, 90, 80, 50, 30, 60, 20 and 10 respectively. The scores 10, 30, 20, 90, 80, 50, 30, 60, 20 and 10 constitute the score set for the large language model to be evaluated.
[0110] In one embodiment, after determining the score for the large language model to be evaluated based on the similarity corresponding to each piece of data to be evaluated in the set of data to be evaluated, it further comprises: The model identifier of the large language model to be evaluated, any to-be-evaluated data, the score for the large language model to be evaluated, the evaluation progress for the large language model to be evaluated, the evaluation state for the large language model to be evaluated, and the set identifier of the test data set are displayed on the evaluation interface.
[0111] Specifically, for example, the evaluation interface is as shown in the following table: Figure 5 As shown in the table, the evaluation option "task id" represents the evaluation task identifier, the evaluation option "rtx" represents the user identifier, the evaluation option "Venus service ID" represents the ID of the user-defined large language model, and the evaluation option "base model" represents the general large language model. The user inputs specific data into the above evaluation options, clicks the button "query", or the user directly clicks the button "query all GPT score data", and the test set, server id, base model, task progress (%), status, result url, create time, update time, etc. can be displayed on the evaluation interface. When the user clicks the button "reset", the evaluation options "task id", "rtx", "Venus service ID", and "base model" are emptied, and the user can re-input specific data into the above evaluation options. When the user clicks the button "refresh", the test set, server id, base model, task progress (%), status, result url, create time, update time, etc. are updated and displayed on the evaluation interface. The test set is represented by the set identifier of the test data set, for example, sstest2, sstest, etc. The server id represents the ID of the user-defined large language model, and the base model is, for example, model A, model B, etc. If the base model is selected for evaluation and no user-defined large language model is selected for evaluation, the server id is 0. The task progress (%) represents the progress of the evaluation task, i.e., the evaluation progress for the large language model to be evaluated. For example, 100% indicates that the progress of the evaluation task is 100%. The status represents the evaluation state, for example, evaluation start, evaluation in progress, evaluation completion, etc., i.e., the evaluation state for the large language model to be evaluated. For example, DONE indicates that the evaluation state is evaluation completion. The result url represents the evaluation result, which includes to-be-evaluated data and the score for the large language model to be evaluated. The create time represents the time point when the evaluation task starts, and the update time represents the time point when the evaluation task ends.
[0112] For example, `result url` represents the evaluation result, which includes test results and evaluation results. The evaluation result is the score for the large language model being evaluated. Test results are, for example: { "instruction":"", "input":"Please describe the scene and multiple characters in the image in detail, using JSON format. Only the final JSON should be provided. [Output Format]:{"Environment":"","Characters":[{"Name":"","Gender":"","Clothing":"","Action":"","Expression":""},{"Name":"","Gender":"","Clothing":"","Action":"","Expression":""}]}", "url":"https: / / smart-ai-production-1258344701.cos.ap-guangzhou.myqcloud.com / public_access / v0017poawc4_32138.jpg", "params":"", "output":"{\n\"Environment\":\"Outdoors, tree background, blurred natural environment\",\n\"Character\":[\n{\n\"Name\":\"\",\n \"Gender\":\"Female\",\n\"Clothing\":\"Traditional Chinese clothing, black cloak, ancient-style hair ornament\",\n"Action\":\"Standing, facing the camera\",\n"Expression\":\"Melancholy, seemingly speaking or expressing emotions\"\n}\n]\n}" } It should be noted that instruction represents the script content, such as a brief introduction to Journey to the West; input represents the first test data (question); url represents the address of an image; params represents the role played by the large language model to be evaluated; and output represents the data to be evaluated (the predicted answer to the question), that is, the output of the large language model to be evaluated.
[0113] Applying the embodiments of this disclosure has at least the following beneficial effects: The first preset prompt word is used to instruct the to-be-evaluated large language model to reply to the first test data (question), and the second preset prompt word is used to instruct the target large language model to determine the similarity between the to-be-evaluated data (estimated answer) and the second test data (actual answer), and score the to-be-evaluated data (estimated answer) based on the similarity, that is, score the to-be-evaluated large language model, and obtain the evaluation result of the to-be-evaluated large language model, thereby realizing the automatic evaluation of the to-be-evaluated large language model. Through automatic evaluation, different users evaluate the same to-be-evaluated large language model, and the same evaluation result is obtained, ensuring the consistency of the evaluation result and improving the accuracy of the evaluation result, thereby improving the evaluation efficiency of the to-be-evaluated large language model.
[0114] In order to better understand the method provided by the embodiments of the present disclosure, the scheme of the embodiments of the present disclosure will be further described below in combination with examples of specific application scenarios.
[0115] In one embodiment, as shown in Figure 1 The evaluation system specially designed for large language model evaluation can effectively solve the problems existing in the evaluation method for large language models in the prior art. The advantages embodied by the evaluation system are as follows: (1) The evaluation system provides a general evaluation method: a centralized evaluation system can provide a standardized evaluation method, and users no longer need to write and maintain independent scripts for each large language model. The evaluation system can support multiple types of large language models, thereby ensuring the consistency and comparability of the evaluation of the same large language model.
[0116] (2) The evaluation system performs visual operation for evaluation: through a user-friendly front-end interface (for example, the evaluation interface shown in Figure 4 and Figure 5 ), a user can quickly set test parameters and start an evaluation task; the front-end interface is a visual interface, which can help users quickly understand and configure test parameters, and can also intuitively display evaluation results, which can include performance charts, error analysis, etc. In this way, non-technical personnel can also understand the evaluation results.
[0117] (3) The evaluation system prevents repeated evaluation: the evaluation system can display the evaluation status of each model to avoid repeating the same evaluation; in this way, not only the valuable hardware resources are saved, but also the use efficiency of human resources is improved; the evaluation system can set permissions and roles to ensure the coordination and efficiency of the evaluation work.
[0118] It should be noted that the evaluation system can become a powerful assistant for algorithm users, providing an efficient and convenient platform for comprehensive and in-depth performance evaluation of large language models; through a series of precise tests and evaluations, the algorithm user can thoroughly understand the performance limits of the large language model, which helps the algorithm user more accurately understand the strengths and weaknesses of the large language model; such accurate understanding becomes a solid foundation for optimizing algorithm design.
[0119] The functional characteristics of the evaluation system include: (1) Test data set upload: users can upload test data sets to cloud storage (as shown in Figure 4 ), which will serve as the cornerstone of evaluating the performance of large language models; multiple data formats are supported to ensure that different types of test data sets can be effectively processed.
[0120] (2) Large language model selection and customization: the evaluation system provides a rich model library, allowing users to select appropriate large language models for evaluation based on their needs; to meet the needs of advanced users, it also supports directly specifying specific large language models through their IDs (as shown in Figure 4 ) and customizing prompts to more accurately control evaluation conditions.
[0121] (3) Custom prompt: to more accurately evaluate the performance of large language models in different scenarios, users can customize prompts (as shown in Figure 4 ), which can simulate various application scenarios in the real world and make the evaluation results more valuable.
[0122] The operation process of the evaluation system is as follows: (1) Evaluation interface display: an intuitive evaluation interface (user interface) is displayed to the user, which reduces the difficulty of user operation.
[0123] (2) Upload data set: in the evaluation interface as shown in Figure 4 , the user uploads the evaluation data set; the evaluation system provides a simple upload interface, allowing users to upload evaluation data sets through drag-and-drop or file browsing.
[0124] (3) Large language model selection: after uploading the evaluation data set, the user can select the large language model to be evaluated on the evaluation interface as shown in Figure 4 ; if the user has specific model requirements, they can also directly specify the large language model by entering its ID.
[0125] (4) Evaluation start: after everything is ready, the user only needs to click on Figure 4The "start test" button in the evaluation interface shown, the evaluation system will automatically initiate the evaluation of the selected large language model; during the evaluation process, the user can monitor the evaluation progress in real time, and after the evaluation is completed, view the detailed evaluation report (evaluation result).
[0126] It should be noted that through the functional characteristics of the evaluation system and the operation process of the evaluation system, the evaluation system can provide a comprehensive and efficient evaluation experience of the large language model for the algorithm user, and the algorithm user can use the evaluation system to verify the performance of the large language model, thereby optimizing the algorithm.
[0127] By applying the embodiments of the present disclosure, at least the following beneficial effects are achieved: The evaluation system provides a general evaluation framework that can adapt to various large language models, simplifies the user's workflow and ensures the standardization of large language model evaluation; the evaluation system emphasizes ease of operation, reduces the technical threshold through an intuitive visual interface, so that non-professionals can easily evaluate large language models; in addition, as a unified evaluation platform, the evaluation system avoids repeated evaluation work, saves hardware resources and human resources, and provides a cost-effective evaluation solution; the evaluation system has the following advantages of universality, ease of operation and resource saving: (1) Universality The evaluation system adopts an innovative general evaluation framework, and the evaluation system can adapt to and evaluate various different models; this comprehensive evaluation system greatly simplifies the user's workflow, and the user no longer needs to write and maintain complex evaluation scripts for each model; in this way, the standardization and consistency of the evaluation are ensured, and the evaluation efficiency of the large language model to be evaluated is also improved.
[0128] (2) Ease of operation In order to further improve the user experience, the evaluation system particularly emphasizes the convenience of operation; by introducing an intuitive visual interface, the user can easily complete a complex evaluation task with a few clicks; this design makes the front-end page (evaluation interface) the only requirement for interacting with the evaluation system, greatly reducing the technical threshold, so that non-professionals can also evaluate large language models.
[0129] (3) Resource saving As a unified evaluation platform, the evaluation system effectively avoids repeated work in the evaluation process, thereby saving valuable hardware resources and human resources; in this way, not only unnecessary expenses are reduced, but also the overall evaluation efficiency is improved, so that resources can be more reasonably and efficiently allocated and used.
[0130] In one specific application scenario embodiment, for example, in the evaluation scenario of a large language model, refer to Figure 6, a processing flow of an evaluation method is shown, as Figure 6 The processing flow of the evaluation method provided by the embodiments of the present disclosure includes the following steps: S601, the front-end platform in the evaluation system generates evaluation instruction information for the large language model to be evaluated, and the front-end platform sends the evaluation instruction information for the large language model to be evaluated to the evaluation background service in the evaluation system.
[0131] Specifically, for example, as Figure 3 indicated, the front-end platform in the evaluation system generates evaluation instruction information for the large language model to be evaluated, that is, the evaluation system obtains evaluation instruction information for the large language model to be evaluated; the front-end platform sends the evaluation instruction information for the large language model to be evaluated to the evaluation background service in the evaluation system, and the evaluation background service receives the evaluation instruction information for the large language model to be evaluated. The evaluation interface of the front-end platform in the evaluation system, for example, as Figure 4 indicated, for example, as Figure 4 indicated, the user clicks the button "Start Test", that is, initiates the evaluation task for the large language model to be evaluated, and the front-end platform generates the evaluation instruction information for the large language model to be evaluated.
[0132] S602, the evaluation background service in the evaluation system receives the evaluation instruction information for the large language model to be evaluated, and obtains a test data set based on the evaluation instruction information.
[0133] Specifically, for example, as Figure 3 indicated, the evaluation background service in the evaluation system analyzes and processes the evaluation instruction information for the large language model to be evaluated to obtain the model identifier of the large language model to be evaluated, the cloud storage address of the test data set, etc.; the evaluation background service obtains the test data set from the cloud based on the cloud storage address of the test data set, that is, downloads the test data set, and the cloud is, for example, the cloud storage shown in Figure 3 .
[0134] It should be noted that, for example, as Figure 3 indicated, the test data is uploaded to the cloud storage through the front-end platform in the evaluation system; for example, as Figure 4 indicated, the evaluation option "Upload Test Data Set" allows the user to upload a custom test data set to the cloud, that is, the user uploads the test data set to the cloud storage through the front-end platform for storage.
[0135] S603, the evaluation background service in the evaluation system determines any evaluation data in the evaluation data set based on any first test data in the test data set and a preset first prompt word, and replies to the question represented by any first test data through the large language model to be evaluated.
[0136] Specifically, for example, as shown in Figure 3 The single test task construction includes: determining a test task corresponding to any first test data in the original test data set based on the model identifier of the large language model to be evaluated and any first test data in the test data set; and the test task includes the first test data and the model identifier of the large language model to be evaluated.
[0137] For example, as shown in Figure 3 The execution of the test task and the test result acquisition includes: calling the large language model to be evaluated from the preset large language model set based on the model identifier of the large language model to be evaluated included in the test task, and determining the evaluation data corresponding to any first test data by replying to the question represented by any first test data through the large language model to be evaluated based on the preset first prompt word and any first test data included in the test task, and the evaluation data set includes the evaluation data corresponding to any test data.
[0138] S604, the evaluation background service in the evaluation system determines the similarity between any evaluation data in the evaluation data set and any second test data in the test data set through the target large language model based on any evaluation data in the evaluation data set, any second test data in the test data set, and the preset second prompt word.
[0139] Specifically, any evaluation data in the evaluation data set and any second test data in the test data set are written into the preset second prompt word to obtain an updated second prompt word; and the updated second prompt word is input into the target large language model to determine the similarity between any evaluation data and any second test data through similarity calculation processing.
[0140] For example, as shown in Figure 3 The test result re-splicing includes: writing any evaluation data in the evaluation data set and any second test data in the test data set into the preset second prompt word to obtain an updated second prompt word.
[0141] The target large language model is, for example, GPT4, as shown in Figure 3 The GPT4 scoring includes: inputting the updated second prompt word into the target large language model to determine the similarity between any evaluation data and any second test data through similarity calculation processing; and determining the score for the large language model to be evaluated based on the similarity corresponding to each evaluation data in the evaluation data set.
[0142] S605, the evaluation background service in the evaluation system determines the score for the large language model to be evaluated based on the similarity corresponding to each evaluation data in the evaluation data set.
[0143] Specifically, for example, the set of data to be evaluated includes 10 pieces of data to be evaluated, and the similarity corresponding to the 10 pieces of data to be evaluated is 0.1, 0.3, 0.2, 0.9, 0.8, 0.5, 0.3, 0.6, 0.2 and 1 respectively. The score value range is [0, 100], and the scores corresponding to the similarities 0.1, 0.3, 0.2, 0.9, 0.8, 0.5, 0.3, 0.6, 0.2 and 1 are 10, 30, 20, 90, 80, 50, 30, 60, 20 and 10 respectively. That is, the scores of the large language model to be tested are 10, 30, 20, 90, 80, 50, 30, 60, 20 and 10.
[0144] S606, the evaluation result is uploaded to the cloud for storage by the evaluation background service in the evaluation system.
[0145] Specifically, the evaluation result includes a test result and an evaluation result, the test result includes any data to be evaluated, and the evaluation result includes a score of the large language model to be tested. As shown in Figure 3 The evaluation result is uploaded to the cloud storage for storage by the evaluation background service in the evaluation system.
[0146] S607, the evaluation result is displayed through the evaluation interface by the front-end platform in the evaluation system.
[0147] Specifically, for example, the model identifier of the large language model to be tested, any data to be evaluated, the score of the large language model to be tested, the evaluation progress of the large language model to be tested, the evaluation state of the large language model to be tested, and the set identifier of the test data set are displayed through the evaluation interface; wherein the evaluation result includes any data to be evaluated and the score of the large language model to be tested.
[0148] By applying the embodiments of the present disclosure, the following beneficial effects are achieved: The evaluation system is a platform specially designed to evaluate and analyze the performance of large language models. The user interface (evaluation interface) of the evaluation system is intuitive and easy to use, supporting visual operations, allowing users to easily evaluate and manage large language models. In addition, the evaluation system has advanced scoring functions for large language models, which can automatically quantify the output (to-be-evaluated data) of large language models and provide objective and accurate performance indicators for large language models. At the technical level, the core of the evaluation system is the overall architecture, which makes the evaluation system stable and flexible, and can adapt to the evaluation needs of large language models of various sizes and complexities. The user interface (evaluation interface) not only provides clear information display and smooth user experience, but also allows users to operate efficiently without deep understanding of underlying technical details. The evaluation system is a powerful auxiliary tool that encapsulates complex technical details behind a friendly user interface (evaluation interface), allowing large language model developers and researchers to focus more on innovation and optimization of large language models. The entire evaluation process of large language models can be visualized on the platform, such as evaluation progress and evaluation results.
[0149] The embodiments of the present disclosure also provide an evaluation device, a structural schematic diagram of which is shown as Figure 7 The evaluation device 70 includes a first processing module 701, a second processing module 702, a third processing module 703, and a fourth processing module 704.
[0150] The first processing module 701 is configured to obtain a test data set, the test data set including a plurality of first test data and second test data corresponding to each of the first test data, the first test data being used to represent a question, and the second test data being used to represent an actual answer corresponding to the question represented by the first test data. The second processing module 702 is configured to reply to the question represented by any first test data in the test data set based on the any first test data and a preset first prompt word through a large language model to be evaluated, to determine a to-be-evaluated data set corresponding to the test data set, the to-be-evaluated data set including a plurality of to-be-evaluated data corresponding to the plurality of first test data in the test data set, the to-be-evaluated data corresponding to the any first test data being used to represent a predicted answer corresponding to the question represented by the any first test data, and the preset first prompt word including a first task indication for the large language model to be evaluated, the first task indication being used to indicate the large language model to be evaluated to output the to-be-evaluated data corresponding to the any first test data. The third processing module 703 is configured to determine, by the target large language model, similarity between any to-be-evaluated data in the to-be-evaluated data set and any second test data in the test data set based on the any to-be-evaluated data, the any second test data, and a preset second prompt word, the preset second prompt word including a second task instruction for the target large language model, the second task instruction being used to instruct the target large language model to output the similarity between the any to-be-evaluated data and the any second test data. The fourth processing module 704 is configured to determine a score for the large language model under test based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set.
[0151] In an embodiment, the first processing module 701 is specifically configured to: obtain evaluation instruction information for the large language model under test, the evaluation instruction information including a model identifier of the large language model under test and a storage address of a cloud end of the test data set; obtain the test data set from the cloud end based on the storage address of the cloud end.
[0152] In an embodiment, the first processing module 701 is specifically configured to: in response to an evaluation start operation on an evaluation option in an evaluation interface, initiate an evaluation task for the large language model under test based on a preset query rate per second, and generate evaluation instruction information for the large language model under test; the evaluation option in the evaluation interface includes the model identifier of the large language model under test and the storage address of the cloud end.
[0153] In an embodiment, the second processing module 702 is specifically configured to: determine a test task corresponding to any first test data in the original test data set based on the any first test data and the model identifier of the large language model under test in the test data set; determine any to-be-evaluated data in the to-be-evaluated data set by having the large language model under test reply to a question represented by the any first test data based on the test task and a preset first prompt word.
[0154] In an embodiment, the second processing module 702 is specifically configured to: invoke the large language model under test from a preset large language model set based on the model identifier of the large language model under test included in the test task, and determine to-be-evaluated data corresponding to any first test data by having the large language model under test reply to a question represented by the any first test data based on the preset first prompt word and the any first test data included in the test task, the to-be-evaluated data set including to-be-evaluated data corresponding to any test data.
[0155] In an embodiment, the second processing module 702 is specifically configured to: write any first test data included in the test task into the preset first prompt word to obtain an updated first prompt word; input the updated first prompt word into the large language model to be evaluated, reply to the question represented by any first test data, and determine the to-be-evaluated data corresponding to any first test data.
[0156] In an embodiment, the third processing module 703 is specifically configured to: write any to-be-evaluated data in the to-be-evaluated data set and any second test data in the test data set into the preset second prompt word to obtain an updated second prompt word; input the updated second prompt word into the target large language model, and determine the similarity between any to-be-evaluated data and any second test data through similarity calculation processing.
[0157] In an embodiment, the fourth processing module 704 is specifically configured to: based on the similarity corresponding to each to-be-evaluated data in the to-be-evaluated data set, convert the to-be-evaluated data through the target large language model to determine a score corresponding to each similarity; based on the score corresponding to each similarity, determine a score for the large language model to be evaluated.
[0158] In an embodiment, the fourth processing module 704 is further configured to: display, through the evaluation interface, at least one of the model identifier of the large language model to be evaluated, any to-be-evaluated data, the score for the large language model to be evaluated, the evaluation progress of the large language model to be evaluated, the evaluation state of the large language model to be evaluated, and the set identifier of the test data set.
[0159] By applying the embodiments of the present disclosure, the following beneficial effects are achieved: The test data set is obtained, the test data set includes a plurality of first test data and second test data corresponding to each first test data, the first test data is used to represent a question, and the second test data is used to represent an actual answer corresponding to the question represented by the first test data; based on any first test data in the test data set and a preset first prompt word, the question represented by any first test data is replied by the large language model to be evaluated based on the preset first prompt word, and a data set to be evaluated corresponding to the test data set is determined, the data set to be evaluated includes a plurality of data to be evaluated corresponding to the plurality of first test data in the test data set, and the data to be evaluated corresponding to any first test data is used to represent an estimated answer corresponding to the question represented by any first test data, the preset first prompt word includes a first task instruction for the large language model to be evaluated, and the first task instruction is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data; based on any data to be evaluated in the data set to be evaluated, any second test data in the test data set and a preset second prompt word, the similarity between any data to be evaluated and any second test data is determined by a target large language model, the preset second prompt word includes a second task instruction for the target large language model, and the second task instruction is used to instruct the target large language model to output the similarity between any data to be evaluated and any second test data; based on the similarity corresponding to each data to be evaluated in the data set to be evaluated, a score for the large language model to be evaluated is determined; in this way, the large language model to be evaluated is instructed by the preset first prompt word to reply to the first test data (question) to obtain the data to be evaluated (estimated answer), the similarity between the data to be evaluated (estimated answer) and the second test data (actual answer) is determined by the target large language model based on the preset second prompt word, and the data to be evaluated (estimated answer) is scored based on the similarity, that is, the large language model to be evaluated is scored to obtain the evaluation result of the large language model to be evaluated, thereby realizing the automatic evaluation of the large language model to be evaluated; through automatic evaluation, different users evaluate the same large language model to be evaluated, and the same evaluation result will be obtained, ensuring the consistency of the evaluation result and improving the accuracy of the evaluation result, thereby improving the evaluation efficiency of the large language model to be evaluated.
[0160] The electronic device provided by the embodiment of the present disclosure has the structure as shown in the schematic structural diagram of the electronic device Figure 8 as shown, Figure 8The electronic device 4000 shown includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 can also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception, etc. It should be noted that the transceiver 4004 is not limited to one in actual application, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure.
[0161] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in combination with the disclosure. The processor 4001 can also be a combination of computing functions, such as one or more microprocessor combinations, combinations of DSP and microprocessor, etc.
[0162] The bus 4002 can include a channel for transmitting information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect, peripheral component interconnect) bus or an EISA (Extended Industry Standard Architecture, extended industry standard architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the middle, but it does not mean that there is only one bus or only one type of bus.
[0163] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disc, a laser disc, an optical disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium, other magnetic storage device, or any other medium that can be used to carry or store computer programs and that can be accessed by a computer, without limitation.
[0164] The memory 4003 is configured to store a computer program for implementing the embodiments of the present disclosure, and the processor 4001 is configured to control the execution of the computer program stored in the memory 4003. The processor 4001 is configured to execute the computer program stored in the memory 4003 to implement the steps of the foregoing method embodiments.
[0165] The electronic device includes, but is not limited to, a server, etc.
[0166] According to the embodiments of the present disclosure, at least the following beneficial effects can be achieved: The test data set is obtained, the test data set includes a plurality of first test data and second test data corresponding to each first test data, the first test data is used to represent a question, and the second test data is used to represent an actual answer corresponding to the question represented by the first test data; based on any first test data in the test data set and a preset first prompt word, the question represented by any first test data is replied to by the large language model to be evaluated, and a data set to be evaluated corresponding to the test data set is determined, the data set to be evaluated includes a plurality of data to be evaluated corresponding to the plurality of first test data in the test data set, and the data to be evaluated corresponding to any first test data is used to represent an estimated answer corresponding to the question represented by any first test data, the preset first prompt word includes a first task instruction for the large language model to be evaluated, and the first task instruction is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data; based on any data to be evaluated in the data set to be evaluated, any second test data in the test data set and a preset second prompt word, a similarity between any data to be evaluated and any second test data is determined by a target large language model, the preset second prompt word includes a second task instruction for the target large language model, and the second task instruction is used to instruct the target large language model to output the similarity between any data to be evaluated and any second test data; based on the similarity corresponding to each data to be evaluated in the data set to be evaluated, a score for the large language model to be evaluated is determined; in this way, the large language model to be evaluated is instructed by the preset first prompt word to reply to the first test data (question), and the data to be evaluated (estimated answer) is obtained, the target large language model is instructed by the preset second prompt word to determine the similarity between the data to be evaluated (estimated answer) and the second test data (actual answer), and the data to be evaluated (estimated answer) is scored based on the similarity, that is, the large language model to be evaluated is scored, and an evaluation result for the large language model to be evaluated is obtained, thereby realizing automatic evaluation of the large language model to be evaluated; through automatic evaluation, different users evaluate the same large language model to be evaluated, and the same evaluation result will be obtained, ensuring the consistency of the evaluation result and improving the accuracy of the evaluation result, thereby improving the evaluation efficiency of the large language model to be evaluated.
[0167] The embodiment of the present disclosure provides a computer readable storage medium, and a computer program is stored on the computer readable storage medium. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiment can be realized.
[0168] The embodiment of the present disclosure also provides a computer program product, including a computer program, and the computer program is executed by a processor to realize the steps and corresponding contents of the foregoing method embodiment.
[0169] It should be understood that although the various operation steps in the flowcharts of the embodiments of the present disclosure are indicated by arrows, the implementation order of the steps is not limited to the order indicated by the arrows. Unless otherwise specified herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders as required. In addition, part or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on the actual implementation scenario. Part or all of these sub-steps or stages can be executed at the same time, and each of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured as required, and the embodiments of the present disclosure do not limit this.
[0170] The above is only an optional implementation of some implementation scenarios of the present disclosure, and it should be pointed out that, for ordinary skilled persons in the technical field, other similar implementation means based on the technical idea of the present disclosure without departing from the technical concept of the present disclosure also belong to the protection scope of the embodiments of the present disclosure.
Claims
1. An assessment method characterized by, include: Obtain a test data set, which includes multiple first test data and second test data corresponding to each first test data. The first test data is used to represent a problem, and the second test data is used to represent the actual answer corresponding to the problem represented by the first test data. Based on any first test data in the test data set and a preset first prompt word, the large language model to be evaluated answers the question represented by the first test data, thereby determining the set of data to be evaluated corresponding to the test data set. The set of data to be evaluated includes multiple data to be evaluated corresponding to multiple first test data in the test data set. The data to be evaluated corresponding to any first test data is used to represent the predicted answer corresponding to the question represented by the first test data. The preset first prompt word includes a first task instruction for the large language model to be evaluated. The first task instruction is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data. Based on any data to be evaluated in the dataset to be evaluated, any second test data in the dataset to be tested, and a preset second prompt word, the similarity between the data to be evaluated and the second test data is determined by a target large language model. The preset second prompt word includes a second task instruction for the target large language model. The second task instruction is used to instruct the target large language model to output the similarity between the data to be evaluated and the second test data. Based on the similarity of each data point in the dataset to be evaluated, a score is determined for the large language model to be evaluated.
2. The method of claim 1, wherein, The acquisition of the test data set includes: Obtain evaluation instruction information for the large language model to be evaluated, the evaluation instruction information including the model identifier of the large language model to be evaluated and the cloud storage address of the test data set; The test data set is obtained from the cloud based on the storage address in the cloud.
3. The method of claim 2, wherein, The process of obtaining assessment instruction information for the large language model to be assessed includes: In response to the evaluation start operation for the evaluation options in the evaluation interface, an evaluation task is initiated for the large language model to be evaluated based on a preset query rate per second, and evaluation instruction information for the large language model to be evaluated is generated. The evaluation interface includes evaluation options such as the model identifier of the large language model to be evaluated and the cloud storage address.
4. The method of claim 2, wherein, The step involves determining the set of data to be evaluated corresponding to the test data set by answering the question represented by any first test data and a preset first prompt word using a large language model to be evaluated, based on any first test data in the test data set. This includes: Based on any first test data in the test data set and the model identifier of the large language model to be evaluated, determine the test task corresponding to any first test data in the original test data set; Based on the test task and the preset first prompt words, the large language model to be evaluated answers the question represented by any of the first test data, and determines any data to be evaluated in the set of data to be evaluated.
5. The method of claim 4, wherein, Based on the test task and preset first prompt words, the process involves answering the question represented by any of the first test data using the large language model to be evaluated, thereby determining any data to be evaluated in the dataset to be evaluated, including: Based on the model identifier of the large language model to be evaluated included in the test task, the large language model to be evaluated is called from the preset large language model set. Based on the preset first prompt word and any first test data included in the test task, the large language model to be evaluated answers the question represented by the any first test data, and determines the data to be evaluated corresponding to the any first test data. The set of data to be evaluated includes the data to be evaluated corresponding to the any test data.
6. The method of claim 5, wherein, The process involves using a preset first prompt word and any first test data included in the test task to answer the question represented by any first test data through the large language model to be evaluated, thereby determining the data to be evaluated corresponding to any first test data, including: Write any of the first test data included in the test task into a preset first prompt word to obtain the updated first prompt word; The updated first prompt word is input into the large language model to be evaluated. By answering the question represented by any first test data, the data to be evaluated corresponding to any first test data is determined.
7. The method of claim 1, wherein, The step of determining the similarity between any data to be evaluated and any second test data in the dataset to be evaluated, based on any data to be evaluated in the dataset to be evaluated, any second test data in the dataset to be tested, and a preset second prompt word, using a target large language model, includes: Write any data to be evaluated from the dataset to be evaluated and any second test data from the dataset to be tested into a preset second prompt word to obtain an updated second prompt word; The updated second prompt word is input into the target large language model, and the similarity between any data to be evaluated and any second test data is determined through similarity calculation.
8. The method of claim 1, wherein, The step of determining a score for the large language model to be evaluated based on the similarity of each data point in the dataset to be evaluated includes: Based on the similarity of each data point in the dataset to be evaluated, the target large language model is used for conversion processing to determine the score corresponding to each similarity. Based on the scores corresponding to each similarity level, a score is determined for the large language model to be evaluated.
9. The method of claim 1, wherein, After determining the score for the large language model to be evaluated based on the similarity corresponding to each data point in the dataset to be evaluated, the method further includes: The evaluation interface displays at least one of the following: the model identifier of the large language model to be evaluated, any data to be evaluated, the score of the large language model to be evaluated, the evaluation progress of the large language model to be evaluated, the evaluation status of the large language model to be evaluated, and the set identifier of the test data set.
10. An evaluation device, characterized by include: A first processing module is used to acquire a test data set, the test data set including multiple first test data and second test data corresponding to each first test data, the first test data being used to characterize a problem, and the second test data being used to characterize the actual answer corresponding to the problem characterized by the first test data; The second processing module is used to respond to the question represented by any first test data in the test data set and a preset first prompt word through a large language model to be evaluated, and to determine the set of data to be evaluated corresponding to the test data set. The set of data to be evaluated includes multiple data to be evaluated corresponding to multiple first test data in the test data set. The data to be evaluated corresponding to any first test data is used to represent the estimated answer corresponding to the question represented by any first test data. The preset first prompt word includes a first task instruction for the large language model to be evaluated. The first task instruction is used to instruct the large language model to be evaluated to output the data to be evaluated corresponding to any first test data. The third processing module is used to determine the similarity between any data to be evaluated and any second test data in the dataset to be evaluated, any second test data in the dataset to be tested, and a preset second prompt word through a target large language model. The preset second prompt word includes a second task instruction for the target large language model. The second task instruction is used to instruct the target large language model to output the similarity between the data to be evaluated and any second test data. The fourth processing module is used to determine the score for the large language model to be evaluated based on the similarity of each data to be evaluated in the dataset to be evaluated.
11. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, is arranged to perform the method of any one of claims 1 to 10. The processor executes the computer program to implement the steps of the method according to any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-9.