Model performance evaluation method, device, equipment and storage medium
Patent Information
- Application Number
- CN202510229365.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2026-08-28
AI Technical Summary
但是,这种方法耗时长且容易出错,评估效率低,人力成本较高,而且难以应对大规模的模型评估需求
[0024]The model performance evaluation method, apparatus, device, and storage medium provided in this disclosure acquire test data, which includes multiple test questions and related options for each test question. Based on each test question and its related options, prompt information is constructed. Each test question, its related options, and the prompt information are sent to a target model to be evaluated. The target model generates an answer to each test question. The target model returns its answer to each test question. Based on the answer to each test question and its corresponding standard answer, the performance of the target model is evaluated. Compared to existing technologies, this disclosure automates the generation of prompt information by constructing prompt information based on each test question and its related options. The performance evaluation of the target model based on the answer to each test question and its corresponding standard answer allows for automated model performance evaluation, reducing labor costs and improving evaluation efficiency.
Smart Images

Figure CN122654577A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a model performance evaluation method, apparatus, device, and storage medium. Background Technology
[0002] With the development of science and technology, models have been gradually applied to various fields. To verify the quality of model training, it is necessary to evaluate the model's performance.
[0003] Existing model performance evaluation methods require human intervention throughout the process. This involves manually writing questions and options, then having professionals read the model's responses and manually extract the answers for evaluation. However, this method is time-consuming, error-prone, inefficient, and labor-intensive, and it struggles to meet the demands of large-scale model evaluation. Summary of the Invention
[0004] To address the aforementioned technical issues, this disclosure provides a model performance evaluation method, apparatus, device, and storage medium that can automatically evaluate model performance, reduce labor costs, and improve evaluation efficiency.
[0005] In a first aspect, embodiments of this disclosure provide a model performance evaluation method, the method comprising:
[0006] Acquire test data, which includes multiple test questions and related options for each test question;
[0007] Based on each test question in the test data and the relevant options for each test question, construct prompt information;
[0008] Each test question, its related options, and the prompt information are sent to the target model to be evaluated, so that the target model can generate the answer to each test question.
[0009] Receive the answer result for each test question returned by the target model;
[0010] The performance of the target model is evaluated based on the answers to each test question and the standard answer corresponding to each test question.
[0011] Secondly, embodiments of this disclosure provide a model performance evaluation apparatus, the apparatus comprising:
[0012] The acquisition module is used to acquire test data, which includes multiple test questions and related options for each test question.
[0013] The module is used to construct prompt information based on each test question in the test data and the relevant options for each test question;
[0014] The sending module is used to send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated, so that the target model can generate the answer result for each test question;
[0015] A receiving module is used to receive the answer results for each test question returned by the target model;
[0016] The evaluation module is used to evaluate the performance of the target model based on the answer results of each test question and the standard answer corresponding to each test question.
[0017] Thirdly, embodiments of this disclosure provide an electronic device, including:
[0018] Memory;
[0019] Processor; and
[0020] Computer programs;
[0021] The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.
[0022] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method as described in the first aspect.
[0023] Fifthly, embodiments of this disclosure also provide a computer program product comprising a computer program or instructions that, when executed by a processor, implement the method described in the first aspect.
[0024] The model performance evaluation method, apparatus, device, and storage medium provided in this disclosure acquire test data, which includes multiple test questions and related options for each test question. Based on each test question and its related options, prompt information is constructed. Each test question, its related options, and the prompt information are sent to a target model to be evaluated. The target model generates an answer to each test question. The target model returns its answer to each test question. Based on the answer to each test question and its corresponding standard answer, the performance of the target model is evaluated. Compared to existing technologies, this disclosure automates the generation of prompt information by constructing prompt information based on each test question and its related options. The performance evaluation of the target model based on the answer to each test question and its corresponding standard answer allows for automated model performance evaluation, reducing labor costs and improving evaluation efficiency. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0026] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 A flowchart of the model performance evaluation method provided in the embodiments of this disclosure;
[0028] Figure 2 A flowchart of a model performance evaluation method provided in another embodiment of this disclosure;
[0029] Figure 3 A flowchart of a model performance evaluation method provided in another embodiment of this disclosure;
[0030] Figure 4 This is a schematic diagram of the overall process of the model performance evaluation method provided in the embodiments of this disclosure;
[0031] Figure 5 This is a schematic diagram of the structure of the model performance evaluation device provided in the embodiments of this disclosure;
[0032] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0033] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0034] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0035] Existing model performance evaluation methods require human intervention throughout the process. This involves manually writing questions and options, then having professionals read the model's responses and manually extract the answers for evaluation. However, this method is time-consuming, error-prone, inefficient, and labor-intensive, and it struggles to meet the demands of large-scale model evaluation.
[0036] To address this issue, this disclosure provides a model performance evaluation method, which will be described below with reference to specific embodiments.
[0037] Figure 1 This is a flowchart illustrating the model performance evaluation method provided in this embodiment. The method is executed by an electronic device, which can be a client, specifically a tablet computer, laptop computer, or personal computer. This method can be applied to scenarios involving model performance evaluation.
[0038] It is understood that the model performance evaluation method provided in this disclosure can also be applied in other scenarios.
[0039] The following is about Figure 1 The model performance evaluation method shown is introduced. This method can be applied to electronic devices. Taking a client as an example, the specific steps of the method are as follows:
[0040] S101. Obtain test data, which includes multiple test questions and related options for each test question.
[0041] In this step, the user pre-stores the test data in the client's local storage. The client retrieves the test data from the local storage. The test data includes multiple test questions and their corresponding options. In some embodiments, the test data may also be stored on a server, and the client retrieves the test data from the server; this is not limited.
[0042] S102. Construct prompt information based on each test question in the test data and the relevant options for each test question.
[0043] In this step, after acquiring the test data, the client constructs prompts based on the test questions and their related options, reducing manual intervention and improving efficiency. Prompts are essentially injected instructions and a templated structure for answering questions. By learning from these prompts, the model can think about questions and output content according to a pre-defined approach.
[0044] In some embodiments, S102 may include, but is not limited to, S1021, S1022, and S1023:
[0045] S1021. Determine whether the current sample is being evaluated based on the user configuration information.
[0046] S1022. If the current sample is being evaluated, then construct prompt information based on each test question and the relevant options for each test question.
[0047] S1023. If the current sample is not being evaluated, read the sample data and construct a prompt message based on the sample data, each test question, and the relevant options for each test question.
[0048] In this embodiment, the first sample evaluation refers to a zero-sample evaluation, i.e., there is no example sample data. The client determines whether it is currently in the first sample evaluation stage based on the user configuration information. The configuration information indicates whether learning samples are needed to construct prompt information; that is, the configuration information can indicate whether learning samples are needed or not, and is configured according to user needs. If it is currently the first sample evaluation stage (i.e., there are no example samples), the client constructs prompt information based on each test question and its related options. If it is not currently the first sample evaluation stage (i.e., there is example sample data, for example, five sample data points, i.e., a five-sample evaluation), the client reads the sample data and constructs prompt information based on the sample data, each test question, and its related options. Optionally, the sample data includes sample questions, sample options, and sample prompt information.
[0049] S103. Send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated, so that the target model can generate the answer result for each test question.
[0050] In this step, after the prompt information is constructed, the client will send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated. The target model will then generate the answer result for each test question.
[0051] S104. Receive the answer results for each test question returned by the target model.
[0052] In this step, the target model returns the answer to each test question to the client, and the client receives the answer to each test question returned by the target model.
[0053] In some embodiments, S104 may include, but is not limited to, S1041 and S1042:
[0054] S1041. Obtain the answer result for each test question returned by the target model based on the target interface;
[0055] S1042. Store the answer result of each test question into a preset output file.
[0056] In this embodiment, a target interface is pre-set for data interaction between the client and the target model. The client can obtain the answer results for each test question returned by the target model based on the target interface, and further store the answer results for each test question in a preset output file for subsequent processing.
[0057] S105. Evaluate the performance of the target model based on the answer results of each test question and the standard answer corresponding to each test question.
[0058] In this step, the client evaluates the performance of the target model based on the answers to each test question and the corresponding standard answer. Specifically, the client compares the answers to each test question with the corresponding standard answer to determine if the answers are correct. Furthermore, it can calculate the accuracy of the target model and evaluate its performance based on the accuracy. Figure 4 As shown, since the answers to each test question are stored in a preset output file, the client can first check if the output file exists. If it exists, it means that data has already been processed, and the client reads the identifier (ID) of the processed data. If it doesn't exist, it means that no data has been processed, and all data needs to be processed. Further, the data is processed in batches, and the results are output to a specified file. Simultaneously, the client extracts the standard answers. When extracting the standard answers, it first checks if the standard answers exist in the output path. If they exist, an output path is created, and the standard answers are read from that path. If they don't exist, the question IDs and answers are extracted from the original test data (JSON file), written to a new file, and the standard answers are read from that file.
[0059] This disclosure embodiment acquires test data, including multiple test questions and related options for each test question. Based on each test question and its related options, it constructs prompt information and sends each test question, its related options, and the prompt information to a target model to be evaluated. The target model then generates an answer to each test question. The embodiment receives the answer to each test question returned by the target model and evaluates the performance of the target model based on the answer to each test question and its corresponding standard answer. Compared to existing technologies, this disclosure embodiment automatically generates prompt information by constructing prompt information based on each test question and its related options. The performance evaluation of the target model based on the answer to each test question and its corresponding standard answer is automated, reducing labor costs and improving evaluation efficiency.
[0060] Figure 2 A flowchart of a model performance evaluation method provided in another embodiment of this disclosure is shown below. Figure 2 As shown, the method includes the following steps:
[0061] S201. Obtain test data, which includes multiple test questions and related options for each test question.
[0062] Specifically, the implementation process and principle of S201 and S101 are the same, and will not be repeated here.
[0063] S202. Construct prompt information based on each test question in the test data and the relevant options for each test question.
[0064] Specifically, the implementation process and principle of S202 and S102 are the same, and will not be repeated here.
[0065] S203. Create a session connection between the current client and the target model, and obtain the session identifier.
[0066] In this step, request headers are pre-set, and the target API URL is defined. Next, a session connection is created between the current client and the target model, and the session identifier (session UUID) is obtained. After the session connection is established, the current client and the target model can conduct a session.
[0067] S204. Send a test message to the target model to perform the test.
[0068] After obtaining the session identifier, the client will send a test message to the target model to test whether the client and the target model can interact with each other normally.
[0069] In some embodiments, S204 includes, but is not limited to, S2041 and S2042:
[0070] S2041. Determine whether the current client has successfully connected to the target model based on the response status code;
[0071] S2042. If the current client successfully connects to the target model, a test message is sent to the target model to perform the test.
[0072] In this embodiment, the client checks the response status code and determines whether the current client has successfully connected with the target model based on the response status code. If the current client has successfully connected with the target model, a test message is sent to the target model for testing.
[0073] In some embodiments, if the current client fails to connect to the target model, an error message is printed and returned.
[0074] S205. When the test response message returned by the target model is received, the test is complete.
[0075] In this step, after the client sends a test message to the target model, the target model receives the test message and returns a test response message to the client. If the client receives the test response message returned by the target model, it indicates that the test is complete. In some embodiments, if the client does not receive the test response message returned by the target model, it indicates that the test has failed.
[0076] S206. Call the target interface to send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated.
[0077] In this step, the client can call the target interface to send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated.
[0078] S207. Receive the answer results for each test question returned by the target model.
[0079] Specifically, the implementation process and principle of S207 and S104 are the same, and will not be repeated here.
[0080] S208. Extract the answer results for each test question to obtain the extraction results corresponding to each test question.
[0081] After obtaining the answers to each test question, the client further extracts the answers to each test question, resulting in an extraction result corresponding to each test question. The answers are automatically extracted from the model-generated responses and then subjected to subsequent evaluation processing, reducing human error and improving the consistency and reliability of the results.
[0082] S209. Evaluate the performance of the target model based on the extraction results corresponding to each test question and the standard answer corresponding to each test question.
[0083] In this step, the client evaluates the performance of the target model based on the extraction results for each test question and the corresponding standard answer. Specifically, the client can compare the extraction results for each test question with the corresponding standard answer to determine whether the extraction results are correct. Furthermore, it can calculate the accuracy of the target model and evaluate its performance based on the accuracy, providing support for model development and optimization.
[0084] This embodiment of the disclosure acquires test data, which includes multiple test questions and related options for each test question. Based on each test question and its related options in the test data, prompt information is constructed. Further, a session connection is created between the current client and the target model, and a session identifier is obtained. A test message is sent to the target model for testing. When a test response message is received from the target model, the test is complete. The target interface is then called to send each test question, its related options, and the prompt information to the target model to be evaluated. Next, the answer results for each test question returned by the target model are received, and the answer results for each test question are extracted to obtain the extraction results corresponding to each test question. Then, the performance of the target model is evaluated based on the extraction results corresponding to each test question and the standard answer corresponding to each test question. Through this method, this embodiment of the disclosure can automatically extract answers from the model's returned answer results for each test question, improving automation capabilities, enabling rapid processing of large amounts of data, and significantly improving evaluation efficiency.
[0085] Figure 3 A flowchart of a model performance evaluation method provided in another embodiment of this disclosure is shown below. Figure 3 As shown, the method includes the following steps:
[0086] S301. Obtain test data, which includes multiple test questions and related options for each test question.
[0087] Specifically, the implementation process and principle of S301 and S101 are the same, and will not be repeated here.
[0088] S302. Construct prompt information based on each test question in the test data and the relevant options for each test question.
[0089] Specifically, the implementation process and principle of S302 and S102 are the same, and will not be repeated here.
[0090] S303. Send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated, so that the target model can generate the answer result for each test question.
[0091] Specifically, the implementation process and principle of S303 and S103 are the same, and will not be repeated here.
[0092] S304. Receive the answer results for each test question returned by the target model.
[0093] Specifically, the implementation process and principle of S304 and S104 are the same, and will not be repeated here.
[0094] S305. Extract the answer results for each test question to obtain the extraction results corresponding to each test question.
[0095] Specifically, the implementation process and principle of S305 and S208 are the same, and will not be repeated here.
[0096] S306. For each test question, if there is an option in the relevant options that matches the test question, then that option is determined to be the test answer for each test question.
[0097] In this step, for any extracted result corresponding to a test question, the matching degree between the extracted result and each related option is calculated. If there is an option among the related options of the test question that matches the extracted result corresponding to the test question, then that option is determined as the test answer for each test question. In some embodiments, if the matching degree between the extracted result and the related option is greater than a preset threshold, then the extracted result is considered to match the related option.
[0098] S307. If there is no option in the relevant options of the test question that matches the extraction result corresponding to the test question, then the test answer corresponding to the test question is randomly determined from the relevant options of the test question, and the test question is marked.
[0099] In this step, if there is no option in the relevant options of the test question that matches the extracted result, it indicates that the test answer is rather ambiguous and requires manual review. In this case, the test answer corresponding to the test question is randomly determined from the relevant options of the test question, and the test question is marked. The tester will review the marked test questions.
[0100] S308. Obtain the standard answer corresponding to each test question from the target path, compare the test answer corresponding to each test question with the standard answer corresponding to each test question, and count the number of correct answers.
[0101] like Figure 4 As shown, the client extracts the standard answers. When extracting the standard answers, it first checks if the standard answers exist in the output path. If they exist, it creates the output path and reads the standard answers from it. If they don't exist, it extracts the question ID and answer from the original test data (JSON file), writes them to a new file, and reads the standard answers from that file. Furthermore, it compares the test answer for each test question with the standard answer for each test question to count the number of correct answers.
[0102] In some embodiments, the client determines whether the test answer for each test question has been compared with the standard answer for each test question. If it is determined that the comparison for each test question has been completed, the number of correct answers is counted.
[0103] S309. Based on the number of correct answers and the total number of test questions, calculate the accuracy of the target model.
[0104] In some embodiments, after counting the number of correct answers, the accuracy of the target model can be calculated based on the number of correct answers and the total number of test questions. Specifically, the accuracy of the target model is obtained by dividing the number of correct answers by the total number of test questions.
[0105] This embodiment of the disclosure acquires test data, which includes multiple test questions and related options for each test question. Based on each test question and its related options in the test data, prompt information is constructed. Each test question, its related options, and the prompt information are then sent to a target model to be evaluated, so that the target model generates an answer to each test question. Further, the answer to each test question returned by the target model is received, and the answer to each test question is extracted to obtain an extraction result corresponding to each test question. Next, for each extraction result corresponding to a test question, if there is an option among the related options that matches the extraction result, that option is determined as the test answer for each test question. If there is no option among the related options that matches the extraction result, the test answer for that test question is randomly determined from the related options, and the test question is marked. The standard answer for each test question is obtained from the target path, and the test answers for each test question are compared with the standard answers for each test question to count the number of correct answers. Then, based on the number of correct answers and the total number of test questions, the accuracy of the target model is calculated. Compared with the prior art, the embodiments of this disclosure construct prompt information based on each test question and its related options in the test data. This can automatically generate prompt information, calculate the accuracy of the target model based on the number of correct answers and the total number of test questions, and automatically perform model performance evaluation, reducing labor costs and improving evaluation efficiency.
[0106] Figure 5 This is a schematic diagram of the structure of the model performance evaluation device provided in the embodiments of this disclosure. The model performance evaluation device can be a client as described in the above embodiments, or it can be a component or part within the client. The model performance evaluation device provided in the embodiments of this disclosure can execute the processing flow provided in the embodiments of the model performance evaluation method, such as... Figure 5As shown, the model performance evaluation device 50 includes: an acquisition module 51, a construction module 52, a sending module 53, a receiving module 54, and an evaluation module 55. The acquisition module 51 acquires test data, which includes multiple test questions and related options for each test question. The construction module 52 constructs prompt information based on each test question and its related options in the test data. The sending module 53 sends each test question, its related options, and the prompt information to the target model to be evaluated, so that the target model generates an answer to each test question. The receiving module 54 receives the answer to each test question returned by the target model. The evaluation module 55 evaluates the performance of the target model based on the answer to each test question and the corresponding standard answer.
[0107] Optionally, when the construction module 52 constructs prompt information based on each test question in the test data and the related options of each test question, it is specifically used to: determine whether the current period is the first sample evaluation based on the user configuration information; if the current period is the first sample evaluation, construct prompt information based on each test question and the related options of each test question; if the current period is not the first sample evaluation, read the sample data and construct prompt information based on the sample data, each test question, and the related options of each test question.
[0108] Optionally, when the sending module 53 sends each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated, it is specifically used to: create a session connection between the current client and the target model and obtain a session identifier; send a test message to the target model to perform the test; when the test response message returned by the target model is received, the test is completed; and call the target interface to send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated.
[0109] Optionally, when the sending module 53 sends a test message to the target model for testing, it is specifically used to: determine whether the current client has successfully connected with the target model based on the response status code; if the current client has successfully connected with the target model, then send a test message to the target model for testing.
[0110] Optionally, when the receiving module 54 receives the answer result for each test question returned by the target model, it is specifically used to: obtain the answer result for each test question returned by the target model based on the target interface; and store the answer result for each test question in a preset output file.
[0111] Optionally, after receiving the answer results for each test question returned by the target model, the device 50 further includes: an extraction module 56; the extraction module 56 is used to extract the answer results for each test question to obtain the extraction result corresponding to each test question;
[0112] When the evaluation module 55 evaluates the performance of the target model based on the answer results of each test question and the standard answer corresponding to each test question, it is specifically used to: evaluate the performance of the target model based on the extraction results corresponding to each test question and the standard answer corresponding to each test question.
[0113] Optionally, when the evaluation module 55 evaluates the performance of the target model based on the extraction results corresponding to each test question and the standard answers corresponding to each test question, it specifically performs the following: for each extraction result corresponding to a test question, if there is an option in the relevant options of the test question that matches the extraction result corresponding to the test question, then that option is determined as the test answer corresponding to each test question; if there is no option in the relevant options of the test question that matches the extraction result corresponding to the test question, then the test answer corresponding to the test question is randomly determined from the relevant options of the test question, and the test question is marked; the standard answers corresponding to each test question are obtained from the target path, and the test answers corresponding to each test question are compared with the standard answers corresponding to each test question to count the number of correct answers; based on the number of correct answers and the total number of test questions, the accuracy of the target model is calculated.
[0114] Figure 5 The model performance evaluation device shown in the embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.
[0115] This disclosure also provides a model performance evaluation system, which can be applied to various scenarios requiring model capability evaluation, providing support for model development and optimization. For example, in the field of natural language processing, this system can be used to evaluate the retrieval capabilities of a large language model. The system can automatically generate questions and options based on a given dataset, and the model needs to output the correct answers. The system calculates the accuracy rate based on the model's answers, helping developers understand the model's strengths and weaknesses.
[0116] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. See below for details. Figure 6 It shows a schematic diagram of a structure suitable for implementing the electronic device 600 in the embodiments of this disclosure. Figure 6The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0117] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603 to implement the model performance evaluation method as described in the embodiments of this disclosure. The RAM 603 also stores various programs and data required for the operation of electronic device 600. The processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0118] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0119] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the model performance evaluation method as described above. In such embodiments, the computer program can be downloaded and installed from a network via communication device 609, or installed from storage device 608, or installed from ROM 602. When the computer program is executed by processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0120] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0121] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0122] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0123] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0124] Acquire test data, which includes multiple test questions and related options for each test question;
[0125] Based on each test question in the test data and the relevant options for each test question, construct prompt information;
[0126] Each test question, its related options, and the prompt information are sent to the target model to be evaluated, so that the target model can generate the answer to each test question.
[0127] Receive the answer result for each test question returned by the target model;
[0128] The performance of the target model is evaluated based on the answers to each test question and the standard answer corresponding to each test question.
[0129] Optionally, when one or more of the above-described procedures are executed by the electronic device, the electronic device may also perform other steps described in the above embodiments.
[0130] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0132] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0133] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0134] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0135] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0136] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0137] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for evaluating model performance, characterized in that, The method includes: Acquire test data, which includes multiple test questions and related options for each test question; Based on each test question in the test data and the relevant options for each test question, construct prompt information; Each test question, its related options, and the prompt information are sent to the target model to be evaluated, so that the target model can generate the answer to each test question. Receive the answer result for each test question returned by the target model; The performance of the target model is evaluated based on the answers to each test question and the standard answer corresponding to each test question.
2. The method according to claim 1, characterized in that, The step of constructing prompt information based on each test question in the test data and the relevant options for each test question includes: Determine whether the current sample is being evaluated based on the user's configuration information; If the current sample is being evaluated, then prompt information is constructed based on each test question and the relevant options for each test question; If the current sample is not being evaluated, then the sample data is read, and a prompt message is constructed based on the sample data, each test question, and the relevant options for each test question.
3. The method according to claim 1, characterized in that, Sending each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated includes: Create a session connection between the current client and the target model, and obtain the session identifier; Send test messages to the target model to perform the test; The test is complete when the test response message returned by the target model is received. The target interface is invoked to send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated.
4. The method according to claim 3, characterized in that, The step of sending test messages to the target model for testing includes: Determine whether the current client has successfully connected to the target model based on the response status code; If the current client successfully connects to the target model, a test message is sent to the target model to perform the test.
5. The method according to claim 1, characterized in that, The step of receiving the answer results for each test question returned by the target model includes: Based on the target interface, obtain the answer result for each test question returned by the target model; The answer to each test question is stored in a preset output file.
6. The method according to claim 1, characterized in that, After receiving the answer results for each test question returned by the target model, the method further includes: The answer results for each test question are extracted to obtain the extraction results corresponding to each test question; The evaluation of the target model's performance based on the answers to each test question and the corresponding standard answer includes: The performance of the target model is evaluated based on the extraction results for each test question and the standard answer for each test question.
7. The method according to claim 6, characterized in that, The evaluation of the target model's performance based on the extraction results for each test question and the standard answer for each test question includes: For each test question, if there is an option in the relevant options that matches the test question, then that option is determined to be the test answer for each test question. If there is no option in the relevant options of the test question that matches the extraction result corresponding to the test question, then the test answer corresponding to the test question is randomly determined from the relevant options of the test question, and the test question is marked. Obtain the standard answer corresponding to each test question from the target path, compare the test answer corresponding to each test question with the standard answer corresponding to each test question, and count the number of correct answers; The accuracy of the target model is calculated based on the number of correct answers and the total number of test questions.
8. A model performance evaluation device, characterized in that, The device includes: The acquisition module is used to acquire test data, which includes multiple test questions and related options for each test question. The module is used to construct prompt information based on each test question in the test data and the relevant options for each test question; The sending module is used to send each test question, the relevant options for each test question, and the prompt information to the target model to be evaluated, so that the target model can generate the answer result for each test question; A receiving module is used to receive the answer results for each test question returned by the target model; The evaluation module is used to evaluate the performance of the target model based on the answer results of each test question and the standard answer corresponding to each test question.
9. An electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.