Large-model performance assessment method and related device

By constructing structured text descriptions and using multimodal large models and visual pre-trained model annotations, the problem of lack of hierarchy in large model evaluation in existing technologies is solved, and more accurate performance evaluation is achieved.

WO2025208909A1PCT designated stage Publication Date: 2025-10-09HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/137365
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-30
Filing Date
2024-12-06
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

In the existing technology, when evaluating large models by converting the information of test images into text descriptions, there is a lack of hierarchy, which makes it impossible to accurately evaluate the performance of the large model to be tested.

Method used

Structured text descriptions are used, including information on multiple dimensions of the test material. Multimodal large models and visual pre-training models are used for annotation to construct detailed structured text descriptions, which are then corrected through large language models to obtain the true value of the evaluation and improve the accuracy of the evaluation.

Benefits of technology

Through multi-dimensional structured text descriptions, we can more comprehensively mine key information and details in the test materials and improve the accuracy of large model performance evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024137365_09102025_PF_FP_ABST
    Figure CN2024137365_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a large-model performance assessment method and a related device. The method comprises: acquiring structured text description of a test material, wherein the structured text description comprises information of the test material in a plurality of dimensions; inputting the structured text description and a test question into a large language model, so as to acquire an evaluation true value; inputting the test material and the test question into a large model under test, so as to acquire an answer under test of the test question; and inputting the evaluation true value, the test question and the answer under test into the large language model, so as to acquire an assessment result for the answer under test.
Need to check novelty before this filing date? Find Prior Art

Description

A performance evaluation method for large models and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 30, 2024, with application number 202410385861.X and application name “A performance evaluation method for a large model and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence, and in particular to a performance evaluation method for a large model and related equipment. Background Art

[0003] Large models (such as Chat GPT) have become a major highlight in AI technology due to their exceptional performance and wide application. These large models not only excel in natural language processing tasks but also demonstrate enormous potential in areas such as image processing, video understanding, and autonomous driving. With the widespread adoption of large model technology, building a high-quality, high-precision, and credible large model capability assessment platform has become increasingly important. This platform can help model developers and users more accurately understand the performance and applicability of models.

[0004] In the existing technology, by converting the information of the test image into a text description and obtaining the inference results of the large model to be tested on the test image, the large language model judges the inference results of the large model to be tested based on the above text description, thereby achieving quantifiable automatic evaluation.

[0005] However, the above text description only contains rough information about the test image and lacks hierarchy, which is not conducive to the understanding of large language models, and thus makes it impossible to accurately evaluate the performance of the large model to be tested. Summary of the Invention

[0006] This application provides a large model performance evaluation method and related equipment for accurately evaluating the performance of the large model to be tested.

[0007] The first aspect of this application provides a performance evaluation method for a large model:

[0008] Obtain test materials and test questions, where test questions are questions raised based on the test materials. Obtain a structured text description of the test materials, where the structured text description includes information on multiple dimensions of the test materials. Input the structured text description and the test questions into the large language model to obtain the true evaluation value output by the large language model, where the true evaluation value includes information related to the test questions in the test materials, or is a standard answer to the test questions; input the test materials and the test questions into the large model to be tested to obtain the answer to the test question to be tested output by the large model to be tested; input the true evaluation value, the test questions, and the answer to be tested into the large language model to obtain the evaluation result of the large language model on the answer to be tested based on the true evaluation value and the test questions.

[0009] In this application, since the structured text description includes information on multiple dimensions of the test material, it is more conducive to mining and reflecting more key information and details in the test material, and the description is more hierarchical, which is more conducive to the understanding of the large language model, thereby obtaining accurate evaluation truth values ​​and improving the accuracy of performance evaluation.

[0010] In one possible implementation, a structured text description is tested against a large language model to identify errors in the structured text description. These errors include inconsistencies between information in multiple dimensions. The errors in the structured text description are corrected, and the corrected structured text description and a test question are input into the large language model to obtain the true value of the evaluation output by the large language model.

[0011] In this application, the structured text description is also corrected based on a large language model to further improve the accuracy of performance evaluation.

[0012] In one possible implementation, the test material is an image of a traffic scene, and the information in multiple dimensions includes scene-level information, object-level information, and area-level information. The scene-level information includes an overall description of the traffic scene, the object-level information includes a description of each object in each object category in the traffic scene, and the area-level information includes a description of the target area in the traffic scene.

[0013] In one possible implementation, obtaining the structured text description of the test material is specifically as follows:

[0014] Input the test material into the multimodal large model to obtain scene-level information output by the multimodal large model. Obtain multiple object-annotated images, where the object-annotated images are obtained by annotating each object in the same object category in the test material. The objects annotated in different object-annotated images belong to different object categories. Input multiple object-annotated images into the multimodal large model in sequence to obtain object-level information output by the multimodal large model. Obtain region-annotated images, where the region-annotated images are obtained by annotating the target region in the test material. Input the region-annotated images and the position coordinates of the target region into the multimodal large model to obtain region-level information output by the multimodal large model.

[0015] In one possible implementation, the object categories include vehicles, road users, and roadblocks.

[0016] In one possible implementation, the test material is input into the visual pre-training model to obtain the annotation information output by the visual pre-training model. The annotation information includes annotations of objects in the test material. Object annotation images and region annotation images are obtained based on the annotation information.

[0017] In this application, a visual pre-training model is introduced in the process of constructing a structured text description to identify various objects in the test material, thereby improving the automation level and accuracy of constructing the structured text description.

[0018] In a possible implementation, the test material is an image of a geometry problem, and the information in multiple dimensions includes geometric element information and geometric relationship information. The geometric element information includes the geometric elements appearing in the geometry problem, and the geometric relationship information includes the geometric relationship between the geometric elements.

[0019] The second aspect of this application provides a performance evaluation method for a large model:

[0020] Obtain test materials and test questions, where the test questions are questions posed based on the test materials. Obtain a structured text description of the test materials, which includes information on multiple dimensions of the test materials. Input the test materials and test questions into the large model to be tested to obtain the answer to the test question output by the large model to be tested. Input the structured text description, test question, and answer to be tested into the large language model to obtain the large language model's evaluation result of the answer to be tested based on the structured text description and test question.

[0021] In this application, since the structured text description includes information on multiple dimensions of the test material, it is more conducive to mining and reflecting more key information and details in the test material, and the description is more hierarchical, which is more conducive to the understanding of the large language model, thereby improving the accuracy of performance evaluation.

[0022] In one possible implementation, the test material is an image of a traffic scene, and the information in multiple dimensions includes scene-level information, object-level information, and area-level information. The scene-level information includes an overall description of the traffic scene, the object-level information includes a description of each object in each object category in the traffic scene, and the area-level information includes a description of the target area in the traffic scene.

[0023] In one possible implementation, obtaining the structured text description of the test material is specifically as follows:

[0024] Input the test material into the multimodal large model to obtain scene-level information output by the multimodal large model. Obtain multiple object-annotated images, where the object-annotated images are obtained by annotating each object in the same object category in the test material. The objects annotated in different object-annotated images belong to different object categories. Input multiple object-annotated images into the multimodal large model in sequence to obtain object-level information output by the multimodal large model. Obtain region-annotated images, where the region-annotated images are obtained by annotating the target region in the test material. Input the region-annotated images and the position coordinates of the target region into the multimodal large model to obtain region-level information output by the multimodal large model.

[0025] In one possible implementation, the object categories include vehicles, road users, and roadblocks.

[0026] In one possible implementation, the test material is input into the visual pre-training model to obtain the annotation information output by the visual pre-training model. The annotation information includes annotations of objects in the test material. Object annotation images and region annotation images are obtained based on the annotation information.

[0027] In a possible implementation, the test material is an image of a geometry problem, and the information in multiple dimensions includes geometric element information and geometric relationship information. The geometric element information includes the geometric elements appearing in the geometry problem, and the geometric relationship information includes the geometric relationship between the geometric elements.

[0028] A third aspect of the present application provides a computing device, including:

[0029] The acquisition unit is used to acquire test materials and test questions, where the test questions are questions raised based on the test materials.

[0030] The acquisition unit is further configured to acquire a structured text description of the test material, where the structured text description includes information of multiple dimensions of the test material.

[0031] The processing unit is used to input the structured text description and the test question into the large language model to obtain the evaluation truth value output by the large language model, where the evaluation truth value includes information related to the test question in the test material, or is a standard answer to the test question.

[0032] The processing unit is further used to input the test materials and test questions into the large model to be tested, so as to obtain the answers to the test questions output by the large model to be tested.

[0033] The processing unit is further used to input the evaluation truth value, the test question and the answer to be tested into the large language model to obtain the evaluation result of the large language model on the answer to be tested based on the evaluation truth value and the test question.

[0034] In one possible implementation,

[0035] The processing unit is further configured to test the structured text description based on the large language model to determine description errors in the structured text description, where the description errors include contradictions between information in multiple dimensions.

[0036] The processing unit is further used to correct description errors in the structured text description.

[0037] The processing unit is specifically used to input the corrected structured text description and test questions into the large language model to obtain the evaluation truth value output by the large language model.

[0038] In one possible implementation, the test material is an image of a traffic scene, and the information in multiple dimensions includes scene-level information, object-level information, and area-level information. The scene-level information includes an overall description of the traffic scene, the object-level information includes a description of each object in each object category in the traffic scene, and the area-level information includes a description of the target area in the traffic scene.

[0039] In one possible implementation,

[0040] The acquisition unit is specifically used to input the test material into the multimodal large model to obtain the scene-level information output by the multimodal large model.

[0041] The acquisition unit is specifically used to acquire multiple object-annotated images. The object-annotated images are obtained by annotating each object in the same object category in the test material. The objects annotated in different object-annotated images belong to different object categories.

[0042] The acquisition unit is specifically used to input multiple object-annotated images into the multimodal large model in sequence to obtain object-level information output by the multimodal large model.

[0043] The acquiring unit is specifically used to acquire a region-annotated image, where the region-annotated image is obtained by annotating a target region in a test material.

[0044] The acquisition unit is specifically used to input the region annotation image and the position coordinates of the target region into the multimodal large model to obtain the region-level information output by the multimodal large model.

[0045] In one possible implementation, the object categories include vehicles, road users, and roadblocks.

[0046] In one possible implementation,

[0047] The processing unit is further configured to input the test material into the visual pre-training model to obtain annotation information output by the visual pre-training model, wherein the annotation information includes annotations of objects in the test material;

[0048] The object annotated image and the region annotated image are obtained based on the annotation information.

[0049] In a possible implementation, the test material is an image of a geometry problem, and the information in multiple dimensions includes geometric element information and geometric relationship information. The geometric element information includes the geometric elements appearing in the geometry problem, and the geometric relationship information includes the geometric relationship between the geometric elements.

[0050] In a fourth aspect, the present application provides a computing device, including a processor and a memory, wherein the processor is configured to execute instructions stored in the memory so that a computing device cluster executes the method in the aforementioned first aspect.

[0051] A fifth aspect of the present application provides a computer program product comprising instructions, which, when executed by a computing device, causes the computing device to execute the method in the aforementioned first aspect.

[0052] In a sixth aspect, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method in the aforementioned first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] FIG1 is a schematic diagram of the system architecture in this application;

[0054] FIG2 is another schematic diagram of the system architecture in this application;

[0055] FIG3 is a flow chart of a performance evaluation method for a large model in this application;

[0056] FIG4 is a schematic diagram of the test material in this application;

[0057] FIG5 is a schematic diagram of obtaining scene-level information in this application;

[0058] FIG6 is a schematic diagram of an object annotation image in this application;

[0059] 7-9 are schematic diagrams of obtaining object-level information in this application;

[0060] FIG10 is a schematic diagram of a region-annotated image in this application;

[0061] FIG11 is a schematic diagram of obtaining regional level information in this application;

[0062] FIG12 is a schematic diagram of evaluating responses based on a large language model in this application;

[0063] FIG13 is another flow chart of the performance evaluation method for a large model in this application;

[0064] FIG14 is a schematic diagram of evaluating responses based on a large language model in this application;

[0065] FIG15 is another schematic diagram of the test material in this application;

[0066] FIG16a is a schematic diagram of obtaining the true evaluation value based on the large language model in this application;

[0067] FIG16 b is a schematic diagram of obtaining the answers of the large model to be tested to the test questions in this application;

[0068] FIG17 is a schematic diagram showing the effect of the performance evaluation method of a large model in this application;

[0069] FIG18 is a schematic diagram of a structure of a computing device in this application;

[0070] FIG19 is another schematic diagram of the structure of the computing device in this application. DETAILED DESCRIPTION

[0071] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art will appreciate that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0072] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0073] Please refer to Figure 1, which is a schematic diagram of the deployment of an evaluation system disclosed in this application. The performance evaluation method for large models in this application is executed by the evaluation system. The evaluation system can be deployed in a cloud environment. A cloud environment is an entity that uses basic resources to provide cloud services to users in a cloud computing model. The cloud environment includes a cloud data center and a cloud service platform. The cloud data center includes a large number of basic resources (including computing resources, storage resources, and network resources) owned by the cloud service provider. The computing resources included in the cloud data center can be a large number of computing devices (such as servers). The evaluation system can be deployed independently on a server or virtual machine in the cloud data center. The evaluation system can also be distributed across multiple servers in the cloud data center, or distributed across multiple virtual machines in the cloud data center, or distributed across servers and virtual machines in the cloud data center. As shown in Figure 1, the evaluation system is abstracted by the cloud service provider on the cloud service platform as an evaluation cloud service and provided to users. After the user purchases the cloud service on the cloud service platform (pre-charge is possible and settlement is based on the final resource usage), the cloud environment uses the evaluation system deployed in the cloud data center to provide the evaluation cloud service to the user. The evaluation system can also be deployed in an edge environment, which refers to a data center or a collection of edge computing devices that are closer to the user. The edge environment includes one or more edge computing devices. The evaluation system can be deployed independently on an edge computing device, or the evaluation system can be distributedly deployed on multiple edge servers, or distributedly deployed on multiple edge stations with computing power, or distributedly deployed on edge servers and edge stations with computing power. In addition, the evaluation system can also be deployed in other environments, such as a cluster of terminal computing devices. The evaluation system can be a software system that runs on computing devices such as servers. Please refer to Figure 2, which is a deployment diagram of another evaluation system disclosed in this application. As shown in Figure 2, the evaluation system provided in this application can also be distributedly deployed in different environments. The evaluation system provided in this application can be logically divided into multiple parts, each part having different functions. The various parts in the evaluation system can be deployed in any two or three environments of the terminal computing device, edge environment and cloud environment. Terminal computing devices include: terminal servers, smart phones, laptops, tablets, personal desktop computers, smart cameras, etc. The edge environment is an environment that includes a collection of edge computing devices that are close to the terminal computing devices. Edge computing devices include: edge servers, edge stations with computing power, etc. The various parts of the evaluation system deployed in different environments or devices work together to realize the evaluation function. It should be understood that this application does not restrictively divide which parts of the evaluation system are deployed in what environment. In actual application, it can be adaptively deployed according to the computing power of the terminal computing device, the resource possession of the edge environment and the cloud environment, or the specific application requirements.Of course, the evaluation system can also be deployed in a non-cloud environment as a toolkit to help users perform performance evaluation of large models.

[0074] Please refer to Figure 3. The following describes the process of the performance evaluation method of the large model in this application:

[0075] 301. Obtain test materials and test questions, where the test questions are questions raised based on the test materials;

[0076] The evaluation system first obtains test materials and test questions for testing the large model to be tested. The test materials can be multimodal information, such as images, audio, etc., and the test questions are questions based on the test materials. This application can accurately evaluate the performance of the large model to be tested when facing autonomous driving tasks. Please refer to Figure 4. For example, the test material is an image of a traffic scene, and the test question is: "What is the driving recommendation in this traffic scene?"

[0077] 302. Obtain a structured text description of the test material, where the structured text description includes information on multiple dimensions of the test material;

[0078] The evaluation system will obtain information on multiple dimensions of the test material based on the multimodal large language model (MLLM). This information includes scene-level information, object-level information, and area-level information. Among them, the scene-level information includes an overall description of the traffic scene, the object-level information includes a description of each object in each object category in the traffic scene, and the area-level information includes a description of the target area in the traffic scene.

[0079] The following describes how the evaluation system obtains the above information:

[0080] 1. Obtaining scene-level information

[0081] As shown in Figure 5, the evaluation system directly inputs the aforementioned test material into the multimodal large model and instructs it to describe the information in the image as detailed as possible. The output of the multimodal large model is scene-level information. For example, "During the daytime, the vehicle is driving forward on a city road. A green bus occupies the lane ahead. A silver SUV is slightly ahead of the vehicle on the left, and the vehicle needs to observe whether it will change lanes. Meanwhile, there are some cyclists on the right side of the road. There is a red and white fence on the right side of the road. The traffic light ahead is currently green."

[0082] 2. Get object-level information

[0083] The evaluation system inputs the above test material into the visual pre-training model, and the visual pre-training model identifies and labels each object in the test material (such as vehicles, roadblocks, etc.). Afterwards, the evaluation system obtains multiple object-labeled images based on the labeling performed by the visual pre-training model. The object-labeled images are obtained by labeling each object in the same object category in the test material, and the objects labeled in different object-labeled images belong to different object categories. For example, please refer to Figure 6. The evaluation system obtains object-labeled image 1, object-labeled image 2, and object-labeled image 3, wherein object-labeled image 1 is obtained by labeling each vehicle in the test material, object-labeled image 2 is obtained by labeling each road user in the test material, and object-labeled image 3 is obtained by labeling each roadblock in the test material. Please refer to Figure 7. The evaluation system inputs object-labeled image 1 into the multimodal large model and instructs the multimodal large model to describe the information in the labeling box in the image. The output of the multimodal large model is the description of each object in the object category of vehicle. For example: "Vehicle 1 Description: A green bus is directly in front of the ego vehicle, occupying its lane. Vehicle 1 Description: Buses are large, slow-moving vehicles that may block the ego vehicle's path and limit its view ahead. Vehicle 2 Description: A silver SUV is to the left of the ego vehicle, slightly ahead of the adjacent lane. Vehicle 2 Description: The SUV may change lanes or maintain its current trajectory, which the ego vehicle needs to monitor to safely change lanes or turn."

[0084] Please refer to Figure 8. The evaluation system inputs the object annotation image 2 into the multimodal large model and instructs the multimodal large model to describe the information in the annotation box in the image. The output of the multimodal large model is the description of each object in the road user object category, for example: "Road User 1 Description: A cyclist appears to the right of the vehicle, traveling parallel to the car. Road User 1 Explanation: Cyclists can behave unpredictably, so the vehicle must be prepared to respond if they change direction or enter the lane."

[0085] Please refer to Figure 9. The evaluation system inputs the object annotation image 3 into the multimodal large model and instructs the multimodal large model to describe the information in the annotation box in the image. The output of the multimodal large model is the description of each object in the object category of roadblocks, for example: "Fence 1 Description: There is a red and white fence on the right side of the road, which demarcates a part of the road. Fence 1 Description: These fences indicate that there may be construction zones or lane closures, which require caution and may require a change in normal driving mode."

[0086] The evaluation system combines the above descriptions of vehicles, road users, and roadblocks to form object-level information.

[0087] 3. Obtain regional level information

[0088] The evaluation system obtains a region-annotated image based on the annotation performed by the visual pre-training model. The region-annotated image is obtained by annotating the target area in the test material. The target area can be the bounding box of an object in the image. For example, please refer to Figure 10. The above-mentioned target area can be, for example, the bounding box of a bus. Please refer to Figure 11. The evaluation system inputs the region-annotated image into the multimodal large model, provides the position coordinates of the target area, and instructs the multimodal large model to describe the information in the annotated box in the image. The output of the multimodal large model is regional-level information, such as: "A large bus designed to carry multiple passengers. The vehicle is directly ahead, indicating that the autonomous vehicle needs to maintain a safe following distance and be prepared to respond accordingly when the bus stops or slows down, especially when passengers may get on and off."

[0089] The evaluation system stitches together the information from the above multiple dimensions to obtain a structured text description.

[0090] The above describes a method in which the evaluation system obtains structured text descriptions on its own. In another possible implementation method, users can also upload test materials, test questions, and corresponding structured text descriptions to the evaluation system on their own.

[0091] In one possible implementation, the evaluation system can also input the structured text description into a large language model, which analyzes the structured text description to determine whether there are any errors, such as inconsistencies between information in different dimensions. The evaluation system then corrects the errors in the structured text description, either manually or automatically.

[0092] 303. Input the structured text description and the test question into the large language model to obtain the true evaluation value output by the large language model;

[0093] The evaluation system inputs structured text descriptions and test questions into a large language model, which then outputs a ground-truth evaluation, which is the standard answer to the test question. In this example, the test question is, "What are the driving recommendations for this traffic scenario?" The corresponding ground-truth evaluation is, "The vehicle should maintain a safe distance from the bus, prepare for the bus's possible stop, monitor any lane-changing attempts by the SUV on the left, and be aware of the cyclist on the right. Since the traffic light is currently green, the vehicle can continue driving, but should exercise caution and slow down, preparing for possible changes in the traffic light, changes in the cyclist's behavior, and possible lane narrowing due to the barrier."

[0094] 304. Input the test material and the test questions into the large model to be tested to obtain the answers to the test questions output by the large model to be tested;

[0095] The evaluation system inputs test materials and test questions into the large model to be tested, and the large model to be tested outputs answers to the test questions, such as: "Keep a safe distance from the bus, pay attention to the SUV changing lanes on the left and the cyclist on the right. When the traffic light is green, continue driving but slow down and be prepared for traffic light changes, cyclists' behavior and possible lane narrowing."

[0096] 305. Input the true value of the evaluation, the test question, and the answer to be tested into the large language model to obtain an evaluation result of the large language model on the answer to be tested based on the true value of the evaluation and the test question.

[0097] Please refer to Figure 12. The evaluation system inputs the test questions, the answers of the large model to be tested, and the evaluation truth values ​​into the large language model (such as the plain text version of GPT-4), and instructs the large language model to analyze the correctness of the answers of the large model to be tested based on the test questions and the evaluation truth values, and gives a corresponding score. In this embodiment, the analysis result of the large language model is, for example: "The answers of the model to be tested cover most of the key safety factors, such as maintaining a safe distance from the bus, monitoring the lane change intention of the SUV on the left, paying attention to the cyclist on the right, the status of the traffic light, and slowing down and preparing to respond to changes in traffic conditions, demonstrating an understanding of traffic safety rules. However, the answer of the model to be tested is slightly lacking in the richness and detail of expression. For example, it does not explicitly mention "preparing for possible stops of the bus", which is an important safety consideration. Overall, the model accurately identifies and responds to most key safety factors, but lacks some details and depth. Score (0-10 points): 7 points." In this application, the analysis results of the large language model can be in json format and can be directly parsed.

[0098] It should be understood that the above performance evaluation process can be repeated, and multiple sets of test materials and test questions can be used to evaluate the large model to be tested, and the final total score of the large model to be tested can reflect the performance of the large model to be tested.

[0099] In this application, since the structured text description includes information on multiple dimensions of the test material, it is more conducive to mining and reflecting more key information and details in the test material, and the description is more hierarchical, which is more conducive to the understanding of the large language model, thereby obtaining accurate evaluation truth values ​​and improving the accuracy of performance evaluation.

[0100] Please refer to FIG13 , which describes another process of the performance evaluation method for a large model in this application:

[0101] 1301. Obtain test materials and test questions, where the test questions are questions raised based on the test materials;

[0102] The evaluation system first obtains test materials and test questions for testing the large model to be tested. Similarly, please refer to Figure 4. The test material is, for example, an image of a traffic scene, and the test question is: "What is the driving recommendation in this traffic scene?"

[0103] 1302. Obtain a structured text description of the test material, where the structured text description includes information on multiple dimensions of the test material;

[0104] The evaluation system will obtain information of multiple dimensions of the test material based on the multimodal large model. The information of multiple dimensions includes scene-level information, object-level information and area-level information, which is similar to that described in the embodiment shown in Figure 3 above and will not be repeated here.

[0105] 1303. Input the test materials and test questions into the large model to be tested to obtain the answers to the test questions output by the large model to be tested;

[0106] The evaluation system inputs test materials and test questions into the large model under test, which then outputs answers to the test questions. For example, "Keep a safe distance from buses, watch out for SUVs changing lanes on the left and cyclists on the right. When the traffic light is green, continue driving but slow down and be prepared for traffic light changes, cyclists' behavior, and possible lane narrowing."

[0107] 1304. Input the structured text description, the test question, and the answer to be tested into the large language model to obtain an evaluation result of the large language model on the answer to be tested based on the structured text description and the test question.

[0108] Please refer to Figure 14. Unlike the embodiment shown in Figure 3, in this embodiment, no true evaluation value is generated. Instead, the test questions, the answers of the large model to be tested to the test questions, and the structured text description are directly input into the large language model (such as the plain text version of GPT-4), and the large language model is instructed to analyze the correctness of the answers to the large model to be tested based on the structured text description and the test questions, and give a corresponding score. It should be understood that the above performance evaluation process can be repeated, and multiple sets of test materials and test questions can be used to evaluate the large model to be tested, and the performance of the large model to be tested can be reflected by the final total score of the large model to be tested.

[0109] In this application, since the structured text description includes information on multiple dimensions of the test material, it is more conducive to mining and reflecting more key information and details in the test material, and the description is more hierarchical, which is more conducive to the understanding of the large language model, thereby improving the accuracy of performance evaluation.

[0110] The above describes the process of applying this application in the autonomous driving scenario. Please refer to steps A01 to A05. The following describes the process of applying this application in the geometry problem analysis scenario:

[0111] A01. Obtain test materials and test questions. Test questions are questions raised based on the test materials.

[0112] Please refer to FIG. 15 . In this embodiment, the test material is an image of a geometry problem. The test question is: “Circle O is the circumcircle of triangle ABC. Given angle ABO=30 degrees, what is the size of angle ACB?”

[0113] A02. Obtain a structured text description of the test material, which includes information on multiple dimensions of the test material.

[0114] The assessment system can also obtain structured description text of geometry problems. This structured description text can be obtained by, for example, professional annotation personnel who have annotated the geometry problems. It includes, but is not limited to, the properties of the graphics (such as lines, circles, polygons, etc.), positional relationships (such as parallel, perpendicular, tangent, etc.), and the application of geometric theorems or formulas. By drawing graphics, writing logical explanation text, and citing relevant geometric theorems or formulas, a set of structured description texts of the geometry problems is constructed. The structured description text includes, for example, geometric element information and geometric relationship information. The geometric element information includes the geometric elements that appear in the geometry problem, such as:

[0115] Circle (O)

[0116] Triangle (A, B, C)

[0117] Triangle (A, O, B)

[0118] Geometric relationship information includes the geometric relationships between geometric elements, such as:

[0119] Point Lies On circle(A,circle(O))

[0120] Indicates that point A is on circle O

[0121] Point Lies On circle(B,circle(0))

[0122] Indicates that point B is on circle O

[0123] Point Lies On circle(c,circle(O))

[0124] Indicates that point C is on circle O

[0125] A03. Input the structured text description and test questions into the large language model to obtain the true evaluation value output by the large language model;

[0126] Refer to Figure 16a. The evaluation system inputs the structured text description and the test question into the large language model and instructs the large language model to describe as much information relevant to the test question as possible based on the structured text description. The output of the large language model is the true evaluation value, for example: "This image consists of a triangle ABC, which is inscribed in a circle, labeled circle 0. Points A, B, and C are the vertices of the triangle, all located on the circumference of the circle. The center of the circle is labeled point 0. The center point O and points A and B form another triangle AOB."

[0127] A04. Input the test materials and test questions into the large model to be tested, so as to obtain the answers to the test questions output by the large model to be tested;

[0128] Please refer to Figure 16b. The evaluation system inputs the test materials and test questions into the large model to be tested. The large model to be tested will output the answer to the test question, for example: "In triangle AOB, OA = OB, angle ABO = 30°: therefore, angle AOB = 180-2×angle ABO = 120°: therefore, angle ACB = 1 / 2×angle AOB = 60°."

[0129] A05. Input the true value of the evaluation, the test question, and the answer to be tested into the large language model to obtain the evaluation result of the large language model on the answer to be tested based on the true value of the evaluation and the test question.

[0130] The evaluation system inputs the test questions, the answers of the large model to be tested to the test questions, and the evaluation truth values ​​into the large language model (such as the plain text version of GPT-4), and instructs the large language model to analyze the correctness of the answers of the large model to be tested based on the test questions and the evaluation truth values, and give a corresponding score. In this embodiment, the analysis result of the large language model is, for example: "First, the large model to be tested points out that in triangle AOB, OA=OB. This is correct because OA and OB are both the radii of circle O. The large model to be tested mentions that angle ABO is 30°, which is consistent with the test question. Next, the large model to be tested calculates the size of angle A0B as 180°-2×angle ABO. This calculation is based on the fact that in the isosceles triangle AOB, the base angles (angle ABO and angle BAO) are equal and the sum of the internal angles of the triangle is 180°. Since angle ABO is 30°, angle A0B is indeed 180°-2×30°=120°. Finally, the large model to be tested claims that angle ACB is half of angle AOB, that is, 60°. This conclusion is based on the inclination angle theorem, that is, the angle on a circle is equal to half of the central angle it corresponds to. In this case, angle AOB is the central angle of angle ACB, so angle ACB is indeed half of angle AOB, that is, 60°. Correctness score: 1 point."

[0131] In this application, since the structured text description includes information on multiple dimensions of the test material, it is more conducive to mining and reflecting more key information and details in the test material, and the description is more hierarchical, which is more conducive to the understanding of the large language model, thereby obtaining accurate evaluation truth values ​​and improving the accuracy of performance evaluation.

[0132] Please refer to steps B01 to B04. The following describes another process used in the geometry problem analysis scenario of this application:

[0133] B01. Obtain test materials and test questions. Test questions are questions raised based on the test materials.

[0134] This step is similar to the aforementioned step A01 and will not be repeated here.

[0135] B02. Obtain a structured text description of the test material, which includes information on multiple dimensions of the test material;

[0136] This step is similar to the aforementioned step A02 and will not be repeated here.

[0137] B03. Input the test materials and test questions into the large model to be tested, so as to obtain the answers to the test questions output by the large model to be tested;

[0138] The evaluation system inputs the test materials and test questions into the large model to be tested, and the large model to be tested will output the answers to the test questions. For example: in triangle AOB, OA=OB, angle ABO=30°: therefore, angle AOB=180-2×angle ABO=120°: therefore, angle ACB=1 / 2×angle AOB=60°.

[0139] B04. Input the structured text description, test questions, and answers to be tested into the large language model to obtain the evaluation results of the large language model on the answers to be tested based on the true value of the evaluation and the test questions.

[0140] Unlike the previous embodiment, in this embodiment, there is no need to obtain the true evaluation value. Instead, the test question, the answer of the large model to be tested to the test question, and the true evaluation value are directly input into the large language model (such as the plain text version of GPT-4), and the large language model is instructed to analyze the correctness of the answer of the large model to be tested based on the test question and the structured text description, and give a corresponding score.

[0141] In this application, since the structured text description includes information on multiple dimensions of the test material, it is more conducive to mining and reflecting more key information and details in the test material, and the description is more hierarchical, which is more conducive to the understanding of the large language model, thereby improving the accuracy of performance evaluation.

[0142] In this application, when the evaluation system evaluates the performance of the model to be tested in different task scenarios (such as autonomous driving, geometry problems), it can automatically select the prompt project applicable to the current task scenario, such as a customized prompt template, thereby optimizing the accuracy of the evaluation.

[0143] Referring to FIG. 17 , based on the method provided in this application, the performance of each large model can be quickly and accurately evaluated in various aspects, including but not limited to general perception, regional perception, driving suggestions, comprehensive capabilities (all), vehicles, vulnerable road users, traffic cones, traffic lights, traffic signs, barriers, and miscellaneous.

[0144] The above describes the method in this application. The following describes the computing device in this application:

[0145] Please refer to FIG. 18 . The computing device 1800 in the present application includes an acquisition unit 1801 and a processing unit 1802 .

[0146] The acquisition unit 1801 is used to acquire test materials and test questions, where the test questions are questions raised based on the test materials.

[0147] The acquiring unit 1801 is further configured to acquire a structured text description of the test material, where the structured text description includes information of multiple dimensions of the test material.

[0148] Processing unit 1802 is used to input the structured text description and the test question into the large language model to obtain the evaluation truth value output by the large language model. The evaluation truth value includes information related to the test question in the test material, or is a standard answer to the test question.

[0149] The processing unit 1802 is further configured to input the test materials and the test questions into the large model to be tested, so as to obtain the answers to the test questions output by the large model to be tested.

[0150] The processing unit 1802 is further configured to input the true evaluation value, the test question, and the answer to be tested into the large language model to obtain an evaluation result of the large language model on the answer to be tested based on the true evaluation value and the test question.

[0151] In one possible implementation,

[0152] The processing unit 1802 is further configured to verify the structured text description according to the large language model to determine description errors in the structured text description, where the description errors include contradictions between information in multiple dimensions.

[0153] The processing unit 1802 is further configured to correct description errors in the structured text description.

[0154] The processing unit 1802 is specifically configured to input the corrected structured text description and the test question into the large language model to obtain a true evaluation value output by the large language model.

[0155] In one possible implementation, the test material is an image of a traffic scene, and the information in multiple dimensions includes scene-level information, object-level information, and area-level information. The scene-level information includes an overall description of the traffic scene, the object-level information includes a description of each object in each object category in the traffic scene, and the area-level information includes a description of the target area in the traffic scene.

[0156] In one possible implementation,

[0157] The acquisition unit 1801 is specifically configured to input the test material into the multimodal large model to obtain scene-level information output by the multimodal large model.

[0158] The acquiring unit 1801 is specifically configured to acquire a plurality of object-annotated images. The object-annotated images are obtained by annotating each object in the same object category in the test material. Objects annotated in different object-annotated images belong to different object categories.

[0159] The acquisition unit 1801 is specifically configured to sequentially input a plurality of object-annotated images into the multimodal large model to obtain object-level information output by the multimodal large model.

[0160] The acquiring unit 1801 is specifically configured to acquire a region-annotated image, where the region-annotated image is obtained by annotating a target region in a test material.

[0161] The acquisition unit 1801 is specifically configured to input the region annotation image and the position coordinates of the target region into the multimodal large model to obtain region-level information output by the multimodal large model.

[0162] In one possible implementation, the object categories include vehicles, road users, and roadblocks.

[0163] In one possible implementation,

[0164] The processing unit 1802 is further configured to input the test material into the visual pre-training model to obtain annotation information output by the visual pre-training model, where the annotation information includes annotations of objects in the test material;

[0165] The object annotated image and the region annotated image are obtained based on the annotation information.

[0166] In a possible implementation, the test material is an image of a geometry problem, and the information in multiple dimensions includes geometric element information and geometric relationship information. The geometric element information includes the geometric elements appearing in the geometry problem, and the geometric relationship information includes the geometric relationship between the geometric elements.

[0167] FIG19 is a schematic diagram of the structure of a computing device provided by the present application, which is used to implement the methods in the aforementioned embodiments. The computing device 1900 may include one or more central processing units (CPUs) 1901 and a memory 1905 , wherein the memory 1905 stores one or more applications or data.

[0168] Memory 1905 may be volatile storage or persistent storage. The program stored in memory 1905 may include one or more modules, each of which may include a series of instruction operations on the server. Furthermore, central processing unit 1901 may be configured to communicate with memory 1905 and execute the series of instruction operations in memory 1905 on computing device 1900. Computing device 1900 may also include one or more power supplies 1902, one or more wired or wireless network interfaces 1903, one or more input / output interfaces 1904, and / or one or more operating systems.

[0169] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0170] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0171] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0172] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0173] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A performance evaluation method for a large model, characterized in that: include: Obtaining test materials and test questions, wherein the test questions are questions raised based on the test materials; Obtaining a structured text description of the test material, where the structured text description includes information of multiple dimensions of the test material; Inputting the structured text description and the test question into a large language model to obtain a true evaluation value output by the large language model, where the true evaluation value includes information related to the test question in the test material, or is a standard answer to the test question; Inputting the test material and the test question into the large model to be tested, so as to obtain the answer to the test question output by the large model to be tested; The evaluation truth value, the test question, and the answer to be tested are input into a large language model to obtain an evaluation result of the large language model on the answer to be tested based on the evaluation truth value and the test question.

2. The method according to claim 1, characterized in that The method further comprises: Testing the structured text description according to the large language model to determine description errors in the structured text description, wherein the description errors include contradictions between the information in the multiple dimensions; Correcting the description error in the structured text description; Inputting the structured text description and the test question into the large language model to obtain the true evaluation value output by the large language model includes: The corrected structured text description and the test question are input into the large language model to obtain the true evaluation value output by the large language model.

3. The method according to claim 1 or 2, characterized in that The test material is an image of a traffic scene, and the information in multiple dimensions includes scene-level information, object-level information, and area-level information. The scene-level information includes an overall description of the traffic scene, the object-level information includes a description of each object in each object category in the traffic scene, and the area-level information includes a description of the target area in the traffic scene.

4. The method according to claim 3, characterized in that The step of obtaining a structured text description of the test material includes: Inputting the test material into the multimodal large model to obtain the scene-level information output by the multimodal large model; Acquire a plurality of object-annotated images, wherein the object-annotated images are obtained by annotating each object of the same object category in the test material, and the objects annotated in different object-annotated images belong to different object categories; Inputting the plurality of object-annotated images into the multimodal large model in sequence to obtain the object-level information output by the multimodal large model; Acquire a region-annotated image, where the region-annotated image is obtained by annotating a target region in the test material; The region annotation image and the position coordinates of the target region are input into the multimodal large model to obtain the region-level information output by the multimodal large model.

5. The method according to claim 4, characterized in that The object categories include vehicles, road users, and roadblocks.

6. The method according to claim 5, characterized in that The method further comprises: Inputting the test material into a visual pre-training model to obtain annotation information output by the visual pre-training model, wherein the annotation information includes annotations of objects in the test material; The object annotated image and the region annotated image are acquired based on the annotation information.

7. The method according to claim 2, characterized in that The test material is an image of a geometry problem, and the information in multiple dimensions includes geometric element information and geometric relationship information. The geometric element information includes the geometric elements appearing in the geometry problem, and the geometric relationship information includes the geometric relationship between the geometric elements.

8. A computing device, characterized in that include: an acquiring unit, configured to acquire test materials and test questions, wherein the test questions are questions raised based on the test materials; The acquiring unit is further configured to acquire a structured text description of the test material, wherein the structured text description includes information of multiple dimensions of the test material; a processing unit, configured to input the structured text description and the test question into a large language model to obtain a true evaluation value output by the large language model, wherein the true evaluation value includes information related to the test question in the test material, or is a standard answer to the test question; The processing unit is further configured to input the test material and the test question into the large model to be tested, so as to obtain an answer to the test question output by the large model to be tested; The processing unit is further configured to input the true evaluation value, the test question, and the answer to be tested into the large language model to obtain an evaluation result of the large language model on the answer to be tested based on the true evaluation value and the test question.

9. The device according to claim 8, characterized in that The processing unit is further configured to verify the structured text description according to the large language model to determine description errors in the structured text description, wherein the description errors include contradictions between the information in the multiple dimensions; The processing unit is further configured to correct the description error in the structured text description; The processing unit is specifically configured to input the corrected structured text description and the test question into the large language model to obtain the true evaluation value output by the large language model.

10. The device according to claim 8 or 9, characterized in that The test material is an image of a traffic scene, and the information in multiple dimensions includes scene-level information, object-level information, and area-level information. The scene-level information includes an overall description of the traffic scene, the object-level information includes a description of each object in each object category in the traffic scene, and the area-level information includes a description of the target area in the traffic scene.

11. The device according to claim 10, characterized in that The acquisition unit is specifically configured to input the test material into the multimodal large model to obtain the scene-level information output by the multimodal large model; The acquisition unit is specifically configured to acquire a plurality of object-annotated images, wherein the object-annotated images are obtained by annotating each object of the same object category in the test material, and the objects annotated in different object-annotated images belong to different object categories; The acquisition unit is specifically configured to sequentially input the plurality of object-annotated images into the multimodal large model to acquire the object-level information output by the multimodal large model; The acquisition unit is specifically configured to acquire a region-annotated image, where the region-annotated image is obtained by annotating a target region in the test material; The acquisition unit is specifically configured to input the region annotation image and the position coordinates of the target region into the multimodal large model to acquire the region-level information output by the multimodal large model.

12. The device according to claim 11, characterized in that The object categories include vehicles, road users, and roadblocks.

13. The device according to claim 12, characterized in that The processing unit is further configured to input the test material into a visual pre-training model to obtain annotation information output by the visual pre-training model, wherein the annotation information includes annotations of objects in the test material; The object annotated image and the region annotated image are acquired based on the annotation information.

14. The device according to claim 9, characterized in that The test material is an image of a geometry problem, and the information in multiple dimensions includes geometric element information and geometric relationship information. The geometric element information includes the geometric elements appearing in the geometry problem, and the geometric relationship information includes the geometric relationship between the geometric elements.

15. A computing device, characterized in that including processor and memory; The processor is configured to execute instructions stored in the memory, so that the computing device performs the method according to any one of claims 1 to 7.

16. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device, the computing device is caused to perform the method according to any one of claims 1 to 7.

17. A computer-readable storage medium, characterized in that The method comprises computer program instructions, which, when executed by a computing device cluster, perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image reasoning question and answer method based on priori knowledge inspired large language model

    CN116595151A

  • Knowledge base retrieval method and device based on large language model self-evaluation and self-feedback

    CN116932717A

  • Assessment method of large language model based on dynamic graph

    CN117609439A

  • Content evaluation method, device and equipment for large model scene and storage medium

    CN117744664A

  • Computer implemented methods for the automated analysis or use of data, including use of a large language model

    WO2023161630A1

Cited By

  • Question and answer method and device based on large model, medium, equipment and program product

    CN121303368A

  • Method, device, equipment and medium for generating test sequence for detecting safety performance of large model

    CN121349883A

  • Evaluation method and device of multi-modal model, equipment, storage medium and product

    CN121478616A

  • Geology big data-oriented agent evaluation method, equipment and product

    CN122220476A