Vehicle-end large model evaluation method and system, electronic equipment and readable storage medium

By acquiring real-vehicle scenario data for scenario simulation and answer comparison, the problem of the inability to fully evaluate the capabilities of large models in existing technologies has been solved, enabling a comprehensive and accurate evaluation of large vehicle-side models.

CN120849293APending Publication Date: 2025-10-28CHONGQING CHANGAN AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511217764.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Current technologies for evaluating large models cannot fully reflect their actual capabilities, especially their reasoning processes.

Method used

By acquiring real-vehicle scenario data, scenario simulations are performed, and the inference answers of the target test model are compared with the scenario answers to obtain test results.

Benefits of technology

It enables comprehensive and accurate evaluation of large vehicle-side models, and can assess their perception and reasoning capabilities in specific scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849293A_ABST
    Figure CN120849293A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle-end large model evaluation method and system, electronic equipment and a readable storage medium. The method comprises the steps of obtaining real vehicle scene data and evaluation data corresponding to the real vehicle scene data; performing scene simulation on the target test model through the real vehicle scene data; obtaining associated scene questions and scene answers in the evaluation data; inputting the scene question into the target test model to obtain a reasoning answer; and comparing the reasoning answer with the scene answer to obtain a test result of the target test model. The target test model is tested through the scene question and the scene answer, so that the perception condition of the target test model for the specific scene can be tested, and the evaluation of the specific reasoning process of the target test model can be realized by constructing the evaluation data corresponding to the real vehicle scene data. Therefore, comprehensive and accurate evaluation of the vehicle-end large model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving, and in particular to a method, system, electronic device, and readable storage medium for evaluating large vehicle models. Background Technology

[0002] With the development of autonomous driving technology, large models are being used more and more widely in the field of autonomous driving, and their capabilities are becoming stronger. However, the current technology for evaluating large models usually involves comparing the large model's control of the vehicle with the expected situation. This evaluation method can only determine whether the large model meets the specific driving requirements, but it cannot evaluate the reasoning process of the large model. Therefore, the evaluation results cannot fully reflect the actual capabilities of the large model. Summary of the Invention

[0003] The main objective of this invention is to propose a method, system, electronic device, and readable storage medium for evaluating large vehicle models, aiming to solve the problem that existing evaluations of large models cannot fully reflect their capabilities.

[0004] To achieve the above objectives, the present invention provides a method for evaluating large vehicle models, the method comprising the following steps:

[0005] Acquire real-vehicle scenario data, and the corresponding evaluation data for the real-vehicle scenario data;

[0006] The target test model is simulated using the real vehicle scenario data.

[0007] Obtain the associated scenario questions and answers from the evaluation data;

[0008] The scenario problem is input into the target test model to obtain the reasoning answer;

[0009] The test results of the target test model are obtained by comparing the reasoning answer with the scenario answer.

[0010] Optionally, the step of simulating the target test model using the real vehicle scenario data includes:

[0011] Receive the evaluation command sent by the client and obtain the target model address in the evaluation command;

[0012] The target model address is placed at the end of the evaluation queue;

[0013] The target test models corresponding to the target model addresses in the evaluation queue are tested sequentially.

[0014] When each target test model is tested, the real vehicle scene data is sent to the target test model through the target model address to simulate the scene.

[0015] Optionally, the evaluation data includes multiple question-answer pairs, each pair including an associated scenario question and scenario answer; comparing the inferred answer with the scenario answer to obtain the test result of the target test model includes:

[0016] For each question-answer pair, the reasoned answer is compared with the scenario answer to obtain a sub-score;

[0017] The test results of the target test model are obtained by combining the sub-scores corresponding to each question and answer pair.

[0018] Optionally, the step of synthesizing the sub-scores corresponding to each of the question-and-answer pairs to obtain the test result of the target test model includes:

[0019] Determine the corresponding assessment type for each question-and-answer pair;

[0020] For each of the aforementioned evaluation types, obtain the corresponding sub-score results for the question-and-answer pairs;

[0021] Determine the type score result corresponding to the evaluation type based on the sub-score results;

[0022] The overall score is obtained by combining the scores from each of the aforementioned types.

[0023] The type score and the total score are used as the test result.

[0024] Optionally, when the evaluation type corresponding to the scenario answer and the scenario question is scenario understanding, comparing the reasoning answer with the scenario answer to obtain the test result of the target test model includes:

[0025] Obtain the reasoning elements and their states in the reasoning answer; obtain the scene elements and their states in the scene answer.

[0026] Determine whether the reasoning element is consistent with the scene element, and determine whether the state of the reasoning element is consistent with the state of the scene element;

[0027] If the reasoning element is consistent with the scene element, and the state of the reasoning element is consistent with the state of the scene element, then the test result is determined to be correct.

[0028] Optionally, when the evaluation type corresponding to the scenario answer and the scenario question is multimodal reasoning, the step of comparing the reasoning answer with the scenario answer to obtain the test result of the target test model includes:

[0029] Obtain the reasoning steps and reasoning decisions in the reasoning answer;

[0030] Obtain the reasoning elements in each reasoning step, and the element functions corresponding to the reasoning elements;

[0031] The reasoning logic, which includes reasoning elements and reasoning functions, is generated based on the reasoning sequence of the reasoning steps.

[0032] A reasoning score is obtained by scoring the reasoning logic and the reasoning decision based on the answer to the scenario.

[0033] The reasoning score is used as the test result.

[0034] To achieve the above objectives, the present invention also provides a vehicle-side large model evaluation system, the vehicle-side large model evaluation system including a server, the server comprising:

[0035] The model inference module is used to obtain real vehicle scene data and corresponding evaluation data from the resource configuration module; and to perform scene simulation on the target test model using the real vehicle scene data.

[0036] The indicator calculation module is used to obtain the scenario problems in the evaluation data;

[0037] The model reasoning module is used to input the scenario problem into the target test model to obtain the reasoning answer;

[0038] The referee module is used to obtain the scenario questions and scenario answers associated with the evaluation data;

[0039] The indicator calculation module is used to call the referee module to compare the reasoning answer with the scenario answer to obtain the test result of the target test model.

[0040] Optionally, the vehicle-side large model evaluation system further includes a client, and the server further includes:

[0041] The evaluation scheduling module is used to receive the evaluation instruction sent by the client and obtain the target model address in the evaluation instruction; put the target model address into the end of the evaluation queue; and test the target test models corresponding to the target model addresses in the evaluation queue in turn.

[0042] The model inference module is also used to send the real vehicle scene data to the target test model through the target model address for scene simulation when each target test model is tested.

[0043] To achieve the above objectives, the present invention also provides an electronic device, the electronic device including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the vehicle-end large model evaluation method as described above.

[0044] To achieve the above objectives, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the vehicle-side large model evaluation method as described above.

[0045] This invention proposes a method, system, electronic device, and readable storage medium for evaluating large vehicle-side models. The method involves acquiring real-vehicle scenario data and corresponding evaluation data; simulating scenarios using the real-vehicle scenario data to evaluate a target test model; obtaining scenario questions and answers associated with the evaluation data; inputting the scenario questions into the target test model to obtain inference answers; and comparing the inference answers with the scenario answers to obtain the test results of the target test model. By simulating scenarios using real-vehicle scenario data, the target test model can be tested in the required scenarios. Furthermore, by testing the target test model using scenario questions and answers, the perception of the target test model in specific scenarios can be tested. Finally, by constructing evaluation data corresponding to the real-vehicle scenario data, the specific reasoning process of the target test model can be evaluated, achieving a comprehensive and accurate evaluation of large vehicle-side models. Attached Figure Description

[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart illustrating the first embodiment of the vehicle-side large model evaluation method of the present invention;

[0049] Figure 2 This is a block diagram of the module structure of the vehicle-mounted large model evaluation system of the present invention;

[0050] Figure 3 This is a schematic diagram of the overall process of the vehicle-side large model evaluation method of the present invention;

[0051] Figure 4 This is a schematic diagram of the module structure of the electronic device of the present invention. Detailed Implementation

[0052] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0053] This invention provides a method for evaluating large vehicle models, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the vehicle-side large model evaluation method of the present invention. The method includes the following steps:

[0054] Step S10: Obtain real vehicle scenario data and the corresponding evaluation data;

[0055] Real-vehicle scenario data refers to relevant data collected during actual vehicle driving. The specific data types included in real-vehicle scenario data can be set based on actual evaluation needs, such as sensor multimodal data, vehicle control signals, and environmental parameters. Sensor multimodal data is a multimodal data stream of real road scenes synchronously collected by vehicle-mounted sensors such as cameras, LiDAR, and millimeter-wave radar during driving. Specifically, sensor multimodal data includes visual images collected by cameras and point clouds collected by LiDAR or millimeter-wave radar. Vehicle control signals are vehicle driving status signals, such as steering angle, acceleration, and navigation information output by map software. Environmental parameters indicate the state of the driving environment, such as lighting and weather.

[0056] Real-world scenario data can include multiple sets of data from different scenarios. When conducting evaluations, one or more sets of real-world scenario data can be selected to evaluate the target test model.

[0057] The evaluation data is constructed based on real-vehicle scenario data and is used to evaluate the target test model.

[0058] The number of scenario questions included in the evaluation data can be set based on actual needs, and the specific scenario questions and answers can be obtained by labeling the actual vehicle scenario data used.

[0059] The system pre-constructs corresponding evaluation data for various real-vehicle scenario data. When it is necessary to test the target test model, it receives the evaluation instruction sent by the client; obtains the evaluation data identifier in the evaluation instruction; and retrieves the real-vehicle scenario data corresponding to the evaluation data identifier and the evaluation data from the evaluation database.

[0060] Step S20: Simulate the target test model using the real vehicle scene data;

[0061] The target test model is the large vehicle-side model that needs to be evaluated.

[0062] After obtaining real vehicle scenario data, scenario simulation can be performed on the target test model. The specific scenario simulation method can be set according to actual needs, such as simulating real vehicle scenarios by injecting real vehicle scenario data into the target test model in the form of fault injection.

[0063] Step S30: Obtain the scenario questions and scenario answers associated with the evaluation data;

[0064] In this embodiment, the evaluation data is set as a set of associated scenario questions and scenario answers; the scenario questions are questions raised for the driving scenarios corresponding to the real vehicle scenario data; for example, for a scenario of an intersection, questions can be asked about the elements in the scenario, driving strategies, etc., such as "How many lanes are there before the intersection?", "What is the driving direction of the middle lane before the intersection?", "What targets need to be paid attention to before the intersection?", "Which lane should this vehicle take?"

[0065] The scenario answers are the answers corresponding to the scenario questions in the driving scenarios corresponding to the actual vehicle scenario data; the scenario questions and scenario answers are set in a one-to-one correspondence; for example, for the scenario question "How many lanes are there before the intersection?", if the actual vehicle scenario data indicates that the intersection contains three lanes, then the corresponding scenario answer can be set to "three lanes"; for the scenario question "What is the driving direction of the middle lane before the intersection?", if the actual vehicle scenario data indicates that the driving direction of the middle lane before the intersection is left turn, then the corresponding scenario answer can be set to "left turn"; for the scenario question "What targets need to be paid attention to before the intersection?", if the actual vehicle scenario data indicates that there is a white stationary vehicle in the second lane from the left, then the corresponding scenario answer can be set to "There is a white vehicle in the second lane from the left".

[0066] Step S40: Input the scenario problem into the target test model to obtain the reasoning answer;

[0067] After the scenario problem is input into the target test model, the target test model infers the scenario problem based on the simulated scenario and outputs the inferred answer to the scenario problem.

[0068] Step S50: Compare the reasoning answer with the scenario answer to obtain the test result of the target test model.

[0069] It is understandable that the inference answer is the output of the target test model based on its perception of the simulated scenario, that is, the inference answer reflects the perception of the target test model; while the scenario answer indicates the standard answer corresponding to the scenario question; therefore, by comparing the inference answer and the scenario answer, we can clarify the accuracy of the target test model's perception of the scenario.

[0070] It should be noted that this application differs from existing technologies that use evaluation methods based on the actual control of the vehicle. Existing evaluation methods detect relevant data when a large model controls the vehicle for autonomous driving, such as speed, legality, and accident occurrences, and compare the detected results with expected results to evaluate the large model. This approach completely ignores the reasoning process of the large model, focusing only on the model's control results over the vehicle. However, this application evaluates the target test model's perception of a specific simulated scenario. Specifically, it asks the target test model scenario-related questions and determines whether the model's output corresponds to the simulated scenario, thereby determining the target test model's perception of the specific scenario. By designing specific evaluation data, it is possible to evaluate the target test model's perception of specific aspects of scenario reasoning, such as the perception of object types in the scenario (e.g., vehicles, pedestrians, lanes, traffic lights); the perception of driving strategies in the scenario (e.g., steering direction, vehicle speed); and the perception of object information in the scenario (e.g., traffic sign text, object color), thus evaluating the specific reasoning process of the target test model.

[0071] By simulating scenarios using real-vehicle scene data, the target test model can be tested in the required scenarios. Furthermore, by testing the target test model through scenario questions and answers, the perception of the target test model in specific scenarios can be tested. By constructing evaluation data corresponding to real-vehicle scene data, the specific reasoning process of the target test model can be evaluated, thus achieving a comprehensive and accurate evaluation of the large vehicle-side model.

[0072] Furthermore, in the second embodiment of the vehicle-side large model evaluation method of the present invention based on the first embodiment of the present invention, step S20 includes the following steps:

[0073] Step S21: Receive the evaluation instruction sent by the client and obtain the target model address in the evaluation instruction;

[0074] Step S22: Place the target model address at the end of the evaluation queue;

[0075] Step S23: Test the target test models corresponding to the target model addresses in the evaluation queue in sequence;

[0076] Step S24: When each of the target test models is being tested, the real vehicle scene data is sent to the target test model through the target model address to simulate the scene.

[0077] Evaluation instructions are commands sent by the client to evaluate the target test model; the evaluation instructions specifically indicate the information of the target test model and related evaluation information.

[0078] The target model address is the address of the target test model.

[0079] The target test model can be connected to the vehicle-side large model evaluation system through the provided interface. After being connected to the vehicle-side large model evaluation system, the target test model will be assigned an address, namely the target model address.

[0080] It is understandable that there can be multiple target test models that need to be evaluated. However, the number of models that the vehicle-side large model evaluation system can support for evaluation at the same time is limited. Therefore, it is necessary to schedule the evaluation order of the target test models. In this embodiment, an evaluation queue is set up, and the target model addresses corresponding to the target test models are sorted in the evaluation queue according to the order in which the evaluation instructions are sent. Users can also set the target test models to be evaluated first on the client. The target test models to be evaluated first can be inserted into the evaluation queue for faster evaluation.

[0081] The number of target test models that can be tested simultaneously in the vehicle-side large model evaluation system can be set according to actual needs; this embodiment takes supporting one target test model at the same time as an example for explanation.

[0082] In practical implementation, the evaluation scheduling model can periodically read the address of the top-ranked target model in the evaluation queue and determine whether there is a target test model currently being tested. If there is a target test model currently being tested, wait for the next address read. If there is no target test model currently being tested, test the target test model corresponding to the top-ranked target model address in the evaluation queue. Specifically, the real vehicle scenario data is sent to the target test model through the target model address for scenario simulation, and subsequent steps are executed.

[0083] In this embodiment, by maintaining an evaluation queue, the testing of multiple target test models can be scheduled, thereby enabling the evaluation of multiple target test models.

[0084] Furthermore, in the third embodiment of the vehicle-side large model evaluation method of the present invention based on the first embodiment, the evaluation data includes multiple sets of question-answer pairs, each set of question-answer pairs including the associated scenario question and the scenario answer; step S50 includes the following steps:

[0085] Step S51: For each question-answer pair, compare the reasoning answer with the scenario answer to obtain a sub-score;

[0086] Step S52: Combine the sub-scores corresponding to each question-and-answer pair to obtain the test results of the target test model.

[0087] A question-and-answer pair consists of a related scenario question and a scenario answer; for example, scenario question: "How many lanes are there at the intersection ahead?" - scenario answer: "Three lanes" forms a question-and-answer pair.

[0088] As the evaluation requirements for large vehicle models become increasingly stringent, it is necessary to set up more comprehensive test items for the evaluation of large vehicle models. In this embodiment, by setting up evaluation data containing multiple sets of question-answer pairs, it is possible to test the target test model's perception of the scene in different aspects through diverse questions, and to comprehensively evaluate the overall perception capability of the target test model.

[0089] The sub-score is the score for a question-answer pair. The final test result of the target test model can be obtained by combining the scores of all question-answer pairs.

[0090] The specific method for combining sub-scores can be set based on actual needs, such as average, weighted average, etc. For example, if the evaluation data contains six question-answer pairs, and the sub-scores corresponding to the question-answer pairs are 100, 1, 1, 100, 100, and 100 respectively, then the test result obtained by averaging is 67.

[0091] Further, step S52 includes the following steps:

[0092] Step S521: Determine the evaluation type corresponding to each question-answer pair;

[0093] Step S522: For each of the evaluation types, obtain the sub-score results of the corresponding question-answer pair;

[0094] Step S523: Determine the type score result corresponding to the evaluation type based on the sub-score result;

[0095] Step S524: Combine the scoring results of each type to obtain the total score;

[0096] Step S525: The type scoring result and the total scoring result are used as the test result.

[0097] Evaluation type is used to indicate the question-answering system's perception ability in a specific direction for the target test model; for example, evaluation type may include scene understanding, multimodal reasoning, and interpretive behavior understanding.

[0098] Scene understanding refers to the target test model's perception of specific elements within a scene, such as environmental understanding, target understanding, and driving risk understanding. Environmental understanding refers to the perception of the state of the driving environment, such as weather conditions and visibility. Target understanding refers to the perception of specific objects within the scene, such as vehicles, pedestrians, and traffic signs. Driving risk understanding refers to the perception of driving risks present in the scene, such as collisions and traffic violations. For this type of assessment, the question-and-answer pairs can be set with specific element conditions. For example, scene questions can ask about the type and state of elements, while scene answers specify the type and state of the elements.

[0099] 1. {"q":"How many lanes are there at the intersection ahead?"}, {"a":"Three lanes"}

[0100] 2. {"q":"What is the direction of travel in the middle lane before the intersection ahead?"}, {"a":"Turn left"}

[0101] 3. {"q":"What targets need to be monitored at the upcoming intersection? What is their state of motion?"}, {"a":"The white vehicle in the second lane from the left is stationary."}

[0102] Where q represents the scenario question and a represents the scenario answer, and so on.

[0103] Multimodal reasoning assesses the target test model's perception of the reasoning logic of control strategies within a scenario. Specifically, it can include a continuous reasoning chain formed by the influence of elements within the scenario, as well as the final decision; for example:

[0104] 1. {"q":"Which lane should this vehicle use?"}, {"a":"The navigation system requires this vehicle to turn left. Information from the telephoto lens ahead shows that the first and second lanes from the left are open for left turns, so this vehicle can use these two lanes. However, since there is a vehicle stationary in the second lane waiting for the red light at the intersection, this vehicle should turn left into the first lane from the left to pass through the intersection more efficiently. Therefore, this vehicle should use the first lane from the left."}

[0105] As can be seen from the above examples, in multimodal reasoning, specific elements within the scene are determined, such as the navigation indicating a left turn, the first and second lanes being left-turn lanes, and a car waiting at a red light in the second lane. Then, the feasibility of driving is determined based on the influence of these elements. Since the navigation indicates a left turn, the vehicle can only use the first and second lanes. However, since there is a car waiting in the second lane, the first lane is more efficient than the second lane. Therefore, the final driving strategy is to use the first lane.

[0106] Explaining behavior comprehension refers to the logical understanding of driving behavior reasoning and decision-making. For this type of assessment, the question-and-answer pair can be designed to suggest setting scenario-based questions that describe driving behavior in natural language, and setting the scenario-based answers to match the driving behavior reasoning and decision-making within that scenario; for example:

[0107] 1. {"q":"Why should this car take the first lane from the left?"}, {"a":"The navigation system requires this car to turn left. The telephoto lens shows that the first and second lanes from the left ahead are open for left turns, so this car can use these two lanes. However, since there is a car stationary in the second lane waiting for the red light at the intersection, this car should turn left into the first lane from the left to pass through the intersection more efficiently. Therefore, this car should take the first lane from the left."}

[0108] 2. {"q":"If I turn right ahead, which lane should I use?"}, {"a":"As seen in the traffic sign information from the telephoto lens, the first and second lanes from the left are for left turns, and the first lane from the right is for right turns. Therefore, if I want to turn right ahead, I need to use the rightmost lane."}

[0109] As can be seen from the above examples, in explaining behavioral understanding, the reasoning logic of driving behavior decisions is specifically tested.

[0110] Different assessment types have different testing tendencies for the target test model. In this application, by setting corresponding question-answer pairs for multiple assessment types, it is possible to test different aspects of the target test model's capabilities. For each assessment type, a type score is obtained based on the sub-scores of the question-answer pairs corresponding to that assessment type, thereby revealing the target test model's perceptual ability in that assessment type. Finally, a total score is obtained based on the type score results corresponding to all assessment types, enabling the assessment of the target test model's overall perceptual ability.

[0111] The specific method for obtaining the type scoring results can be set based on actual needs, such as calculating the average of sub-scores.

[0112] The specific method for obtaining the overall score can be set based on actual needs, such as averaging the type scores or performing a weighted average. For example, users can set the priority of different evaluation types of the target test results through the client. If it is necessary to focus on testing the multimodal reasoning perception ability of the target test model, a higher weight can be set for the multimodal reasoning type score results, while a lower weight can be set for the score results corresponding to other evaluation types, thereby controlling the bias of the evaluation.

[0113] Furthermore, in the fourth embodiment of the vehicle-side large model evaluation method of the present invention based on the first embodiment of the present invention, when the evaluation type corresponding to the scene answer and the scene question is scene understanding, step S50 includes the following steps:

[0114] Step S53: Obtain the reasoning elements and their states in the reasoning answer; obtain the scene elements and their states in the scene answer.

[0115] Step S55: Determine whether the reasoning element is consistent with the scene element, and determine whether the state of the reasoning element is consistent with the state of the scene element;

[0116] Step S55: If the reasoning element is consistent with the scene element, and the state of the reasoning element is consistent with the state of the scene element, then the test result is determined to be correct.

[0117] The reasoning element is an element that exists in the scene indicated in the reasoning answer, and the reasoning element state indicates the specific state of the reasoning element.

[0118] For elements in a scene, the target test model needs to be able to accurately perceive the elements, and the specific content to be perceived needs to be clear about the type and state of the elements.

[0119] Scene elements are the correct elements set in the scene answer, and scene element states are the correct element states set in the scene answer.

[0120] If the reasoning element is consistent with the scene element and the state of the reasoning element is consistent with the state of the scene element, the reasoning answer is considered to be consistent with the scene answer and the reasoning answer is correct; however, if the reasoning element is inconsistent with the scene element or the state of the reasoning element is inconsistent with the state of the scene element, the test result is determined to be incorrect.

[0121] The scoring for whether a specific answer is correct or not can be set based on actual needs. For example, in this embodiment, a percentage system is used. When the reasoning answer is correct, the corresponding sub-score is 100, and when the reasoning answer is incorrect, the corresponding sub-score is 1. Understandably, when multiple items need to be compared are included, scoring can also be based on the number of correct answers. For example, if the reasoning element is inconsistent with the scene element, and the state of the reasoning element is inconsistent with the state of the scene element, the test result is determined to be incorrect and scored as 1. If the reasoning element is inconsistent with the scene element, and the state of the reasoning element is consistent with the state of the scene element, the test result is determined to be incorrect and scored as 50. If the reasoning element is consistent with the scene element, and the state of the reasoning element is inconsistent with the state of the scene element, the test result is determined to be incorrect and scored as 50. For example, taking a specific intersection scenario, the specific information of the intersection scenario includes the image captured by the telephoto camera. There are three lanes to choose from in front of the vehicle. The first and second lanes are for left turns, and the third lane is for right turns. There is a car waiting at a red light in the second lane, and a yellow taxi is driving in the opposite lane. The navigation shows that the vehicle needs to turn left in front. Subsequent multimodal reasoning and behavioral understanding are also based on this scenario and will not be elaborated further.

[0122] The question-and-answer pairs corresponding to scene understanding include:

[0123] 1. {"q":"How many lanes are there at the intersection ahead?"}, {"a":"Three lanes"}

[0124] 2. {"q":"What is the direction of travel in the middle lane before the intersection ahead?"}, {"a":"Turn left"}

[0125] 3. {"q":"What targets need to be monitored at the upcoming intersection? What is their state of motion?"}, {"a":"The white vehicle in the second lane from the left is stationary."}

[0126] The reasoning answers corresponding to scene understanding include:

[0127] 1. {"q":"How many lanes are there at the intersection ahead?"}, {"a":"Three lanes"}

[0128] 2. {"q":"What is the direction of travel in the middle lane before the next intersection?"}, {"a":"Turn right"}

[0129] 3. {"q":"What targets should we pay attention to at the intersection ahead? What is their state of motion?"}, {"a":"The yellow taxi on the left side of this vehicle is moving."}

[0130] By comparing the question-and-answer pairs with the inference answers, we can see that the inference answer corresponding to question-and-answer pair 1 is correct, while the inference answers corresponding to question-and-answer pairs 2 and 3 are incorrect. Therefore, the scores for the three inference answers are 100, 1, and 1, respectively, and the score for the type of scene understanding is 34.

[0131] When the evaluation type corresponding to the scenario answer and the scenario question is multimodal reasoning, step S50 includes the following steps:

[0132] Step S56: Obtain the reasoning steps and reasoning decisions in the reasoning answer;

[0133] Step S57: Obtain the reasoning elements in each reasoning step, and the element functions corresponding to the reasoning elements;

[0134] Step S58: Generate reasoning logic containing reasoning elements and reasoning functions based on the reasoning sequence of the reasoning steps;

[0135] Step S59: Based on the scenario answer, score the reasoning logic and the reasoning decision to obtain a reasoning score;

[0136] Step S510: The reasoning score is used as the test result.

[0137] The reasoning steps are the smallest unit of reasoning in the reasoning answer; for example, if the navigation requires a left turn, then it is determined that the first lane and the second lane can be used.

[0138] Reasoning leads to the final driving decision, such as choosing to drive in the first lane.

[0139] The function of an element indicates the role it plays in the reasoning and decision-making process. For example, the function of navigation is to require a left turn; the function of the first lane, second lane, and third lane is to indicate the direction in which the lane is allowed to turn; and the function of a stationary car in the second lane is to indicate that the second lane is less efficient than the first lane.

[0140] The reasoning order is the order of the reasoning steps in the reasoning answer; for example, first determine the need for a left turn through navigation, then determine the lane that can be entered based on lane attributes, then determine the driving efficiency based on the vehicles in the lane, and finally obtain the lane to enter.

[0141] Generating reasoning logic based on reasoning order, which includes reasoning elements and reasoning functions, can reflect the target test model's perception of the reasoning logic of the control strategy; furthermore, scoring the reasoning logic and reasoning decisions through scenario answers can yield a reasoning score.

[0142] For example, the question-answer pairs corresponding to multimodal reasoning include:

[0143] 4. {"q":"Which lane should this vehicle use?"}, {"a":"The navigation system requires this vehicle to turn left. The telephoto lens shows that the first and second lanes from the left ahead are available for left turns, so this vehicle can use these two lanes. However, since there is a vehicle stationary in the second lane waiting for the red light at the intersection, this vehicle should turn left into the first lane from the left to pass through the intersection more efficiently. Therefore, this vehicle should use the first lane from the left."}

[0144] The inference answers corresponding to multimodal reasoning include:

[0145] 4. {"q":"Which lane should this vehicle use?"}, {"a":"The navigation system requires this vehicle to turn left. The telephoto lens shows that the first and second lanes from the left ahead are available for left turns, so this vehicle can use these two lanes. However, since there is a vehicle stationary in the second lane waiting for the red light at the intersection, this vehicle should turn left into the first lane from the left to pass through the intersection more efficiently. Therefore, this vehicle should use the first lane from the left."}

[0146] By comparing the question-answer pairs with the inferred answers, we can see that the inferred answer for question-answer pair 4 is correct, and the score for the inferred answer is 100; the corresponding multimodal reasoning type score is also 100. It should be noted that since multimodal reasoning evaluates reasoning logic and reasoning strategies, it is difficult to determine the correctness of the inferred answer through direct comparison in scene understanding. In practical applications, a referee model can be set up to compare the question with the inferred answer and obtain the corresponding sub-score. Similarly, for the evaluation types of scene understanding and explanatory behavior understanding, a referee model can also be used for scoring. The specific type of referee model can be set according to actual needs, such as an open-source large language model or a large language model fine-tuned for the autonomous driving field.

[0147] When the evaluation type corresponding to the scenario answer and the scenario question is explanatory behavior understanding, step S50 includes the following steps:

[0148] Step S511: Obtain the reasoning steps in the reasoning answer;

[0149] Step S512: Obtain the reasoning elements in each reasoning step, and the element functions corresponding to the reasoning elements;

[0150] Step S513: Generate reasoning logic containing reasoning elements and reasoning functions based on the reasoning sequence of the reasoning steps;

[0151] Step S514: Score the reasoning logic based on the scenario answer to obtain a reasoning score;

[0152] Step S515: The reasoning score is used as the test result.

[0153] Explaining behavioral understanding can be analogous to multimodal reasoning.

[0154] For example, the question-and-answer pairs corresponding to explaining behavioral understanding include:

[0155] 5. {"q":"Why should this car take the first lane from the left?"}, {"a":"The navigation system requires this car to turn left. The telephoto lens shows that the first and second lanes from the left ahead are open for left turns, so this car can use these two lanes. However, since there is a car stationary in the second lane waiting for the red light at the intersection, this car should turn left into the first lane from the left to pass through the intersection more efficiently. Therefore, this car should take the first lane from the left."}

[0156] 6. {"q":"If I turn right ahead, which lane should I use?"}, {"a":"As seen in the traffic sign information in the telephoto lens, the first and second lanes from the left are for left turns, and the first lane from the right is for right turns. Therefore, if I want to turn right ahead, I need to use the rightmost lane."}

[0157] The reasoning answers corresponding to the explanation of behavior include:

[0158] 5. {"q":"Why should this car take the first lane from the left?"}, {"a":"The navigation system requires this car to turn left. The telephoto lens shows that the first and second lanes from the left ahead are open for left turns, so this car can use these two lanes. However, since there is a car stationary in the second lane waiting for the red light at the intersection, this car should turn left into the first lane from the left to pass through the intersection more efficiently. Therefore, this car should take the first lane from the left."}

[0159] 6. {"q":"If I turn right ahead, which lane should I use?"}, {"a":"As seen in the traffic sign information in the telephoto lens, the first and second lanes from the left are for left turns, and the first lane from the right is for right turns. Therefore, if I want to turn right ahead, I need to use the rightmost lane."}

[0160] By comparing the question-and-answer pairs with the inference answers, it can be seen that the inference answers corresponding to question-and-answer pairs 5 and 6 are correct, and the scores for the inference answers are both 100; the corresponding type score for explaining the behavior comprehension is 100.

[0161] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0163] This application also provides a vehicle-side large model evaluation system for implementing the above-described vehicle-side large model evaluation method, see [link to relevant documentation]. Figure 2 The vehicle-side large model evaluation system includes a server, and the server includes:

[0164] The model inference module is used to obtain real vehicle scene data and corresponding evaluation data from the resource configuration module; and to perform scene simulation on the target test model using the real vehicle scene data.

[0165] The indicator calculation module is used to obtain the scenario problems in the evaluation data;

[0166] The model reasoning module is used to input the scenario problem into the target test model to obtain the reasoning answer;

[0167] The referee module is used to obtain the scenario questions and scenario answers associated with the evaluation data;

[0168] The indicator calculation module is used to call the referee module to compare the reasoning answer with the scenario answer to obtain the test result of the target test model.

[0169] This vehicle-side large-scale model evaluation system simulates target test models using real-vehicle scenario data, enabling testing of the target test models in the required scenarios. Furthermore, it tests the target test models by presenting scenario questions and answers, allowing for the evaluation of the target test models' perception of specific scenarios. By constructing evaluation data corresponding to real-vehicle scenario data, it can evaluate the specific reasoning process of the target test models, achieving a comprehensive and accurate evaluation of vehicle-side large-scale models.

[0170] Furthermore, the vehicle-side large model evaluation system also includes a client, and the server also includes:

[0171] The evaluation scheduling module is used to receive the evaluation instruction sent by the client and obtain the target model address in the evaluation instruction; put the target model address into the end of the evaluation queue; and test the target test models corresponding to the target model addresses in the evaluation queue in turn.

[0172] The model inference module is also used to send the real vehicle scene data to the target test model through the target model address for scene simulation when each target test model is tested.

[0173] Furthermore, the evaluation data includes multiple question-answer pairs, each of which includes the associated scenario question and scenario answer; the model inference module compares the inferred answer with the scenario answer for each question-answer pair to obtain a sub-score; and the test result of the target test model is obtained by combining the sub-scores corresponding to each question-answer pair.

[0174] Furthermore, the model reasoning module determines the evaluation type corresponding to each question-answer pair; for each evaluation type, it obtains the sub-score result of the corresponding question-answer pair; it determines the type score result corresponding to the evaluation type based on the sub-score result; it combines the type score results to obtain the total score result; and it uses the type score result and the total score result as the test result.

[0175] Furthermore, when the evaluation type corresponding to the scenario answer and the scenario question is scenario understanding, the model reasoning module obtains the reasoning elements and the state of the reasoning elements in the reasoning answer, and obtains the scenario elements and the state of the scenario elements in the scenario answer; it determines whether the reasoning elements are consistent with the scenario elements, and determines whether the state of the reasoning elements is consistent with the state of the scenario elements; if the reasoning elements are consistent with the scenario elements, and the state of the reasoning elements is consistent with the state of the scenario elements, then the test result is determined to be correct.

[0176] Furthermore, when the evaluation type corresponding to the scenario answer and the scenario question is multimodal reasoning, the model reasoning module obtains the reasoning steps and reasoning decisions in the reasoning answer; obtains the reasoning elements in each reasoning step, and the element functions corresponding to the reasoning elements; generates reasoning logic containing reasoning elements and reasoning functions based on the reasoning order of the reasoning steps; scores the reasoning logic and the reasoning decisions based on the scenario answer to obtain a reasoning score; and uses the reasoning score as the test result.

[0177] The following combination Figure 2 , Figure 3 The overall implementation principle of the vehicle-side large model evaluation system is explained below:

[0178] 1. Server Setup:

[0179] a. Prepare real vehicle scenario data and corresponding evaluation data and store them in the hardware resource configuration. The hardware resource configuration can set up acceleration modules to provide computing power for other modules, such as evaluation scheduling module, model inference module, judge model, index calculation module, and acceleration modules such as GPU (Graphics Processing Unit).

[0180] b. Deploy the referee module; deploy the referee model on hardware resources and provide the referee model's API (Application Programming Interface) for the indicator calculation module to use.

[0181] c. Model Inference Module: Infers the target test model based on the target model address. Specifically, the model inference module reads the real vehicle scene data and evaluation data from the hardware resource configuration, loads the target test model through the target model address, and performs scene simulation through the real vehicle scene data. It then infers the scene questions from the evaluation data to the target test model to obtain the inference answer, sends the inference answer to the index calculation model, and stores it along with the corresponding evaluation data in the hardware resources.

[0182] d. Index calculation module: Obtain the reasoning answer output by the model reasoning module, and compare the reasoning answer with the scenario answer by calling the judge model through the judge model's API to obtain the sub-score results. Combine the sub-score results to obtain the type score result and the total score result. Store each score result in hardware resources and send it to the evaluation scheduling module.

[0183] e. Evaluation scheduling module; When there is more than one target test model waiting to be tested, the evaluation scheduling module maintains the evaluation queue; according to the queue order, the address of the target model in the queue is provided to the model inference module for inference to evaluate the target test model; after the evaluation is completed, the indicator calculation module is automatically called to calculate the scoring result, and the scoring result of the indicator calculation module is obtained, and the scoring result information and the storage address in the hardware resource configuration are sent to the client.

[0184] 2. Client-side construction.

[0185] a. Evaluation Submission Module: This module provides a user interface where users can input or select the target model address and evaluation data, such as the type and version of the evaluation data. This generates evaluation instructions that are then sent to the server's evaluation scheduling module. The evaluation scheduling module schedules the evaluation according to the evaluation instructions, enabling each module to perform its own function to evaluate the target test model.

[0186] b. Evaluation result display module; used to display the evaluation results. Specifically, it obtains the scoring results information sent by the evaluation scheduling module and their storage address in the hardware resource configuration. Based on the scoring results information and the information obtained from the storage address, it can display specific evaluation-related content:

[0187] i. Type Scoring Results

[0188] ii. Overall score

[0189] iii. Data categorization and display where the reasoning answer differs from the scenario answer; specifically, this can include visualized real-vehicle scenario data, scenario questions, scenario answers, reasoning answers, scoring results, etc., to enable the analysis of error samples and detailed performance analysis of the model.

[0190] Reference Figure 4 In terms of hardware structure, the electronic device may include components such as a communication module 10, a memory 20, and a processor 30. In the electronic device, the processor 30 is connected to both the memory 20 and the communication module 10. The memory 20 stores a computer program, which is executed by the processor 30. When the computer program is executed, it implements the steps of the above-described method embodiments.

[0191] The communication module 10 can connect to external communication devices via a network. The communication module 10 can receive requests from the external communication devices and can also send requests, instructions, and information to the external communication devices. The external communication devices can be other electronic devices, servers, or IoT devices, such as televisions, etc.

[0192] The memory 20 can be used to store software programs and various data. The memory 20 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as acquiring real-world vehicle scenario data), etc.; the data storage area may include a database, and may store data or information created based on system usage. Furthermore, the memory 20 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0193] The processor 30 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 20, and by calling data stored in the memory 20, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. The processor 30 may include one or more processing units; optionally, the processor 30 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 30.

[0194] although Figure 4 Not shown, but the above-described electronic device may further include a circuit control module for connecting to a power supply to ensure the normal operation of other components. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0195] The present invention also proposes a computer-readable storage medium having a computer program stored thereon. The computer-readable storage medium may be... Figure 4 The memory 20 in the electronic device may also be at least one of ROM (Read-Only Memory) / RAM (Random Access Memory), magnetic disk, optical disk, etc. The computer-readable storage medium includes a number of instructions to cause a terminal device with a processor (which may be a television, automobile, mobile phone, computer, server, terminal, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0196] In this invention, the terms "first," "second," "third," "fourth," and "fifth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0197] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0198] Although embodiments of the present invention have been shown and described above, the scope of protection of the present invention is not limited thereto. It is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, and substitutions to the above embodiments within the scope of the present invention, and such changes, modifications, and substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for evaluating large vehicle-mounted models, characterized in that, The vehicle-side large model evaluation method includes: Acquire real-vehicle scenario data, and the corresponding evaluation data for the real-vehicle scenario data; The target test model is simulated using the real vehicle scenario data. Obtain the associated scenario questions and answers from the evaluation data; The scenario problem is input into the target test model to obtain the reasoning answer; The test results of the target test model are obtained by comparing the reasoning answer with the scenario answer.

2. The vehicle-side large model evaluation method as described in claim 1, characterized in that, The step of simulating the target test model using the real vehicle scene data includes: Receive the evaluation command sent by the client and obtain the target model address in the evaluation command; The target model address is placed at the end of the evaluation queue; The target test models corresponding to the target model addresses in the evaluation queue are tested sequentially. When each target test model is tested, the real vehicle scene data is sent to the target test model through the target model address to simulate the scene.

3. The vehicle-side large model evaluation method as described in claim 1, characterized in that, The evaluation data includes multiple question-answer pairs, each of which includes the associated scenario question and scenario answer; The process of comparing the reasoning answer with the scenario answer to obtain the test result of the target test model includes: For each question-answer pair, the reasoned answer is compared with the scenario answer to obtain a sub-score; The test results of the target test model are obtained by combining the sub-scores corresponding to each question and answer pair.

4. The vehicle-side large model evaluation method as described in claim 3, characterized in that, The test results of the target test model obtained by combining the sub-scores corresponding to each of the question-and-answer pairs include: Determine the corresponding assessment type for each question-and-answer pair; For each of the aforementioned evaluation types, obtain the corresponding sub-score results for the question-and-answer pairs; Determine the type score result corresponding to the evaluation type based on the sub-score results; The overall score is obtained by combining the scores from each of the aforementioned types. The type score and the total score are used as the test result.

5. The vehicle-side large model evaluation method as described in claim 1, characterized in that, When the evaluation type corresponding to the scenario answer and the scenario question is scenario understanding, the step of comparing the reasoning answer with the scenario answer to obtain the test result of the target test model includes: Obtain the reasoning elements and their states in the reasoning answer; obtain the scene elements and their states in the scene answer. Determine whether the reasoning element is consistent with the scene element, and determine whether the state of the reasoning element is consistent with the state of the scene element; If the reasoning element is consistent with the scene element, and the state of the reasoning element is consistent with the state of the scene element, then the test result is determined to be correct.

6. The vehicle-side large model evaluation method as described in claim 1, characterized in that, When the evaluation type corresponding to the scenario answer and the scenario question is multimodal reasoning, the step of comparing the reasoning answer with the scenario answer to obtain the test result of the target test model includes: Obtain the reasoning steps and reasoning decisions in the reasoning answer; Obtain the reasoning elements in each reasoning step, and the element functions corresponding to the reasoning elements; The reasoning logic, which includes reasoning elements and reasoning functions, is generated based on the reasoning sequence of the reasoning steps. A reasoning score is obtained by scoring the reasoning logic and the reasoning decision based on the answer to the scenario. The reasoning score is used as the test result.

7. A vehicle-mounted large model evaluation system, characterized in that, The vehicle-mounted large model evaluation system includes a server, and the server includes: The model inference module is used to obtain real vehicle scene data and corresponding evaluation data from the resource configuration module; and to perform scene simulation on the target test model using the real vehicle scene data. The indicator calculation module is used to obtain the scenario problems in the evaluation data; The model reasoning module is used to input the scenario problem into the target test model to obtain the reasoning answer; The referee module is used to obtain the scenario questions and scenario answers associated with the evaluation data; The indicator calculation module is used to call the referee module to compare the reasoning answer with the scenario answer to obtain the test result of the target test model.

8. The vehicle-side large model evaluation system as described in claim 7, characterized in that, The vehicle-mounted large model evaluation system also includes a client, and the server also includes: The evaluation scheduling module is used to receive the evaluation instruction sent by the client and obtain the target model address in the evaluation instruction; put the target model address into the end of the evaluation queue; and test the target test models corresponding to the target model addresses in the evaluation queue in turn. The model inference module is also used to send the real vehicle scene data to the target test model through the target model address for scene simulation when each target test model is tested.

9. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the vehicle-end large model evaluation method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the vehicle-end large model evaluation method as described in any one of claims 1 to 6.