Evaluation method of intelligent cabin model and construction method of cabin index evaluation system

By defining evaluation strategies for preset usage scenarios in the intelligent cockpit model and combining them with multi-dimensional assessments, the problem of low evaluation accuracy in existing technologies is solved, and more accurate model performance evaluation is achieved.

CN120872764APending Publication Date: 2025-10-31CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510975407.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Current technologies for evaluating smart cockpit models primarily focus on testing single functions, neglecting the model's overall performance across multiple scenarios, resulting in low evaluation accuracy.

Method used

By identifying the preset usage scenarios associated with the intelligent cockpit model under test, formulating relevant evaluation strategies, and calling the target cockpit indicator evaluation system, the model performance is evaluated from multiple evaluation dimensions (model application, system application, scenario application, user experience), thus constructing the target cockpit indicator evaluation system.

Benefits of technology

This improved the accuracy of the intelligent cockpit model evaluation, ensuring that the evaluation results are closely related to actual application performance, and enhancing the relevance and practicality of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872764A_ABST
    Figure CN120872764A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an evaluation method of an intelligent cockpit model and a construction method of a cockpit index evaluation system. The evaluation method of the intelligent cockpit model comprises the steps of determining a to-be-tested intelligent cockpit model in an intelligent cockpit system of a vehicle; determining an evaluation strategy of the to-be-tested intelligent cabin model based on a preset use scene associated with the to-be-tested intelligent cabin model; according to the evaluation strategy, evaluating the to-be-tested intelligent cabin model to obtain an evaluation result; calling a target cabin index evaluation system, and evaluating the model performance of the to-be-tested intelligent cabin model from a plurality of evaluation dimensions according to the evaluation result to obtain a performance evaluation result of the to-be-tested intelligent cabin model, the evaluation dimensions at least comprise a model application evaluation dimension, a system application evaluation dimension, a scene application evaluation dimension and a user experience evaluation dimension. The technical problem of low evaluation accuracy of the intelligent cockpit model in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and more specifically, to a method for evaluating an intelligent cockpit model and a method for constructing a cockpit index evaluation system. Background Technology

[0002] With the development and improvement of vehicle intelligence, large-scale models are increasingly being applied in intelligent cockpits. Intelligent cockpits not only integrate advanced infotainment systems but also incorporate various functions such as driver assistance, environmental perception, and personalized services, greatly enhancing the driving experience and safety. Therefore, to ensure that the intelligent cockpit model within the intelligent cockpit system can accurately understand and respond to user needs, thereby providing a smooth, intuitive, and personalized interactive experience, the evaluation of the intelligent cockpit model becomes particularly important.

[0003] In related technologies, the evaluation of intelligent cockpit models mainly focuses on testing single functions, paying attention to only one aspect of the model and ignoring the model's comprehensive performance in multiple scenarios, resulting in the technical problem of low evaluation accuracy of intelligent cockpit models.

[0004] There is currently no good solution to the above problems. Summary of the Invention

[0005] This application provides an evaluation method for intelligent cockpit models to at least address the technical problem of low evaluation accuracy of intelligent cockpit models in related technologies.

[0006] According to one aspect of the embodiments of this application, an evaluation method for an intelligent cockpit model is provided, comprising: determining a test intelligent cockpit model in a vehicle's intelligent cockpit system; determining an evaluation strategy for the test intelligent cockpit model based on a preset usage scenario associated with the test intelligent cockpit model, wherein the evaluation strategy is used to represent the evaluation rules for evaluating the test intelligent cockpit model, and the evaluation strategy is associated with usage requirement information in the preset usage scenario; evaluating the test intelligent cockpit model according to the evaluation strategy to obtain an evaluation result; calling a target cockpit index evaluation system, and evaluating the model performance of the test intelligent cockpit model from multiple evaluation dimensions according to the evaluation result to obtain a performance evaluation result of the test intelligent cockpit model, wherein the multiple evaluation dimensions include at least a model application evaluation dimension, a system application evaluation dimension, a scenario application evaluation dimension, and a user experience evaluation dimension.

[0007] Furthermore, based on the preset usage scenarios associated with the intelligent cockpit model under test, the evaluation strategy for the intelligent cockpit model under test is determined, including: determining the usage requirement information of the intelligent cockpit model under test in the preset usage scenarios based on the preset usage scenarios associated with the intelligent cockpit model under test; obtaining at least one test case based on the usage requirement information, wherein the at least one test case is used to simulate the operational behavior of the passenger in the process of using the intelligent cockpit model under test; and determining the evaluation strategy for the intelligent cockpit model under test based on the at least one test case.

[0008] Furthermore, the target cockpit index evaluation system is invoked, and the model performance of the intelligent cockpit model under test is evaluated from multiple evaluation dimensions according to the evaluation results, so as to obtain the performance evaluation results of the intelligent cockpit model under test. This includes: invoking multiple evaluation modules in the target cockpit index evaluation system, and evaluating the model performance of the intelligent cockpit model under test from multiple evaluation dimensions according to the evaluation results, so as to obtain the performance evaluation results of the intelligent cockpit model under test. Among them, the evaluation module is used to evaluate the model performance of the intelligent cockpit model under test from at least one of the model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension, and the evaluation module corresponds one-to-one with the evaluation dimension.

[0009] Furthermore, multiple evaluation modules within the target cockpit performance evaluation system are invoked to evaluate the model performance of the intelligent cockpit model under test from multiple evaluation dimensions based on the test results, thereby obtaining the performance evaluation results of the intelligent cockpit model under test. This includes: invoking the model application evaluation module within the target cockpit performance evaluation system to evaluate the test results and obtain model application evaluation indicators for the intelligent cockpit model under test, wherein the model application evaluation indicators are used to at least indicate the degree of matching between the intelligent cockpit model under test and the preset usage scenario; and invoking the system application evaluation module within the target cockpit performance evaluation system to evaluate the test results and obtain service capability evaluation indicators for the intelligent cockpit model under test. The evaluation criteria include: service capability assessment indicators to at least indicate the service performance of the intelligent cockpit model under test in the intelligent cockpit system; calling the scenario application assessment module in the target cockpit indicator assessment system to evaluate the test results and obtain the scenario application assessment indicators of the intelligent cockpit model under test, which at least indicate the task completion results of the intelligent cockpit model under test in the preset usage scenario; and calling the user experience assessment module in the target cockpit indicator assessment system to evaluate the test results and obtain the user experience assessment indicators of the intelligent cockpit model under test, which at least indicate the user satisfaction level of the intelligent cockpit model under test in the preset usage scenario.

[0010] According to another aspect of the embodiments of this application, a method for constructing a cockpit index evaluation system is also provided. The method includes: acquiring multiple preset scenario information of a vehicle's intelligent cockpit system, wherein the preset scenario information is used to represent preset usage scenarios of the intelligent cockpit system, and the preset scenario information corresponds one-to-one with the preset usage scenarios; based on the multiple preset scenario information, determining intelligent cockpit models that match the preset usage scenarios respectively, thereby obtaining multiple intelligent cockpit models, wherein the intelligent cockpit models are used to respond to task requests in the preset usage scenarios and generate corresponding response results; acquiring multiple performance evaluation indicators corresponding to the multiple intelligent cockpit models respectively, wherein the performance evaluation indicators are used to evaluate the model performance of the intelligent cockpit models in the preset usage scenarios from at least one of the model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension, and the performance evaluation indicators correspond one-to-one with the evaluation dimensions; and constructing a target cockpit index evaluation system based on the multiple performance evaluation indicators corresponding to the multiple intelligent cockpit models respectively, wherein the target cockpit index evaluation system is used to evaluate the model performance of the aforementioned intelligent cockpit models under test.

[0011] Furthermore, based on multiple preset scenario information, intelligent cockpit models matching the preset usage scenarios are determined respectively, including: based on multiple preset scenario information, determining the usage requirement information corresponding to the preset usage scenarios respectively, obtaining usage requirement information corresponding to multiple preset usage scenarios respectively, wherein the usage requirement information is used to represent the usage requirements in the preset usage scenarios; based on the usage requirement information corresponding to multiple preset usage scenarios respectively, intelligent cockpit models matching the preset usage scenarios are determined respectively.

[0012] Furthermore, based on multiple performance evaluation metrics corresponding to multiple smart cockpit models, a target cockpit metric evaluation system is constructed, including: identifying common performance evaluation metrics and scenario-dependent performance evaluation metrics from the multiple performance evaluation metrics corresponding to multiple smart cockpit models, wherein common performance evaluation metrics are performance evaluation metrics used to evaluate the same performance characteristics of multiple smart cockpit models, and scenario-dependent performance evaluation metrics are performance evaluation metrics used to evaluate the performance characteristics of any smart cockpit model in a corresponding preset usage scenario; fusing common performance evaluation metrics of the same type to obtain fused common performance evaluation metrics; and constructing the target cockpit metric evaluation system based on the fused common performance evaluation metrics and scenario-dependent performance evaluation metrics.

[0013] Furthermore, based on the fused common performance evaluation indicators and scenario-dependent performance evaluation indicators, a target cockpit indicator evaluation system is constructed, including: determining the evaluation modules to which the fused common performance evaluation indicators and scenario-dependent performance evaluation indicators belong in the initial cockpit indicator evaluation system, based on the mapping relationship between the indicator types and evaluation modules of the performance evaluation indicators. Each evaluation module is at least one of the following: a model application evaluation module, a system application evaluation module, a scenario application evaluation module, and a user experience evaluation module. The model application evaluation module is used to evaluate the matching degree between the intelligent cockpit model and the preset usage scenario; the system application evaluation module is used to evaluate the service performance of the intelligent cockpit model in the intelligent cockpit system; the scenario application evaluation module is used to evaluate the task completion results of the intelligent cockpit model in the preset usage scenario; and the user experience evaluation module is used to evaluate the user satisfaction of the intelligent cockpit model in the preset usage scenario. The fused common performance evaluation indicators and scenario-dependent performance evaluation indicators are then assigned to their respective evaluation modules to obtain the target cockpit indicator evaluation system.

[0014] According to another aspect of the embodiments of this application, a vehicle is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.

[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0016] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0018] According to another aspect of the embodiments of this application, a computer program is also provided, which, when executed by a processor, implements the methods of the various embodiments of this application.

[0019] In this embodiment, when evaluating a smart cockpit model under test in a smart cockpit system, the evaluation strategy for the model can be determined based on the preset usage scenarios associated with it. This determined evaluation strategy is closely related to the preset usage scenarios of the smart cockpit model under test. Compared to traditional evaluation methods that use more general test cases and lack a deep understanding of specific scenario requirements, leading to discrepancies between evaluation results and actual application performance, this embodiment evaluates the smart cockpit model under test using an evaluation strategy combined with specific usage scenarios, improving the relevance and practicality of the evaluation. Furthermore, after obtaining the evaluation results, a target cockpit indicator evaluation system can be invoked. Based on the evaluation results, the smart cockpit model under test can be evaluated from multiple dimensions, including model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation. This allows for a more comprehensive evaluation of the overall performance of the smart cockpit model, achieving the technical effect of improving the evaluation accuracy of the smart cockpit model under test, thereby solving the technical problem of low evaluation accuracy of smart cockpit models under test in related technologies. Attached Figure Description

[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0021] Figure 1 This is a flowchart of an evaluation method for an intelligent cockpit model according to an embodiment of this application;

[0022] Figure 2 This is a flowchart of a method for constructing a cabin performance evaluation system according to an embodiment of this application;

[0023] Figure 3 This is a schematic diagram of a target cockpit index evaluation system according to an embodiment of this application;

[0024] Figure 4 This is a flowchart of a method for constructing a target cockpit performance evaluation system according to an embodiment of this application;

[0025] Figure 5 This is a schematic diagram of an evaluation device for a smart cockpit model according to an embodiment of this application;

[0026] Figure 6 This is a schematic diagram of a construction device for a cabin performance evaluation system according to an embodiment of this application; Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] According to an embodiment of this application, an embodiment of an evaluation method for an intelligent cockpit model is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0030] This embodiment provides a method for evaluating a smart cockpit model. Figure 1 This is a flowchart of an evaluation method for an intelligent cockpit model according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps:

[0031] Step S101: Determine the smart cockpit model to be tested in the vehicle's smart cockpit system.

[0032] In the technical solution provided in step S101 of this application, the intelligent cockpit system of the vehicle can be a cockpit integrating multiple intelligent technologies to provide a safe, comfortable, and personalized driving and riding experience. The core of the intelligent cockpit system is to utilize advanced information technology, artificial intelligence, and the Internet of Things to transform the vehicle's cockpit into an intelligent interaction center. The aforementioned intelligent cockpit model under test can be a large-scale pre-trained language model or artificial intelligence model applied to the intelligent cockpit system. This intelligent cockpit model under test is pre-trained on a large amount of language data using deep learning technology, thereby acquiring the ability to understand and generate human natural language. In the intelligent cockpit system, the intelligent cockpit model under test can achieve natural language interaction with passengers, providing a series of services such as navigation, entertainment, and information query, enhancing the riding experience and safety. For example, this large-scale intelligent cockpit model under test can be used for: voice interaction, natural language understanding, multimodal interaction, personalized services, and context awareness.

[0033] In this embodiment, the intelligent cockpit system may include multiple intelligent cockpit models, each suitable for different preset usage scenarios, to meet the user needs in the corresponding scenarios. These preset usage scenarios may include at least: transportation scenarios and entertainment / leisure scenarios. Each preset usage scenario can utilize at least one intelligent cockpit model to fulfill the user's needs.

[0034] Optionally, in the travel scenario, to meet users' basic needs during vehicle operation, such as navigation, traffic information inquiry, trip planning, nearby recommendations, and safe driving, the intelligent cockpit system is equipped with a trip planning intelligent cockpit model, a nearby recommendation intelligent cockpit model, and a safe driving intelligent cockpit model to fulfill user needs in this travel scenario. In the entertainment and leisure scenario, to meet users' entertainment needs during their journey, the intelligent cockpit system is equipped with a variety of entertainment and leisure functions, such as casual conversation, music playback, video viewing, news broadcasts, and spoken dialogue. The intelligent cockpit system is equipped with a general casual conversation intelligent cockpit model, an Artificial Intelligence (AI) news intelligent cockpit model, a spoken dialogue intelligent cockpit model, and a super-recognition intelligent cockpit model to fulfill user needs in this entertainment and leisure scenario. The intelligent cockpit model to be tested can be any intelligent cockpit model corresponding to any of the above-mentioned preset usage scenarios.

[0035] Step S102: Based on the preset usage scenarios associated with the intelligent cockpit model under test, determine the evaluation strategy for the intelligent cockpit model under test.

[0036] In the technical solution provided in step S102 of this application, the aforementioned preset usage scenario is a usage scenario predefined according to the user's usage needs. This preset usage scenario may include at least travel scenarios, entertainment and leisure scenarios, information consultation scenarios, and service booking scenarios, without specific limitations. The aforementioned evaluation strategy is used to represent the evaluation rules for evaluating the intelligent cockpit model under test, and the evaluation strategy is associated with the usage needs information in the preset usage scenario.

[0037] In this embodiment, after determining the smart cockpit model to be tested, the preset usage scenarios associated with the smart cockpit model can be determined based on the smart cockpit model to be tested, and then the evaluation strategy of the smart cockpit model to be tested can be determined based on the user usage needs in the preset usage scenarios.

[0038] For example, there is a pre-defined association between the pre-defined usage scenarios of the intelligent cockpit system and the intelligent cockpit model. Based on this, after determining the intelligent cockpit model to be tested, the pre-defined usage scenarios associated with the intelligent cockpit model can be determined according to the pre-defined association, and then the user's usage needs in those pre-defined usage scenarios can be determined. After determining the user's usage needs, the evaluation strategy for the intelligent cockpit model to be tested can be determined based on the user's usage needs.

[0039] Optionally, in the travel and transportation scenario, user needs may include destination search, route planning, real-time traffic reports, and nearby recommendations. This scenario focuses on the intelligent cockpit model's response speed, accuracy, and ability to interpret complex commands. In the entertainment and leisure scenario, user needs may include music playback requests, video streaming, voice storytelling, and news broadcasts. This scenario focuses on the intelligent cockpit model's content recommendation capabilities, speech synthesis quality, and real-time updates of entertainment information. In the information consultation scenario, user needs may include inquiring about weather forecasts, stock quotes, and sports results. This scenario focuses on the intelligent cockpit model's information retrieval capabilities and data accuracy. In the service booking scenario, user needs may include booking restaurants, hotels, and movie tickets. This scenario focuses on the intelligent cockpit model's service interface call capabilities and transaction security. Based on this, user needs within the pre-defined usage scenarios of the intelligent cockpit model under test can be determined. Then, based on these determined user needs, an evaluation strategy can be developed for the intelligent cockpit model under test, thereby evaluating its performance within the pre-defined usage scenarios.

[0040] Optionally, the evaluation strategy mainly includes the selection of test cases, the execution of the testing process, the definition of evaluation metrics, and the setting of evaluation criteria. The selection of test cases primarily involves obtaining a series of representative test cases based on user needs within the preset usage scenarios of the intelligent cockpit model under test, to comprehensively verify the model performance of the intelligent cockpit model under test in the preset usage scenarios. Evaluation metrics may include response time, recognition accuracy, task completion rate, user satisfaction, etc., to quantify the performance of the intelligent cockpit model under test. The setting of evaluation criteria mainly focuses on the output results of the intelligent cockpit model under test, used to measure the pass rate of the output results. Based on the above methods, an evaluation strategy closely related to user needs in the preset usage scenarios can be formulated, and this evaluation strategy can then be determined as the evaluation strategy for evaluating the intelligent cockpit model under test.

[0041] For example, by deeply analyzing user needs in preset usage scenarios, potential user behavior patterns and information needs can be identified. Then, based on these identified patterns and needs, the most suitable evaluation metrics can be selected or created. For instance, in entertainment and leisure scenarios, evaluation metrics might include the smoothness of audio and video playback, the richness of content, and the level of personalization. One or more test cases are designed for each evaluation metric to ensure comprehensive coverage of the functional boundaries and potential anomalies of the smart cockpit model under test in that scenario.

[0042] In this step, the evaluation strategy for assessing the smart cockpit model is determined by identifying pre-defined usage scenarios associated with the model under test and the user needs within those scenarios. This ensures that the determined evaluation strategy is directly related to users' actual needs and experiences, avoiding blind and ineffective evaluations. Using this evaluation strategy to assess the smart cockpit model allows for a more direct reflection of whether it meets user expectations, thereby driving the optimization of the model and improving the overall user experience.

[0043] Step S103: According to the evaluation strategy, evaluate the intelligent cockpit model to be tested and obtain the evaluation results.

[0044] In the technical solution provided by step S103 of this application, after determining the evaluation strategy through step S102, the intelligent cockpit model to be tested can be evaluated according to the evaluation strategy to obtain the evaluation result. The evaluation result is the data basis for evaluating the model performance of the intelligent cockpit model to be tested in the preset usage scenario. The evaluation result may include the output result of the intelligent cockpit model to be tested, the response time of the intelligent cockpit model to be tested, the operating status of the intelligent cockpit model to be tested, and user feedback, etc., without specific limitations.

[0045] In this embodiment, when evaluating the intelligent cockpit model under test according to the evaluation strategy, the model can be run in a simulated cockpit environment. Then, according to the evaluation strategy, the model is controlled to execute pre-designed test cases, and its output results are collected. These output results can be considered as the evaluation results of the intelligent cockpit model.

[0046] Step S104: Call the target cockpit index evaluation system, and evaluate the model performance of the intelligent cockpit model under test from multiple evaluation dimensions according to the evaluation results, so as to obtain the performance evaluation results of the intelligent cockpit model under test.

[0047] In the technical solution provided in step S104 of this application, the target cockpit index evaluation system is a comprehensive platform pre-designed to evaluate the model performance of the intelligent cockpit model. It covers multiple dimensions of indicators such as model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation, and is used to comprehensively and objectively evaluate the model performance of the intelligent cockpit model in practical applications.

[0048] Optionally, the aforementioned multiple evaluation dimensions include at least model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation. The model application evaluation dimension primarily assesses the matching degree between the intelligent cockpit model and the preset usage scenarios. For example, it uses model evaluation metrics to measure the appropriateness of the intelligent cockpit model selection in the preset usage scenarios and the capability boundaries of the intelligent cockpit model in the associated preset usage scenarios (e.g., whether the intelligent cockpit model can handle the functional requirements of the preset usage scenarios, such as navigation, entertainment, information query, and personalized services). The system application evaluation dimension primarily assesses the service performance of the intelligent cockpit model in the intelligent cockpit system. For example, it assesses the operational stability, resource utilization, and system service capabilities (e.g., anomaly handling capabilities, fault recovery mechanisms, and system response time) of the intelligent cockpit model in the intelligent cockpit system. The scenario application evaluation dimension primarily assesses the task completion results of the intelligent cockpit model in the usage scenarios. For example, it assesses the intelligent cockpit model's ability to handle multi-task and complex instructions, such as planning itineraries and recommending restaurants while considering weather and real-time traffic conditions. The user experience evaluation dimension assesses the user satisfaction level of the intelligent cockpit model in the usage scenarios. For example, user feedback surveys can be used to determine user satisfaction and acceptance of the smart cockpit model.

[0049] In this embodiment, after obtaining the evaluation results through the above step S104, the target cockpit index evaluation system can be called to evaluate the model performance of the intelligent cockpit model under test from multiple evaluation dimensions according to the evaluation results, thereby obtaining the performance evaluation results of the intelligent cockpit model under test, wherein the performance evaluation results are...

[0050] Optionally, the evaluation results of the intelligent cockpit model under test can be input into the target cockpit indicator evaluation system. After receiving the evaluation results, the target cockpit indicator evaluation system can evaluate the model performance of the intelligent cockpit model under test from the dimensions of model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation, and then output the performance evaluation results. The performance evaluation results are used to represent the performance of the intelligent cockpit model under test under the above-mentioned multiple evaluation dimensions, such as the degree of matching with the preset use scenario, the service performance in the intelligent cockpit system, the task completion effect in the preset use scenario, and user satisfaction.

[0051] Optionally, after obtaining the performance evaluation results of the intelligent cockpit model under test, an optimization strategy can be formulated based on the performance evaluation results. Then, the intelligent cockpit model under test can be adjusted according to the optimization strategy to obtain an optimized intelligent cockpit model under test. The optimized intelligent cockpit model under test can then be further evaluated according to the above evaluation method. Through continuous evaluation and optimization, the intelligent cockpit model under test can better serve users, promote the development of intelligent vehicle technology, and improve the overall intelligence level of vehicles.

[0052] In this step, the target cockpit performance evaluation system is invoked to assess the performance of the smart cockpit model under test from multiple evaluation dimensions based on the test results. This results provide crucial information for the development and optimization of the smart cockpit model, reflecting not only its current technical status and user experience level but also providing clear guidance for future performance improvements and product upgrades. Through continuous evaluation and optimization, the smart cockpit model can better serve users and enhance the overall intelligence level of the vehicle.

[0053] In steps S101 to S104 above, when evaluating the intelligent cockpit model under test in the intelligent cockpit system, the evaluation strategy for the intelligent cockpit model under test can be determined based on the preset usage scenarios associated with the model. This determined evaluation strategy is closely related to the preset usage scenarios of the intelligent cockpit model under test. Compared to traditional evaluation methods that use more general test cases and lack a deep understanding of specific scenario requirements, leading to deviations between evaluation results and actual application performance, this embodiment of the application evaluates the intelligent cockpit model under test through an evaluation strategy combined with specific usage scenarios, improving the relevance and practicality of the evaluation. Furthermore, after obtaining the evaluation results, the target cockpit index evaluation system can be invoked. Based on the evaluation results, the intelligent cockpit model under test can be evaluated from multiple evaluation dimensions, such as model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension. This can more comprehensively evaluate the overall performance of the intelligent cockpit model, thereby improving the evaluation accuracy of the intelligent cockpit model under test and solving the technical problem of low evaluation accuracy of the intelligent cockpit model under test in related technologies.

[0054] The evaluation method for the aforementioned intelligent cockpit model in this application will be further described below.

[0055] As an optional implementation, step S102, based on the preset usage scenarios associated with the smart cockpit model under test, determines the evaluation strategy for the smart cockpit model under test, including: determining the usage requirement information of the smart cockpit model under test in the preset usage scenarios based on the preset usage scenarios; obtaining at least one test case based on the usage requirement information, wherein the at least one test case is used to simulate the operation behavior of the passenger in the process of using the smart cockpit model under test; and determining the evaluation strategy for the smart cockpit model under test based on the at least one test case.

[0056] In this embodiment, the aforementioned usage requirement information can be the functional requirements of the smart cockpit model under test by a user in a preset usage scenario. This usage requirement information can be derived through a deep understanding of the user's behavioral patterns in the preset usage scenario and an analysis of the smart cockpit environment. The aforementioned test cases are used to simulate the operational behavior of passengers using the smart cockpit model under test, wherein each test case can represent a type of user requirement or operational behavior.

[0057] Optionally, after determining the preset usage scenarios of the intelligent cockpit module under test, the user demand information within these scenarios can be further determined, and at least one test case can be obtained based on this demand information. For example, for a car-hailing scenario, assuming the user demand is route planning and nearby recommendations, the test case obtained based on this demand information can simulate user operations while driving. For example, the test case could be "Tell me how to get to XX train station" or "Go to the nearby gas station." For an entertainment and leisure scenario, if the user demand is playing music, listening to music radio, or adjusting the volume of media devices, the test case obtained based on this demand information can simulate user requests. For example, the test case could be "Play song XXX" or "Turn up the volume." Based on this, for different preset usage scenarios, at least one test case can be obtained based on the user demand information within those scenarios.

[0058] Optionally, test cases should at least meet the following conditions: test cases should cover all functional boundaries and potential abnormal situations of the intelligent cockpit model under test in the preset scenario; test cases should represent a type of user needs or operational behavior to ensure the general applicability of the evaluation results; test cases should be specific and clear to avoid ambiguity and facilitate the evaluation of the accurate understanding and execution capabilities of the intelligent cockpit model under test.

[0059] Optionally, after obtaining at least one test case, an evaluation strategy for the intelligent cockpit model under test can be determined based on at least one test case.

[0060] For example, a thorough analysis of the acquired test cases is conducted to understand the user intent, scenario characteristics, and technical requirements corresponding to each test case. For instance, a test case for "music playback" might require the smart cockpit model under test to respond quickly, accurately identify song titles, and provide a smooth playback experience. The analyzed user needs and scenario characteristics are mapped to predefined evaluation metrics to determine which metrics are applicable to the current test case. For example, the "music playback" test case might be related to metrics such as "speech recognition accuracy," "naturalness of response," and "music content accuracy." Weights for different evaluation metrics are configured according to the importance of the test case.

[0061] Optionally, after determining the weights of test cases and relevant evaluation metrics, a simulated or actual cockpit environment matching the test cases is created to ensure that the intelligent cockpit model under test faces the same or similar conditions as real-world use during testing. For example, complex road and weather conditions can be set up for testing navigation functions. According to the testing strategy, test cases are executed, and the responses, actions, and feedback of the intelligent cockpit model under test are recorded. During execution, information such as the output results, system logs, and user feedback of the intelligent cockpit model under test is collected.

[0062] In the above embodiments, specific and comprehensive test cases are designed based on the preset usage scenarios of the intelligent cockpit model under test and the user needs in the preset usage scenarios, thereby formulating an effective evaluation strategy to comprehensively evaluate the performance of the intelligent cockpit model under test in various scenarios.

[0063] As an optional implementation, step S104 involves calling the target cockpit index evaluation system and, according to the evaluation results, evaluating the model performance of the intelligent cockpit model under test from multiple evaluation dimensions to obtain the performance evaluation results of the intelligent cockpit model under test. This includes calling multiple evaluation modules in the target cockpit index evaluation system and, according to the evaluation results, evaluating the model performance of the intelligent cockpit model under test from multiple evaluation dimensions to obtain the performance evaluation results of the intelligent cockpit model under test. The evaluation modules are used to evaluate the model performance of the intelligent cockpit model under test from at least one of the following dimensions: model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension. The evaluation modules correspond one-to-one with the evaluation dimensions.

[0064] In this embodiment, the target cockpit performance evaluation system is a comprehensive performance evaluation tool used to systematically analyze and evaluate the various functions and user experience of the intelligent cockpit model within an intelligent cockpit system. This system includes multiple evaluation modules, each focusing on a specific evaluation dimension. These modules work collaboratively to comprehensively evaluate the performance of the intelligent cockpit model under test.

[0065] Optionally, the multiple evaluation modules may include at least a model application evaluation module, a system application evaluation module, a scenario application evaluation module, and a user experience evaluation module. These multiple evaluation modules can be tested and evaluated independently, and also support joint testing across modules to ensure the flexibility and relevance of the evaluation.

[0066] As an optional implementation, multiple evaluation modules in the target cockpit index evaluation system are invoked to evaluate the model performance of the intelligent cockpit model under test from multiple evaluation dimensions according to the evaluation results, thereby obtaining the performance evaluation results of the intelligent cockpit model under test. This includes: invoking the model application evaluation module in the target cockpit index evaluation system to evaluate the evaluation results and obtain the model application evaluation index of the intelligent cockpit model under test, wherein the model application evaluation index is used to at least indicate the degree of matching between the intelligent cockpit model under test and the preset usage scenario; and invoking the system application evaluation module in the target cockpit index evaluation system to evaluate the evaluation results and obtain the service capability evaluation of the intelligent cockpit model under test. The evaluation criteria include: a service capability evaluation criterion to at least indicate the service performance of the intelligent cockpit model under test within the intelligent cockpit system; a scenario application evaluation module in the target cockpit criterion evaluation system to evaluate the test results and obtain scenario application evaluation criteria for the intelligent cockpit model under test, which at least indicate the task completion results of the intelligent cockpit model under test in a preset usage scenario; and a user experience evaluation module in the target cockpit criterion evaluation system to evaluate the test results and obtain user experience evaluation criteria for the intelligent cockpit model under test, which at least indicate the user satisfaction level of the intelligent cockpit model under test in a preset usage scenario.

[0067] In this embodiment, the model application evaluation module is used to evaluate the intelligent cockpit model under test from the perspective of model application evaluation, mainly to evaluate the degree of matching between the intelligent cockpit model under test and the preset usage scenario; the system application evaluation module is used to evaluate the intelligent cockpit model under test from the perspective of system application evaluation, mainly to evaluate the service performance of the intelligent cockpit model under test in the intelligent cockpit system; the scenario application evaluation module is used to evaluate the intelligent cockpit model under test from the perspective of scenario application evaluation, mainly to evaluate the task completion results of the intelligent cockpit model under test in the preset usage scenario; and the user experience module is used to evaluate the intelligent cockpit model under test from the perspective of user experience, mainly to evaluate the user satisfaction of the intelligent cockpit model under test in the preset usage scenario.

[0068] Optionally, after obtaining the evaluation results, multiple evaluation modules in the target cockpit indicator evaluation system can be invoked to evaluate the model performance of the intelligent cockpit model under test. When evaluating the model performance of the intelligent cockpit model under test based on the evaluation results, the multiple evaluation modules can perform the evaluation according to their respective evaluation indicators. For example, the intent recognition accuracy indicator can be determined by dividing the number of accurately predicted intents by the total number of queries; the question-answering accuracy indicator can be determined by dividing the number of correct answers by the total number of queries; and the answer completeness or task completion rate can be determined by dividing the number of completed tasks by the number of task requests. These are merely illustrative examples, and the calculation methods for the evaluation indicators are not limited.

[0069] Optionally, the model application evaluation module is used to evaluate the basic technical and functional capabilities of the intelligent cockpit model under test in a preset usage scenario, covering performance in areas such as speech recognition, natural language processing, and multimodal interaction. The evaluation metrics for this module include, but are not limited to: intent recognition accuracy, question-answering accuracy, answer completeness, and user matching degree. For example, the model application evaluation module can calculate various evaluation metrics based on the evaluation results and the test cases adopted, and determine whether each evaluation metric meets the standards. If it does, it determines that the degree of matching between the intelligent cockpit model under test and the preset usage scenario meets the requirements.

[0070] Optionally, the system application evaluation module is used to evaluate the operation and performance of the intelligent cockpit model under test in the intelligent cockpit system, ensuring the service performance of the intelligent cockpit model under test in the intelligent cockpit system. The evaluation indicators of this system application evaluation module include, but are not limited to: system stability, resource utilization, system response time, and compatibility with other system components. For example, based on the system performance of the intelligent cockpit model under test under different loads and long-term operation in the evaluation results, the system application evaluation module evaluates the service capability of the intelligent cockpit model under test in the intelligent cockpit system.

[0071] Optionally, the scenario application evaluation module is used to evaluate the performance of the intelligent cockpit model under test in specific usage scenarios to determine the task completion results of the intelligent cockpit model under test in preset usage scenarios. The evaluation metrics corresponding to this scenario application evaluation module include, but are not limited to, task completion rate, recommended task accuracy, and recommended location accuracy. For example, the scenario application evaluation module can evaluate the intelligent cockpit model under test's ability to complete tasks, as well as its response and processing capabilities when facing complex tasks, based on the output results of the intelligent cockpit model under test in the test cases and the actual user needs.

[0072] Optionally, a user experience evaluation module is used to assess user satisfaction with the intelligent cockpit model under test based on the quality of its interaction with the user. The evaluation metrics for this module include, but are not limited to: dialogue interaction experience, dialogue coherence, the practicality and richness of dialogue output content, and the accuracy of music and video content. For example, this user experience evaluation module assesses user satisfaction with the intelligent cockpit model under test through user feedback in the evaluation results, such as user ratings and questionnaires.

[0073] Optionally, when the aforementioned model application evaluation module, system application evaluation module, scenario application evaluation module, and user experience evaluation module evaluate the intelligent cockpit model under test using different evaluation metrics, the evaluation results for each evaluation metric can be determined through a preset metric calculation method. For example, as mentioned above, for answer completeness or task completion rate, it can be determined based on the number of completed tasks / the number of task requests of the intelligent cockpit model under test. No specific restrictions are placed on the metric calculation method here. After obtaining the evaluation results for each evaluation metric, it can be determined whether the intelligent cockpit model under test meets the standards in each evaluation dimension based on the evaluation results and the corresponding metric thresholds. For example, by comparing each metric evaluation result with the corresponding metric threshold, if the metric evaluation result is greater than the metric threshold, it is determined that the intelligent cockpit model under test meets the standard for that metric. Then, the metric evaluation results of the performance evaluation metrics for each evaluation dimension are weighted and averaged to obtain the evaluation results corresponding to each evaluation dimension, thereby determining whether the intelligent cockpit model under test meets the standards in each evaluation dimension.

[0074] Optionally, if the smart cockpit model under test fails to meet the standards in one or more evaluation dimensions, it can be adjusted based on the evaluation results. Then, the adjusted smart cockpit model is re-evaluated using the method described above. This iterative optimization improves the model's performance, enabling it to better serve users. By iteratively optimizing each smart cockpit model in the smart cockpit system using the above method, the performance and user satisfaction of the entire smart cockpit system can be improved, thereby promoting the development of intelligent vehicle technology and enhancing the overall intelligence level of vehicles.

[0075] According to an embodiment of this application, an embodiment of a method for constructing a cabin performance evaluation system is also provided. Figure 2 This is a flowchart of a method for constructing a cabin performance evaluation system according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:

[0076] Step S201: Obtain multiple preset scenario information of the vehicle's intelligent cockpit system.

[0077] In the technical solution provided by step S201 of this application, the aforementioned preset scenario information is information pre-set to describe the corresponding preset usage scenario, wherein one preset scenario information is used to describe one preset usage scenario.

[0078] In this embodiment, multiple preset scenario information of the vehicle's smart cockpit can be acquired. This preset scenario information is designed to simulate and cover various situations that may be encountered during daily vehicle use, thereby ensuring that the intelligent system within the cockpit can adapt to the user's needs in different situations. Optionally, the preset scenario information can be descriptive information of preset usage scenarios.

[0079] Step S202: Based on multiple preset scenario information, determine the smart cockpit model that matches the preset usage scenario, and obtain multiple smart cockpit models.

[0080] In the technical solution provided by step S202 of this application, the aforementioned preset usage scenario is the scenario described by the preset scenario information.

[0081] In this embodiment, after obtaining multiple preset scenario information of the vehicle's intelligent cockpit system, intelligent cockpit models matching the multiple preset usage scenarios can be determined according to the preset usage scenarios described by the multiple preset scenarios, resulting in multiple intelligent cockpit models. Among them, each preset usage scenario corresponds to at least one intelligent cockpit model, and the intelligent cockpit model can also be called an intelligent agent model.

[0082] Optionally, when determining the smart cockpit model that matches the preset usage scenarios, a comprehensive analysis can be conducted on the characteristics of multiple preset usage scenarios, the user's usage needs in the corresponding preset usage scenarios, and the types of services that the user expects in the preset usage scenarios. Then, based on the analysis results, the smart cockpit model is determined, and this determined smart cockpit model is used as the smart cockpit model that matches the preset usage scenarios.

[0083] Optionally, intelligent agent models can be constructed based on the usage requirements in preset usage scenarios, and then the constructed intelligent agent models can be determined as intelligent cockpit models that match the preset usage scenarios. For example, the most basic application scenario in an intelligent cockpit system is the car travel scenario, which is mainly used to meet the basic needs of users during vehicle travel, such as navigation, traffic information query, and nearby recommendations, to help users reach their destinations efficiently. These user needs can be realized by constructing intelligent agent models, such as trip planning intelligent agents, nearby recommendation intelligent agents, and safe driving intelligent agents. These intelligent agents can also be called intelligent cockpit models. Based on this, it can be determined that the intelligent cockpit models matching this car travel scenario can at least include: intelligent cockpit models corresponding to trip planning intelligent agents, intelligent cockpit models corresponding to nearby recommendation intelligent agents, and intelligent cockpit models corresponding to safe driving intelligent agents. These intelligent cockpit models can be represented by intelligent agents.

[0084] Optionally, for entertainment and leisure scenarios, primarily designed to meet users' entertainment needs during their journeys, the intelligent cockpit system is equipped with a rich array of entertainment and leisure functions. These include, for example, casual conversation, music playback, video viewing, news broadcasts, and spoken dialogue. Users can select their favorite music and entertainment functions via voice commands or touch controls to enjoy a pleasant in-car experience. These user needs can be fulfilled by constructing intelligent agent models. The intelligent agents that can be constructed can include at least: a general casual conversation intelligent agent, an AI news intelligent agent, a spoken dialogue intelligent agent, and a super-recognition intelligent agent. Based on this, the intelligent cockpit models matching this entertainment and leisure scenario can be determined to include at least: intelligent cockpit models corresponding to general casual conversation intelligent agents, AI news intelligent agents, spoken dialogue intelligent agents, and super-recognition intelligent agents.

[0085] Optionally, when determining the smart cockpit model that matches the preset usage scenario, a model that can meet the usage requirements of the preset usage scenario can be determined from the existing models based on the user's usage needs and the expected results that the user expects the smart cockpit model to achieve, and used as the smart cockpit model that matches the preset usage scenario.

[0086] For example, taking one of several preset usage scenarios as an example, based on the user's usage needs in the preset usage scenario and the expected results that the user expects the smart cockpit model to achieve, a model that can meet the usage needs in the preset usage scenario is determined from the existing language recognition models or AI models, and then the determined model is used as the smart cockpit model that matches the preset usage scenario.

[0087] For example, in a car-hailing scenario, users might need to check weather information for their destination, request recommendations for attractions along the way, play music, and listen to children's stories. Additionally, they need to monitor for driver fatigue and inattention to ensure driving safety. Based on these user needs, candidate models can be evaluated. For instance, the accuracy of speech recognition, semantic understanding, personalized service matching, and the coverage and accuracy of entertainment content recommendations can be assessed. Models that perform well in multi-turn dialogue, entertainment content recommendations, and safety alerts can then be selected as the smart cockpit model that best matches the intended use case. For example, if the first model excels in entertainment content recommendations while the second model leads in safety detection, these two models can be combined to meet the user needs in a car-hailing scenario.

[0088] Optionally, after determining the smart cockpit model that matches the preset usage scenario, if there are multiple smart cockpit models, the user can set the priority of each smart cockpit model to ensure that the user's most important usage needs can be responded to first, thereby improving the user experience.

[0089] Through the above steps, based on the user's needs in the preset usage scenarios, a smart cockpit model that matches the preset usage scenarios is determined. This ensures that the determined smart cockpit model can accurately match the physical needs of the preset usage scenarios, thereby improving the overall user experience of the smart cockpit.

[0090] Step S203: Obtain multiple performance evaluation indicators corresponding to multiple smart cockpit models.

[0091] In the technical solution provided in step S203 of this application, the above-mentioned performance evaluation index is used to evaluate the model performance of the intelligent cockpit model in the preset use scenario from at least one of the model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension. The above-mentioned performance evaluation index corresponds one-to-one with the evaluation dimension.

[0092] In this embodiment, the performance evaluation indicators and their focus will vary for different smart cockpit models, but they will generally revolve around the core dimensions of model application, system application, scenario application, and user experience.

[0093] Optionally, the evaluation metrics for model application are primarily used to assess the technical performance of the intelligent cockpit model, such as whether the intelligent cockpit model selection is correct, i.e., the capability boundaries of the intelligent cockpit model in the preset usage scenarios. These evaluation metrics for model application include, but are not limited to: intent recognition accuracy, question-answering accuracy, speech recognition rate, response completeness / task completion rate, and complex task completion accuracy. These are merely illustrative examples and do not limit the evaluation metrics for model application. Optionally, intent recognition accuracy measures the intelligent cockpit model's ability to recognize user command intent; question-answering accuracy measures the correctness of the intelligent cockpit model's response; speech recognition rate assesses the accuracy of the intelligent cockpit model in recognizing user voice commands; and speech recognition accuracy measures the intelligent cockpit model's speech recognition accuracy. These evaluation metrics determine the degree of matching between the intelligent cockpit model and the preset usage scenarios, i.e., whether the intelligent cockpit model selection is correct. If all evaluation metrics exceed the threshold, the matching degree between the intelligent cockpit model and the preset usage scenarios is considered satisfactory.

[0094] Optionally, the evaluation metrics for the system application dimension are primarily used to assess the operational efficiency and stability of the intelligent cockpit model at the system level. These system application evaluation metrics include, but are not limited to: system response time, resource utilization, system stability, and data security. These are merely illustrative examples and do not limit the scope of the system application evaluation metrics. Optionally, system response time is used to evaluate the time from when the intelligent cockpit model receives an instruction to when it completes execution; resource utilization is used to evaluate the intelligent cockpit model's usage of system resources such as processors and memory during operation; system stability is used to evaluate the system performance of the intelligent cockpit model under long-term operation or high load conditions; and data security is used to evaluate the security measures and privacy protection mechanisms of the intelligent cockpit model when processing user data. Through these evaluation metrics, the service capabilities of the intelligent cockpit model within the intelligent cockpit system can be evaluated at the system application dimension.

[0095] Optionally, the evaluation metrics for the scenario application dimension mainly focus on specific usage scenarios, assessing the ability of the intelligent cockpit model to solve practical problems in the corresponding usage scenarios. Assuming the intelligent cockpit model is a trip planning intelligent cockpit model, the evaluation metrics for assessing its performance from the perspective of scenario application include, but are not limited to: navigation accuracy, location planning accuracy, and route planning accuracy. Assuming the intelligent cockpit model is a surrounding recommendation intelligent cockpit model, the evaluation metrics for assessing its performance from the perspective of scenario application include, but are not limited to: recommendation task accuracy and primary recommendation accuracy. That is, the evaluation metrics for the intelligent cockpit model in the scenario application dimension for different application scenarios can be determined according to the specific application scenario. The evaluation metrics may differ between different models; this is merely an example and does not limit the evaluation metrics for the intelligent cockpit model in the scenario application dimension. These performance metrics can be used to evaluate the task completion effect of the intelligent cockpit model in the preset usage scenarios, thereby determining the model performance of the intelligent cockpit model in the preset usage scenarios.

[0096] Optionally, the evaluation metrics for the user experience dimension are primarily used to assess the user's subjective feelings about interacting with the intelligent cockpit model, i.e., user satisfaction. These user experience evaluation metrics include, but are not limited to, user satisfaction, model usability, personalization level, and emotional interaction capability. These are merely illustrative examples and do not limit the specific user experience evaluation metrics. Optionally, user satisfaction can be assessed through questionnaires or user feedback to evaluate the overall user satisfaction with the intelligent cockpit model. Model usability reflects the intuitiveness and convenience of operating the intelligent cockpit model. Personalization level assesses the intelligent cockpit model's ability to provide personalized services. Emotional interaction capability assesses the intelligent cockpit model's ability to recognize and respond to user emotions. These evaluation metrics can comprehensively assess the user's satisfaction with the intelligent cockpit model in the preset usage scenario.

[0097] Optionally, in addition to the above-mentioned indicator system, evaluation indicators for assessing the intelligent cockpit model under test may also include: content timeliness, robustness of command recognition, multimodal interaction performance, and energy efficiency. Content timeliness is used to evaluate the timeliness and relevance of the information provided by the intelligent cockpit model. Robustness of command recognition is used to evaluate the intelligent cockpit model's ability to recognize commands in noisy environments or under complex commands. Multimodal interaction performance is used to evaluate the intelligent cockpit model's ability to process multimodal inputs such as voice, gestures, and facial expressions. Energy efficiency is used to evaluate the impact of the intelligent cockpit model's operation on the efficiency of the vehicle's battery or engine.

[0098] In this step, each smart cockpit model selects and emphasizes different evaluation metrics based on its main functions and application scenarios. By integrating these metrics, the performance of the smart cockpit model can be comprehensively evaluated, providing data support for subsequent optimization and improvement, and ensuring that the smart cockpit system can provide an efficient, safe, personalized, and user-friendly experience.

[0099] Step S204: Construct a target cockpit indicator evaluation system based on multiple performance evaluation indicators corresponding to multiple smart cockpit models.

[0100] In the technical solution provided in step S204 of this application, the target cockpit index evaluation system is a comprehensive platform for evaluating the model performance of the intelligent cockpit model. It covers multiple dimensions of indicators such as model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation, and is used to comprehensively and objectively evaluate the model performance of the intelligent cockpit model in practical applications.

[0101] In this embodiment, after determining the smart cockpit model that matches each preset usage scenario, multiple smart cockpit models can be obtained. Then, a target cockpit indicator evaluation system can be constructed based on the multiple performance evaluation indicators corresponding to the multiple smart cockpit models.

[0102] Optionally, common performance indicators and scenario dependencies can be identified from multiple performance evaluation indicators corresponding to multiple smart cockpit models. Then, common performance indicators of the same type corresponding to multiple smart cockpit models can be merged, and a target cockpit indicator evaluation system can be constructed based on the merged common performance indicators and scenario-dependent performance evaluation indicators.

[0103] For example, based on the evaluation dimensions to which the fused common performance indicators and scenario-dependent performance evaluation indicators belong, the fused common performance indicators and scenario-dependent performance evaluation indicators can be assigned to the model application evaluation module, system application evaluation module, scenario application evaluation module, and user experience evaluation module in the initial cockpit indicator evaluation system, thus obtaining the target intelligent cockpit evaluation system. Specifically, the model application evaluation module assesses the degree of matching between the intelligent cockpit model and the preset usage scenario; the system application evaluation module assesses the service performance of the intelligent cockpit module in the intelligent cockpit system, such as the operational efficiency and stability of the intelligent cockpit model within the cockpit system; the scenario application evaluation module assesses the task completion results of the intelligent cockpit model in a specific usage scenario; and the user experience evaluation module assesses user satisfaction with the intelligent cockpit model.

[0104] In this step, a target cockpit indicator evaluation system is constructed by using multiple performance evaluation indicators corresponding to multiple smart cockpit models. The resulting target cockpit indicator evaluation system can include multiple evaluation modules. Through these multiple evaluation modules, the performance of the smart cockpit model can be evaluated systematically and scientifically from multiple evaluation dimensions, thereby guiding the optimization and upgrading of the smart cockpit model and ultimately improving the overall performance and user experience of the in-vehicle intelligent system.

[0105] In steps S201 to S204 above, by deeply analyzing information from multiple preset scenarios, the specific needs of each scenario and the expectations of drivers / passengers can be accurately understood, thereby selecting or designing an intelligent cockpit model that highly matches the preset usage scenarios. This precise scenario adaptability ensures that the intelligent cockpit model can provide optimal service in various usage situations, improving the overall applicability and user satisfaction of the intelligent cockpit system. By integrating multiple evaluation indicators corresponding to multiple intelligent cockpit models into the evaluation module of the initial intelligent cockpit model, the target intelligent cockpit model is obtained. The model performance of the intelligent cockpit model under test can be comprehensively evaluated from the dimensions of model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation. This provides a benchmark and direction for the development and performance optimization of intelligent cockpit models, enhances the scalability and flexibility of the evaluation system, and achieves the technical effect of improving the evaluation accuracy of intelligent cockpit models, thereby solving the technical problem of low evaluation accuracy of intelligent cockpit models.

[0106] The construction method of the cabin index evaluation system described above in this application will be further introduced below.

[0107] As an optional implementation, step S202, based on multiple preset scenario information, determines the smart cockpit model that matches the preset usage scenario, including: based on multiple preset scenario information, determining the usage requirement information corresponding to the preset usage scenario, obtaining the usage requirement information corresponding to each of the multiple preset usage scenarios, wherein the usage requirement information is used to represent the usage requirements in the preset usage scenario; and based on the usage requirement information corresponding to each of the multiple preset usage scenarios, determining the smart cockpit model that matches the preset usage scenario.

[0108] In this embodiment, the aforementioned preset scenario information is a detailed description of various driving and riding environments expected to be encountered in the vehicle's intelligent cockpit system. Based on the preset scenario information, usage demand information corresponding to the preset usage scenarios can be determined. For example, in a travel scenario, the user's usage demand information may include navigation, traffic information query, trip planning, and nearby recommendations. In an entertainment and leisure scenario, the user's usage demand information may include casual conversation, music playback, video viewing, news broadcasts, and spoken conversation. This is merely an example and does not limit the user's usage demand information in the preset usage scenarios.

[0109] Optionally, after determining the user needs information corresponding to the preset scenarios, the key functionalities that the intelligent cockpit model must possess in each preset scenario can be identified based on these needs, such as voice recognition, navigation planning, and music recommendation. Then, the technical requirements of existing intelligent cockpit models are evaluated to determine whether they meet these key functionalities, including the model's voice recognition accuracy, multimodal interaction capabilities, and the breadth and depth of its knowledge base. Furthermore, the model's ability to flexibly adapt to the complexity and changes of the preset scenarios is examined; for example, in adverse weather conditions, can the intelligent cockpit model prioritize providing safe driving advice while handling other user needs? The model's ability to provide personalized services based on user preferences and historical interaction data is also assessed; for example, in entertainment and leisure scenarios, can the intelligent cockpit model understand and respond to changes in the user's music preferences? Based on the above analysis, the intelligent cockpit model that best meets the usage requirements of the preset scenarios is selected.

[0110] Optionally, after determining the user needs information corresponding to the preset scenarios, the key functionalities that the intelligent cockpit model must possess in each preset scenario can be identified based on the user needs information, such as voice recognition, navigation planning, and music recommendation. Then, based on the user needs information corresponding to each preset scenario, intelligent agents are constructed to meet the user's needs. For example, in a car-travel scenario, based on user needs such as navigation, traffic information query, trip planning, nearby recommendations, and safe driving, trip planning intelligent agents, nearby recommendation intelligent agents, and safe driving intelligent agents can be constructed, and the intelligent models corresponding to these intelligent agents can be determined as intelligent cockpit models matching the car-travel scenario. Optionally, in an entertainment and leisure scenario, based on user needs such as casual conversation, music playback, video viewing, news broadcasting, and spoken dialogue, casual conversation intelligent agents, AI news intelligent agents, spoken dialogue intelligent agents, and super-recognition intelligent agents can be constructed, and the intelligent models corresponding to these intelligent agents can be determined as intelligent cockpit models matching the entertainment and leisure scenario.

[0111] In this step, by identifying the usage needs information in the preset usage scenarios, a smart cockpit model that matches the preset usage scenarios is determined. Then, by using the determined smart cockpit model, the user's usage needs can be accurately understood and the user's expectations can be met.

[0112] As an optional implementation, step S204 involves constructing a target cockpit indicator evaluation system based on multiple performance evaluation indicators corresponding to multiple smart cockpit models. This includes: identifying common performance evaluation indicators and scenario-dependent performance evaluation indicators from the multiple performance evaluation indicators corresponding to the multiple smart cockpit models, wherein the common performance evaluation indicators are performance evaluation indicators used to evaluate the same performance characteristics of multiple smart cockpit models, and the scenario-dependent performance evaluation indicators are performance evaluation indicators used to evaluate the performance characteristics of any smart cockpit model in a corresponding preset usage scenario; fusing the common performance evaluation indicators of the same type to obtain fused common performance evaluation indicators; and constructing the target cockpit indicator evaluation system based on the fused common performance evaluation indicators and the scenario-dependent performance evaluation indicators.

[0113] In this embodiment, the aforementioned common performance indicators can be cross-scenario indicators used to evaluate the basic functions and performance of the intelligent cockpit model. These indicators are not dependent on specific usage scenarios. For example, intent recognition accuracy, question-answering accuracy, and response completion rate are universally applicable across all scenarios. The aforementioned scenario-dependent performance evaluation indicators are evaluation indicators specific to a particular scenario, reflecting the performance characteristics of the intelligent cockpit model in that specific scenario. For example, fatigue monitoring accuracy in long-distance driving scenarios and rapid response capabilities in emergency rescue scenarios; the evaluation results of these indicators are directly related to the user experience within the specific usage scenario.

[0114] Optionally, after obtaining multiple performance evaluation metrics corresponding to multiple smart cockpit models, universally applicable common performance evaluation metrics and scenario-dependent performance evaluation metrics for specific scenarios can be identified from these metrics. Then, common performance evaluation metrics of the same type are fused to create a more comprehensive evaluation metric, reducing redundancy and improving efficiency. For example, the intent recognition accuracy metrics corresponding to multiple smart cockpit models can be fused to obtain a fused intent recognition accuracy. Subsequently, the fused common metric is combined with the scenario-dependent performance evaluation metrics to jointly construct a target cockpit metric evaluation system. This target cockpit metric evaluation system focuses on both the model's basic capabilities and its performance in specific scenarios.

[0115] In this step, the construction of the target cockpit index evaluation system not only needs to focus on the common performance evaluation indicators of the intelligent cockpit model, but also fully consider the scenario-dependent performance evaluation indicators. The target cockpit index evaluation system constructed in this way can provide a more comprehensive perspective when evaluating the intelligent cockpit model, and evaluate the general performance and professional performance of the intelligent cockpit model in specific scenarios, so as to improve the evaluation accuracy of the intelligent cockpit model.

[0116] As an optional implementation, a target cockpit indicator evaluation system is constructed based on the fused common performance evaluation indicators and scenario-dependent performance evaluation indicators. This includes: determining the evaluation modules to which the fused common performance evaluation indicators and scenario-dependent performance evaluation indicators belong in the initial cockpit indicator evaluation system, based on the mapping relationship between the indicator types and evaluation modules. Each evaluation module is at least one of the following: a model application evaluation module, a system application evaluation module, a scenario application evaluation module, and a user experience evaluation module. The model application evaluation module evaluates the matching degree between the intelligent cockpit model and the preset usage scenario; the system application evaluation module evaluates the service performance of the intelligent cockpit model in the intelligent cockpit system; the scenario application evaluation module evaluates the task completion results of the intelligent cockpit model in the usage scenario; and the user experience evaluation module evaluates the user satisfaction level of the intelligent cockpit model in the usage scenario. The fused common performance evaluation indicators and scenario-dependent performance evaluation indicators are then assigned to their respective evaluation modules to obtain the target cockpit indicator evaluation system.

[0117] In this embodiment, the process of constructing the target cockpit index evaluation system involves mapping the common performance evaluation indexes and scenario-dependent performance evaluation indexes of the intelligent cockpit model to appropriate evaluation modules to ensure the comprehensiveness and relevance of the evaluation.

[0118] Optionally, the initial cockpit performance evaluation system includes multiple evaluation modules, such as a model application evaluation module, a system application evaluation module, a scenario application evaluation module, and a user experience evaluation module. These modules are used to evaluate the performance of the intelligent cockpit model from different dimensions. For example, the model application evaluation module evaluates the degree of matching between the intelligent cockpit model and the preset usage scenario; the system application evaluation module evaluates the service performance of the intelligent cockpit model in the intelligent cockpit system; the scenario application evaluation module evaluates the task completion results of the intelligent cockpit model in the usage scenario; and the user experience evaluation module evaluates the user satisfaction of the intelligent cockpit model in the usage scenario. Based on this, the fused common performance indicators and scenario-dependent performance evaluation indicators can be divided into the above evaluation modules according to the indicator type, so that the above evaluation modules can achieve accurate evaluation of the intelligent cockpit model through the evaluation indicators.

[0119] Optionally, since common performance evaluation metrics such as speech recognition rate and semantic understanding accuracy reflect the general matching degree between the basic technical capabilities of the intelligent cockpit model and its usage scenarios, these common performance metrics correspond to model application types. Based on this, these common performance metrics can be assigned to the model application evaluation module in the initial cockpit metric evaluation system, allowing the model application evaluation module to assess the matching degree between the intelligent cockpit model and the preset usage scenarios based on these evaluation metrics.

[0120] Optionally, since evaluation metrics such as navigation accuracy, topic coverage, recommended task accuracy, and emergency service response speed are closely related to specific scenarios, i.e., these performance evaluation metrics are scenario application types, these evaluation metrics can be classified into the scenario application evaluation module in the initial cockpit metric evaluation system. This allows the scenario application evaluation module to select the corresponding evaluation metrics based on the application scenario of the intelligent cockpit model when evaluating the intelligent cockpit model, so as to evaluate the task completion capability of the intelligent cockpit model in a specific usage scenario.

[0121] Optionally, since evaluation indicators such as stability, response time, and resource utilization are related to the service performance of the intelligent cockpit model in the intelligent cockpit system, that is, the indicator type of these evaluation indicators is the system application type, based on this, these evaluation indicators can be assigned to the system application evaluation module of the initial cockpit indicator evaluation system, so that the system application evaluation module can evaluate the service performance of the intelligent cockpit model in the intelligent cockpit system according to these evaluation indicators.

[0122] Optionally, since evaluation indicators such as satisfaction, ease of use, and emotional interaction capabilities are closely related to user experience, that is, the indicator type of these evaluation indicators is user experience type, these evaluation indicators can be assigned to the user experience evaluation module in the initial cockpit indicator evaluation system, so that the user experience module can evaluate the user satisfaction of the intelligent cockpit model based on these evaluation indicators.

[0123] Optionally, based on the type of performance evaluation indicators, the fused common indicators and scenario-dependent performance evaluation indicators are respectively assigned to the evaluation modules of the initial cockpit indicator evaluation system to form a comprehensive indicator system, namely, the target cockpit indicator evaluation system. Then, when evaluating the intelligent cockpit model under test, the target cockpit indicator evaluation system can be called. The target cockpit indicator evaluation system can then evaluate various aspects of the intelligent cockpit model under test in a targeted manner based on its multiple internal evaluation modules, thereby obtaining an overview of the model performance of the intelligent cockpit model under test.

[0124] Optionally, Figure 3This is a schematic diagram of a target cabin performance evaluation system according to an embodiment of this application, such as... Figure 3 As shown, a structured view of the target cockpit performance evaluation system is presented, designed to illustrate the components of the system and the relationships between them. Figure 3 As shown, the target cockpit performance evaluation system 300 includes a model application evaluation module 301, a system application evaluation module 302, a scenario application evaluation module 303, and a user experience evaluation module 304. Each evaluation module contains several key indicators used to comprehensively measure and optimize the performance and effectiveness of the intelligent cockpit model. Specifically, the model application evaluation module primarily uses model application evaluation indicators to measure whether the selected intelligent cockpit model is appropriate and optimal, as well as the performance boundaries of the intelligent cockpit model in various usage scenarios. The system application evaluation module is mainly used to evaluate the stability of the intelligent cockpit model within the intelligent cockpit system and its system service capabilities. The scenario application evaluation module and the user experience evaluation module primarily reflect the final effectiveness of the intelligent cockpit model, evaluating user experience and the completion of user tasks in various scenarios.

[0125] In this embodiment, by precisely matching indicators with modules, the target cockpit indicator evaluation system can more accurately identify and quantify the advantages and disadvantages of intelligent cockpit models in different scenarios. Common indicators ensure the assessment of the basic capabilities of the intelligent cockpit model under test, while scenario-dependent performance evaluation indicators focus on the model's performance in specific situations. The evaluation results directly point to potential improvement directions for the model. Based on feedback from different evaluation modules, the development team can specifically adjust the model to improve its performance in areas such as speech recognition, system response, task execution, or user experience. Through continuous testing and optimization, this accelerates the development of intelligent cockpit technology, thereby promoting the high-quality development of the intelligent vehicle industry. In other words, the construction of the target cockpit indicator evaluation system achieves a comprehensive and accurate evaluation of the performance of intelligent cockpit models. This not only promotes the technical optimization of the models themselves but also ensures that the intelligent cockpit system can meet diverse user needs, providing higher levels of safety and user experience, thus driving the high-quality development of the intelligent vehicle industry.

[0126] The above technical solutions of the present application embodiments will be further illustrated below with reference to preferred embodiments of the present invention.

[0127] With the intelligent transformation of the automotive industry, smart cockpits have become a key area for enhancing the driving experience. Traditional smart cockpits, however, are limited in function, relying primarily on pre-set rules and limited datasets, making it difficult to fully understand and respond to diverse and complex user needs. In recent years, the development of large-scale models has enabled smart cockpits to handle more complex tasks and provide more personalized and high-quality services. However, evaluating the performance of these large-scale models in the cockpit environment has become a major challenge for technological development. Currently, the industry lacks a systematic and standardized evaluation system, especially for large-scale models in areas such as multimodal interaction and scene adaptability, where evaluation standards are still immature.

[0128] Based on this, this application provides an evaluation method for intelligent cockpit models. According to the preset usage scenarios of the intelligent cockpit model under test and the user needs within those scenarios, an evaluation strategy is determined. Then, the evaluation strategy is used to evaluate the large-scale intelligent cockpit model under test, obtaining evaluation results. Next, a target cockpit indicator evaluation system is invoked. Based on the evaluation results, the model performance of the intelligent cockpit model under test is evaluated from multiple dimensions, including model application, system application, scenario application, and user experience. The resulting evaluation results comprehensively characterize the model performance of the intelligent cockpit model under test. In other words, this application constructs a complete evaluation framework consisting of model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation, ensuring that the evaluation covers all key aspects of the large-scale intelligent cockpit model. The intelligent cockpit model evaluation method provided in this application can provide clear guidance for the development and optimization of intelligent cockpit models, ensuring that they meet user needs in different usage scenarios, thereby improving the overall performance and user experience of the intelligent cockpit.

[0129] The following section will further describe the construction method of the target cockpit index evaluation system using a preferred embodiment.

[0130] Figure 4 This is a flowchart illustrating a method for constructing a target cockpit performance evaluation system according to an embodiment of this application. Figure 4 As shown, the method includes the following steps:

[0131] Step S401: Determine multiple usage scenarios for the smart cockpit.

[0132] In this embodiment, multiple potential use cases in the smart cockpit are identified and defined, providing specific contexts for subsequent indicator design.

[0133] Optionally, based on users' behavior patterns, needs, and interaction methods in the vehicle, the use scenarios of the smart cockpit can be divided. For example, the use scenarios of the smart cockpit can be divided into: travel scenarios, entertainment and leisure scenarios, etc. There are no specific restrictions here. Then, a detailed list of smart cockpit use scenarios can be formed.

[0134] Step S402: Determine the usage requirements information in multiple usage scenarios.

[0135] In this embodiment, after identifying multiple usage scenarios for the smart cockpit, the user's usage needs and goals in each usage scenario are analyzed in depth to determine the usage requirement information corresponding to each usage scenario.

[0136] Optionally, for each use case, user research should be conducted to collect user feedback. Simultaneously, industry trends and technological capabilities should be considered to clarify user needs in each use case. For example, user needs in a transportation scenario might include at least: navigation, traffic information queries, trip planning, and nearby recommendations. User needs in an entertainment and leisure scenario might include at least: casual conversation, music playback, video viewing, news broadcasts, and spoken conversation. This is merely an example and does not limit the specific user needs for each use case.

[0137] Step S403: Determine the smart cockpit models that match multiple usage scenarios.

[0138] In this embodiment, after determining the usage requirements information in each usage scenario through the above step S402, a smart cockpit model matching multiple usage scenarios can be determined respectively.

[0139] Optionally, for travel-related scenarios, since user needs in this scenario may include navigation, traffic information queries, trip planning, nearby recommendations, and safe driving, intelligent cockpits can be constructed based on these needs. These intelligent agents can then fulfill the user's requirements in the scenario. For example, trip planning, nearby recommendation, and safe driving agents can be constructed to meet the user's needs in the scenario, and the corresponding intelligent models for each agent can be identified as the intelligent cockpit models matching the travel-related scenario. Optionally, key strategies can be formulated for each intelligent cockpit model, and key indicators can be developed based on these strategies to evaluate and analyze whether the intelligent cockpit models can complete the task requests in the scenario.

[0140] Optionally, for entertainment and leisure scenarios, since user needs in such scenarios may include: casual conversation, music playback, video viewing, news broadcasting, and spoken dialogue, intelligent agents can be constructed based on these needs. These agents can then fulfill the user's needs in the scenario. For example, casual conversation agents, AI news agents, spoken dialogue agents, and super-recognition agents can be constructed to meet the user's needs in the scenario. The corresponding intelligent models of these agents can then be defined as intelligent cockpit models that match the entertainment and leisure scenario.

[0141] Step S404: Determine the performance evaluation indicators corresponding to the smart cockpit models that match multiple usage scenarios.

[0142] In this embodiment, multiple evaluation metrics are defined for the smart cockpit models in various usage scenarios to measure the performance of the corresponding smart cockpit models.

[0143] Optionally, based on the usage scenarios and user needs of the intelligent cockpit model, quantifiable evaluation metrics should be determined for the model. These metrics should cover multiple evaluation dimensions, including model application, system application, scenario application, and user experience. Specifically, the model application evaluation dimension primarily assesses the degree of matching between the intelligent cockpit model and the usage scenario; the system application evaluation dimension primarily assesses the service performance of the intelligent cockpit model within the intelligent cockpit system; the scenario application evaluation dimension primarily assesses the task completion effectiveness of the intelligent cockpit model in the usage scenario; and the user experience evaluation dimension primarily assesses user satisfaction with the intelligent cockpit model.

[0144] Optionally, the following are multiple performance evaluation indicators provided in the embodiments of this application, as well as the indicator definitions, indicator calculation methods, test cases that can be used when evaluating the cockpit model using performance evaluation indicators, and judgment criteria for performance evaluation indicators.

[0145] Intent recognition accuracy is defined as a performance evaluation metric that measures the model's ability to correctly recognize user command intents. The metric is calculated as: number of accurately predicted intents / total number of queries. A test case that can be used to evaluate the cockpit model using this metric is: "I'm hungry, I want to find a nearby Sichuan restaurant," where the expected intent is: Sichuan restaurant recommendation. The criteria for judging this performance metric are: can the model recognize the user's intent when the user's question has no specific task or service in mind? Can it effectively recognize intents when user input is ambiguous or uncertain?

[0146] Question answering accuracy is defined as a performance evaluation metric that measures a model's ability to correctly answer general task requests, requiring both correct responses and logical reasoning. The metric is calculated as: number of accurate answers / number of queries. A test case that can be used to evaluate a cockpit model using this metric could be: "How much is a train ticket from XX to XX?" The intended purpose of this test case is to provide accurate price information. The criteria for judging this performance evaluation metric are: the model's responses should accurately complete the task or answer the question, requiring both correct responses and logical reasoning.

[0147] Response completeness / task completion rate is defined as a performance evaluation metric that measures the model's ability to complete all user-requested tasks. The calculation method for this performance evaluation metric is: number of completed tasks / number of task requests. A test case that can be used to evaluate the cockpit model using this performance evaluation metric could be: "I'm going on a business trip to XX tomorrow. Can you tell me what the weather is like there, and what clothes I should wear?" The expected response for this test case would be: "The weather forecast for XX tomorrow shows cloudy skies, with temperatures between 10 and 15 degrees Celsius. We suggest you wear a light jacket with a long-sleeved shirt or a thin sweater underneath to cope with the temperature difference." The criteria for judging this performance evaluation metric are: the model should successfully complete all requested tasks, weather temperature information, and clothing recommendations.

[0148] User matching degree is defined as the ability of a model to provide personalized responses that take into account the user's identity. The calculation method for this performance evaluation metric is subjective scoring, using a multi-round scoring system, and calculating the average. A test case that can be used to evaluate the cockpit model using this performance evaluation metric could be: "I am a university student and want to know about nearby libraries." The judgment criterion for this performance evaluation metric is: the model's response should provide personalized suggestions based on the user's identity.

[0149] The accuracy rate for completing complex tasks is defined as a performance evaluation metric that measures the model's ability to correctly complete complex task requests. The metric is calculated as: number of correctly completed complex tasks / number of task requests. A test case that can be used to evaluate the cockpit model using this metric could be: "Please help me plan a weekend family trip, including attraction selection, dining arrangements, and transportation methods." The criterion for judging this performance evaluation metric is that the model should accurately complete the planning of all tasks.

[0150] Topic coverage is defined as a performance evaluation metric that measures whether a model's responses cover the topics inquired about by the user. The metric is calculated as: number of queries matching the topic / total number of queries. A test case that can be used to evaluate the cockpit model using this metric could be: "Discuss your views on artificial intelligence and its applications in the medical field." The criterion for judging this performance evaluation metric is that the model's responses should cover all mentioned topics.

[0151] Data accuracy is defined as a performance evaluation metric that measures the correctness of the data portion of a model's response. The metric is calculated as: number of accurate responses / number of queries. A test case that can be used to evaluate a cockpit model using this metric could be: "What is the GDP of location XX?". The criteria for judging this performance evaluation metric are: the data provided by the model should be consistent with actual data, and the accuracy of the data portion of the response should be verified.

[0152] Recommendation task accuracy is defined as a performance evaluation metric that measures whether the model's recommended tasks match the user's requests. The metric is calculated as: (Number of accurate recommended tasks / Total number of requested tasks). A test case that can be used to evaluate the cockpit model using this metric could be: "Recommend a nearby Sichuan restaurant with good reviews." The criterion for this performance evaluation metric is: the recommended restaurant should meet the user's needs; in this case, recommending a Sichuan restaurant would suffice.

[0153] Recommendation location accuracy is defined as a performance evaluation metric that measures whether the model's recommended locations match the user's request. The metric is calculated as: number of accurately recommended locations / total number of requests. A test case that can be used to evaluate the cockpit model using this metric is: "Find the nearest hospital to me." The criteria for judging this performance evaluation metric are as follows: If the user provides a specific range, and if more than 95% of the recommended locations in the accurate recommendation list are correct, then the model is considered to have made an accurate recommendation. If the user does not provide a range but only gives a vague range, such as "nearby," then any recommended hospital within a 5-kilometer radius of the user's current location is considered a correct recommendation.

[0154] Practicality is defined as a performance metric that measures the usability and operability of a model's response. The calculation method for this performance evaluation metric is subjective scoring, using a multi-round scoring system, and calculating the average. A test case that can be used to evaluate a cockpit model using this performance evaluation metric is: "How to set up the vehicle's navigation system." The judgment criteria for this performance evaluation metric are: the response should be practical and operable, taking into account the impact of factors such as operating hours, weather, and traffic conditions, and thus possess operability.

[0155] Richness is defined as a performance evaluation metric that measures the richness of the model's response content. The calculation method for this metric is a multi-round scoring system, calculating the average. A test case that can be used to evaluate the cockpit model using this metric could be: "In XX location, besides the beach, what other fun places are there nearby?" The judgment criterion for this performance evaluation metric is: whether the content of the response or recommendation is rich.

[0156] Location planning accuracy is defined as a performance evaluation metric that measures the accuracy of locations and routes planned by the model. The metric is calculated as: number of user expectations met / number of inquiries. A test case that can be used to evaluate the cockpit model using this metric could be: "I want to spend two days in or around XX location; please plan my itinerary." The criteria for judging this performance evaluation metric include: the degree to which the intelligently planned attractions or locations match the user's actual interests or needs, and whether location errors occur.

[0157] Route planning accuracy is defined as a performance evaluation metric that measures the accuracy of the locations and routes planned by the model. The corresponding calculation method for this performance evaluation metric is: number of user expectations met / number of queries. A test case that can be used to evaluate the cockpit model using this performance evaluation metric could be: "What is the best driving route from my home to the office?" The judgment criteria for this performance evaluation metric are: measuring the degree of closeness between the planned route and the actual optimal route (the route on the map) (comparing distance, congestion, etc.).

[0158] Content timeliness is defined as a performance evaluation metric that measures the timeliness of the model's response content. The calculation method for this performance evaluation metric is as follows: a multi-round scoring system is used to calculate the average. When using this performance evaluation metric to evaluate the cockpit model, a possible test case is: "What are the latest technology news?" The judgment criterion for this performance evaluation metric is: the provided content should be up-to-date.

[0159] The dialogue interaction experience is defined as a performance evaluation metric that measures a user's ability to engage in in-depth conversations. The metric is calculated using a multi-turn scoring system, with the average score calculated. When evaluating a cockpit model using this metric, test cases can include multi-turn dialogues, such as asking about the weather followed by clothing suggestions. The criteria for judging this performance evaluation metric are: the dialogue should be fluent and have depth.

[0160] Dialogue coherence is defined as a performance evaluation metric that measures a model's ability to maintain contextual memory during dialogue. The metric is calculated using a multi-turn scoring system, with the average score calculated. Test cases that can be used to evaluate a cockpit model using this metric include multi-turn dialogues that maintain contextual consistency. The criterion for judging this performance evaluation metric is that the dialogue should maintain contextual coherence.

[0161] Music content accuracy is defined as a performance evaluation metric that measures the accuracy of the content in the model's responses. The metric is calculated as: number of correct responses / number of queries. A test case that can be used to evaluate the cockpit model using this metric could be: "What is the latest album by artist XX?" The criterion for judging this performance evaluation metric is that the response content should be accurate.

[0162] Film and television content accuracy is defined as a performance evaluation metric that measures the accuracy of the content in the model's responses. The calculation method for this performance evaluation metric is: number of correct responses / number of queries. A test case that can be used to evaluate the cockpit model using this performance evaluation metric could be: "I remember a story about a group of over a hundred people who revolted in the mountains, resisted, and eventually surrendered and were granted amnesty. What story was that again?" The judgment criterion for this performance evaluation metric is: the response content should be accurate and error-free.

[0163] Text content accuracy is defined as a performance evaluation metric that measures the accuracy of the content in the model's responses. The metric is calculated as: number of correct responses / number of queries. When using this metric to evaluate the cockpit model, a test case can be: "What is the outcome of mission XX?" The criterion for judging this performance evaluation metric is: the response content should be accurate and error-free.

[0164] Knowledge breadth is defined as a performance evaluation metric that measures whether a model's knowledge across multiple domains can cover user queries from multiple domains. The calculation method for this performance evaluation metric is subjective scoring. Test cases that can be used to evaluate the cockpit model using this performance evaluation metric include: "What are the applications of artificial intelligence in the medical field?" and "What are some rhetorical devices used in literature?" The judgment criterion for this performance evaluation metric is: the response should provide correct answers across multiple domains.

[0165] Knowledge depth is defined as a performance metric that measures a model's ability to mine knowledge. The calculation method for this performance evaluation metric is subjective scoring. A test case that can be used to evaluate the cockpit model using this performance evaluation metric is: "Explaining the basic principles of quantum computing." The judgment criterion for this performance evaluation metric is: the response should demonstrate a deep level of knowledge.

[0166] Cross-domain integration capability is defined as a performance evaluation metric that measures a model's ability to integrate knowledge from multiple domains. The metric is calculated as: number of accurately associated data points / total number of tests. A test case that can be used to evaluate a cockpit model using this metric could be: "Combining historical and geographical knowledge, explain why XX is called the Mother River." The criterion for judging this performance evaluation metric is: the response should accurately integrate knowledge from multiple domains.

[0167] The above are merely examples. More performance evaluation metrics can be set. When evaluating the smart cockpit model, matching performance evaluation metrics can be selected from the above performance evaluation metrics based on the usage scenario of the smart cockpit model to evaluate the model performance.

[0168] Step S405: Based on the performance evaluation indicators corresponding to the intelligent cockpit models that match multiple usage scenarios, construct a target cockpit indicator evaluation system.

[0169] In this embodiment, multiple performance evaluation indicators corresponding to multiple smart cockpit models are abstracted and unified to construct a target cockpit coordinate evaluation system.

[0170] Optionally, common performance evaluation indicators and scenario-dependent evaluation indicators are identified from the multiple performance evaluation indicators corresponding to multiple smart cockpit models. Then, referring to the description in step S204 above, a target cockpit indicator evaluation system is constructed based on the common performance evaluation indicators and scenario-dependent evaluation indicators. Further details are omitted here.

[0171] The target cockpit index evaluation system constructed through the above steps S401 to S405 can comprehensively and accurately evaluate the intelligent cockpit model from multiple evaluation dimensions. This not only promotes the technical optimization of the model itself, but also ensures that the intelligent cockpit system can meet the diverse needs of users, provide a higher level of safety and user experience, thereby promoting the high-quality development of the intelligent vehicle industry.

[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0173] According to an embodiment of this application, an embodiment of an evaluation device for a smart cockpit model is provided. It should be noted that the device can be used to perform the above-described evaluation method for a smart cockpit model.

[0174] Figure 5 This is a schematic diagram of an evaluation device for a smart cockpit model according to an embodiment of this application. Figure 5 As shown, the evaluation device 500 of the smart cockpit model may include: a first determining unit 501, a second determining unit 502, a testing unit 503, and an evaluation unit 504.

[0175] The first determining unit 501 is used to determine the smart cockpit model to be tested in the vehicle's smart cockpit system.

[0176] The second determining unit 502 is used to determine the evaluation strategy of the intelligent cockpit model under test based on the preset usage scenarios associated with the intelligent cockpit model under test. The evaluation strategy is used to represent the evaluation rules for evaluating the intelligent cockpit model under test, and the evaluation strategy is associated with the usage requirement information in the usage scenarios.

[0177] The evaluation unit 503 is used to evaluate the intelligent cockpit model under test according to the evaluation strategy and obtain the evaluation results.

[0178] Evaluation unit 504 is used to call the target cockpit index evaluation system and evaluate the model performance of the intelligent cockpit model under test from multiple evaluation dimensions according to the evaluation results, so as to obtain the performance evaluation results of the intelligent cockpit model under test. Among them, the multiple evaluation dimensions include at least the model application evaluation dimension, the system application evaluation dimension, the scenario application evaluation dimension, and the user experience evaluation dimension.

[0179] Optionally, the second determining unit 502 is further configured to: determine the usage requirement information of the intelligent cockpit model under test in the preset usage scenario based on the preset usage scenario associated with the intelligent cockpit model under test; obtain at least one test case based on the usage requirement information, wherein the at least one test case is used to simulate the operation behavior of the passenger in the process of using the intelligent cockpit model under test; and determine the evaluation strategy of the intelligent cockpit model under test based on the at least one test case.

[0180] Optionally, the evaluation unit 504 is further configured to: call multiple evaluation modules in the target cockpit index evaluation system, and evaluate the model performance of the intelligent cockpit model under test from multiple evaluation dimensions according to the evaluation results, so as to obtain the performance evaluation results of the intelligent cockpit model under test. The evaluation modules are used to evaluate the model performance of the intelligent cockpit model under test from at least one of the following dimensions: model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension. The evaluation modules correspond one-to-one with the evaluation dimensions.

[0181] Optionally, the evaluation unit 504 is further configured to: call the model application evaluation module in the target cockpit indicator evaluation system to evaluate the evaluation results and obtain the model application evaluation index of the intelligent cockpit model under test, wherein the model application evaluation index is used to at least indicate the degree of matching between the intelligent cockpit model under test and the preset usage scenario; call the system application evaluation module in the target cockpit indicator evaluation system to evaluate the evaluation results and obtain the service capability evaluation index of the intelligent cockpit model under test, wherein the service capability evaluation index is used to at least indicate the service performance of the intelligent cockpit model under test in the intelligent cockpit system; call the scenario application evaluation module in the target cockpit indicator evaluation system to evaluate the evaluation results and obtain the scenario application evaluation index of the intelligent cockpit model under test, wherein the scenario application evaluation index is used to at least indicate the task completion result of the intelligent cockpit model under test in the usage scenario; and call the user experience evaluation module in the target cockpit indicator evaluation system to evaluate the evaluation results and obtain the user experience evaluation index of the intelligent cockpit model under test, wherein the user experience evaluation index is used to at least indicate the user satisfaction level of the intelligent cockpit model under test in the usage scenario.

[0182] In the evaluation device for the intelligent cockpit model of this embodiment, when evaluating the intelligent cockpit model under test in the intelligent cockpit system, the evaluation strategy for the intelligent cockpit model under test can be determined based on the preset usage scenarios associated with the intelligent cockpit model under test. This determined evaluation strategy is closely related to the preset usage scenarios of the intelligent cockpit model under test. Compared to traditional evaluation methods that use more general test cases and lack in-depth understanding of specific scenario requirements, leading to deviations between evaluation results and actual application performance, this embodiment evaluates the intelligent cockpit model under test through an evaluation strategy combined with specific usage scenarios, improving the relevance and practicality of the evaluation. Furthermore, after obtaining the evaluation results, the target cockpit index evaluation system can be invoked. Based on the evaluation results, the intelligent cockpit model under test can be evaluated from multiple evaluation dimensions, including model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation. This allows for a more comprehensive evaluation of the overall performance of the intelligent cockpit model, achieving the technical effect of improving the evaluation accuracy of the intelligent cockpit model under test, thereby solving the technical problem of low evaluation accuracy of intelligent cockpit models under test in related technologies.

[0183] According to an embodiment of this application, an embodiment of a construction apparatus for a cabin index evaluation system is provided. It should be noted that the apparatus can be used to execute the above-described construction method for the cabin index evaluation system.

[0184] Figure 6 This is a schematic diagram of a construction device for a cabin performance evaluation system according to an embodiment of this application. Figure 6As shown, the construction device 600 of the cabin index evaluation system may include: a first acquisition unit 601, a third determination unit 602, a second acquisition unit 603, and a construction unit 604.

[0185] The first acquisition unit 601 is used to acquire multiple preset scenario information of the vehicle's intelligent cockpit system. The preset scenario information is used to represent the preset usage scenarios of the intelligent cockpit system, and the preset scenario information corresponds one-to-one with the preset usage scenarios.

[0186] The third determining unit 602 is used to determine intelligent cockpit models that match the preset usage scenarios based on multiple preset scenario information, thereby obtaining multiple intelligent cockpit models. The intelligent cockpit models are used to respond to task requests in the preset usage scenarios and generate corresponding response results.

[0187] The second acquisition unit 603 is used to acquire multiple performance evaluation indicators corresponding to multiple smart cockpit models. The performance evaluation indicators are used to evaluate the model performance of the smart cockpit model in a preset use scenario from at least one of the model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension. The performance evaluation indicators correspond one-to-one with the evaluation dimensions.

[0188] The construction unit 604 is used to construct a target cockpit index evaluation system based on multiple performance evaluation indicators corresponding to multiple smart cockpit models. The target cockpit index evaluation system is used to evaluate the model performance of the aforementioned smart cockpit models under test.

[0189] Optionally, the third determining unit 602 is further configured to: determine the usage requirement information corresponding to the preset usage scenario based on multiple preset scenario information, thereby obtaining the usage requirement information corresponding to the multiple preset usage scenarios respectively, wherein the usage requirement information is used to represent the usage requirements in the preset usage scenario; and determine the intelligent cockpit model that matches the preset usage scenario based on the usage requirement information corresponding to the multiple preset usage scenarios respectively.

[0190] Optionally, the construction unit 604 is further configured to: identify common performance evaluation indicators and scenario-dependent performance evaluation indicators from multiple performance evaluation indicators corresponding to multiple smart cockpit models, wherein the common performance evaluation indicators are performance evaluation indicators used to evaluate the same performance characteristics of multiple smart cockpit models, and the scenario-dependent performance evaluation indicators are performance evaluation indicators used to evaluate the performance characteristics of any smart cockpit model in a corresponding preset usage scenario; fuse the common performance evaluation indicators of the same type to obtain the fused common performance evaluation indicators; and construct a target cockpit indicator evaluation system based on the fused common performance evaluation indicators and the scenario-dependent performance evaluation indicators.

[0191] Optionally, the construction unit 604 is further configured to: determine the evaluation modules to which the fused common performance evaluation indicators and scenario-dependent performance evaluation indicators belong in the initial cockpit indicator evaluation system, based on the mapping relationship between the indicator types of performance evaluation indicators and evaluation modules. The evaluation modules are at least one of the following: model application evaluation module, system application evaluation module, scenario application evaluation module, and user experience evaluation module. The model application evaluation module is used to evaluate the matching degree between the intelligent cockpit model and the preset usage scenario; the system application evaluation module is used to evaluate the service performance of the intelligent cockpit model in the intelligent cockpit system; the scenario application evaluation module is used to evaluate the task completion results of the intelligent cockpit model in the usage scenario; and the user experience evaluation module is used to evaluate the user satisfaction level of the intelligent cockpit model in the usage scenario. The fused common performance evaluation indicators and scenario-dependent performance evaluation indicators are then assigned to their respective evaluation modules to obtain the target cockpit indicator evaluation system.

[0192] In the construction apparatus of the cockpit index evaluation system in this embodiment, by deeply analyzing information from multiple preset scenarios, the specific needs of each scenario and the expectations of the driver / passenger can be accurately understood, thereby selecting or designing an intelligent cockpit model that highly matches the preset usage scenarios. This precise scenario adaptability ensures that the intelligent cockpit model can provide optimal service in various usage situations, improving the overall applicability and user satisfaction of the intelligent cockpit system. By integrating multiple evaluation indicators corresponding to multiple intelligent cockpit models into the evaluation module of the initial intelligent cockpit model, the target intelligent cockpit model is obtained. The model performance of the intelligent cockpit model under test can be comprehensively evaluated from the dimensions of model application evaluation, system application evaluation, scenario application evaluation, and user experience evaluation. This provides a benchmark and direction for the development and performance optimization of intelligent cockpit models, enhances the scalability and flexibility of the evaluation system, and achieves the technical effect of improving the evaluation accuracy of intelligent cockpit models, thereby solving the technical problem of low evaluation accuracy of intelligent cockpit models.

[0193] Embodiments of this application also provide a vehicle, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods described in various embodiments of this application when it runs.

[0194] Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.

[0195] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.

[0196] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of this application.

[0197] Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of this application.

[0198] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0199] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0200] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0201] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0202] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0203] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for evaluating an intelligent cockpit model, characterized in that, include: Identify the intelligent cockpit model to be tested in the vehicle's intelligent cockpit system; Based on the preset usage scenarios associated with the smart cockpit model under test, an evaluation strategy for the smart cockpit model under test is determined. The evaluation strategy is used to represent the evaluation rules for evaluating the smart cockpit model under test, and the evaluation strategy is associated with the usage requirement information in the preset usage scenarios. The intelligent cockpit model under test is evaluated according to the evaluation strategy, and the evaluation results are obtained. The target cockpit index evaluation system is invoked, and the model performance of the intelligent cockpit model under test is evaluated from multiple evaluation dimensions according to the evaluation results, so as to obtain the performance evaluation result of the intelligent cockpit model under test. The multiple evaluation dimensions include at least the model application evaluation dimension, the system application evaluation dimension, the scenario application evaluation dimension, and the user experience evaluation dimension.

2. The method according to claim 1, characterized in that, Based on the preset usage scenarios associated with the intelligent cockpit model under test, the evaluation strategy for the intelligent cockpit model under test is determined, including: Based on the preset usage scenarios associated with the intelligent cockpit model under test, determine the usage requirement information of the intelligent cockpit model under test in the preset usage scenarios; Based on the usage requirements information, at least one test case is obtained, wherein the at least one test case is used to simulate the operational behavior of a passenger using the smart cockpit model under test. Based on the at least one test case, determine the evaluation strategy for the smart cockpit model under test.

3. The method according to claim 1, characterized in that, The target cockpit performance evaluation system is invoked, and the model performance of the intelligent cockpit model under test is evaluated from multiple evaluation dimensions according to the evaluation results, to obtain the performance evaluation results of the intelligent cockpit model under test, including: Multiple evaluation modules in the target cockpit index evaluation system are invoked to evaluate the model performance of the intelligent cockpit model under test from multiple evaluation dimensions according to the evaluation results, thereby obtaining the performance evaluation result of the intelligent cockpit model under test. The evaluation modules are used to evaluate the model performance of the intelligent cockpit model under test from at least one of the model application evaluation dimension, the system application evaluation dimension, the scenario application evaluation dimension, and the user experience evaluation dimension, and the evaluation modules correspond one-to-one with the evaluation dimensions.

4. The method according to claim 3, characterized in that, Multiple evaluation modules in the target cockpit performance evaluation system are invoked, and the model performance of the intelligent cockpit model under test is evaluated from multiple evaluation dimensions according to the evaluation results, to obtain the performance evaluation results of the intelligent cockpit model under test, including: The model application evaluation module in the target cockpit index evaluation system is invoked to evaluate the evaluation results and obtain the model application evaluation index of the smart cockpit model under test. The model application evaluation index is used to at least indicate the degree of matching between the smart cockpit model under test and the preset usage scenario. The system application evaluation module of the target cockpit index evaluation system is invoked to evaluate the evaluation results and obtain the service capability evaluation index of the intelligent cockpit model under test. The service capability evaluation index is used to at least indicate the service performance of the intelligent cockpit model under test in the intelligent cockpit system. The scenario application evaluation module in the target cockpit index evaluation system is invoked to evaluate the evaluation results and obtain the scenario application evaluation index of the intelligent cockpit model under test. The scenario application evaluation index is used to at least indicate the task completion result of the intelligent cockpit model under test in the preset usage scenario. The user experience evaluation module in the target cockpit index evaluation system is invoked to evaluate the evaluation results and obtain the user experience evaluation index of the smart cockpit model under test. The user experience evaluation index is used to at least indicate the user satisfaction level of the smart cockpit model under test in the preset usage scenario.

5. A method for constructing a cabin performance evaluation system, characterized in that, include: The system acquires multiple preset scenario information of the vehicle's intelligent cockpit system, wherein the preset scenario information is used to represent preset usage scenarios of the intelligent cockpit system, and the preset scenario information corresponds one-to-one with the preset usage scenarios; Based on multiple preset scenario information, intelligent cockpit models that match the preset usage scenarios are determined respectively, resulting in multiple intelligent cockpit models. The intelligent cockpit models are used to respond to task requests in the preset usage scenarios and generate corresponding response results. Multiple performance evaluation metrics are obtained for each of the multiple intelligent cockpit models. The performance evaluation metrics are used to evaluate the model performance of the intelligent cockpit model in the preset use scenario from at least one of the model application evaluation dimension, system application evaluation dimension, scenario application evaluation dimension, and user experience evaluation dimension. The performance evaluation metrics correspond one-to-one with the evaluation dimensions. Based on the multiple performance evaluation indicators corresponding to the multiple intelligent cockpit models, a target cockpit indicator evaluation system is constructed, wherein the target cockpit indicator evaluation system is used to evaluate the model performance of the intelligent cockpit model under test in the method of any one of claims 1 to 4.

6. The method according to claim 5, characterized in that, Based on multiple preset scenario information, intelligent cockpit models matching the preset usage scenarios are determined respectively, including: Based on multiple preset scenario information, usage requirement information corresponding to each preset usage scenario is determined, thereby obtaining usage requirement information corresponding to each preset usage scenario, wherein the usage requirement information is used to represent the usage requirements in the preset usage scenario; Based on the usage demand information corresponding to the multiple preset usage scenarios, a smart cockpit model matching the preset usage scenario is determined.

7. The method according to claim 5, characterized in that, Based on the various performance evaluation indicators corresponding to the various intelligent cockpit models, a target cockpit indicator evaluation system is constructed, including: From the multiple performance evaluation metrics corresponding to the multiple smart cockpit models, common performance evaluation metrics and scenario-dependent performance evaluation metrics are identified. The common performance evaluation metrics are performance evaluation metrics used to evaluate the same performance characteristics of the multiple smart cockpit models, and the scenario-dependent performance evaluation metrics are performance evaluation metrics used to evaluate the performance characteristics of any smart cockpit model in the corresponding preset use scenario. The common performance evaluation indicators with the same indicator type are fused to obtain the fused common performance evaluation indicators; Based on the fused common performance evaluation indicators and the scenario-dependent performance evaluation indicators, the target cockpit indicator evaluation system is constructed.

8. The method according to claim 7, characterized in that, Based on the fused common performance evaluation metrics and the scenario-dependent performance evaluation metrics, the target cockpit metric evaluation system is constructed, including: Based on the mapping relationship between the performance evaluation index types and evaluation modules, the fused common performance evaluation index and the scenario-dependent performance evaluation index are respectively determined to belong to the evaluation module in the initial cockpit index evaluation system. The evaluation module is at least one of the following: model application evaluation module, system application evaluation module, scenario application evaluation module, and user experience evaluation module. The model application evaluation module is used to evaluate the matching degree between the intelligent cockpit model and the preset usage scenario. The system application evaluation module is used to evaluate the service performance of the intelligent cockpit model in the intelligent cockpit system. The scenario application evaluation module is used to evaluate the task completion results of the intelligent cockpit model in the preset usage scenario. The user experience evaluation module is used to evaluate the user satisfaction level of the intelligent cockpit model in the preset usage scenario. The fused common performance evaluation index and the scenario-dependent performance evaluation index are respectively assigned to their respective evaluation modules to obtain the target cockpit index evaluation system.

9. A vehicle, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 4, or performs the method according to any one of claims 5 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the method of any one of claims 1 to 4, or to perform the method of any one of claims 5 to 8.