Generation method of vehicle function test instruction, readable storage medium and program product

By strengthening the training of human feedback and multiple reward models, the vehicle-machine functional test instruction generation model is optimized, and the traditional test efficiency and insufficient coverage are solved, and more efficient and accurate vehicle-machine functional test is achieved.

CN120162233APending Publication Date: 2025-06-17ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510321768.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Traditional vehicle computer functional testing relies on manual manual, is inefficient, error-prone, and test coverage is difficult to guarantee. Models based on labeled data cannot fully capture human complex preferences and subjective feelings, resulting in increased testing costs and reduced coverage.

Method used

A method for generating vehicle-machine functional test instructions is proposed. By obtaining the original input data, inputting the original test instructions to generate the model, using reinforcement learning human feedback (RLHF) and multiple reward models to train, optimize the parameters of the test instruction generation model, and generate target test instructions.

Benefits of technology

It effectively reduces testing costs, improves test coverage and accuracy, can better capture human subjectivity and complex preferences, and generates test instructions that are more in line with actual user behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162233A_ABST
    Figure CN120162233A_ABST
Patent Text Reader

Abstract

One or more embodiments of the present specification provide a method for generating a vehicle function test instruction, a readable storage medium and a program product, the method comprising: for a target vehicle of a to-be-tested vehicle function, obtaining original input data corresponding to the vehicle function; inputting the original input data into an original test instruction generation model, and obtaining an original test instruction output by the original test instruction generation model; training a plurality of reward models based on reinforcement learning human feedback RLHF and the original test instruction, wherein each reward model evaluates the original test instruction based on different evaluation dimensions of the vehicle function; parameters of the original test instruction generation model are optimized according to evaluation results output by the multiple reward models respectively, so that a target test instruction generation model is obtained, and the target test instruction generation model is used for outputting a target test instruction for the vehicle function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of software testing, and in particular, to a method for generating vehicle infotainment system function test instructions, a readable storage medium, and a program product. Background Art

[0002] With the development of intelligent vehicle technology, the vehicle infotainment system it carries has begun to be upgraded from traditional independent electronic device components to an intelligent cockpit integrated with a large number of functions such as navigation, entertainment, and voice. Due to the complexity of its system, relevant enterprises need to conduct a large number of tests on various vehicle infotainment system functions to ensure their reliability and user experience. However, traditional vehicle infotainment system function tests usually rely on manual testing, which has the defects of low efficiency, easy to make mistakes, and difficult to guarantee test coverage. Moreover, as the number of vehicle infotainment system functions increases, these limitations become more prominent.

[0003] In related technologies, a pre-trained large model's parameters are usually fine-tuned based on supervised learning using a labeled dataset to make the model adapt to specific automotive scenarios, and then automated tests are performed on the vehicle infotainment system functions in that scenario. However, this method has a high demand for a large amount of high-quality labeled data, and obtaining labeled data is often costly and time-consuming, resulting in an increase in the cost of testing. At the same time, a model that only learns based on labeled data cannot fully capture the complex preferences and subjective feelings of humans, resulting in difficulty in covering the actual behaviors of real users when using vehicle infotainment system functions, and thus reducing the test coverage and accuracy. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide the following technical solutions:

[0005] According to a first aspect of one or more embodiments of this specification, a method for generating vehicle infotainment system function test instructions is proposed, and the method includes:

[0006] For a target vehicle of a vehicle infotainment system function to be tested, obtain the original input data corresponding to the vehicle infotainment system function; and input the original input data into an original test instruction generation model, and obtain the original test instructions output by the original test instruction generation model;

[0007] Train multiple reward models respectively based on Reinforcement Learning from Human Feedback (RLHF) and the original test instructions, and each reward model evaluates the original test instructions respectively based on different evaluation dimensions of the vehicle infotainment system function;

[0008] Optimize the parameters of the original test instruction generation model according to the evaluation results respectively output by the multiple reward models to obtain a target test instruction generation model, and the target test instruction generation model is used to output target test instructions for the vehicle infotainment system function.

[0009] Optionally, training multiple reward models respectively based on the Reinforcement Learning from Human Feedback (RLHF) technology and the original test instructions includes:

[0010] For each of the multiple reward models, obtain the feedback results of humans for the original test instructions under the respective evaluation dimensions corresponding to this reward model, and train this reward model according to the feedback results.

[0011] Optionally, the training of this reward model according to the feedback results includes:

[0012] Extract target features from the original input data and the original test instructions;

[0013] Use the target features and the feedback results as training data to train this reward model.

[0014] Optionally, optimizing the parameters of the original test instruction generation model according to the evaluation results respectively output by the multiple reward models includes:

[0015] Obtain the evaluation results respectively output by each of the multiple reward models;

[0016] Input the target expectations into the minimax function respectively, and iteratively optimize the parameters of the original test instruction generation model until the target expectations are less than the expectation threshold, where the target expectations are used to represent the expectations of the differences between each evaluation result and other evaluation results.

[0017] Optionally, the method further includes:

[0018] For the target vehicle of the in-vehicle system function to be tested, obtain the target input data corresponding to the in-vehicle system function; input the target input data into the target test instruction generation model;

[0019] When the target test instruction generation model generates a to-be-processed test instruction according to the target input data, determine the expected anomalies in the in-vehicle system function test scenario corresponding to the target test instruction through Self-Consistency Thought Chain (SeCoT);

[0020] Determine corresponding additional test instructions according to the expected anomalies, and combine the additional test instructions with the to-be-processed test instruction to generate the target test instruction.

[0021] Optionally, the target test instruction generation model is associated with a memory module, and the memory module includes a short-term memory module and a long-term memory module; the method further includes:

[0022] Obtain the data to be memorized, where the data to be memorized includes the feedback data corresponding to the test instruction, and the feedback data includes functional feedback data for characterizing the vehicle state during the execution of the target test instruction by the target vehicle and / or user emotional feedback data for characterizing the emotional state of the user during the execution of the target test instruction by the target vehicle;

[0023] According to the data type of the data to be memorized, store the data to be memorized as historical data in the short-term memory module or the long-term memory module, and the historical data is used to feedback the target test instruction generation model to generate the target test instruction.

[0024] Optionally, the method further includes:

[0025] Obtain the historical data in the short-term memory module and the long-term memory module;

[0026] When the memory elimination condition is met, eliminate at least a part of the historical data from the obtained historical data according to the memory elimination rule.

[0027] Optionally, the method further includes:

[0028] Obtain the expected execution result of the target test instruction, and determine the actual execution result of the target test instruction according to the feedback data;

[0029] Adjust the memory elimination rule according to the expected execution result and the actual execution result.

[0030] Optionally, the method further includes:

[0031] When the data types of the first historical data in the short-term memory module and the second historical data in the long-term memory module match each other, determine the incremental data of the first historical data compared with the second historical data, and store the incremental data in the long-term memory module.

[0032] Optionally, the corresponding relationship between the historical data and the test scenario is maintained in the memory module; the method further includes:

[0033] Determine the target test scenario in which the target vehicle is located according to the in-vehicle computer state of the target vehicle;

[0034] Query the historical data corresponding to the target test scenario from the memory module according to the corresponding relationship, so as to preferentially input it into the target test instruction generation model.

[0035] According to the second aspect of one or more embodiments of this specification, a device for generating a vehicle-mounted function test instruction is proposed, and the device includes:

[0036] An original test instruction acquisition unit, configured to obtain, for a target vehicle of a vehicle head unit function to be tested, original input data corresponding to the vehicle head unit function; and input the original input data into an original test instruction generation model, and obtain an original test instruction output by the original test instruction generation model;

[0037] An original test instruction evaluation unit, configured to train a plurality of reward models respectively based on reinforcement learning from human feedback (RLHF) and the original test instructions, and each reward model evaluates the original test instructions respectively based on different evaluation dimensions of the vehicle head unit function;

[0038] A target test instruction generation model acquisition unit, configured to optimize parameters of the original test instruction generation model according to evaluation results respectively output by the plurality of reward models, so as to obtain a target test instruction generation model, and the target test instruction generation model is configured to output a target test instruction for the vehicle head unit function.

[0039] According to a third aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0040] According to a fourth aspect of one or more embodiments of the present specification, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0041] In the technical solution provided in this specification, after obtaining the original input data corresponding to the target vehicle for the in-vehicle infotainment (IVI) function to be tested, the original input data can be input into the original test instruction generation model, and the original test instructions output by the model can be obtained. Then, multiple reward models can be trained respectively based on Reinforcement Learning from Human Feedback (RLHF) and the obtained original test instructions. Each reward model can evaluate the original test instructions based on different evaluation dimensions of the IVI function, and the parameters of the original test instruction generation model can be optimized according to the evaluation results respectively output by the multiple reward models to obtain the optimized target test instruction generation model, which is then used to output the target test instructions for the IVI function. Among them, based on RLHF, there is no need to prepare a large amount of labeled data in advance, but relevant feedback is gradually obtained during the training process of the multiple reward models, effectively reducing the workload in the traditional data annotation process and helping to reduce the test cost of the IVI function. In addition, the evaluation results of RLHF are derived from human feedback, and the multiple reward models evaluate the original test instructions from different dimensions of the IVI function. In this way, the target test instruction generation model optimized based on these evaluation results can output target test instructions that not only conform to the subjectivity and flexibility of real humans but also meet the multi-dimensional evaluation requirements, increasing the test coverage and accuracy for the IVI function. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0043] Figure 1 is a schematic diagram of the architecture of a system for generating test instructions for an in-vehicle infotainment (IVI) function provided by an exemplary embodiment;

[0044] Figure 2 is a schematic flowchart of a method for generating test instructions for an in-vehicle infotainment (IVI) function provided by an exemplary embodiment;

[0045] Figure 3 is a schematic flowchart of another method for generating test instructions for an in-vehicle infotainment (IVI) function provided by an exemplary embodiment;

[0046] Figure 4 is a schematic diagram of the architecture of a system for testing an in-vehicle infotainment (IVI) function provided by an exemplary embodiment;

[0047] Figure 5It is a schematic flowchart of a process for executing a target test instruction provided by an exemplary embodiment;

[0048] Figure 6 It is a schematic flowchart of a data storage method based on a short-term memory module and a long-term memory module provided by an exemplary embodiment;

[0049] Figure 7 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment;

[0050] Figure 8 It is a schematic structural diagram of a generating device for a vehicle-mounted function test instruction provided by an exemplary embodiment. Detailed Description of the Invention

[0051] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention.

[0052] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0053] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0054] Figure 1 It is a schematic architecture diagram of a generating system for a vehicle-mounted function test instruction provided by an exemplary embodiment. As Figure 1 shown, the system may at least include a target vehicle 12, an original test instruction generation model 14, a reward model set 16, and a target test instruction generation model 18.

[0055] The target vehicle 12 is the carrier for the in-vehicle infotainment (IVI) functions to be tested. For specific IVI functions, by interacting with the target vehicle, the original input data corresponding to the IVI functions can be obtained. These data reflect the specific conditions and operating conditions of the IVI functions in the actual scenario, such as the version information of the IVI system, the sensor data of the vehicle, such as vehicle speed, acceleration, etc., and the regular operation habit data of the user, etc., providing a necessary basis for the generation of subsequent test instructions.

[0056] The present invention does not limit the IVI functions to be tested, which can be any function supported by the IVI system of the target vehicle 12. Exemplarily, it can be IVI functions such as navigation function, voice command function, in-vehicle entertainment function, Bluetooth connection function, air conditioning and seat adjustment function, IVI system compatibility and stability function, communication function, etc.

[0057] The original test instruction generation model 14 is a pre-trained large language model (LLM). Specifically, this model has been pre-trained on a large scale of general data and has powerful language understanding and generation capabilities. In the scenario of generating test instructions for IVI functions, it can receive the original input data obtained from the target vehicle, and use the knowledge and patterns learned from its own pre-training to convert the original input data into original test instructions. Due to its pre-training characteristics, it can quickly generate test instructions with a certain logic and feasibility based on the above original input data, but may not be optimized for specific evaluation dimensions of the IVI functions.

[0058] The reward model set 16 consists of multiple reward models such as reward model 1, reward model 2... reward model n, where n is a positive integer greater than 2. Each reward model can select a machine learning model such as a neural network model as the basic architecture according to actual needs, and after initializing the parameters of the corresponding reward model, it can be trained based on the reinforcement learning human feedback (RLHF) and the original test instructions generated by the original test instruction generation model 14. Each trained reward model can evaluate the original test instructions for different evaluation dimensions of the IVI functions. These evaluation dimensions can be set according to the characteristics of the IVI functions and test requirements. Each reward model calculates the scores of the original test instructions on their respective corresponding evaluation dimensions as the evaluation results, providing feedback information for the subsequent optimization of the original test instruction generation model.

[0059] The target test instruction generation model 18 is a large language model obtained by optimizing the parameters of the original test instruction generation model 14 based on the evaluation results output by the reward model set. It inherits the basic capabilities of the original test instruction generation model and is adjusted and optimized for specific evaluation dimensions of in-vehicle infotainment (IVI) functions. Therefore, the target test instruction generation model 18 can output target test instructions that better meet the test requirements. These instructions perform well in multiple evaluation dimensions and can test the IVI functions more accurately and comprehensively.

[0060] It can be understood that the above-mentioned generation system of IVI function test instructions can be regarded as an AI large model intelligent agent (AI Model Agent), that is, a generation proxy module driven by an LLM (Large Language Model) with the above-mentioned original test instruction generation model 14 and the above-mentioned target test instruction generation model 18. Among them, the generation proxy refers to a program or entity that can generate new content or make decisions based on given input information using specific technologies and algorithms. The generation proxy module can be regarded as a physical unit or logical unit carrying the above-mentioned generation proxy. In the scenario of IVI function testing, this AI large model intelligent agent can generate corresponding original test instructions based on the original input data related to the IVI function obtained from the target vehicle. At the same time, it can also generate target test instructions by virtue of the powerful language understanding and generation capabilities of the LLM to simulate the behavior operations of real users on the IVI system.

[0061] Figure 2 It is a schematic flowchart of a method for generating IVI function test instructions provided by an exemplary embodiment, as Figure 2 shown. The above method includes:

[0062] Step S202, for the target vehicle of the IVI function to be tested, obtain the original input data corresponding to the IVI function; and input the original input data into the original test instruction generation model, and obtain the original test instructions output by the original test instruction generation model.

[0063] When the user hopes to perform automated testing on the IVI function to be tested of the above target vehicle based on the above AI large model intelligent agent, first, the original input data corresponding to the above IVI function can be obtained. The above data can at least include the following types of data:

[0064] I. IVI system data: The IVI system data can further include the version information of the IVI system, such as the operating system version number, application version number, etc. These information can be used to determine the compatibility and applicability of the above original input data; it can also include the device logs of the target vehicle, such as the positioning logs of the current navigation path, voice interaction logs, screen operation logs, etc. These logs can be used to determine the operation results of various operations performed by the target vehicle.

[0065] II. Driving data: The vehicle driving data collected by motion sensors, such as the vehicle speed, acceleration, steering angle, braking state, etc. of the target vehicle. These data can reflect the state of the vehicle during historical and actual operation processes, providing historical and real-time bases for the generation of the original input data.

[0066] III. User preference data: Data used to describe the user's regular operation habits, such as the navigation routes commonly used by the user, music playlists, voice modes, air conditioner settings, etc. These data can help the original test instruction generation model generate more user-habit-compliant original test instructions.

[0067] IV. Environmental data: Data collected by environmental sensors such as cameras and thermohygrometers on the target vehicle, such as the weather conditions, road conditions, traffic signals, map information, etc. of the route where the target vehicle is driving. These data can be used to generate test instructions related to specific environmental conditions.

[0068] The types and specific categories of the above original input data can be determined according to the types of the above in-vehicle system functions. The specific method can be determined based on the preset relationship between the in-vehicle system functions and the original input data, that is, the in-vehicle system of the target vehicle can, through predefined rules or mapping relationships, clarify which types of original input data are required for a certain in-vehicle system function. Taking the automatic parking system, an in-vehicle system function, as an example, the software version and sensor status in the in-vehicle system data, the vehicle speed and steering angle in the driving data, the camera images and ultrasonic sensor data in the environmental data can be used as the original input data corresponding to this in-vehicle system function; or it can be determined according to the similarity degree between the in-vehicle system functions and the original input data, that is, the in-vehicle system of the target vehicle, based on machine learning models such as clustering algorithms and similarity calculations, or a rule engine, analyzes the similarity between each in-vehicle system function and the original input data, and dynamically infers which types of original input data are required for a certain in-vehicle system function. Taking the vehicle navigation, an in-vehicle system function, as an example, the system can, by analyzing the similarity, use the positioning logs in the in-vehicle system data, the map information and road condition information in the environmental data, etc., data that are similar to the vehicle navigation function in name or effect, as the original input data.

[0069] The obtained original input data can be input into the original test instruction generation model. As mentioned above, this model, as a pre-trained LLM, can understand the semantics of the original input data and generate original test instructions based on its pre-trained knowledge. Specifically, during the generation process, the above-mentioned original test instruction generation model can consider the context information of the original input data, and combine the specific requirements of in-vehicle functions and the prompt words (Prompts) previously input by testers of in-vehicle functions into the original test instruction generation model to generate test instructions with certain logic and feasibility. For example, if the original input data is the original input data corresponding to the navigation function, the model can generate a series of original test instructions related to the navigation scenario, such as "navigation route planning with location A set as the navigation destination", "re-planning other navigation routes during navigation", or "positioning accuracy in weak signal areas" and corresponding test steps. In addition, the above-mentioned original test instructions can also include test conditions and expected results, etc., which are not limited in this specification.

[0070] Of course, after obtaining the original input data and before inputting it into the original test instruction generation model, preprocessing operations such as data cleaning, data formatting, and data augmentation can also be performed on the original input data to ensure the quality and consistency of the original input data. Among them, the above-mentioned data cleaning can include operations such as removing noise data, filling in missing values, and correcting incorrect data; the above-mentioned data formatting can be converting the data into a format suitable for input into the original test instruction generation model; the above-mentioned data augmentation can include increasing the diversity and richness of the data by means of data interpolation or data synthesis to improve the generalization ability of the model. This is not limited in this specification.

[0071] Regarding the user preference data in the above original input data, the solution in this specification can also analyze and extract corresponding user preference features therefrom, so as to construct a corresponding personalized profile for the users of the target vehicle, enabling the original test instruction generation model to generate more original test instructions that conform to a user's behavior based on the personalized profile. Similarly, the above personalized profile can also be applied to the above target test instruction generation model to achieve the same effect. Among them, the personalized profile can store the above user preference features in the form of data tables, string texts, etc. in a database. Taking the data table as an example, the personalized profile can include: user preference identity (ID), which is used to uniquely identify user preferences; music preference, which is used to reflect the user's preference degree for different types of music such as pop and classical; voice command preference, which is used to reflect the voice commands preferred by the user; fragrance preference, which is used to record the user's preferences for different fragrance types such as floral and woody; self-driving assistant preference, which is used to indicate whether the user enables or prefers to use the self-driving assistant function; entertainment usage frequency, which is used to reflect the user's usage frequency in the cockpit entertainment domain; recent usage time, which is used to record the time when the user last used the cockpit entertainment function of the target vehicle; personalized settings, which are used to store user personalized settings such as volume and display brightness; remarks, which are used to record other non-sensitive information, such as preference features that cannot be matched to other user preference data. Among them, the number and format of user preference features are not limited in this specification, and testers of in-vehicle infotainment functions can define user preference features in the personalized profile according to actual test requirements. In addition, the above personalized profile can also be centrally maintained in the user profile module. This user profile module can be deployed as a sub-module of the generation proxy module, such as in the original test instruction generation model and the target test instruction generation model, or as shown in Figure 4 shown, as a physical unit or logical unit independent of the above generation proxy module, which is not limited in this specification.

[0072] Step S204: Train multiple reward models based on Reinforcement Learning from Human Feedback (RLHF) and the original test instructions respectively, and each reward model evaluates the original test instructions based on different evaluation dimensions of the in-vehicle infotainment function.

[0073] After the original test instructions output by the original test instruction generation model, they can be combined with RLHF to train the reward models that have been pre-emailed and initialized, so that each trained reward model can evaluate the original test instructions based on different evaluation dimensions for the in-vehicle infotainment function to be tested.

[0074] The above evaluation dimensions may include: accuracy, rationality, time efficiency, resource consumption, maintainability, fault discovery ability, security, and compatibility. Specifically, accuracy refers to evaluating whether the original test instructions conform to the scope of capabilities supported by the in-vehicle system of the target vehicle. For example, whether the input original test instructions can be executed and whether they conform to the call logic of the in-vehicle application programming interface (API); rationality refers to evaluating the integrity and rationality of the original test instructions. For example, whether the test operations in the original test instructions comprehensively cover different test scenarios and whether edge test situations are considered; time efficiency refers to evaluating the time cost required when executing the original test instructions. For example, whether the execution time of the test operations in the original test instructions is greater than the preset operation duration; resource consumption refers to evaluating the consumption degree of the in-vehicle system resources, such as CPU, memory, storage, etc. during the execution of the original test instructions; maintainability refers to evaluating aspects such as the clarity of the test steps, the scalability of the test data, and the modifiability of the test script in the original test instructions; fault discovery ability refers to evaluating whether the original test instructions can effectively discover potential faults and errors in the in-vehicle system, such as anomalies, function deficiencies, data errors, etc. in the in-vehicle system; security refers to evaluating whether the original test instructions consider the security of the in-vehicle system. For example, whether it involves the protection of sensitive information and whether it may trigger security vulnerabilities; and compatibility refers to evaluating whether the original test instructions are applicable to in-vehicle systems of different models and different versions to ensure that the in-vehicle functions can be effectively tested in different hardware and software environments.

[0075] For the training process of each of the multiple reward models, it may include obtaining, for the evaluation dimension corresponding to the reward model, the feedback results of humans on the above original test instructions under this evaluation dimension, and training the reward model according to the feedback results. Among them, the so-called humans may refer to human experts or testers in the field of in-vehicle function testing. They can give feedback on the original test instructions based on their own experience and professional knowledge. Specifically, the above humans can generate feedback results through real-time annotation, enabling the reward model to better meet the requirements of humans for in-vehicle function testing and providing reliable feedback for optimizing the original test instruction generation model. In addition, these feedback results can be further divided into multiple feedback sub-results according to different evaluation dimensions.

[0076] Suppose there are two reward models: one is an accuracy evaluation model for evaluating the accuracy of the original test instructions based on the accuracy dimension of the in-vehicle system functions, and the other is a rationality evaluation model for evaluating the rationality of the original test instructions based on the rationality dimension of the in-vehicle system functions. For the same original test instruction, the system can respectively obtain two feedback sub-results provided by human experts based on the accuracy dimension and the rationality dimension. The feedback sub-result of the accuracy dimension can include whether the original test instruction can effectively trigger the expected behavior of the in-vehicle system functions, while the feedback sub-result of the rationality dimension can include whether the original test instruction comprehensively covers different test scenarios. These feedback sub-results will be used as training data to train the reward models corresponding to the evaluation dimensions respectively. In addition, to further improve the evaluation accuracy of the reward models, the system can also obtain the feedback results of human experts on additional original test instructions based on a single evaluation dimension. For example, for the accuracy evaluation model, the system can specifically collect the scores and annotations of human experts on multiple test instructions in terms of the accuracy dimension. These additional feedback data can be used for targeted training of the accuracy evaluation model, enabling it to more accurately identify and evaluate the accuracy of the test instructions. In this way, the reward models can continuously optimize their performance in specific evaluation dimensions, thereby providing more reliable feedback information for the parameter optimization of the original test instruction generation model. In summary, the training process of multiple reward models in this specification relies on high-quality human feedback data, and through multi-dimensional and multi-round targeted training, gradually improves their evaluation ability and accuracy. This training mechanism based on human expert feedback ensures that the reward models can fully understand the complex requirements of the in-vehicle system functions and provide strong support for generating high-quality target test instructions.

[0077] The process of training any of the above reward models based on the feedback results can be further understood as extracting target features from the original input data and the original test instructions, and then using the above target features and the above feedback results as training data to train the reward model, so that the reward model can better learn the association between the original test instruction features and human evaluations and improve its evaluation accuracy. Among them, the features extracted from the original input data and the original test instructions can include text features such as the word vector representation, semantic features, and correlation metrics of the text. For example, the original input data and the original test instructions are converted into word vectors, and correlation features are obtained by calculating vector similarity and other methods. At the same time, the extracted features can be combined with the above feedback results, such as the scores or evaluation information feedback by humans, to form training data together. For example, the word vector features of the original test instructions, the correlation scores with the original input data, etc. are used as inputs, and the feedback sub-results provided by humans based on different evaluation dimensions are used as labels. For example, a score of 5 is given for the accuracy of the original test instructions, and a score of 3 is given for the rationality. In addition, to use the above training data, appropriate machine learning algorithms, such as the training method of neural networks, can be used to train the reward model, so that the reward model finally learns how to predict the human evaluation based on the corresponding evaluation dimensions according to the features of the original test instructions. After multiple iterative trainings, it can accurately predict the satisfaction of humans with different responses.

[0078] In addition, the above RLHF can more efficiently optimize the reward model in the model fine-tuning stage based on the method of contrast learning. For example, the original test instruction generation model can generate multiple test instructions based on the input data, and then human experts compare these test instructions respectively and give preference scores or rankings based on different evaluation dimensions. Using the method of contrast learning, the reward model can learn how to score candidate outputs according to human preferences. This specification does not limit this.

[0079] Step S206, optimize the parameters of the original test instruction generation model according to the evaluation results respectively output by the multiple reward models to obtain a target test instruction generation model, and the target test instruction generation model is used to output target test instructions for the in-vehicle function.

[0080] After training multiple reward models, the parameters of the above-mentioned original test instruction generation model can be optimized by using the evaluation results respectively output by each reward model. Among them, the above-mentioned evaluation results can include both the evaluation results for the original test instructions and the evaluation results for other test instructions different from the above-mentioned original test instructions. These evaluation results can be presented in the form of scores or probabilities, etc. For example, the accuracy evaluation result is a score of 0.85 or a probability of 85%, the security evaluation result is a score of 0.7 or a probability of 70%, etc. Among them, the numerical size of the evaluation result of the reward model is positively correlated with the evaluation of the original test instruction by the reward model in the corresponding evaluation dimension.

[0081] For the different evaluation results of multiple reward models, this specification can balance the differences between multiple evaluation dimensions to optimize the parameters of the original test instruction generation model, and then generate more comprehensive and higher-quality target test instructions, improving the effect of in-vehicle infotainment system function testing.

[0082] In one embodiment, the evaluation results respectively output by each of multiple reward models can be obtained. Then, the target expectations can be input into the Minimax Function respectively, and the parameters of the original test instruction generation model can be iteratively optimized until the target expectations are less than the expectation threshold. Among them, the target expectations can be used to represent the expectations of the differences between each evaluation result and other evaluation results, and are used to represent the degree of difference between multiple evaluation results. In other words, if the scores of all evaluation results are close, the target expectations are small, indicating a high consistency between the evaluation results; if the scores of the evaluation results vary greatly, the target expectations are large, indicating a large difference between the evaluation results. The above Minimax Function can minimize the maximum difference between multiple evaluation results, thereby balancing the performance of different evaluation dimensions. The above iterative optimization process is to adjust the parameters of the original test instruction generation model through optimization methods such as gradient descent, so that the target test instructions generated by the optimized target test instruction generation model have more balanced scores in multiple evaluation dimensions. Suppose an original test instruction generation model A generates an original test instruction, and the evaluation results of the corresponding three reward models are as follows: accuracy evaluation model: 9 / 10; rationality evaluation model: 6 / 10; time efficiency evaluation model: 5 / 10. Among them, each evaluation result respectively includes the current score and the score upper limit of the above original test instruction in the corresponding evaluation dimension. For example, the evaluation result "9 / 10" of the above accuracy evaluation model can be understood as that the model gives a score of 9 to the above original test instruction, and 10 is the score upper limit. The higher the score, the more the above original test instruction meets the requirements of the corresponding evaluation dimension. At this time, assuming that the expectation threshold is 2, the target expectations of the three evaluation results are [(9 - 6) + (9 - 6) + (6 - 5)] / 3 ≈ 2.67, which is greater than the expectation threshold 2. It is judged that the target expectations are large, and there is a large difference between the evaluation results, that is, the accuracy is high, but the rationality and time efficiency are low.

[0083] Through the Minimax Function and iterative optimization, the original test instruction generation model adjusts its parameters to generate new test instructions, which can make the scores more balanced. The above Minimax Function is as follows:

[0084]

[0085] Among them, θ is the parameter of the original test instruction generation model, πθ is the original test instruction output by the original test instruction generation model, and its content is determined by θ. R1 and R2 are the scoring functions corresponding to two reward models RM1 and RM2 respectively. R1(π θ ) - R2(π θ ) represents the difference between R1(π θ ) and R2(π θ) The difference between the scoring results of these two scoring functions for the original test instructions. E[R1(π θ ) - R2(π θ )] represents the expectation of the above difference. The goal of the entire optimization function can be divided into a maximization part and a minimization part. These two parts will be introduced separately. For the maximization part, the purpose is to find the test instruction π θ that maximizes the scoring difference between the two reward models, so as to identify the worst-case scenario, that is, the scenario with the largest scoring difference between the two reward models; for the minimization part, adjust the parameter θ so that in the worst-case scenario (i.e., when the scoring difference is the largest), the difference is as small as possible. Its purpose is to improve the robustness of the original test instruction generation model and avoid performing poorly in certain evaluation dimensions.

[0086] It should be noted that the above min-max function only involves two reward models for the convenience of understanding the solution. When the number of reward models is greater than two, the above min-max function can be further extended to:

[0087]

[0088] where, R i and R j can be the scoring functions of any two reward models among the m reward models, i, j ∈ {1, 2, 3..., m}. Assuming m is 3, then in the above extended function, R1, R2, and R3 score each test instruction, and after calculating all possible scoring differences, two reward model scores with the largest differences can be selected as the parameters of the min-max function, or operations such as summing or averaging the differences between the scores of each reward model and the scores of all other reward models can also be performed to be used as the parameters of the min-max function. This will not be elaborated in detail in this specification.

[0089] In summary, by applying the above min-max function to the optimization process of the parameters of the original test instruction generation model for multiple reward models, the expected difference in the scores of multiple reward models under different test processes is minimized as much as possible. This ensures that the optimized target test instruction generation model can perform well in multiple evaluation dimensions such as accuracy and rationality, and avoids the situation of poor performance in some evaluation dimensions. At the same time, it can also ensure that the output of the target test instructions of the target test instruction generation model is within the range of the in-vehicle infotainment system's functional capabilities, effectively improving the coverage and rationality of the in-vehicle infotainment system test process. Still taking the above original test instruction generation model A as an example, the target test instructions for the in-vehicle infotainment system functions output by the target test instruction generation model optimized based on this model can have more balanced scores for different reward models. For example: accuracy: 8 / 10; rationality: 7 / 10; time efficiency: 7 / 10. At this time, the target expectation of the three evaluation results is [(8 - 7) + (8 - 7) + (7 - 6)] / 3 ≈ 1, which is less than the expectation threshold of 2. The smaller target expectation indicates a smaller difference between the evaluation results, and the test instructions perform evenly in multiple dimensions.

[0090] After obtaining the target test instruction generation model, this specification can also optimize the inference process of the target test instruction generation model to conform to the in-vehicle infotainment system functions to be tested.

[0091] In one embodiment, for a target vehicle with a vehicle head unit function to be tested, target input data corresponding to the vehicle head unit function can be obtained, and the target input data can be input into the target test instruction generation model. Meanwhile, when the target test instruction generation model generates a to-be-processed test instruction according to the target input data, the expected anomaly in the vehicle head unit function test scenario corresponding to the target test instruction can be determined through a Self-Consistency Chain-of-Thought (SeCoT); the corresponding additional test instruction can be determined according to the expected anomaly, and the additional test instruction can be combined with the to-be-processed test instruction to generate the target test instruction. Among them, the target input data can be either the original input data or other input data corresponding to the vehicle head unit function and different from the original input data. As a comprehensive technology combining Chain-of-Thought (CoT) and Self-Consistency, the SeCoT can optimize the performance of the LLM in complex reasoning tasks. In this embodiment, the target test instruction generation model can initially generate a to-be-processed test instruction. Meanwhile, the model can achieve Self-Prediction based on SeCoT, generating different reasoning paths to anticipate possible expected anomalies in the vehicle head unit function test process. The results include but are not limited to the anomaly type, error information provided by the vehicle head unit system, and interaction feedback records with the user, etc. Then, based on the expected anomaly, additional test instructions are further determined in the target test instruction generation model and combined with the to-be-processed test instruction. There is a corresponding causal logical relationship between the expected anomaly and the additional test instruction. For example, when the expected anomaly is "the user uses a voice command to modify the navigation path on the highway, and the voice recognition fails due to the noise reduction algorithm of the vehicle head unit microphone", then the target test instruction generation model can specifically generate "voice command test in a noisy environment" as an additional test instruction and combine it with the to-be-processed test instruction, thereby further covering the key test scenarios of the to-be-processed test instruction and improving the comprehensiveness of the test.

[0092] Of course, the premise of the above embodiments is that the historical test data for the in-vehicle infotainment system functions does not involve the test data of the above-mentioned additional test instructions. After obtaining the expected anomaly, the solution in this specification can also actively compare the historical test data based on the Self-Reflection of SeCoT to determine whether the expected anomaly covers the anomaly situation in the historical test data. If so, it means that the expected anomaly has been tested before, so there is no need to generate additional test instructions, and the test instructions to be processed can be directly used as the target test instructions; if not, it means that the expected anomaly has not been tested before, so additional test instructions can be generated and combined with the above-mentioned test instructions to be processed. Specifically, the target test instruction generation model can introduce a memory module to maintain the historical test data. The memory module can adopt a hash table, a database, such as a relational database or a non-relational database, a tree structure, such as a binary search tree, a prefix tree, a graph structure, such as a directed acyclic graph, a knowledge graph, or one or more of these data storage structures. Taking the structure of a hash table combined with a linked list as an example, the anomaly type can be used as the hash key to quickly locate the relevant anomaly data linked list. Among them, the linked list node can be used to store the specific anomaly description information and the corresponding historical test data. The storage location of each hash table can be connected to a linked list to avoid the hash collision that may occur in the hash table. For example, when there is an expected anomaly belonging to the anomaly type of "Bluetooth connection anomaly", a hash value can be calculated through a hash function, and then quickly locate the corresponding storage location in the hash table. Suppose in this linked list, the first node stores the anomaly description information of a Bluetooth connection anomaly that occurred when the device temperature was too high, including the location, time, environmental status, device status, interface prompts displayed on the in-vehicle infotainment system, function error logs, etc., as well as the historical test data for this anomaly, such as test steps, test tools, and the data contained in the above anomaly description information; the second node stores another Bluetooth connection anomaly caused by different reasons and related anomaly descriptions and historical test data. In this way, the expected anomaly can be compared with the historical test data contained in each node in the linked list according to the anomaly description information.

[0093] When the SeCoT self-reflection module compares the expected anomalies with the historical test data in the memory module, two methods can be adopted: exact matching and fuzzy matching. Among them, exact matching means that if the new expected anomaly is exactly the same as the existing anomaly description information in the memory module, the relevant historical test data can be directly obtained for reference and analysis. For example, if the new expected anomaly is "audio playback stuttering in low-temperature environment" and there is also the anomaly description information of "audio playback stuttering in low-temperature environment" in the memory module, the corresponding historical test data can be quickly located by obtaining the relevant linked list nodes. Fuzzy matching means using a fuzzy matching algorithm to analyze the similarity of key information in the anomaly description information. For example, the new expected anomaly is "Bluetooth connection interruption under specific network fluctuations", while there is the anomaly description information of "abnormal Bluetooth connection caused by unstable network" in the memory module. Through the fuzzy matching algorithm, key information such as "network", "Bluetooth connection", and "abnormality" is extracted, and the similarity between the two is calculated. When the similarity reaches a preset threshold, it can be determined that the two anomalies are similar. In this way, even if the anomaly expressions are different, relevant historical test data can be found for reference, which helps to more comprehensively judge whether additional test instructions need to be generated, improving the efficiency and accuracy of testing.

[0094] Those skilled in the art can understand that in addition to generating the target test instruction by combining the additional test instruction with the test instruction to be processed, the target test instruction can also be regenerated through a target test instruction generation model in this specification, including the test instruction to be processed and the target test instruction mentioned above, so as to reduce the generation process of the target test instruction and improve the generation efficiency.

[0095] In addition, the above-mentioned processes of predicting the future and self-reflection of SeCoT can be generated based on the same or different modules respectively. For example, a future prediction module and a self-reflection module, where any module can be used as a sub-module of the generation proxy module and deployed in, for example, the original test instruction generation model and the target test instruction generation model, or can be used as a physical unit or logical unit independent of the above-mentioned generation proxy module. This specification does not limit this.

[0096] In summary, by combining the prediction of the future and self-reflection of SeCoT in the scenario of generating test instructions for in-vehicle infotainment system functions, this application can optimize the original test instructions to make them more in line with the actual needs of in-vehicle infotainment system function testing, improve the test coverage and effectiveness, and minimize the risks caused by untested functions to the greatest extent.

[0097] Next, taking the in-vehicle infotainment system test as an example, combined with Figure 3 This paper introduces a method for generating test instructions for in-vehicle infotainment system functions based on RLHF, prediction of the future and self-reflection of SeCoT. The method includes the following steps:

[0098] Step S302: Generate original test instructions based on the original input data.

[0099] In one embodiment, a tester of the in-vehicle infotainment (IVI) system functions can establish a connection between a test device equipped with a target test instruction generation model, such as a tablet computer or a laptop computer, and the IVI system of the target vehicle, so as to use this test device as a generation proxy module to collect the original input data for the IVI system functions to be tested of the target vehicle. Suppose the original input data includes: the version number of the navigation application, and the sensor logs; driving data including vehicle speed and steering angle, environmental data including road conditions and weather conditions; and the following personalized profiles of users:

[0100]

[0101]

[0102] Suppose the IVI system recommends a new route and destination to the above-mentioned generation proxy module, then the original test instruction generation model corresponding to this generation proxy module can generate original test instructions based on the above-mentioned personalized profiles and the feedback of the memory module, so as to decide whether to accept or reject the recommended route according to the original test instructions, and select an appropriate navigation mode and navigation behavior during the navigation process.

[0103] Before officially generating the original test instructions, corresponding prompt words can also be configured for the original test instruction generation model to ensure that the original test instructions are more in line with the actual situation. For example, the following prompt words can be generated:

[0104] "You are an intelligent agent simulating user behavior, making decisions based on the following information: the personalized characteristics of the user: {the characteristic information of each user preference data in the personalized profile, such as activity level, common navigation routes, etc.}, the context conversation history in the memory module: {the relevant historical records stored in the memory module}. Your goal is to simulate the behavior of real users and cover various usage scenarios.

[0105] Please perform operations according to the following steps:

[0106] 1. Receive the recommended content or instructions of the current IVI system.

[0107] 2. Analyze the current situation and generate original test instructions based on the personalized characteristics of the user and the context conversation history in the memory module.

[0108] 3. Execute the original test instructions, such as selecting a certain function, giving feedback, executing a voice command, etc.

[0109] 4. Record each operation and emotional reaction to generate a behavior log.

[0110] 5. Conduct reflection and feedback at appropriate times to optimize subsequent original test instructions.

[0111] Please note that in each step of the operation, consider the user's emotions and thinking and give reasonable feedback. For example, if the user encounters difficulties during navigation, provide assistance and adjust the navigation route. If the user expresses dissatisfaction with a certain music playlist, recommend other playlists that they may be interested in.

[0112] Based on the above input data, the target test instruction generation model can generate an original test instruction or an instruction set composed of multiple consecutive original test instructions. For example, "Test the navigation path planning function: Plan the optimal path from the company, the daily work location, to the home, the residential location", and "Test the voice command response: In an environment where the urban traffic is busy and the vehicle speed is low, change the navigation destination through voice commands". However, at this time, the generated original test instructions are likely to be difficult to cover all possible special cases and complex scenarios due to the extreme complexity of the real-world scenarios.

[0113] Step S304, apply RLHF constraints to screen test instructions for accuracy and reasonableness.

[0114] In one embodiment, the generated original test instructions can be input into the generation agent module together with the in-vehicle infotainment (IVI) function API logic, such as a clearly listed support path algorithm list. Suppose the reward model corresponding to the generation agent module includes an accuracy reward model and a rationality evaluation model trained based on Reinforcement Learning from Human Feedback (RLHF). Among them, for the accuracy reward model, each original test instruction can be scored according to the evaluation dimension of accuracy. This model strictly checks whether there is a situation of calling an unsupported algorithm in the original test instruction based on the above IVI function API logic. For example, in the original test instruction related to path planning, if the Dijkstra algorithm that the IVI does not support is called, the accuracy reward model can accurately identify it. For those instructions that do not conform to the IVI API logic and have a score lower than the pre-set threshold, the parameters of the original test instruction generation model can be adjusted and optimized to eliminate them, so that the optimized test instruction generation model, that is, the target test instruction generation model, outputs target test instructions that meet the above accuracy. For the rationality reward model, each original test instruction can be scored based on RLHF and the evaluation dimension of rationality, so as to determine whether the original test instruction considers the edge issues in the corresponding test scenario. For example, when the original test instruction requires testing the IVI navigation function, whether it requires continuously inputting multiple complex navigation instructions within a short period of time, such as setting a long-distance destination in sequence within 10 seconds, and then quickly modifying the waypoints and adjusting the preferred route type, etc., so as to judge whether the navigation application of the IVI can correctly identify and process these instructions without freezing, crashing or incorrect execution; whether the user frequently switches different map display modes during navigation, such as satellite view, 3D view, ordinary map view, etc., and at the same time performs voice interaction operations, so as to judge whether the above navigation application can run stably and does not affect the normal use of the navigation function, etc. If so, the rationality evaluation model can give a higher score, otherwise a lower score; for those instructions that do not consider the edge issues and have a score lower than the pre-set threshold, the parameters of the original test instruction generation model can be adjusted and optimized to eliminate them, so that the optimized test instruction generation model, that is, the target test instruction generation model, outputs target test instructions that meet the above rationality requirements.

[0115] Those skilled in the art can understand that the evaluation results output by the reward model in this step can be used not only to optimize the parameters of the above original test instruction generation model, but also to screen out a part of the target test instructions whose evaluation results meet the pre-set threshold from all the target test instructions during the inference process of the target test instruction generation model.

[0116] Step S306, predict potential abnormal scenarios based on SeCoT.

[0117] In one embodiment, during the process that the previous-step target test instruction generation model infers the target test instruction, the test instruction generated by the target test instruction generation model can be used as the to-be-processed test instruction, and input into the future prediction module of the in-vehicle system together with the historical test data in the memory module, such as the relevant record about "voice recognition fails in low-speed environment", so as to start the SeCoT inference chain, and thereby deeply predict the possible abnormal situations during the test execution process by means of multi-path inference analysis. During the prediction process, the model can identify potential abnormal problems. For example, it is found that when the vehicle is driving at a high speed, strong wind noise may interfere with the normal recognition of voice commands, resulting in the commands not being accurately transmitted and executed. Based on this analysis, the above model can generate a detailed expected abnormal list, which can include, for example, the expected abnormality of "voice command recognition fails in high-speed scenarios", and at the same time output the test scenario description closely associated with it.

[0118] Step S308, based on the SeCoT reflection, supplement the test scenario to generate additional test instructions.

[0119] In one embodiment, the expected abnormal list generated by the above model can be input into the self-reflection module, and this module can comprehensively compare it with the historical test data in the historical test database. Suppose the historical test data is stored in a hash table structure. If there is a record of "high-speed scenario voice test" that is exactly the same as the current expected abnormality in the historical test data, it means that this scenario has been fully tested and there is no need to supplement new tests. However, if there is only a record related to "low-speed noise test" in the historical test data, through the fuzzy matching algorithm, the system will determine that the "high-speed noise test" scenario needs to be supplemented currently. Based on this judgment result, the target test instruction generation model can generate additional test instructions, such as "voice command response test in a noise environment (simulating strong wind noise when driving at a high speed)".

[0120] Step S310, generate the target test instruction based on the to-be-processed test instruction and the additional test instruction.

[0121] In one embodiment, based on the above to-be-processed test instruction and additional test instruction, a target test instruction set including both of them can be generated. Compared with the original to-be-processed test instruction, the target test instruction set not only includes the original basic path planning test and the test of modifying the destination by low-speed voice command, but also covers the voice command test in a noise environment in high-speed scenarios, making the test process more scientific and comprehensive, and capable of effectively detecting the performance of the in-vehicle navigation system in various complex scenarios.

[0122] It should be noted that, in addition to the above historical test data, the memory module in step S302 of the above embodiment may also include other data during the test process of the in-vehicle function, and this specification will further introduce the memory module.

[0123] Taking the in-vehicle infotainment system function test system shown Figure 4 as an example, the system as a whole includes a proxy module 42, an instruction execution module 44, an in-vehicle infotainment system 46, a memory module 48, and a user profile module 410. Among them, the target test instruction generation model in the above-mentioned proxy module 42 can generate target test instructions according to the personalized profile in the user profile module 410 and the content in the memory module. The target test instructions can be sent to the instruction execution module 44 and executed. During the execution process, the instruction execution module 44 can further send the target test instructions to the in-vehicle infotainment system 46. The in-vehicle infotainment system 46 can then respond to the execution of the above-mentioned target test instructions by the instruction execution module 44 and feedback the execution result. The memory module 48 can obtain a series of operation records of the above-mentioned target test instructions from the in-vehicle infotainment system 46 and store them as the content in the above-mentioned historical test data. At the same time, the memory module 48 or the in-vehicle infotainment system 46 can summarize and reflect on it and store it in the form of a report or the like, or provide it to the proxy module 42 so that the target test instruction generation model can generate new target test instructions. In addition, any one of the instruction execution module 44, the memory module 48, and the user profile module 410 can be used as a sub-module of the proxy module 42 and deployed in, for example, the original test instruction generation model and the target test instruction generation model, or can be used as a physical unit or a logical unit of the in-vehicle infotainment system independent of the above-mentioned proxy module, and this specification does not limit this.

[0124] Of course, the above-mentioned in-vehicle infotainment system 46 can also obtain a series of operation records of the above-mentioned target test instructions from the instruction execution module 44 under specific circumstances. Below, on the Figure 4 basis, combined with Figure 5 , taking the process of executing the target test instruction as an example, the above-mentioned specific circumstances will be further introduced. The steps of this process are as follows.

[0125] Step S502, the target test instruction generation model generates multiple target test instructions and sends them to the instruction execution module.

[0126] In one embodiment, the target test instruction generation model in the proxy module can generate multiple target test instructions for different aspects of the in-vehicle infotainment system functions based on various input data, including in-vehicle infotainment system data, driving data, user preference data, and environmental data, by using its optimized algorithms and parameters. These instructions cover various functional scenarios of the in-vehicle infotainment system, including path planning and real-time traffic condition update tests in the navigation function; audio playback format compatibility and video playback smoothness tests in the multimedia function; Bluetooth connection stability and in-vehicle phone call dialing and answering tests in the communication function. After generation, these target test instructions can be sent to the instruction execution module.

[0127] Step S504, the instruction execution module determines whether the target test instruction passes the accuracy verification.

[0128] In one embodiment, although the target test instruction is scored and screened through, for example, the above accuracy reward model, considering the differences in the training degree of the accuracy reward model, there is still a risk of anomalies caused by the in-vehicle system executing unsupported target test instructions. Therefore, the instruction execution module can further perform accuracy verification on the target test instruction. The instruction execution module can first check whether the target test instruction conforms to the instruction specifications and interface protocols supported by the in-vehicle system. For example, check whether the in-vehicle function APIs called in the instruction exist and whether the parameter settings are correct. For the target test instruction of the navigation function, it will be confirmed whether the format of the navigation destination set in the instruction is correct and whether it can be recognized by the in-vehicle system. At the same time, the instruction execution module will also compare with the function list of the in-vehicle system to determine whether the function involved in the instruction is within the actual function range of the in-vehicle system. If the target test instruction passes the above accuracy verification, if it passes, step S506 is executed; otherwise, step S508 is executed.

[0129] Step S506, send the target test instruction to the in-vehicle system.

[0130] In one embodiment, when the target test instruction passes the accuracy verification, the instruction execution module sends it to the in-vehicle system. After receiving the instruction, the in-vehicle system will call the corresponding function modules and hardware resources according to the requirements of the instruction and start executing the test operation. When executing the target test instruction for navigation path planning, the in-vehicle system will start the navigation software, calculate the path using the start and end point information in the instruction, map data, and path planning algorithms, and display the planning result on the screen. When executing the audio playback test instruction, the in-vehicle system will call the audio playback module, read and play the audio file in the specified format.

[0131] Step S508, input the verification information into the memory module and reflect on it in combination with historical data.

[0132] In one embodiment, if the target test instruction fails the accuracy verification, the instruction execution module may input the verification information, including the reason for the failed verification, the specific content of the target test instruction, etc., into the memory module. After receiving this information, the memory module will reflect on it in combination with the stored historical test data, that is, it will analyze the similarities and differences between the current situation of the failed verification and the similar situations in history, and determine whether it is caused by a new in-vehicle system version, hardware update or other factors. If it is found that after the in-vehicle system software is updated, the calling method of some APIs has changed, resulting in the original correct target test instruction failing the verification, the memory module will record this change and provide a reference for the generation and optimization of subsequent test instructions.

[0133] Step S510, the memory module transmits the reflection result to the generation proxy module to regenerate the target test instruction.

[0134] In one embodiment, after the memory module completes the reflection, it can transmit the reflection result to the target test instruction generation model in the generation proxy module. The target test instruction generation model adjusts its own parameters and algorithms according to these reflection results. It will refer to the information provided by the memory module, re-analyze the input data, and generate target test instructions that are more in line with the current state and accuracy requirements of the in-vehicle system. If the memory module feedbacks that the implementation method of a certain function has changed after the in-vehicle system is updated, the target test instruction generation model will adjust the test instruction generation logic for this function to ensure that the newly generated target test instruction can accurately and effectively test the in-vehicle function.

[0135] The above memory module can be associated with the above target test instruction generation model. Meanwhile, the memory module can include two types of modules, namely, a Short-Term Memory (STM) module and a Long-Term Memory (LTM) module, thus constituting a hierarchical storage structure. When the above generation agent module obtains the data to be memorized, it can store the data to be memorized as historical data in the above short-term memory module or the above long-term memory module according to the data type of the data to be memorized. Among them, the data to be memorized can include the feedback data corresponding to the above target test instruction, and the feedback data includes functional feedback data for characterizing the vehicle state during the execution of the target test instruction by the target vehicle, such as an operation log recording whether the vehicle air conditioner temperature adjustment meets the expected temperature of the temperature adjustment test instruction, whether the seat position meets the expected position of the seat adjustment test instruction, etc., and / or user emotional feedback data for characterizing the emotional state of the user during the execution of the target test instruction by the target vehicle, such as an emotional log recording the perceived navigation difficulty, satisfaction, etc. of the user for the navigation test instruction in the navigation test scenario. The data type of the above data to be memorized can be divided according to the specific content or the relationship with the in-vehicle system functions. For example, the above feedback data is a type of data that can be used to characterize the feedback of the target vehicle and the user on the execution of the target test instruction. In addition, the data type of the data to be memorized can also include the in-vehicle system data, driving data, user preference data, and environmental data in the above original input data.

[0136] It is worth mentioning that the feedback data is not the same as the above historical test data. The feedback data reflects the actual operation of the in-vehicle system in the test scenario and the relevant feedback of the user. These feedback data will be stored in the memory module and become part of the historical test data. Historical test data refers to the general term for various information related to previous in-vehicle system function tests. When new feedback data is generated, it can be compared and analyzed with the historical test data.

[0137] Those skilled in the art can understand that since the above historical data can be used to feedback the target test instruction generation model to generate target test instructions, and the above feedback data can reflect the execution effect of the target test instruction, the memory module can, when the continuous execution time of the generation agent module is greater than the preset duration or a specific instruction is output, summarize the relevant operation logs and emotional logs in the historical data, analyze the trends in these logs, identify the successfully executed and failed instructions from them, generate a reflection report for introducing and clarifying effective instructions and instructions that need to be improved, and finally this reflection report can be transmitted to the generation agent module to adjust the target test instruction generation model so that the target test instructions generated subsequently can not only meet the expected effect but also provide a higher emotional experience for the user.

[0138] For the data to be memorized with different data types, testers of in-vehicle functions can preset the default storage locations for each data type. For example, important but less real-time data such as the above-mentioned feedback data, driving data, and environmental data are all stored in the above-mentioned short-term memory module, while data with high importance or data that needs continuous analysis, such as in-vehicle system data, user preference data, reflection reports, and behavior pattern data, are all stored in the above-mentioned long-term memory module, so as to realize multi-modal dimension management of the data to be memorized through a hierarchical storage structure and improve the management efficiency of the data to be memorized. Among them, the so-called behavior pattern data includes the in-vehicle function information to be tested for characterizing the function application of the in-vehicle function during the test process and the interaction mode information for characterizing the interaction mode defined between the in-vehicle system and users, instructions, etc. when the corresponding in-vehicle function is executed. Further, different subtypes can be further divided for each data type based on preset rules to achieve the ability to store the same data type in the short-term memory module and the long-term memory module respectively. In addition, the data to be memorized can be selectively stored in the short-term memory module or the long-term memory module according to the order of creation time. For example, if there is a piece of feedback data that is continuously stored in the short-term memory module within a preset time period after creation, such as within one month, then it can be stored from the short-term memory module to the long-term memory module after one month. The storage methods of the above data to be memorized are not limited in this specification, and testers of in-vehicle functions can adopt one method or multiple methods according to actual test requirements.

[0139] In addition, there are mainly two ways to obtain the above-mentioned user emotion feedback data. One is to simulate and generate feedback data through an agent; the other is to collect the emotion feedback of real users during actual use by means of offline gray-box testing, etc. Combining multiple methods can effectively improve the authenticity and reliability of user emotion feedback data. The specific data acquisition method can be flexibly selected according to the actual situation, and this specification does not limit it.

[0140] Regarding the process of storing new data to be memorized in the above memory module, the solution of this specification can screen and update the corresponding memory module according to the changes and differences in historical data between different memory modules, so as to optimize the storage efficiency of the memory module.

[0141] In one embodiment, when the data types of the first historical data in the short-term memory module match those of the second historical data in the long-term memory module, the incremental data of the first historical data compared to the second historical data can be determined, and the incremental data can be stored in the long-term memory module. For example, when the target test instruction generation model has already output a target test instruction including the task of "navigation-play music-adjust volume", and the feedback data corresponding to the target test instruction is successively stored in the short-term memory module and the long-term memory module, if a new target test instruction adds a test step of "pause music" and stores the new target test instruction and its corresponding feedback data in the short-term memory module, then the historical data that are both feedback data in the short-term memory module and the long-term memory module can be compared, and it can be determined by comparing the test steps in the two feedback data that only "pause music" is the newly added test step. Then, when storing the new target test instruction from the short-term memory module to the long-term memory module, the feedback data corresponding to the newly added test step can be stored in the long-term memory module as incremental data, thereby avoiding redundant data storage and improving the storage efficiency of the memory module.

[0142] As the vehicle head unit function test process continues to progress, the historical data in the memory module is constantly updated. The method in this application can provide a corresponding elimination strategy for the memory module in the above hierarchical storage structure to reduce the useless data in the memory module.

[0143] In one embodiment, the memory module can obtain the historical data in the short-term memory module and the long-term memory module, and under the condition of meeting the memory elimination condition, eliminate at least a part of the historical data from the obtained historical data according to the memory elimination rule, thereby optimizing the storage efficiency of the memory module and ensuring the validity of the stored data. Among them, the memory elimination condition can include:

[0144] 1. Historical data with high test coverage and low function priority is preferentially eliminated. For the historical test data in the historical data, if the test coverage of a certain part of the data related to the vehicle head unit function is already very high, it means that this function has been fully covered in multiple tests, and at the same time the priority of this function is relatively low, such as some infrequently used vehicle head unit auxiliary functions. In this case, in order to optimize the storage of the memory module, such historical data should be preferentially eliminated. For example, a niche function for a specific vehicle model display in the vehicle head unit system has been fully covered in multiple tests, and this function has little impact on the core experience of most users using the vehicle head unit, then the historical data related to it meets this memory elimination condition.

[0145] The above function priorities can be either fixed values of preset configurations or dynamically adjusted based on test scenarios. For example, in the navigation test scenario, the priority of historical data related to the path can be increased, and in the voice assistant test, the priority of historical data related to speech recognition can be increased, so as to ensure that the test data most relevant to the current scenario is retained within the limited storage space.

[0146] In addition, the expected execution result of the above target test instruction can be obtained, and the actual execution result of the above target test instruction can be determined according to the feedback data above. At the same time, the memory elimination rule can be adjusted according to the above expected execution result and the above actual execution result. Among them, the adjustment object can be the priority of the historical test data corresponding to the above target test instruction. If there is a substantial difference between the above expected execution result and the above actual execution result, such as the navigation path planning time exceeding the limit or the voice command recognition failing, it means that problems have emerged in the relevant in-vehicle infotainment system functions during the test. At this time, the priority content of the corresponding historical test data can be increased based on the feedback data, so that the memory module pays more attention to these data, prompting the target test instructions generated subsequently to be better optimized for these problems, thereby improving the accuracy and comprehensiveness of the in-vehicle infotainment system function test and realizing the optimization of the overall performance of the memory module.

[0147] It can be understood that the adjustment of the priority of the test scenario focuses on retaining the most relevant data in a timely manner according to the current test scenario. The short-term memory module stores data with low importance and strong real-time nature, and it needs to be adjusted flexibly according to the scenario to quickly adapt to the test requirements. Therefore, the above memory elimination conditions are more applicable to the data elimination of the short-term memory module than the long-term memory module.

[0148] 2. Data that has not been accessed for more than a certain preset active duration is preferentially eliminated. The memory module can record the time when each historical data was last accessed. When the time since a certain historical data was last accessed exceeds the preset active duration, it indicates that the data has not been used for a long time, and its reference value for the current in-vehicle infotainment system function test may have decreased. For example, if the preset active duration is three months, and a certain test data about a function of an early version of the in-vehicle infotainment system has not been accessed again for more than three months since the last access, then this data meets this elimination condition and should be preferentially eliminated to release storage resources. It can be understood that because the data stored in the long-term memory module needs to have more long-term value and stability, data that has not been accessed for more than the preset active duration can be considered to have a reduced long-term reference value. The short-term memory module is more concerned with recent and current relevant data, is relatively less sensitive to the active duration, and focuses more on immediate and short-term test requirements. Therefore, the above memory elimination conditions are more applicable to the data elimination of the long-term memory module than the short-term memory module.

[0149] 3. Historical data with low correlation in repeated storage is preferentially eliminated. During the process of storing historical data in the memory module, for repeatedly stored data, it can be further analyzed and found that its correlation with the current in-vehicle system function test is relatively low. For example, some repeated data generated in the early test environment, while the version and functions of the current in-vehicle system have changed significantly, resulting in little guiding significance of these historical data for the current test. At this time, these repeated historical data with low correlation should be preferentially eliminated. It can be understood that considering that whether it is the short-term or long-term memory module, the repeated historical data with low correlation with the current in-vehicle system function test will occupy valuable storage resources in the memory module and have little guiding significance for the current test. Eliminating them can effectively optimize storage and improve test efficiency. Therefore, the above memory elimination conditions apply to both the short-term memory module and the long-term memory module for data elimination.

[0150] When providing the historical data in the memory module to the generation agent module, the input order of the historical data can also affect the target test instruction. The solution in this specification can preferentially provide the historical data with strong correlation with the above test scenario to the target test instruction generation model in combination with the test scenario of the target vehicle.

[0151] In an embodiment, the memory module can obtain the in-vehicle state of the target vehicle from the original input data, such as vehicle speed, external temperature, driving mode, etc. Then, it can determine the target test scenario in which the target vehicle is located according to the in-vehicle state of the above target vehicle, and query the historical data corresponding to the target test scenario from the memory module according to the corresponding relationship between the historical data maintained in the memory module and the test scenario, so as to preferentially input the above target test instruction generation model. Among them, the target test scenario can be characterized by a scenario label. For example, when the current vehicle speed of the target vehicle is higher than the preset speed threshold and the external temperature is lower than the preset low temperature threshold, it is determined that the target vehicle is in cold weather and is driving at high speed. Therefore, a conventional scenario label such as "cold weather" or a composite scenario label combining multiple conventional scenario labels such as "cold weather + high-speed driving" can be generated. In particular, when multiple conventional scenario labels in the composite scenario label respectively correspond to different historical data, the historical data with more coincidence corresponding times can be preferentially input into the above target test instruction generation model. To ensure that the target test instruction output by the target test instruction generation model is more in line with the current state of the target vehicle, avoid misjudgment or missed judgment caused by the mismatch between the test instruction and the vehicle state, improve the reliability and effectiveness of the test results; at the same time, reduce the number of times of readjustment and testing required due to the inapplicability of the test instruction, save test time and cost, and improve test efficiency.

[0152] The following combines Figure 6 , to introduce the data storage process based on the short-term memory module and the long-term memory module. This process includes the following steps:

[0153] Step S602, the generation agent module generates target test instructions.

[0154] In one embodiment, the generation agent module first conducts an in-depth analysis of the test target, such as clarifying whether it is a navigation function test or a voice assistant function test, etc. Assuming that the test target this time is a navigation function test, the target test instruction generation model in the generation agent module can generate target test instructions for in-vehicle functions based on various input data, including in-vehicle system data, driving data, user preference data, and environmental data, such as "Plan a navigation route from the current location to the specified shopping mall" and "Simulate a signal interruption scenario during navigation and observe the navigation recovery situation", etc.

[0155] Step S604, the generation agent module reads the memory module.

[0156] In one embodiment, the generation agent module can obtain real-time multimodal data from the short-term memory module. These data cover in-vehicle system data (such as the real-time operating status of the navigation software and map data update situation), driving data (such as the current vehicle speed and driving direction), environmental data (such as the real-time weather condition and road condition information), and user interaction data (such as user input instructions and operation behaviors, etc.). Based on these real-time data, task assistance information is provided for the current navigation test scenario. For example, according to the real-time road condition information, it is prompted whether it is necessary to adjust the navigation route to avoid congested sections.

[0157] Step S606, obtain relevant historical data from the long-term memory module based on the current test scenario.

[0158] In one embodiment, the generation agent module marks the current test scenario to generate corresponding scenario tags, such as "high-speed driving + sunny weather". Then, historical data related to this scenario tag is extracted from the long-term memory module. For example, in the previous scenarios of high-speed driving and sunny weather, the optimized navigation route strategy of the in-vehicle system or the energy-saving strategy of the air-conditioning system, etc. These historical patterns are used as references to provide more comprehensive information support for the current test.

[0159] Step S608, perform dynamic incremental updates on the short-term memory module and the long-term memory module.

[0160] In one embodiment, the generation agent module can screen the newly added target test instructions to remove redundant and invalid data. Store the memory data corresponding to the screened target test instructions in the short-term memory module. For example, during the navigation test, newly generated real-time navigation route adjustment information and user feedback information on navigation instructions. At the same time, determine the incremental data and store it in the long-term memory module by comparing the short-term memory module and the long-term memory module.

[0161] Step S610: Adjust the memory elimination rule according to the actual execution result and the expected execution result of the target test instruction.

[0162] In one embodiment, the generation proxy module obtains the expected execution result of the target test instruction, which is the ideal result set according to the function requirements of the in-vehicle system and the test target when generating the test instruction. At the same time, determine the actual execution result of the target test instruction based on channels such as the feedback of the in-vehicle system and user feedback. Compare the two. If there is a difference, for example, in the in-vehicle call function test, the expected execution result is a clear and noise-free call quality, while there is obvious noise in the actual execution result, indicating that there is a problem with the in-vehicle call function. At this time, the generation proxy module adjusts the memory elimination rule according to this difference. The specific adjustment object can be the priority of the historical test data corresponding to the target test instruction. For the historical test data related to the above call quality problem, increase its priority so as to pay more attention to these data subsequently, and prompt the generation proxy module to better optimize for such problems when generating the target test instruction in the future.

[0163] Step S612: Update the short-term memory module and the long-term memory module.

[0164] In one embodiment, the generation proxy module updates the short-term memory module and the long-term memory module according to the adjusted memory elimination rule and other relevant information in the feedback loop. For example, when comparing the actual execution result with the expected execution result in step S610, the analysis data and optimization suggestions for the reason of the deviation are used. In the short-term memory module, delete the data with low relevance and low priority to the current test scenario, and release the storage resources. In the long-term memory module, eliminate the irrelevant or low-priority data, such as some outdated multimedia playback strategies or call optimization strategies that are no longer applicable to the current in-vehicle system version, and at the same time store the optimized algorithms, strategies, etc. in the long-term memory module.

[0165] Step S614: Form an optimized memory module.

[0166] In one embodiment, after the above update operation of the generation proxy module, the short-term memory module and the long-term memory module form a more efficient and accurate memory structure. This optimized memory structure better supports the test tasks performed by the subsequent generation proxy module.

[0167] Figure 7 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 7, at the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710. Of course, it may also include other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 702 reads the corresponding computer program from the non-volatile memory 710 into the memory 708 and then runs it. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, but can also be hardware or logical devices.

[0168] Corresponding to the embodiments of the method for generating the in-vehicle function test instructions described above, this specification also provides an embodiment of a device for generating in-vehicle function test instructions. Please refer to Figure 8 ; The device includes:

[0169] An original test instruction acquisition unit 802, configured to obtain, for a target vehicle with an in-vehicle function to be tested, the original input data corresponding to the in-vehicle function; and input the original input data into an original test instruction generation model, and obtain the original test instruction output by the original test instruction generation model.

[0170] An original test instruction evaluation unit 804, configured to train a plurality of reward models respectively based on reinforcement learning human feedback (RLHF) and the original test instructions, and each reward model evaluates the original test instructions respectively based on different evaluation dimensions of the in-vehicle function.

[0171] A target test instruction generation model acquisition unit 806, configured to optimize the parameters of the original test instruction generation model according to the evaluation results respectively output by the plurality of reward models, so as to obtain a target test instruction generation model, and the target test instruction generation model is used to output target test instructions for the in-vehicle function.

[0172] Optionally, the original test instruction evaluation unit 804 is specifically configured to:

[0173] For each of the plurality of reward models, obtain the feedback results of humans on the original test instructions under the respective evaluation dimensions corresponding to the reward model, and train the reward model according to the feedback results.

[0174] Optionally, the original test instruction evaluation unit 804 is specifically configured to:

[0175] Extract target features from the original input data and the original test instructions;

[0176] Use the target feature and the feedback result as training data to train the reward model.

[0177] Optionally, the target test instruction generation model acquisition unit 806 is specifically configured to:

[0178] Obtain the evaluation results respectively output by each reward model in the multiple reward models;

[0179] Input the target expectation into the minimax function respectively, and iteratively optimize the parameters of the original test instruction generation model until the target expectation is less than the expectation threshold, where the target expectation is used to characterize the expectation of the difference between each evaluation result and other evaluation results.

[0180] Optionally, the device further includes:

[0181] A test anomaly prediction unit, configured to obtain target input data corresponding to a target vehicle for a vehicle-mounted function to be tested; input the target input data into the target test instruction generation model;

[0182] When the target test instruction generation model generates a to-be-processed test instruction according to the target input data, determine the expected anomaly in the vehicle-mounted function test scenario corresponding to the target test instruction through the Self-Consistency based Chain of Thought (SeCoT);

[0183] Determine corresponding additional test instructions according to the expected anomaly, and combine the additional test instructions with the to-be-processed test instruction to generate the target test instruction.

[0184] Optionally, the target test instruction generation model is associated with a memory module, and the memory module includes a short-term memory module and a long-term memory module; the device further includes:

[0185] A data memory unit, configured to obtain data to be memorized, where the data to be memorized includes feedback data corresponding to the target test instruction, and the feedback data includes function feedback data for characterizing the vehicle state during the execution of the target test instruction by the target vehicle and / or user emotion feedback data for characterizing the user's emotion state during the execution of the target test instruction by the target vehicle;

[0186] Store the data to be memorized as historical data in the short-term memory module or the long-term memory module according to the data type of the data to be memorized, where the historical data is used to feedback the target test instruction generation model to generate the target test instruction.

[0187] Optionally, the device further includes:

[0188] A historical data elimination unit, configured to obtain historical data in the short-term memory module and the long-term memory module;

[0189] When the memory elimination condition is satisfied, at least a part of the obtained historical data is eliminated according to the memory elimination rule.

[0190] Optionally, the device further includes:

[0191] A memory elimination rule adjustment unit, configured to obtain the expected execution result of the target test instruction, and determine the actual execution result of the target test instruction according to the feedback data;

[0192] Adjust the memory elimination rule according to the expected execution result and the actual execution result.

[0193] Optionally, the device further includes:

[0194] An incremental data storage unit, configured to determine incremental data of the first historical data compared with the second historical data when the data types of the first historical data in the short-term memory module and the second historical data in the long-term memory module match each other, and store the incremental data into the long-term memory module.

[0195] Optionally, the corresponding relationship between the historical data and the test scenario is maintained in the memory module; the device further includes:

[0196] A historical data query unit, configured to determine the target test scenario where the target vehicle is located according to the in-vehicle state of the target vehicle;

[0197] Query historical data corresponding to the target test scenario from the memory module according to the corresponding relationship, so as to preferentially input the target test instruction generation model.

[0198] Based on the same concept as the above method, this specification further provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor runs the executable instructions to implement the steps of the method as described in any one of the above embodiments.

[0199] Based on the same concept as the above method, this specification further provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.

[0200] Based on the same concept as the above method, this specification further provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.

[0201] While the present invention includes many specific implementation details, these should not be construed as limiting the scope of any invention or the scope of what is claimed, but rather are mainly used to describe the features of specific embodiments of a particular invention. Certain features described in multiple embodiments of the present invention can also be implemented in combination in a single embodiment. On the other hand, the various features described in a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may operate in certain combinations as described above and were even initially claimed as such, one or more features from the claimed combination can in some cases be removed from that combination, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.

[0202] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or sequentially, or that all illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0203] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the drawings are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0204] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for generating a vehicle function test instruction, characterized in that: The method comprises: For a target vehicle with a vehicle function to be tested, original input data corresponding to the vehicle function is obtained; and the original input data is input into an original test instruction generation model, and an original test instruction output by the original test instruction generation model is obtained; Based on the reinforcement learning human feedback RLHF and the original test instruction, a plurality of reward models are respectively trained, and each reward model evaluates the original test instruction based on different evaluation dimensions of the vehicle computer function; Optimize the parameters of the original test instruction generation model according to the evaluation results respectively output by the multiple reward models to obtain a target test instruction generation model, wherein the target test instruction generation model is used to output a target test instruction for the vehicle computer function.

2. The method according to claim 1, characterized in that The reinforcement learning based human feedback RLHF technology and the original test instructions respectively train multiple reward models, including: For each reward model among the multiple reward models, obtain the feedback results of humans on the original test instruction under each evaluation dimension corresponding to the reward model, and train the reward model according to the feedback results.

3. The method according to claim 2, characterized in that The step of training the reward model according to the feedback result includes: Extracting target features from the original input data and the original test instructions; The target feature and the feedback result are used as training data to train the reward model.

4. The method according to claim 1, characterized in that The optimizing the parameters of the original test instruction generation model according to the evaluation results respectively output by the multiple reward models comprises: Obtaining evaluation results outputted by each of the multiple reward models; The target expectations are input into the minimax function respectively, and the parameters of the original test instruction generation model are iteratively optimized until the target expectation is less than the expected threshold, and the target expectation is used to characterize the expectation of the difference between each evaluation result and other evaluation results.

5. The method according to claim 1, characterized in that The method further comprises: For a target vehicle of a vehicle-machine function to be tested, obtaining target input data corresponding to the vehicle-machine function; inputting the target input data into the target test instruction generation model; In the case where the target test instruction generation model generates a test instruction to be processed according to the target input data, an expected anomaly in the vehicle function test scenario corresponding to the target test instruction is determined through a self-consistency-based thought chain SeCoT; A corresponding additional test instruction is determined according to the expected exception, and the additional test instruction is combined with the to-be-processed test instruction to generate the target test instruction.

6. The method according to claim 1, characterized in that The target test instruction generation model is associated with a memory module, and the memory module includes a short-term memory module and a long-term memory module; the method further includes: Acquire data to be memorized, wherein the data to be memorized includes feedback data corresponding to the target test instruction, wherein the feedback data includes functional feedback data for characterizing a vehicle state during the target vehicle executing the target test instruction and / or user emotional feedback data for characterizing an emotional state of a user during the target vehicle executing the target test instruction; The data to be memorized are stored as historical data in the short-term memory module or the long-term memory module according to the data type of the data to be memorized, and the historical data are used to feed back the target test instruction generation model to generate the target test instruction.

7. The method according to claim 6, characterized in that The method further comprises: Acquiring historical data in the short-term memory module and the long-term memory module; When the memory elimination condition is met, at least a portion of the historical data is eliminated from the acquired historical data according to the memory elimination rule.

8. The method according to claim 7, characterized in that The method further comprises: Acquire the expected execution result of the target test instruction, and determine the actual execution result of the target test instruction according to the feedback data; The memory elimination rule is adjusted according to the expected execution result and the actual execution result.

9. The method according to claim 6, characterized in that The method further comprises: When the data types of the first historical data in the short-term memory module and the second historical data in the long-term memory module match each other, the incremental data of the first historical data compared to the second historical data is determined and the incremental data is stored in the long-term memory module.

10. The method according to claim 6, characterized in that The memory module maintains a correspondence between the historical data and the test scenario; the method further includes: Determining a target test scenario in which the target vehicle is located according to the vehicle computer status of the target vehicle; According to the corresponding relationship, historical data corresponding to the target test scenario is queried from the memory module to give priority to inputting the target test instruction generation model.

11. A device for generating a vehicle function test instruction, characterized in that: The device comprises: An original test instruction acquisition unit is used to acquire original input data corresponding to the vehicle function to be tested for a target vehicle; and input the original input data into an original test instruction generation model, and acquire the original test instruction output by the original test instruction generation model; An original test instruction evaluation unit, used for training a plurality of reward models based on reinforcement learning human feedback RLHF and the original test instruction, each reward model evaluates the original test instruction based on different evaluation dimensions of the vehicle computer function; A target test instruction generation model acquisition unit is used to optimize the parameters of the original test instruction generation model according to the evaluation results output by the multiple reward models respectively, so as to obtain a target test instruction generation model, wherein the target test instruction generation model is used to output a target test instruction for the vehicle computer function.

12. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

13. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Remote control management system and method applied to cleaning equipment

    CN120928758A