An intelligent cockpit multi-modal test method, system, device and medium

By generating and analyzing multimodal dynamic noise using generative adversarial networks, the problem of single modality in existing intelligent cockpit testing methods is solved. This enables comprehensive evaluation of the cockpit system and robust testing of multimodal interactions, improving testing accuracy and efficiency.

CN120669675BActive Publication Date: 2026-04-07CHONGQING VEHICLE TEST & RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing smart cockpit testing methods mainly focus on functional testing and performance evaluation of a single modality, making it difficult to comprehensively evaluate the overall performance of the cockpit, especially in terms of effectively detecting the mutual influence and potential conflicts between different modalities.

Method used

Generative adversarial networks are used to generate multimodal dynamic noise to interfere with cockpit command signals. The noise is then identified and analyzed by a motion detection system to generate evaluation results. A multi-objective optimization algorithm is used to calculate test priorities, and a multimodal testing system is constructed.

Benefits of technology

It enables comprehensive simulation of the cockpit system, improves testing accuracy and efficiency, effectively evaluates the overall performance and robustness of multimodal interaction systems, and identifies potential system problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure REF-OBJ-1772610601840-000001
    Figure REF-OBJ-1772610601840-000001
  • Figure REF-OBJ-1772610601840-000002
    Figure REF-OBJ-1772610601840-000002
  • Figure REF-OBJ-1772610601840-000003
    Figure REF-OBJ-1772610601840-000003
Patent Text Reader

Abstract

The application provides a kind of intelligent cockpit multi-modal test method, system, equipment and medium, comprising: obtaining the original state information of cockpit and cockpit instruction signal, then based on original state information, using generative adversarial network to generate multi-modal dynamic noise;Multi-modal dynamic noise is used to interfere with cockpit instruction signal, and multi-modal interference signal is generated;Action detection system is constructed, and the multi-modal interference signal is identified and analyzed using the action detection system, to generate evaluation results.The application solves the problem that the test evaluation of the test method in the prior art has a single mode and it is difficult to evaluate the overall performance of the cockpit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, system, device and medium for multimodal testing of intelligent cockpits. Background Technology

[0002] With the rapid development of smart cockpit technology, multimodal interaction systems have become an indispensable part of modern automobiles. These systems integrate various interaction methods such as voice control, touch operation, and gesture recognition, aiming to provide drivers and passengers with a more convenient, safe, and personalized user experience. However, as the complexity of these systems increases, ensuring their stability and reliability in various complex environments becomes increasingly challenging.

[0003] Existing smart cockpit testing methods primarily focus on functional testing and performance evaluation of a single modality. For example, testing of voice recognition systems is usually conducted in ideal laboratory environments, making it difficult to fully simulate the various noise interferences that may be encountered during actual driving. Similarly, touchscreen testing often ignores the impact of external factors such as vehicle vibration and changes in lighting. This fragmented testing approach makes it difficult to assess the overall performance of the cockpit, especially failing to effectively detect the interactions and potential conflicts between different modalities. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a multimodal testing method, system, equipment, and medium for intelligent cockpits, which solves the problem that existing testing methods suffer from single-modality testing and evaluation, making it difficult to assess the overall performance of the cockpit.

[0005] According to an embodiment of the present invention, a multimodal testing method for a smart cockpit includes:

[0006] The system acquires the raw state information and cockpit command signals of the cockpit, and then uses a generative adversarial network to generate multimodal dynamic noise based on the raw state information.

[0007] Multimodal dynamic noise is used to interfere with cockpit command signals, generating multimodal interference signals;

[0008] A motion detection system is constructed and used to identify and analyze multimodal interference signals, generating evaluation results.

[0009] Preferably, the original state information includes an initial state diagram of the cockpit and an initial environmental background diagram;

[0010] The cockpit command signals include voice signals, touch signals, and gesture signals.

[0011] Preferably, after obtaining the original state information, the original state information is subjected to noise filtering and grayscale processing with a weight ratio of R:G:B=0.299:0.587:0.114.

[0012] Preferably, the interference of multimodal dynamic noise with cockpit command signals includes sudden noise interference, illumination change interference, and image distortion interference;

[0013] Before identifying and analyzing multimodal interference signals, it is necessary to convert the multimodal interference signals into a unified format input signal diagram.

[0014] Preferably, before interfering with the cockpit command signal using multimodal dynamic noise, the interference intensity on the cockpit command signal is calculated based on the evaluation results of the previous identification analysis:

[0015]

[0016] Where R represents the interference intensity, Accuracy represents the previous evaluation result, NormDifficulty represents the difficulty of the previous test, and Diversity represents the diversity of the current test scenarios. , It is the weighting coefficient.

[0017] Preferably, before using the motion detection system to identify and analyze multimodal interference signals, a test strategy needs to be formulated. The method for formulating the test strategy is as follows:

[0018] Calculate adaptive test metrics for each cockpit command signal based on the evaluation results of the previous test.

[0019] The test importance of each cockpit command signal is calculated based on the preset test conditions according to the adaptive test index value corresponding to each cockpit command signal.

[0020] The mutual information algorithm is used to calculate the correlation between different modes in multimodal dynamic noise;

[0021] Based on the importance and relevance of the tests, a multi-objective optimization algorithm is used to calculate the test priority of each cockpit command signal.

[0022] Preferably, the formula for calculating the correlation between different modes is as follows:

[0023]

[0024] in, It is the joint probability distribution of mode A and mode B. and These are the marginal probability distributions for mode A and mode B, respectively.

[0025] On the other hand, according to embodiments of the present invention, a smart cockpit multimodal testing system is also provided, which uses the above-described smart cockpit multimodal testing method, including:

[0026] The acquisition module is used for raw status information and cockpit command signals;

[0027] A noise simulation module is used to generate multimodal dynamic noise based on the original state information using a generative adversarial network.

[0028] The interference module is used to interfere with cockpit command signals using multimodal dynamic noise to generate multimodal interference signals.

[0029] An evaluation module is used to identify and analyze multimodal interference signals and generate evaluation results.

[0030] On the other hand, according to an embodiment of the present invention, a computer is also provided, including at least one processor and a memory, the memory storing a computer program configured to be executed by the processor to implement the above-described intelligent cockpit multimodal testing method.

[0031] On the other hand, according to embodiments of the present invention, a storage medium is also provided, which is a computer-readable storage medium storing a computer program that can be executed by one or more processors to implement the above-described intelligent cockpit multimodal testing method.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] This invention utilizes generative adversarial networks (GANs) to generate multimodal interactive dynamic noise based on the original state information, interfering with the acquired cockpit command signals. The interfered cockpit command signals are then analyzed and evaluated in depth. Through multimodal interference, this invention comprehensively simulates various noise interferences that may be encountered during actual driving, improving the accuracy of cockpit testing. Attached Figure Description

[0034] Figure 1 This is a flowchart of the multimodal testing process according to an embodiment of the present invention. Detailed Implementation

[0035] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0036] like Figure 1 As shown in the figure, this invention proposes a multimodal testing method for intelligent cockpits, including:

[0037] The system acquires the raw state information and cockpit command signals of the cockpit, and then uses a generative adversarial network to generate multimodal dynamic noise based on the raw state information.

[0038] Since different environments have different interference noises, manually setting interference noise directly often leads to bias. Therefore, this invention generates noise directly based on the original state information of the cockpit. This original state information includes the initial state diagram of the smart cockpit and the initial environmental background. In addition, cockpit command information, including voice signals, touch signals and gesture signals, will also be collected.

[0039] This invention preferably employs a high dynamic range (HDR) camera to capture the initial state image and initial environmental background, including the multimodal input devices. The use of an HDR camera ensures high-quality images are obtained under varying lighting conditions, which is crucial for subsequent environmental simulation. Furthermore, this invention incorporates voice recognition sensors, touch devices, and gesture recognition sensors within the intelligent cockpit to collect voice signals, touch signals, and gesture signals, respectively. This multimodal signal acquisition method enables this invention to comprehensively evaluate the interactive performance of the intelligent cockpit system.

[0040] This invention first trains a generative adversarial network (GAN) based on the acquired initial state diagram and initial environmental background. Generative adversarial networks are powerful generative models that can simulate complex environmental changes, greatly improving the realism and diversity of test scenarios.

[0041] Before training the GAN, the initial state image and initial environment background are processed, including noise filtering and grayscale processing. This preprocessing step can improve training efficiency and model performance. Noise filtering uses methods such as median filtering or Gaussian filtering, and the specific parameters can be adjusted according to the image features. Grayscale processing can use the weighted average method, and the weight ratio is usually R:G:B=0.299:0.587:0.114.

[0042] After training, GANs are used to generate multimodal interactive dynamic noise to simulate environmental patterns. This dynamic noise includes not only traditional sound noise, but also multimodal noise such as visual interference and tactile feedback. For example, GANs can simulate sudden noises in a smart cockpit, such as a sudden increase in the volume of the car audio system or passengers' conversations; they can also simulate changes in lighting, such as sudden changes in light at tunnel entrances and exits or dappled shadows from trees.

[0043] For noise simulation, this invention employs a GAN model based on acoustic features. This model not only considers the spectral characteristics of noise but also simulates its temporal variations. For example, it can simulate instantaneous noise during vehicle startup, wind noise during driving, and even vibration noise under different road conditions. To ensure the realism of the simulation, this invention uses a large amount of audio data collected in real-world scenarios during GAN training, including noise samples from different vehicle models and under different driving conditions.

[0044] In simulating illumination changes, the method of this invention employs a GAN model based on image-to-image transformation. This model can generate cockpit interior images under different illumination conditions, including but not limited to scenarios such as direct sunlight, light variations in tunnels, and nighttime driving. To improve the accuracy of the simulation, this method introduces an illumination estimation module into the GAN's discriminator to evaluate the illumination realism of the generated images. The illumination estimation module uses a deep learning-based illumination estimation algorithm, whose loss function... It can be represented as:

[0045] ,

[0046] in, and These represent the estimated lighting map and the actual lighting map, respectively. Represents the gradient operator. This is the balance coefficient, which typically takes a value between 0.1 and 1.

[0047] In addition, slight deformations are applied to the images generated by GAN to create minor disturbances. This deformation operation simulates the slight jitter or deformation that may occur during actual use. The deformation can be achieved through elastic transformation, and its mathematical expression is as follows:

[0048] ,

[0049] in, and These represent the original image and the deformed image, respectively. and It is a deformable field, which can be generated using a Gaussian-filtered random field. The intensity of the deformation is usually controlled within the range of 1-5 pixels to ensure that the interference is subtle enough to still affect the system.

[0050] Multimodal dynamic noise is used to interfere with cockpit command signals, generating multimodal interference signals;

[0051] After generating multimodal dynamic noise, multimodal interference signals are used to interfere with the speech signal, touch signal, and gesture signal respectively, generating corresponding multimodal interference signals (a total of 3). Then, the 3 different multimodal interference signals are converted into a unified format input signal graph. This unification process is beneficial for subsequent comprehensive analysis. Then, the input signal graph is screened and calibrated to improve signal quality.

[0052] Construct a motion detection system, use the motion detection system to identify and analyze multimodal interference signals, and generate evaluation results;

[0053] The action detection system employs a deep learning model based on an attention mechanism. The core of this model is a multi-head self-attention layer, the computation of which can be represented as follows:

[0054] ,

[0055] in, These represent the query, key, and value matrices, respectively. This refers to the dimension of the key vector. The multi-head attention mechanism allows the model to focus on different parts of the input simultaneously, thereby improving robustness to complex perturbations. It also introduces a reinforcement learning-driven test path adaptation algorithm to adjust the perturbation intensity. This adaptive mechanism enables the testing process to dynamically adjust the difficulty based on the system's actual performance, thus providing a more comprehensive evaluation of the system's robustness.

[0056] Within the reinforcement learning framework, this invention defines a state space. This includes the current type and intensity of interference, as well as the system's response; action space. This includes increasing or decreasing the intensity of various disturbances; reward function. It can be defined based on the system's performance, for example, it can take the following form:

[0057] ,

[0058] Where R represents the interference intensity, Accuracy represents the previous evaluation result, NormDifficulty represents the difficulty of the previous test, and Diversity represents the diversity of the current test scenarios. , It is a weighting coefficient used to balance the importance of different factors. Based on the evaluation results of the previous test, it adjusts the interference intensity of the cockpit command signal in this test to avoid distortion of the multimodal interference signal due to excessively high or low intensity.

[0059] Before using a motion detection system to identify and analyze multimodal interference signals, a testing strategy needs to be developed. The method for developing a testing strategy is as follows:

[0060] (1) Calculate adaptive test metrics based on interaction failures. Here, interaction failure refers to the output result of the action detection system in the previous test not matching the expectation, that is, the identified result is not the original signal. This may be due to misidentification caused by interference or response delay, etc. The calculation of adaptive test metrics adopts a comprehensive scoring mechanism, which can be expressed as:

[0061] ,

[0062] in, Indicates the interaction failure rate. Indicates the length of the test path. This indicates the diversity of test scenarios. , , It is the weighting coefficient.

[0063] (2) Set interaction priority. This priority determines the order of different interaction modes and test scenarios. Calculate the test importance S of each test according to the following formula to sort each test during the testing process:

[0064] ,

[0065] in, It is the adaptive testing metric mentioned earlier. Indicates test coverage. This is the time required for testing. It is a balance coefficient. In this way, reinforcement learning models can find the optimal balance between testing efficiency and comprehensiveness.

[0066] (3) The influence of noise on the correlation between different modes was analyzed.

[0067] This method uses mutual information to quantify the correlation between different modes. The mutual information between mode A and mode B can be expressed as:

[0068] ,

[0069] Where I(A;B) represents the correlation between mode A and mode B. It is the joint probability distribution of mode A and mode B. and These are the marginal probability distributions of mode A and mode B, respectively. By comparing the mutual information under noisy and noise-free conditions, the impact of noise on modal correlation can be quantified.

[0070] (4) Generate a test priority list that considers not only the performance of a single modality but also the synergistic effect of multiple modalities. Preferably, the list is generated using a multi-objective optimization algorithm, and the objective function can be expressed as:

[0071] ,

[0072] in, These represent different optimization objectives, such as accuracy, robustness, and response time. By utilizing the concept of Pareto optimality, this method can find the best balance among multiple objectives, assign priorities to each test, and generate a test priority list accordingly.

[0073] Finally, based on multiple metrics such as interaction failure rate and test path length, the evaluation results of the multimodal interference signal are output. This comprehensive evaluation adopts a weighted scoring mechanism, which can be expressed as:

[0074] ,

[0075] in, For comprehensive scoring, Indicates the first One evaluation indicator, These are the corresponding weights. These metrics may include, but are not limited to: average recognition accuracy, minimum recognition accuracy, average response time, and anti-interference capability. The weights can be adjusted according to the specific application scenario requirements.

[0076] In addition, extreme interference signals can be generated based on the evaluation results. Then, the motion detection system can be used to identify and analyze the extreme interference signals and generate analysis results under extreme conditions. This allows for a comprehensive evaluation of the cockpit's anti-interference capabilities and failure modes.

[0077] On the other hand, embodiments of the present invention also provide a smart cockpit multimodal testing system, which uses the above-described smart cockpit multimodal testing method, including:

[0078] The acquisition module is used for raw status information and cockpit command signals;

[0079] A noise simulation module is used to generate multimodal dynamic noise based on the original state information using a generative adversarial network.

[0080] The interference module is used to interfere with cockpit command signals using multimodal dynamic noise to generate multimodal interference signals.

[0081] An evaluation module is used to identify and analyze multimodal interference signals and generate evaluation results.

[0082] On the other hand, embodiments of the present invention also provide a computer, including at least one processor and a memory, the memory storing a computer program configured to be executed by the processor to implement the above-described intelligent cockpit multimodal testing method.

[0083] On the other hand, embodiments of the present invention also provide a storage medium, which is a computer-readable storage medium, and stores a computer program on the storage medium. The computer program can be executed by one or more processors to implement the above-described intelligent cockpit multimodal testing method.

[0084] To verify the superiority of the testing method and system of this invention, a set of simulation experiments were designed. The experimental environment simulated an intelligent cockpit system, including three interaction modes: voice control, touch operation, and gesture recognition. The simulation conditions are as follows:

[0085] 1. Hardware environment: A virtual intelligent cockpit is built, equipped with a high-fidelity audio system, a 10.1-inch touch screen and a depth camera.

[0086] 2. Software Platform: The CARLA car simulator is used to simulate the driving environment, and OpenAI Gym is used to build the reinforcement learning environment.

[0087] 3. Interference factors: The simulation includes road noise (60-80dB), light variation (100-10000 lux) and vehicle vibration (0.5-2g).

[0088] The performance of the method of the present invention (example) was compared with that of two existing methods (Comparative Example 1 and Comparative Example 2):

[0089] Example: The multimodal interaction robustness adaptive testing method of the present invention is adopted.

[0090] Comparative Example 1: Traditional static test case method.

[0091] Comparative Example 2: Single-modal adaptive testing method.

[0092] The test indicators and their detection methods are as follows:

[0093] 1. Test coverage: Use code coverage tools to calculate the percentage of system functionality being tested.

[0094] 2. Fault detection rate: The percentage of known faults that are successfully detected is calculated by artificially implanting them.

[0095] 3. Testing efficiency: Record the time required to complete a full test.

[0096] 4. Multimodal collaborative performance: Evaluate the accuracy when multiple interaction methods are used simultaneously.

[0097] 5. Extreme case handling capability: Test the accuracy of the system response under the most severe conditions.

[0098] The test results are shown in the table below:

[0099]

[0100] Analysis and Discussion:

[0101] 1. Test Coverage: The method of this invention achieves a high coverage rate of 95%, far exceeding that of traditional methods (78%) and single-modal methods (85%). This is due to the fact that this method can dynamically generate diverse test scenarios, more comprehensively covering various possible usage situations.

[0102] 2. Fault Detection Rate: The 92% fault detection rate of this method is significantly higher than the other two methods. This indicates that the adaptive testing strategy of this invention can more effectively locate potential problems in the system, especially hidden faults in complex interactive scenarios.

[0103] 3. Testing Efficiency: This method can complete comprehensive testing in just 24 hours, which is 3 times more efficient than traditional methods and 2 times more efficient than single-modal methods. This demonstrates the advantage of reinforcement learning-driven adaptive test path generation in this invention, enabling rapid location of key test points.

[0104] 4. Multimodal collaborative performance: Our method achieves 88% significantly higher performance than other methods in this key metric. This highlights the unique advantages of this invention in evaluating and optimizing the overall performance of multimodal interaction systems.

[0105] 5. Extreme Case Handling Capability: This method retains 82% of its processing capability even under extreme conditions, far exceeding other methods. This proves that the strategy of generating the "worst-case" test samples in this invention can indeed effectively evaluate the system's extreme performance.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multimodal testing method for an intelligent cockpit, characterized in that: include: The system acquires the original state information and cockpit command signals of the cockpit, wherein the original state information includes the initial state map of the cockpit and the initial environmental background map, and the cockpit command signals include voice signals, touch signals and gesture signals; then, based on the original state information, a generative adversarial network is used to generate multimodal dynamic noise. Before using multimodal dynamic noise to interfere with cockpit command signals, the interference intensity must be calculated based on the evaluation results of the previous identification analysis. in, For interference intensity, This indicates the results of the previous assessment. This indicates the difficulty level of the previous test. This indicates the diversity of scenarios being tested. , , These are weighting coefficients; Multimodal dynamic noise is used to interfere with cockpit command signals, generating multimodal interference signals. The interference with cockpit command signals by multimodal dynamic noise includes sudden noise interference, illumination change interference, and image distortion interference. Before identifying and analyzing the multimodal interference signals, the multimodal interference signals need to be converted into input signal images of a unified format. Before using a motion detection system to identify and analyze multimodal interference signals, a testing strategy needs to be developed. The method for developing the testing strategy is as follows: Calculate adaptive test metrics for each cockpit command signal based on the evaluation results of the previous test. The test importance of each cockpit command signal is calculated based on the preset test conditions according to the adaptive test index value corresponding to each cockpit command signal. The mutual information algorithm is used to calculate the correlation between different modes in multimodal dynamic noise. The formula for calculating the correlation between different modes is as follows: in, For modality and modality Relevance It is modal and modality The joint probability distribution, and They are modal and modality The marginal probability distribution; Based on the importance and relevance of the tests, a multi-objective optimization algorithm is used to calculate the test priority of each cockpit command signal. A motion detection system is used to identify and analyze multimodal interference signals and generate evaluation results.

2. The intelligent cockpit multimodal testing method as described in claim 1, characterized in that: After obtaining the original state information, noise filtering and grayscale processing with a weight ratio of R:G:B=0.299:0.587:0.114 are performed on the original state information.

3. A multimodal testing system for an intelligent cockpit, characterized in that: The system uses a multimodal testing method for a smart cockpit as described in any one of claims 1-2, comprising: The acquisition module is used to acquire raw status information and cockpit command signals; A noise simulation module is used to generate multimodal dynamic noise based on the original state information using a generative adversarial network. The interference module is used to interfere with cockpit command signals using multimodal dynamic noise to generate multimodal interference signals. An evaluation module is used to identify and analyze multimodal interference signals and generate evaluation results.

4. A computer, characterized in that: It includes at least one processor and a memory, the memory storing a computer program configured to be executed by the processor to implement a smart cockpit multimodal testing method according to any one of claims 1-2.

5. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, on which a computer program is stored. The computer program can be executed by one or more processors to implement a smart cockpit multimodal testing method as described in any one of claims 1-2.

Citation Information

Patent Citations

  • Building crack detection algorithm based on multiple modes

    CN116563262A

  • Simulation test method and device for intelligent cockpit

    CN119089802A