Intelligent cabin multi-mode test method, system and device and medium

By generating multimodal dynamic noise and motion detection system identification analysis through generative adversarial networks, the problem of single modality in existing smart cockpit testing methods is solved, and a comprehensive evaluation of the cockpit system and optimization of multimodal interaction are achieved, thereby improving the accuracy and efficiency of the test.

CN120669675AActive Publication Date: 2025-09-19CHONGQING VEHICLE TEST & RES INST CO LTD

Patent Information

Application Number
CN202510821117.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing smart cockpit testing methods mainly focus on functional testing and performance evaluation of a single mode, which makes it difficult to comprehensively evaluate the overall performance of the cockpit, especially unable to effectively detect the mutual influence and potential conflicts between different modes.

Method used

A generative adversarial network is used to generate multimodal dynamic noise to interfere with the cockpit command signal. The multimodal dynamic noise is combined with the motion detection system for identification and analysis to generate evaluation results. Through a multi-objective optimization algorithm and a reinforcement learning-driven test strategy, various noise interferences in the actual driving process are fully simulated.

Benefits of technology

It improves the accuracy and comprehensiveness of cockpit testing, can better evaluate the stability and reliability of the cockpit in complex environments, significantly improves the fault detection rate and testing efficiency, and optimizes the overall performance of the multimodal interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 6NQ3VBES88WBKT2WPJPLQOSMBQJJBYX7EL9EYKUD
    Figure 6NQ3VBES88WBKT2WPJPLQOSMBQJJBYX7EL9EYKUD
  • Figure HMJ56DUXZVTSEBPJJPAHXXL6DQEHFLY7DEVNKYVC
    Figure HMJ56DUXZVTSEBPJJPAHXXL6DQEHFLY7DEVNKYVC
  • Figure HW1OE4DDTPJANC8BD2PDYJTXXKDO6FCVRADM7DPX
    Figure HW1OE4DDTPJANC8BD2PDYJTXXKDO6FCVRADM7DPX
Patent Text Reader

Abstract

The invention provides an intelligent cockpit multi-modal test method, system and device and a medium, and the method comprises the steps: obtaining the original state information of a cockpit and a cockpit instruction signal, and generating multi-modal dynamic noise through employing a generative adversarial network based on the original state information; performing interference on the cabin instruction signal by using the multi-modal dynamic noise to generate a multi-modal interference signal; and constructing an action detection system, and performing identification analysis on the multi-mode interference signal by using the action detection system to generate an evaluation result. The problems that in the prior art, testing and evaluation modes of a testing method are single, and the overall performance of the cabin is difficult to evaluate are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multimodal testing method, system, equipment, and medium for an intelligent cockpit. Background Art

[0002] With the rapid development of smart cockpit technology, multimodal interactive systems have become an integral part of modern vehicles. These systems integrate multiple interaction methods, including voice control, touch operation, and gesture recognition, aiming to provide drivers and passengers with a more convenient, safe, and personalized experience. However, as these systems grow in complexity, ensuring their stability and reliability in diverse environments becomes increasingly challenging.

[0003] Existing smart cockpit testing methods primarily focus on functional testing and performance evaluation of a single modality. For example, voice recognition system testing is typically conducted in an idealized laboratory environment, which fails to fully simulate the various noise interferences encountered during actual driving. Similarly, touchscreen testing often overlooks the impact of external factors such as vehicle vibration and lighting changes. This fragmented testing approach makes it difficult to evaluate the overall performance of the cockpit, especially when effectively detecting interactions and potential conflicts between different modalities. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides a multi-modal testing method, system, equipment and medium for an intelligent cockpit, which solves the problem that the test evaluation of the test method in the existing technology has a single modality and it is difficult to evaluate the overall performance of the cockpit.

[0005] According to an embodiment of the present invention, a multimodal testing method for an intelligent cockpit includes: Obtain the original state information and cockpit command signals of the cockpit, and then use the generative adversarial network to generate multimodal dynamic noise based on the original state information; Use multimodal dynamic noise to interfere with cockpit command signals to generate multimodal interference signals; Build a motion detection system and use it to identify and analyze multimodal interference signals to generate evaluation results.

[0006] Preferably, the original state information includes an initial state image of the cockpit and an initial environment background image; The cockpit command signal includes a voice signal, a touch signal and a gesture signal.

[0007] Preferably, after the original state information is acquired, noise filtering and grayscale processing with a weight ratio of R:G:B=0.299:0.587:0.114 are performed on the original state information.

[0008] Preferably, the multimodal dynamic noise interfering with the cockpit command signal includes sudden noise interference, illumination change interference, and image deformation interference; Before identifying and analyzing multimodal interference signals, the multimodal interference signals need to be converted into an input signal graph in a unified format.

[0009] Preferably, before using multimodal dynamic noise to interfere with the cockpit command signal, the interference intensity of the cockpit command signal is calculated based on the evaluation result of the previous recognition analysis: Among them, R is the interference intensity, Accuracy is the last evaluation result, NormDifficulty is the last test difficulty, Diversity is the diversity of the current test scene, 、 is the weight coefficient.

[0010] Preferably, before using the motion detection system to identify and analyze the multimodal interference signal, a test strategy needs to be formulated. The method for formulating the test strategy is as follows: Calculate the adaptive test index of each cockpit command signal based on the evaluation results of the previous test; Calculate the test importance of each cockpit command signal according to the preset test conditions of the adaptive test index value corresponding to each cockpit command signal; The mutual information algorithm is used to calculate the correlation between different modes in multimodal dynamic noise; According to the test importance and relevance, a multi-objective optimization algorithm is used to calculate the test priority of each cockpit command signal.

[0011] Preferably, the calculation formula for the correlation between different modalities is as follows: in, is the joint probability distribution of mode A and mode B, and are the marginal probability distributions of mode A and mode B respectively.

[0012] On the other hand, according to an embodiment of the present invention, a smart cockpit multimodal testing system is also provided. The system uses the above-mentioned smart cockpit multimodal testing method, including: An acquisition module, which is used for original status information and cockpit command signals; A noise simulation module, configured to generate multimodal dynamic noise based on original state information using a generative adversarial network; an interference module, configured to interfere with the cockpit command signal using multimodal dynamic noise to generate a multimodal interference signal; The evaluation module is used to identify and analyze the multimodal interference signal and generate an evaluation result.

[0013] On the other hand, according to an embodiment of the present invention, a computer is further provided, comprising at least one processor and a memory, wherein the memory stores a computer program, and the computer program is configured to be executed by the processor to implement the above-mentioned smart cockpit multimodal testing method.

[0014] On the other hand, according to an embodiment of the present invention, a storage medium is further provided. The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. The computer program can be executed by one or more processors to implement the above-mentioned smart cockpit multimodal testing method.

[0015] Compared with the prior art, the present invention has the following beneficial effects: The present invention uses a generative adversarial network to generate multimodal interactive dynamic noise based on the original state information, interferes with the collected cockpit command signal, and conducts in-depth analysis and evaluation of the interfered cockpit command signal. Through multimodal interference, it can fully simulate various noise interferences that may be encountered during actual driving, thereby improving the accuracy of cockpit testing. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a multimodal testing flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The technical solutions of the present invention are further described below with reference to the accompanying drawings and embodiments.

[0018] like Figure 1 As shown, an embodiment of the present invention proposes a multimodal testing method for an intelligent cockpit, including: Obtain the original state information and cockpit command signals of the cockpit, and then use the generative adversarial network to generate multimodal dynamic noise based on the original state information; Since different environments have different interference noises, directly setting the interference noise manually is often biased. Therefore, the present invention generates noise directly based on the original state information of the cockpit. This original state information includes the initial state diagram of the smart cockpit and the initial environmental background. In addition, cockpit command information is also collected, including voice signals, touch signals and gesture signals.

[0019] The present invention preferably uses a high-dynamic-range camera to capture the initial state image of the multimodal input device and the initial environmental background. This ensures high-quality images under varying lighting conditions, which is crucial for subsequent environmental simulations. Furthermore, the present invention employs voice recognition sensors, touch devices, and gesture recognition sensors within the smart cockpit to collect voice, touch, and gesture signals, respectively. This multimodal signal acquisition method enables the present invention to comprehensively evaluate the interactive performance of the smart cockpit system.

[0020] The present invention first trains a generative adversarial network (GAN) based on the acquired initial state diagram and initial environmental background. The generative adversarial network is a powerful generative model that can simulate complex environmental changes, greatly improving the authenticity and diversity of test scenarios.

[0021] Before training GAN, the initial state image and the initial environment background will be processed, including noise filtering and grayscale processing. This preprocessing step can improve training efficiency and model performance. Among them, noise filtering uses methods such as median filtering or Gaussian filtering. The specific parameters can be adjusted according to the image characteristics. Grayscale processing can use the weighted averaging method, and the weight ratio is usually R:G:B=0.299:0.587:0.114.

[0022] After training, GANs are used to generate multimodal interactive dynamic noise that simulates environmental patterns. This dynamic noise includes not only traditional acoustic noise but also multimodal noise such as visual interference and tactile feedback. For example, GANs can simulate sudden noises within a smart cockpit, such as the sudden increase in volume of the car's audio system or the chatter of passengers. They can also simulate lighting changes, such as sudden changes in light at tunnel entrances and exits or dappled shadows from trees.

[0023] To simulate noise, the present invention uses a GAN model based on acoustic features. This model not only considers the spectral characteristics of noise but also simulates its temporal variations. For example, it can simulate the transient noise during vehicle startup, wind noise during driving, and even vibration noise under different road conditions. To ensure the authenticity of the simulation, the present invention uses a large amount of audio data collected in real scenes when training the GAN, including noise samples from different vehicle models and driving conditions.

[0024] In terms of simulating illumination changes, the method of the present invention adopts a GAN model based on image-to-image conversion. This model can generate cabin interior images under different lighting conditions, including but not limited to direct sunlight, changing light in tunnels, and night driving. To improve the accuracy of the simulation, the method introduces an illumination estimation module into the GAN discriminator to evaluate the illumination authenticity of the generated images. The illumination estimation module adopts an illumination estimation algorithm based on deep learning, and its loss function is It can be expressed as: , in, and Represent the estimated illumination map and the real illumination map respectively, represents the gradient operator, is the balance coefficient, which is usually between 0.1 and 1.

[0025] In addition, the GAN-generated images are slightly deformed to create a slight disturbance. This deformation operation simulates the slight jitter or deformation that may occur during actual use. The deformation can be achieved through elastic transformation, which is mathematically expressed as follows: , in, and represent the original image and the deformed image respectively, and It is a deformation field that can be generated by a Gaussian filtered random field. The strength of the deformation is usually controlled in the range of 1-5 pixels to ensure that the interference is subtle enough but still has an impact on the system.

[0026] Use multimodal dynamic noise to interfere with cockpit command signals to generate multimodal interference signals; After generating multimodal dynamic noise, multimodal interference signals are used to interfere with the voice signal, touch signal, and gesture signal respectively to generate corresponding multimodal interference signals (a total of three). The three different multimodal interference signals are then converted into input signal graphs in a unified format. This unified processing facilitates subsequent comprehensive analysis. The input signal graph is then screened and calibrated to improve signal quality.

[0027] Build a motion detection system, use it to identify and analyze multimodal interference signals, and generate evaluation results; The action detection system uses a deep learning model based on the attention mechanism. The core of this model is a multi-head self-attention layer, and its calculation process can be expressed as follows: , in, denote query, key, and value matrices respectively, is the dimension of the key vector. The multi-head attention mechanism allows the model to focus on different parts of the input simultaneously, thereby improving robustness to complex interference. A reinforcement learning-driven test path adaptation algorithm is also introduced to adjust interference intensity. This adaptive mechanism enables the test process to dynamically adjust difficulty based on the actual system performance, thereby more comprehensively evaluating the system's robustness.

[0028] In the reinforcement learning framework, the present invention defines a state space , including the current interference type, intensity and system response; action space Including increasing or decreasing the intensity of various types of interference; reward function It is defined based on the performance of the system, for example, it can take the following form: , Among them, R is the interference intensity, Accuracy is the last evaluation result, NormDifficulty is the last test difficulty, Diversity is the diversity of the current test scene, 、 It is a weight coefficient used to balance the importance of different factors. Based on the evaluation results of the previous test, the interference intensity of the cockpit command signal in this test is adjusted to avoid distortion of the multimodal interference signal caused by excessively high or low intensity.

[0029] Before using the motion detection system to identify and analyze multimodal interference signals, a test strategy must be developed. The method for developing the test strategy is as follows: (1) Calculate the adaptive test index based on the interaction failure. The interaction failure here means that the output result of the motion detection system in the previous test is inconsistent with the expectation, that is, the recognized result is not the original signal, which may be due to misrecognition or response delay caused by interference. The calculation of the adaptive test index adopts a comprehensive scoring mechanism, which can be expressed as: , in, represents the interaction failure rate, represents the test path length, Indicates the diversity of test scenarios, 、 、 is the weight coefficient.

[0030] (2) Set the interaction priority. This priority determines the priority of different interaction modes and test scenarios. Calculate the test importance S of each test according to the following formula to sort each test in the test process: , in, It is the adaptive test indicator mentioned above. Indicates the test coverage, is the time required for the test, is the balance coefficient. In this way, the reinforcement learning model can find the best balance between test efficiency and comprehensiveness.

[0031] (3) Analyze the impact of noise on the correlation between modes This method uses mutual information to quantify the correlation between different modes. For mode A and mode B, their mutual information can be expressed as: , Among them, I (A; B) is the correlation between mode A and mode B, is the joint probability distribution of mode A and mode B, and are the marginal probability distributions of mode A and mode B, respectively. By comparing the mutual information in the presence and absence of noise, the effect of noise on modal correlation can be quantified.

[0032] (4) Generate a test priority list that not only considers the performance of a single modality but also takes into account the synergistic effects of multiple modalities. Preferably, the list is generated using a multi-objective optimization algorithm, and the objective function can be expressed as: , in, Represents different optimization objectives, such as accuracy, robustness, response time, etc. Through the concept of Pareto optimal solution, this method can find the best balance between multiple objectives, assign a priority to each test, and generate a test priority list according to the priority.

[0033] Finally, based on multiple indicators such as interaction failure rate and test path length, the evaluation results of the multimodal interference signal are output. This comprehensive evaluation uses a weighted scoring mechanism and can be expressed as: , in, For the comprehensive rating, Indicates the evaluation indicators, is the corresponding weight. These indicators may include but are not limited to: average recognition accuracy, minimum recognition accuracy, average response time, anti-interference ability, etc. The weight setting can be adjusted according to the needs of specific application scenarios.

[0034] In addition, extreme interference signals can be generated based on the evaluation results, and then the motion detection system can be used to identify and analyze the extreme interference signals and generate analysis results under extreme environments to comprehensively evaluate the cockpit's anti-interference capabilities and failure modes.

[0035] On the other hand, an embodiment of the present invention further provides a smart cockpit multimodal testing system, which uses the above-mentioned smart cockpit multimodal testing method, including: An acquisition module, which is used for original status information and cockpit command signals; A noise simulation module, configured to generate multimodal dynamic noise based on original state information using a generative adversarial network; an interference module, configured to interfere with the cockpit command signal using multimodal dynamic noise to generate a multimodal interference signal; The evaluation module is used to identify and analyze the multimodal interference signal and generate an evaluation result.

[0036] On the other hand, an embodiment of the present invention further provides a computer, comprising at least one processor and a memory, wherein the memory stores a computer program, and the computer program is configured to be executed by the processor to implement the above-mentioned smart cockpit multimodal testing method.

[0037] On the other hand, an embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium and stores a computer program. The computer program can be executed by one or more processors to implement the above-mentioned smart cockpit multimodal testing method.

[0038] To verify the superiority of the test method and system of the present invention, a set of simulation experiments was designed. The experimental environment simulated an intelligent cockpit system, including three interaction modes: voice control, touch operation, and gesture recognition. The simulation conditions are as follows: 1. Hardware environment: Build a virtual intelligent cockpit equipped with a high-fidelity audio system, a 10.1-inch touch screen, and a depth camera.

[0039] 2. Software Platform: Use the CARLA car simulator to simulate the driving environment, and OpenAI Gym to build a reinforcement learning environment.

[0040] 3. Interference factors: The simulation includes road noise (60-80dB), lighting changes (100-10,000 lux) and vehicle vibration (0.5-2g).

[0041] The performance of the method of the present invention (Example) was compared with two existing methods (Comparative Example 1 and Comparative Example 2): Embodiment: The multimodal interaction robustness adaptive testing method of the present invention is adopted.

[0042] Comparative Example 1: Traditional static test case method.

[0043] Comparative Example 2: Single modal adaptive testing method.

[0044] The test indicators and their detection methods are as follows: 1. Test coverage: Use code coverage tools to calculate the proportion of system functions that are tested.

[0045] 2. Fault detection rate: Artificially implant known faults and calculate the proportion of successful detections.

[0046] 3. Test efficiency: Record the time required to complete a comprehensive test.

[0047] 4. Multimodal collaboration performance: evaluates the accuracy when multiple interaction methods are used simultaneously.

[0048] 5. Extreme situation handling capability: Test the accuracy of system response under the most severe conditions.

[0049] The test results are shown in the following table: Analysis and discussion: 1. Test coverage: The proposed method achieved a high coverage of 95%, far exceeding traditional methods (78%) and single-modality methods (85%). This is due to its ability to dynamically generate diverse test scenarios, more comprehensively covering a wide range of possible use cases.

[0050] 2. Fault Detection Rate: The 92% fault detection rate of this method is significantly higher than that of the other two methods. This shows that the adaptive testing strategy of this invention can more effectively locate potential system problems, especially hidden faults in complex interaction scenarios.

[0051] 3. Test Efficiency: This method can complete a comprehensive test in just 24 hours, which is three times more efficient than traditional methods and twice as efficient as single-modality methods. This demonstrates the advantage of the reinforcement learning-driven adaptive test path generation in this invention, which can quickly locate key test points.

[0052] 4. Multimodal Collaboration Performance: On this key metric, our method achieved an 88% score, significantly outperforming other methods. This highlights the unique advantages of our method in evaluating and optimizing the overall performance of multimodal interaction systems.

[0053] 5. Extreme Performance: Our method maintains 82% of its performance under extreme conditions, far exceeding other methods. This demonstrates that our strategy of generating “worst-case” test samples is indeed effective in evaluating the system’s extreme performance.

[0054] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A multimodal testing method for an intelligent cockpit, characterized by: include: Obtain the original state information and cockpit command signals of the cockpit, and then use the generative adversarial network to generate multimodal dynamic noise based on the original state information; Use multimodal dynamic noise to interfere with cockpit command signals to generate multimodal interference signals; Build a motion detection system and use it to identify and analyze multimodal interference signals to generate evaluation results.

2. The multimodal testing method for an intelligent cockpit according to claim 1, wherein: The original state information includes an initial state map of the cockpit and an initial environment background map; The cockpit command signal includes a voice signal, a touch signal and a gesture signal.

3. The multimodal testing method for an intelligent cockpit according to claim 2, wherein: After obtaining the original state information, the original state information is subjected to noise filtering and grayscale processing with a weight ratio of R:G:B=0.299:0.587:0.

114.

4. The multimodal testing method for an intelligent cockpit according to claim 2, wherein: Multimodal dynamic noise interferes with cockpit command signals, including sudden noise interference, illumination change interference, and image deformation interference; Before identifying and analyzing multimodal interference signals, the multimodal interference signals need to be converted into an input signal graph in a unified format.

5. The multimodal testing method for an intelligent cockpit according to claim 1, wherein: Before using multimodal dynamic noise to interfere with the cockpit command signal, the interference intensity on the cockpit command signal must be calculated based on the evaluation results of the previous identification analysis: Among them, R is the interference intensity, Accuracy is the last evaluation result, NormDifficulty is the last test difficulty, Diversity is the diversity of the current test scene, 、 is the weight coefficient.

6. The multimodal testing method for a smart cockpit according to claim 1, wherein: Before using the motion detection system to identify and analyze multimodal interference signals, a test strategy must be developed. The method for developing the test strategy is as follows: Calculate the adaptive test index of each cockpit command signal based on the evaluation results of the previous test; Calculate the test importance of each cockpit command signal according to the preset test conditions of the adaptive test index value corresponding to each cockpit command signal; The mutual information algorithm is used to calculate the correlation between different modes in multimodal dynamic noise; According to the test importance and relevance, a multi-objective optimization algorithm is used to calculate the test priority of each cockpit command signal.

7. The multimodal testing method for an intelligent cockpit according to claim 6, characterized in that: The calculation formula for the correlation between different modes is as follows: Among them, I (A; B) mode A and mode B correlation is is the joint probability distribution of mode A and mode B, and are the marginal probability distributions of mode A and mode B respectively.

8. A multimodal testing system for an intelligent cockpit, characterized by: The system uses a smart cockpit multimodal testing method according to any one of claims 1 to 7, comprising: An acquisition module, which is used for original status information and cockpit command signals; A noise simulation module, configured to generate multimodal dynamic noise based on original state information using a generative adversarial network; an interference module, configured to interfere with the cockpit command signal using multimodal dynamic noise to generate a multimodal interference signal; The evaluation module is used to identify and analyze the multimodal interference signal and generate an evaluation result.

9. A computer, characterized in that: The system comprises at least one processor and a memory, wherein the memory stores a computer program, and the computer program is configured to be executed by the processor to implement the multimodal testing method for an intelligent cockpit according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, on which a computer program is stored. The computer program can be executed by one or more processors to implement the multimodal testing method for an intelligent cockpit as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Direct time sequence action detection method based on relaxation transformation decoder

    CN114821379A

  • Test method for generating application classification model evaluation based on generative adversarial network image and text data

    CN115758096A

  • Building crack detection algorithm based on multiple modes

    CN116563262A

  • Cabin voice test system and method, electronic equipment and readable storage medium

    CN116665713A

  • Multi-mode cabin atmosphere control method and device, electronic equipment and storage medium

    CN118409537A

Cited By

  • Real-time voice transcription anti-interference test system based on environmental noise simulation

    CN122313950A