Cockpit Display and Control Interaction Test System Based on Multimodal Fusion

The multimodal fusion cockpit display and control interaction test system solves the shortcomings of traditional test systems in multimodal signal processing, signal acquisition adaptability and test scenario construction, and realizes efficient and accurate testing of the interactive performance of cockpit display and control equipment.

CN120653501BActive Publication Date: 2025-10-28NORTHWESTERN POLYTECHNICAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511171292.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-10-28
Estimated Expiration
2045-08-21

Smart Images

  • Figure CN120653501B_ABST
    Figure CN120653501B_ABST
Patent Text Reader

Abstract

This invention relates to the field of cockpit display and control interaction testing technology, and discloses a cockpit display and control interaction testing system based on multimodal fusion. The system includes an interaction data input module for acquiring basic parameters such as the device model and screen size of the cockpit display and control device under test; a multimodal signal acquisition module for acquiring visual, auditory, and tactile signals from the cockpit environment based on the basic parameters; a fusion feature processing module for performing feature fusion processing on the multimodal signals to obtain interaction feature fusion data; a test scenario construction module for generating standardized test scenarios based on the interaction feature fusion data; and an interaction performance verification module for inputting the standardized test scenarios into the device under test to perform verification tests and transmitting the results to a display terminal for presentation. This system achieves the acquisition and fusion of multimodal signals, can generate test scenarios that closely resemble reality, and improves the accuracy and efficiency of cockpit display and control device interaction performance testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cockpit display and control interaction testing technology, specifically a cockpit display and control interaction testing system based on multimodal fusion. Background Technology

[0002] With the rapid development of aerospace, automotive, and other fields, cockpit display and control equipment, as a key interface for human-computer interaction, directly affects the safety, convenience, and comfort of operation. However, traditional cockpit display and control interaction testing systems have many shortcomings and are unable to meet the testing needs of modern complex cockpit environments.

[0003] Traditional testing systems often only collect and process signals of a single modality, such as visual or auditory signals, while ignoring the synergistic effects between multimodal signals. This single-modal testing approach cannot fully reflect the interactive performance of cockpit display and control equipment in actual use, resulting in significant limitations in test results. For example, during actual driving, drivers not only need to acquire information visually but also interact using multiple senses such as hearing and touch. Single-modal testing cannot simulate this complex interactive scenario.

[0004] Furthermore, traditional testing systems lack flexibility and adaptability in signal acquisition. Different models of cockpit display and control equipment have different basic parameters, such as device model, screen size, touch sensitivity, and voice recognition threshold. Traditional systems struggle to automatically adjust their signal acquisition strategies based on these parameters, resulting in low-quality acquired signals and affecting the accuracy of test results. For example, for devices with different screen sizes, traditional systems may use the same visual signal acquisition range, failing to accurately capture all interactive information on the screen.

[0005] In terms of test scenario construction, traditional systems typically use fixed test scenarios, which cannot be dynamically adjusted based on actual interaction characteristics and data integration. These fixed test scenarios differ significantly from the actual cockpit environment, making the test results unable to accurately reflect the device's performance in real-world use. For example, factors such as light intensity, background noise levels, and operating temperature in the actual cockpit environment are constantly changing, and traditional systems cannot simulate the impact of these dynamically changing environmental factors on the device's interactive performance.

[0006] Traditional testing systems lack effective algorithms and models for multimodal feature fusion processing, making it difficult to deeply fuse signals from multiple modalities such as vision, hearing, and touch. This results in the inability to extract comprehensive and accurate interaction feature fusion data. Consequently, subsequent test scenario construction and interaction performance verification testing lack reliable basis, leading to low credibility of test results. The interaction performance verification testing process in traditional systems lacks intelligence and automation, requiring extensive manual operations and data processing. This is not only inefficient but also prone to introducing human error. For example, after inputting the test scenario into the device under test, manual observation and recording of verification test results are necessary. This method is not only time-consuming and labor-intensive but may also result in inaccurate records due to human factors.

[0007] Traditional cockpit display and control interaction testing systems have significant shortcomings in multimodal signal processing, signal acquisition adaptability, test scenario construction, feature fusion algorithms, and test process automation. There is an urgent need for a cockpit display and control interaction testing system based on multimodal fusion to solve these problems and improve the accuracy, comprehensiveness, and efficiency of cockpit display and control device interaction performance testing. Summary of the Invention

[0008] The purpose of this invention is to provide a cockpit display and control interaction test system based on multimodal fusion to solve the problems mentioned in the background art.

[0009] To achieve the above objectives, the present invention provides the following technical solution: a cockpit display and control interaction test system based on multimodal fusion, the system comprising:

[0010] The interactive data input module is used to obtain the basic parameters of the cockpit display and control device under test input by the user. The basic parameters include device model, screen size, touch sensitivity and voice recognition threshold.

[0011] The multimodal signal acquisition module is used to acquire visual, auditory, and tactile signals from the cockpit environment based on the basic parameters of the cockpit display and control device under test.

[0012] The feature fusion processing module is used to perform multimodal feature fusion processing on the visual signals, auditory signals and tactile signals to obtain interactive feature fusion data;

[0013] The test scenario construction module is used to generate standardized test scenarios based on the interaction feature fusion data.

[0014] The interactive performance verification module is used to input the standardized test scenario into the cockpit display and control device under test to perform interactive performance verification test, and transmit the verification test results to the display terminal for presentation.

[0015] Preferably, the multimodal signal acquisition module includes:

[0016] An environmental perception unit is used to extract configuration information of visual signal acquisition devices, auditory signal acquisition devices, and tactile signal acquisition devices from the sensor array.

[0017] The signal type classification unit is used to extract the functional descriptions of each signal acquisition device in the configuration information to obtain a set of signal type descriptions.

[0018] The signal feature encoding unit is used to perform feature encoding on the basic parameters of the cockpit display and control device under test and each signal type description in the set of signal type descriptions to obtain a set of device parameter encoding vectors and signal type encoding vectors;

[0019] A multimodal correlation analysis unit is used to perform multimodal correlation analysis on the set of device parameter encoding vectors and signal type encoding vectors to obtain device-signal correlation optimized encoding vectors;

[0020] The signal acquisition and control unit is used to generate acquisition strategies for the visual, auditory, and tactile signals based on the device-signal correlation optimization encoding vector.

[0021] Preferably, the multimodal correlation analysis unit includes:

[0022] The signal type encoding feature aggregation subunit is used to perform hierarchical mapping-based feature aggregation processing on the set of signal type encoding vectors to obtain a guiding template;

[0023] A cross-domain collaborative coding subunit is used to perform cross-domain collaborative coding on the set of device parameter coding vectors and signal type coding vectors based on the guiding template to obtain the device-signal association optimized coding vector.

[0024] Preferably, the signal type encoding feature aggregation subunit includes:

[0025] The signal type coding hierarchical mapping second-level subunit is used to perform hierarchical mapping processing on each signal type coding vector in the set of signal type coding vectors using a mapping matrix to obtain a set of signal type coding vectors after hierarchical transformation.

[0026] The signal type encoding matrix arrangement second-level sub-unit is used to arrange the set of signal type encoding vectors after the hierarchical transformation into a matrix to obtain the signal type encoding feature matrix;

[0027] The signal type coding matrix extremum extraction secondary sub-unit is used to extract the extrema of the signal type coding vector after each level transformation in the signal type coding feature matrix to obtain the key vector of the signal type coding feature matrix as the guiding template.

[0028] Preferably, the signal type coding level mapping second-level subunit includes:

[0029] The signal type encoding vector is multiplied by the mapping matrix and then added to the mapping bias vector positionally to obtain the signal type encoding vector after hierarchical transformation.

[0030] Preferably, the cross-domain cooperative coding subunit includes:

[0031] The device parameter hierarchical mapping second-level subunit is used to perform hierarchical mapping processing on the device parameter encoding vector using a query matrix and a value matrix to obtain a device parameter query vector and a device parameter value vector.

[0032] The template-guided heterogeneous conversion coding secondary subunit is used to combine the device parameter query vector, the device parameter value vector, the signal type coding vector after transformation at each level in the signal type coding feature matrix, and the guiding template input into the template-guided heterogeneous conversion structure to obtain a sequence of device-signal cross-domain collaborative coding vectors.

[0033] The position mean calculation subunit is used to calculate the position mean vector of the sequence of device-signal cross-domain cooperative coding vectors to obtain the device-signal association optimized coding vector.

[0034] Preferably, the template-guided heterogeneous conversion coding secondary subunit includes:

[0035] After calculating the product between the device parameter query vector and the transpose of the signal type encoding vector after hierarchical transformation, the resulting device-signal association feature matrix is ​​divided by the magnitude of the key vector of the signal type encoding feature matrix according to its position to obtain the device-signal association weight matrix.

[0036] The device-signal association weight matrix is ​​input into a normalization function for processing to obtain the device-signal association normalized weight matrix;

[0037] After multiplying the device-signal association normalized weight matrix with the key vector of the signal type coding feature matrix, the resulting feature vector is multiplied by the device parameter value vector at position to obtain the device-signal cross-domain collaborative coding vector.

[0038] Preferably, the signal acquisition and control unit includes:

[0039] The device-signal association optimization encoding vector is input into the decision-maker-based signal acquisition strategy recommendation module to obtain the acquisition strategies for the visual, auditory, and tactile signals;

[0040] Based on the acquisition strategy, the acquisition range of the visual, auditory, and tactile signals is determined.

[0041] Preferably, the test scenario construction module includes:

[0042] The scene parameter extraction unit is used to extract scene parameters such as ambient light intensity, background noise decibel value, and operating temperature range from the interaction feature fusion data;

[0043] A scene template matching unit is used to match a basic test template corresponding to the scene parameters from a preset scene template library;

[0044] The scene detail filling unit is used to fill in the touch response latency, voice command recognition accuracy and interface switching smoothness parameters of the basic test template based on the interaction feature fusion data to obtain the standardized test scene.

[0045] Preferably, the fusion feature processing module includes:

[0046] A visual feature extraction subunit is used to perform edge detection and color space conversion on the visual signal to obtain a visual feature vector;

[0047] An auditory feature extraction unit is used to perform frequency analysis and time-domain waveform extraction on the auditory signal to obtain an auditory feature vector;

[0048] The tactile feature extraction unit is used to identify the pressure distribution and count the contact duration of the tactile signal to obtain a tactile feature vector;

[0049] A multimodal feature alignment unit is used to perform timestamp synchronization and dimension unification processing on the visual feature vector, auditory feature vector and tactile feature vector to obtain an aligned feature vector set;

[0050] The feature fusion unit is used to perform weighted summation on the aligned feature vector set to obtain the interactive feature fusion data.

[0051] Compared with the prior art, the beneficial effects of the present invention are:

[0052] The system acquires basic parameters of the cockpit display and control equipment under test through an interactive data input module, providing accurate data for subsequent testing. Based on these basic parameters, the multimodal signal acquisition module extracts configuration information of various signal acquisition devices from the sensor array, classifies and encodes signal types, and obtains optimized device-signal correlation encoding vectors through multimodal correlation analysis. This generates targeted acquisition strategies, ensuring more accurate and comprehensive acquisition of visual, auditory, and tactile signals, thus improving the quality and efficiency of signal acquisition.

[0053] The fusion feature processing module processes the acquired multimodal signals, extracts feature vectors of each modality through operations such as edge detection, frequency analysis, and pressure distribution recognition, performs alignment processing such as timestamp synchronization and dimension unification, and finally obtains interactive feature fusion data by weighted summation, realizing deep fusion of multimodal features and providing reliable data support for the construction of subsequent test scenarios.

[0054] The test scenario construction module extracts scene parameters such as ambient light intensity and background noise decibels from the interaction feature fusion data, matches them with basic test templates in the preset scene template library, and fills in details such as touch response latency to generate standardized test scenarios. This makes the test scenarios closer to the actual cockpit environment and improves the authenticity and effectiveness of the test.

[0055] The interactive performance verification module inputs standardized test scenarios into the device under test to perform verification tests and transmits the results to the display terminal for presentation. This achieves automation and intelligence in the testing process, reduces manual operation and human error, and improves testing efficiency and the accuracy of results.

[0056] This system utilizes multimodal fusion technology to comprehensively consider the synergistic effects of visual, auditory, and tactile signals, enabling a more complete reflection of the interactive performance of cockpit display and control equipment. Simultaneously, the system can automatically adjust signal acquisition strategies and construct test scenarios based on the equipment's fundamental parameters, improving its adaptability and flexibility. Furthermore, the various modules of the system collaborate to form a complete testing process. From signal acquisition to feature fusion, scenario construction, and performance verification, each step has been meticulously designed and optimized, ensuring the efficiency and reliability of the entire testing process and providing strong technical support for the research and development and optimization of cockpit display and control equipment. Attached Figure Description

[0057] Figure 1 This is a schematic diagram illustrating the working principle of the cockpit display and control interaction test system based on multimodal fusion as described in this invention.

[0058] Figure 2 This is a flowchart of the multimodal signal acquisition module.

[0059] Figure 3 A flowchart for cross-domain collaborative coding subunits;

[0060] Figure 4 Detailed diagram of template-guided heterogeneous conversion encoding;

[0061] Figure 5 The flowchart for the feature processing module. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] Please see Figures 1-5 This invention relates to a cockpit display and control interaction test system based on multimodal fusion. The system includes: an interaction data input module, a multimodal signal acquisition module, a fusion feature processing module, a test scenario construction module, and an interaction performance verification module. The specific implementation is as follows:

[0064] The interactive data input module acquires the basic parameters of the cockpit display and control device under test, including device model, screen size, touch sensitivity, and voice recognition threshold. The multimodal signal acquisition module, based on the basic parameters of the cockpit display and control device under test, acquires visual, auditory, and tactile signals from the cockpit environment. The fusion feature processing module performs multimodal feature fusion processing on the visual, auditory, and tactile signals to obtain interactive feature fusion data. The test scenario construction module generates standardized test scenarios based on the interactive feature fusion data. The interactive performance verification module inputs the standardized test scenarios into the cockpit display and control device under test to perform interactive performance verification tests and transmits the verification test results to the display terminal for presentation.

[0065] Example 1: The multimodal signal acquisition module includes an environmental perception unit, a signal type classification unit, a signal feature encoding unit, a multimodal correlation analysis unit, and a signal acquisition control unit. The environmental perception unit operates by establishing a data interaction channel with the sensor array. It parses the raw configuration information data stream output by the sensor array using a preset protocol, filtering out configuration parameters related to the visual, auditory, and tactile signal acquisition devices. These parameters include physical attributes such as camera resolution, microphone sampling rate, and pressure sensor array distribution. This information is then structured and stored to form a device configuration information database.

[0066] After obtaining the device configuration information output by the environmental perception unit, the signal type classification unit performs semantic parsing on the functional description field of each signal acquisition device. Using keyword extraction technology in natural language processing, it identifies core functional keywords such as "image capture," "audio recording," and "pressure sensing" from the functional description text. Devices with the same functional keywords are grouped into the same signal type, thus constructing a signal type description set that includes categories such as visual signal type, auditory signal type, and tactile signal type. Each element in the set corresponds to an abstract functional description of a specific device.

[0067] The signal feature encoding unit's processing is divided into two parts. First, for the basic parameters of the cockpit display and control device under test, including device model, screen size, touch sensitivity, and voice recognition threshold, these parameters are first standardized. The device model is converted into a one-hot encoded vector, and numerical parameters such as screen size, touch sensitivity, and voice recognition threshold are normalized and mapped to the [0,1] interval, ultimately combining to form a multi-dimensional device parameter encoding vector. For each signal type description in the signal type description set, a word embedding model is used to convert the text description into a low-dimensional dense semantic vector. Through model training, the semantic relationships between different signal types are captured, generating the corresponding signal type encoding vector. All signal type encoding vectors constitute a set.

[0068] After receiving the device parameter encoding vector and signal type encoding vector set, the multimodal correlation analysis unit first performs feature aggregation processing on the signal type encoding vector set. Through hierarchical mapping, each signal type encoding vector is decomposed at different semantic levels, capturing its feature representations in different dimensions such as the basic function layer and application scenario layer. These hierarchical features are then aggregated to form a guiding template that can characterize the core features of the entire signal type set. Subsequently, based on this guiding template, cross-domain co-coding is performed on the device parameter encoding vector and signal type encoding vector set. Through matrix operations and feature transformations, the device parameter features and signal type features are correlated and mapped in the same feature space, generating a device-signal correlation optimized encoding vector that reflects the matching relationship between the device and the signal. This vector contains the correlation weight information between device characteristics and signal acquisition requirements.

[0069] After acquiring the device-signal association optimized encoding vector, the signal acquisition control unit inputs it into the decision-maker-based signal acquisition strategy recommendation module. This module contains a pre-trained decision model. Based on the device characteristics and signal association weights in the encoding vector, combined with a pre-defined signal acquisition rule base, the model generates acquisition strategies for visual, auditory, and tactile signals through logical reasoning and weight calculation. The acquisition strategy includes parameter settings such as signal acquisition frequency, sampling accuracy, and trigger conditions. Then, based on the generated acquisition strategy, the acquisition range for each signal type is further determined, such as the acquisition field of view for visual signals, the sound pickup range for auditory signals, and the sensing area range for tactile signals. Through coordinate mapping and parameter configuration, the abstract acquisition strategy is transformed into specific device control commands and sent to the corresponding signal acquisition devices, achieving precise control of the signal acquisition process. The entire multimodal signal acquisition module, through the collaborative work of its various units, achieves intelligent and precise acquisition of multimodal signals based on the device's basic parameters, providing high-quality raw data for subsequent multimodal feature fusion and test scenario construction.

[0070] Example 2: The multimodal correlation analysis unit includes a signal type coding feature aggregation subunit and a cross-domain cooperative coding subunit. The signal type coding feature aggregation subunit performs hierarchical mapping-based feature aggregation processing on the set of signal type coding vectors to obtain a guiding template. The workflow of the signal type coding hierarchical mapping second-level subunit is as follows: for each signal type coding vector in the set of signal type coding vectors... Using a preset mapping matrix The hierarchical mapping process involves multiplying the signal type encoding vector by the mapping matrix and then multiplying it by the mapping bias vector. By performing positional addition, the signal type encoding vector after hierarchical transformation is obtained. Its mathematical expression is:

[0071] ;

[0072] In the formula, Indicates the first Each signal type encoding vector has a dimension of . It contains semantic feature information about the signal type; For dimension The mapping matrix is ​​used to map the original signal type encoding vector to a new feature space, where... The feature dimensions after mapping; For dimension The mapping bias vector is used to adjust the feature offset after mapping; This is the signal type encoding vector after hierarchical transformation, with dimension . This transformation enables the decomposition and reconstruction of signal type characteristics at different levels.

[0073] The signal type encoding matrix arrangement subunit arranges the set of signal type encoding vectors after hierarchical transformation into a matrix. Assuming the hierarchical transformation yields... Signal type encoding vector Each vector has a dimension of Then these vectors are arranged in rows to form a Signal type coding feature matrix This matrix stores the feature information of the signal type encoding vectors after all levels of transformation in a structured form, which facilitates subsequent feature extraction and analysis.

[0074] Signal type encoding matrix extremum extraction second-level sub-unit for signal type encoding feature matrix The process involves extracting the extreme values ​​of the signal type encoding vectors after each level of transformation within the matrix. Specifically, for the matrix... For each column (i.e., each feature dimension), find the maximum and minimum values ​​in that column, forming a two-dimensional extremum vector. Then, combine the extremum vectors of all columns in order to obtain a dimension... Signal type encoding feature matrix key vector This vector, serving as a guiding template, encapsulates key feature information from the signal type encoding feature matrix, and can characterize the core feature distribution of the entire signal type set.

[0075] The cross-domain cooperative coding subunit, based on a guiding template, performs cross-domain cooperative coding on the set of device parameter coding vectors and signal type coding vectors to obtain device-signal correlation optimized coding vectors. Device parameter coding vectors The system contains basic parameter features of the cockpit display and control equipment under test, such as equipment model and screen size, which are encoded into a multi-dimensional feature vector. In the cross-domain collaborative coding process, the equipment parameter coding vector and the signal type coding feature matrix are first correlated and mapped in the same feature space. Using a guiding template as a bridge, cross-domain fusion of equipment parameter features and signal type features is achieved. This ensures that the generated device-signal correlation optimization coding vector reflects the matching relationship between equipment characteristics and signal acquisition requirements, providing crucial feature correlation basis for subsequent signal acquisition strategy generation. The entire multimodal correlation analysis unit, through the collaborative work of the signal type coding feature aggregation subunit and the cross-domain collaborative coding subunit, achieves effective fusion and correlation analysis of equipment parameter features and signal type features. This provides important feature processing support for the precise control of the multimodal signal acquisition module, ensuring that the acquisition strategy can better adapt to the characteristics of the cockpit display and control equipment under test, improving the quality and efficiency of signal acquisition. During the processing, key features are gradually extracted from the signal type encoding vector set through operations such as hierarchical mapping, matrix arrangement, and extreme value extraction. These features are then cross-domain collaboratively encoded with device parameter features to form an optimized encoding vector containing device-signal correlation information, laying a solid feature foundation for subsequent processing stages of the entire testing system.

[0076] Example 3: The cross-domain collaborative coding subunit includes a device parameter hierarchical mapping subunit, a template-guided heterogeneous conversion coding subunit, and a position mean calculation subunit. The device parameter hierarchical mapping subunit works by using a query matrix and a value matrix to perform hierarchical mapping processing on the device parameter coding vector. Device parameter coding vector It is a multi-dimensional vector obtained by standardizing and encoding the basic parameters of the cockpit display and control equipment under test (such as equipment model, screen size, touch sensitivity, voice recognition threshold, etc.), and its dimension is _____. Query matrix The dimension is Value matrix The dimension is ,in and These are the dimensions of the mapped query vector and value vector, respectively. The device parameter encoding vector is then generated through matrix multiplication. With query matrix Multiply to obtain the device parameter query vector. Its expression is Similarly, the device parameter encoding vector AND-value matrix Multiplying them yields a vector of device parameter values. The expression is This process enables the decomposition of device parameter features into different feature subspaces, providing a foundation for subsequent cross-domain interaction with signal type features.

[0077] Template-guided heterogeneous conversion encoding secondary subunit will query the device parameter vector. Device parameter value vector Signal type encoding vectors after transformation at each level in the signal type encoding feature matrix ( , (number of signal types) and boot template The input is a template-guided heterogeneous transformation structure. First, the device parameter query vector is calculated. Signal type encoding vector after hierarchical transformation transpose vector The product of these two elements yields the device-signal correlation feature matrix. Its dimensions are ( (where is the dimension of the signal type encoding vector after hierarchical transformation), the expression is: Next, the device-signal association feature matrix is... Encode the key vector of the feature matrix by dividing by the signal type based on position. Length of the module The device-signal correlation weight matrix is ​​obtained. ,Right now ,in" "" indicates element-wise division. This step normalizes the associated features by guiding the template's modulus, enhancing the stability and comparability of the features.

[0078] Subsequently, the device-signal association weight matrix was used. The input is processed using a normalization function (such as the softmax function) to obtain the device-signal correlation normalized weight matrix. The softmax function maps the elements of the weight matrix to... An interval is defined such that the sum of elements in the same row is 1, thus yielding the relative importance weights of each feature dimension. Specifically, for a matrix... Each element in Its normalized elements are .

[0079] Next, the device-signal association normalized weight matrix will be used. Key vectors of the signal type encoding feature matrix Multiplying them together yields the eigenvectors. ,Right now This feature vector integrates the association weight information between the device parameter query vector and the signal type feature, and then combines it with the device parameter value vector. Perform positional dot product processing to obtain the device-signal cross-domain cooperative coding vector. The expression is ,in" "" indicates element-wise multiplication. Through this series of operations, cross-domain fusion of device parameter features and signal type features is achieved, enabling the generated encoding vector to simultaneously reflect the semantic relationship between device characteristics and signal type.

[0080] The position mean calculation is the sequence of device-signal cross-domain co-coding vectors for the second-level sub-unit. Processing is performed on each encoded vector. Its dimensions are (Consistent with the dimension of the device parameter value vector), the mean of all encoded vectors is calculated according to their corresponding dimensions based on their position. Specifically, for the first... Dimensions ( ), calculate all The sum of the elements of each encoded vector along that dimension, divided by The mean vector is obtained. The element ,in For the first Encoded vectors The The final mean vector consists of [number] elements. This is the device-signal correlation optimized coding vector, whose dimension is... This vector integrates the cross-domain collaborative coding results of all signal types and device parameters, and condenses the overall correlation characteristics between devices and signals, providing a comprehensive and representative feature representation for subsequent signal acquisition and control.

[0081] The entire cross-domain collaborative coding subunit process decomposes the device parameter coding vector into query vectors and value vectors through hierarchical mapping. Using a guiding template as a bridge, it performs correlation calculations, weight normalization, and feature fusion between device parameter features and signal type features within a heterogeneous transformation structure. Finally, it obtains the optimized coding vector through position mean calculation. This process fully considers the semantic relationship between device characteristics and signal types. Through matrix operations and vector operations, features from different modalities and sources are mapped to the same feature space for collaborative processing. This ensures that the generated coding vector accurately reflects the matching relationship between the device and the signal, providing crucial feature support for the generation of multimodal signal acquisition strategies. This allows the signal acquisition process to more accurately adapt to the parameter characteristics of the cockpit display and control equipment under test, improving the targeting and effectiveness of signal acquisition. During processing, the mathematical operations in each step strictly follow the basic rules of linear algebra. Through reasonable matrix dimension design and operational logic, the rationality and accuracy of feature transformation and fusion are guaranteed, providing high-quality feature input for subsequent modules of the entire testing system.

[0082] Example 4: The workflow of the signal acquisition control unit involves inputting the device-signal association optimization encoding vector into the signal acquisition strategy recommendation module based on the decision-maker. This module has a pre-set decision model constructed based on the association between device parameters and signal types. Taking a certain model of cockpit display and control equipment as an example, its device parameter encoding vector includes a screen size of 10.1 inches, a normalized touch sensitivity parameter of 0.85, and a voice recognition threshold parameter of -35dB. In the device-signal association optimization encoding vector, the visual signal association weight component is 0.42, the auditory signal association weight component is 0.38, and the tactile signal association weight component is 0.2. Based on these weight components and combined with the pre-set acquisition rules, the decision-maker determines that the priority of visual signal acquisition is higher than that of auditory and tactile signals.

[0083] When generating the visual signal acquisition strategy, considering the device's large screen size and high touch sensitivity, the decision model recommends a high image acquisition resolution, such as 1920×1080 pixels, and a frame rate of 30 frames per second, to ensure clear capture of touch operation details on the screen. Simultaneously, based on the spatial feature components of the visual signal in the device-signal association optimization encoding vector, the acquisition viewing angle is determined to be 120 degrees horizontally and 80 degrees vertically, ensuring coverage of the entire screen display area. For auditory signal acquisition, since the device's speech recognition threshold is -35dB, meaning it can recognize voice commands even in low-noise environments, the decision model recommends setting the microphone sampling rate to 44.1kHz, adjusting the pickup gain to a medium level, and setting the pickup range to a 3-meter radius area centered on the device, effectively acquiring voice commands while minimizing environmental noise interference.

[0084] The tactile signal acquisition strategy is generated based on the device's touch sensitivity parameter, which, after normalization, is 0.85, indicating that the device is relatively sensitive to touch pressure. Therefore, the decision model recommends setting the pressure sensor's sampling frequency to 100Hz, adjusting the sensing accuracy to a higher level, and ensuring the sensing area covers the entire touchscreen surface to accurately capture touch operations of varying pressure levels. After generating the above acquisition strategy, the signal acquisition control unit further needs to determine the acquisition range of each signal. Taking visual signals as an example, the acquisition viewing angle is converted into specific camera installation positions and angle parameters through coordinate mapping. For example, the camera is installed 1.5 meters directly in front of the device, with a horizontal lens deflection angle of 0 degrees and a vertical lens deflection angle of 5 degrees, ensuring that the acquired image completely covers the screen.

[0085] For auditory signals, the microphone installation position must meet the requirement of a 3-meter radius pickup range, based on the pickup range setting. Two microphones can be evenly distributed around the device, located at the upper left and upper right of the device, 2.5 meters from the center of the device, to achieve omnidirectional voice acquisition. The acquisition range of tactile signals directly corresponds to the touch area of ​​the screen. By configuring the scanning range of the pressure sensor array, it is ensured that every pixel of the screen is covered. For example, for a 10.1-inch screen, the pressure sensor array is set to a 100×100 sensing point matrix, with each sensing point corresponding to a small area on the screen, thereby achieving pressure distribution acquisition across the entire touch surface.

[0086] In the test scenario construction module, the scenario parameter extraction unit extracts scenario parameters from the interaction feature fusion data. Assume the interaction feature fusion data includes information such as ambient light intensity of 500 lux, background noise level of 40 dB, and operating temperature range of 20-30 degrees Celsius. Based on these parameters, the scenario template matching unit searches for a matching basic test template in a preset scenario template library. For example, an ambient light intensity of 500 lux represents medium lighting conditions, a background noise level of 40 dB represents a quiet environment, and an operating temperature of 20-30 degrees Celsius represents a normal operating temperature. Based on this, the basic test template numbered SC-003 is matched, which corresponds to a normal functional test scenario under medium lighting and quiet conditions.

[0087] The scene detail filling unit fills in parameters to the basic test template based on interaction feature fusion data. This interaction feature fusion data may include information such as an average touch response latency of 80 milliseconds, a voice command recognition accuracy of 95%, and a frame rate of 60 frames per second in the interface switching smoothness parameter. These parameters are filled into the corresponding fields of the basic test template; for example, 80 milliseconds for touch response latency, 95% for voice command recognition accuracy, and 60 frames per second for interface switching smoothness, thus forming a complete standardized test scenario. Specifically, this scenario includes: conducting touch response latency tests under 500 lux illumination and 40 dB noise conditions, recording response times under different pressure touch operations; conducting voice command recognition tests, verifying recognition accuracy using command sets containing different keywords; and conducting interface switching tests, monitoring frame rate changes when switching between different functional modules to ensure a smoothness standard of 60 frames per second.

[0088] Throughout the entire implementation, the signal acquisition control unit generates targeted acquisition strategies and ranges through device-signal association optimization coding vectors, ensuring that the acquired multimodal signals accurately reflect the device characteristics. The test scenario construction module extracts scenario parameters and fills in details based on interaction feature fusion data, generating standardized test scenarios that conform to the actual application environment. Taking a specific device model and parameters as an example, from the generation of signal acquisition strategies and the determination of acquisition ranges to the extraction of test scenario parameters, template matching, and detail filling, each step closely revolves around the actual characteristics of the device and the interaction data, making the entire testing process more targeted and effective, and providing reliable scenario and data support for subsequent interaction performance verification tests. Through this specific implementation method, a complete process from device parameters to signal acquisition control and then to test scenario construction is realized, ensuring that the test system can accurately evaluate the interaction performance of the cockpit display and control equipment.

[0089] Example 5: The fusion feature processing module includes a visual feature extraction subunit, an auditory feature extraction subunit, a tactile feature extraction subunit, a multimodal feature alignment unit, and a feature fusion unit. Taking a cockpit display and control device test as an example, the visual feature extraction subunit processes the visual signal of the screen display captured by the camera. Assuming the acquired original visual signal is a 1920×1080 pixel RGB image, the edge detection algorithm is first used to identify the outlines of controls in the screen interface, such as the edge features of interactive elements like buttons and sliders. The Canny operator is used to extract the edges of the image, highlighting the boundary information of the interface elements. Next, a color space conversion is performed, converting the RGB image to the HSV color space, separating the hue, saturation, and brightness components to facilitate the extraction of color features of the interface elements. For example, the hue range of the red button is extracted to form the color feature components in the visual feature vector, ultimately obtaining a multidimensional visual feature vector containing edge outlines and color information.

[0090] The auditory feature extraction subunit processes the auditory signals of voice commands captured by the microphone. Assuming the acquired voice signal is audio data with a sampling rate of 44.1 kHz, frequency analysis is first performed. The time-domain signal is converted to a frequency-domain signal using a Fast Fourier Transform (FFT) to extract characteristic frequency components from the voice command, such as the typical frequency ranges of initials and finals in Mandarin Chinese. Simultaneously, time-domain waveform extraction is performed to identify the start and end points of the voice signal and calculate parameters such as the duration and fundamental frequency. For example, for the voice command "Open navigation," the energy distribution characteristics of 200-800 Hz in the frequency domain and the waveform characteristics of the command lasting 1.5 seconds in the time domain are extracted to form an auditory feature vector containing both frequency and time-domain information.

[0091] The tactile feature extraction subunit processes the tactile signals from touch operations acquired by the pressure sensor. Assuming the touch signals acquired by the pressure sensor array are 100×100 pressure distribution data, pressure distribution recognition is first performed to determine the location and pressure magnitude of the touch point. For example, a touch operation at screen coordinates (200, 300) generates a pressure value of 0.5N. After Gaussian smoothing to eliminate noise, a clear pressure distribution matrix is ​​obtained. Simultaneously, contact duration statistics are performed, recording the press and release times of the touch operation and calculating the contact duration, for example, 300 milliseconds. The pressure distribution and contact duration information are quantized and encoded to form a tactile feature vector, which contains feature components such as touch location, pressure value, and contact duration.

[0092] The multimodal feature alignment unit performs timestamp synchronization and dimensionality unification processing on visual, auditory, and tactile feature vectors. For example, in a certain interactive operation, the acquisition timestamp of the visual feature vector is 10.234 seconds, the timestamp of the auditory feature vector is 10.236 seconds, and the timestamp of the tactile feature vector is 10.235 seconds. A linear interpolation method is used to align the timestamps of the auditory and tactile feature vectors to the visual feature vector's timestamp of 10.234 seconds, ensuring consistency in the temporal dimension of the multimodal features. Regarding dimensionality unification, assuming the visual feature vector has a dimension of 128, the auditory feature vector has a dimension of 64, and the tactile feature vector has a dimension of 32, the dimensions of the auditory and tactile feature vectors are expanded to 128 through zero-padding or linear transformation, making all three have the same dimension, forming an aligned feature vector set.

[0093] The feature fusion unit performs a weighted summation of the aligned feature vector set to obtain interactive feature fusion data. Assume the aligned visual feature vectors are... The auditory feature vector is The tactile feature vector is Based on the association weights of device parameters and signal types, the weights for visual features are determined to be 0.5, auditory features 0.3, and tactile features 0.2. The three feature vectors are multiplied by their respective weights and then summed to obtain the interaction feature fusion data. This fused data integrates multimodal interaction features. For example, in the interaction of "touching the screen and issuing voice commands", the fused data simultaneously includes the visual features of screen interface changes, the auditory features of voice commands, and the tactile features of touch pressure, comprehensively representing the multimodal information of this interaction.

[0094] Taking a specific interaction scenario as an example, when a tester performs the operation of "clicking the volume adjustment button on the screen and saying 'volume up'" on the cockpit display and control device, the visual feature extraction subunit captures the edge features and color features of the interface changes before and after the button click, forming a visual feature vector; the auditory feature extraction subunit extracts the frequency and temporal features of the "volume up" voice command, forming an auditory feature vector; and the tactile feature extraction subunit records the pressure distribution and contact duration when the button is clicked, forming a tactile feature vector. The multimodal feature alignment unit synchronizes the timestamps of the three feature vectors to the same moment the operation occurs and unifies the dimensions. The feature fusion unit performs a weighted summation of the three feature vectors according to preset weights to obtain interactive feature fusion data containing visual, auditory, and tactile features. This data comprehensively describes the multimodal information of this interaction operation, providing rich feature inputs for subsequent test scenario construction and interaction performance verification.

[0095] The entire fusion feature processing module, through the collaborative work of its sub-units, achieves feature extraction, alignment, and fusion of multimodal signals, converting raw visual, auditory, and tactile signals into representative interactive feature fusion data. Based on specific operational examples, and starting from the processing flow of each sub-unit, the module details the conversion process from raw signals to fused features, ensuring the accuracy and comprehensiveness of feature fusion. This provides high-quality feature data support for subsequent functional modules of the cockpit display and control interaction test system, enabling the system to more comprehensively analyze and evaluate the interactive performance of the equipment.

[0096] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0097] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A cockpit display and control interaction test system based on multimodal fusion, characterized in that, include: The interactive data input module is used to obtain the basic parameters of the cockpit display and control device under test input by the user. The basic parameters include device model, screen size, touch sensitivity and voice recognition threshold. The multimodal signal acquisition module is used to acquire visual, auditory, and tactile signals from the cockpit environment based on the basic parameters of the cockpit display and control device under test. The feature fusion processing module is used to perform multimodal feature fusion processing on the visual signals, auditory signals and tactile signals to obtain interactive feature fusion data; The test scenario construction module is used to generate standardized test scenarios based on the interaction feature fusion data. The interactive performance verification module is used to input the standardized test scenario into the cockpit display and control device under test to perform interactive performance verification test, and transmit the verification test results to the display terminal for presentation. The multimodal signal acquisition module includes: An environmental perception unit is used to extract configuration information of visual signal acquisition devices, auditory signal acquisition devices, and tactile signal acquisition devices from the sensor array. The signal type classification unit is used to extract the functional descriptions of each signal acquisition device in the configuration information to obtain a set of signal type descriptions. The signal feature encoding unit is used to perform feature encoding on the basic parameters of the cockpit display and control device under test and each signal type description in the set of signal type descriptions to obtain a set of device parameter encoding vectors and signal type encoding vectors; The multimodal correlation analysis unit includes a signal type coding feature aggregation subunit and a cross-domain cooperative coding subunit, which are used to perform multimodal correlation analysis on the set of device parameter coding vectors and signal type coding vectors to obtain device-signal correlation optimized coding vectors; The signal type encoding feature aggregation subunit is used to receive the device parameter encoding vector and the signal type encoding vector set, perform feature aggregation processing on the signal type encoding vector set, decompose each signal type encoding vector at different semantic levels through hierarchical mapping, capture its feature representation at different dimensions of the basic function layer and application scenario layer, and then aggregate these hierarchical features to form a guiding template that can characterize the core features of the entire signal type set. The cross-domain collaborative coding subunit is used to perform cross-domain collaborative coding on the set of device parameter coding vectors and signal type coding vectors based on the guiding template. Through matrix operations and feature transformations, the device parameter features and signal type features are correlated and mapped in the same feature space to generate a device-signal correlation optimization coding vector that can reflect the matching relationship between the device and the signal. This vector contains the correlation weight information between the device characteristics and the signal acquisition requirements. The signal acquisition and control unit is used to generate acquisition strategies for the visual, auditory, and tactile signals based on the device-signal correlation optimization encoding vector.

2. The cockpit display and control interaction test system based on multimodal fusion according to claim 1, characterized in that, The signal type encoding feature aggregation subunit includes: The signal type coding hierarchical mapping second-level subunit is used to perform hierarchical mapping processing on each signal type coding vector in the set of signal type coding vectors using a mapping matrix to obtain a set of signal type coding vectors after hierarchical transformation. The signal type encoding matrix arrangement second-level sub-unit is used to arrange the set of signal type encoding vectors after the hierarchical transformation into a matrix to obtain the signal type encoding feature matrix; The signal type coding matrix extremum extraction secondary sub-unit is used to extract the extrema of the signal type coding vector after each level transformation in the signal type coding feature matrix to obtain the key vector of the signal type coding feature matrix as the guiding template.

3. The cockpit display and control interaction test system based on multimodal fusion according to claim 2, characterized in that, The signal type coding level mapping second-level subunit includes: The signal type encoding vector is multiplied by the mapping matrix and then added to the mapping bias vector positionally to obtain the signal type encoding vector after hierarchical transformation.

4. The cockpit display and control interaction test system based on multimodal fusion according to claim 3, characterized in that, The cross-domain collaborative coding subunit includes: The device parameter hierarchical mapping second-level subunit is used to perform hierarchical mapping processing on the device parameter encoding vector using a query matrix and a value matrix to obtain a device parameter query vector and a device parameter value vector. The template-guided heterogeneous conversion coding secondary subunit is used to combine the device parameter query vector, the device parameter value vector, the signal type coding vector after transformation at each level in the signal type coding feature matrix, and the guiding template input into the template-guided heterogeneous conversion structure to obtain a sequence of device-signal cross-domain collaborative coding vectors. The position mean calculation subunit is used to calculate the position mean vector of the sequence of device-signal cross-domain cooperative coding vectors to obtain the device-signal association optimized coding vector.

5. The cockpit display and control interaction test system based on multimodal fusion according to claim 4, characterized in that, The template-guided heterogeneous conversion coding secondary subunit includes: After calculating the product between the device parameter query vector and the transpose of the signal type encoding vector after hierarchical transformation, the resulting device-signal association feature matrix is ​​divided by the magnitude of the key vector of the signal type encoding feature matrix according to its position to obtain the device-signal association weight matrix. The device-signal association weight matrix is ​​input into a normalization function for processing to obtain the device-signal association normalized weight matrix; After multiplying the device-signal association normalized weight matrix with the key vector of the signal type coding feature matrix, the resulting feature vector is multiplied by the device parameter value vector at position to obtain the device-signal cross-domain collaborative coding vector.

6. The cockpit display and control interaction test system based on multimodal fusion according to claim 5, characterized in that, The signal acquisition and control unit includes: The device-signal association optimization encoding vector is input into the decision-maker-based signal acquisition strategy recommendation module to obtain the acquisition strategies for the visual, auditory, and tactile signals; Based on the acquisition strategy, the acquisition range of the visual, auditory, and tactile signals is determined.

7. The cockpit display and control interaction test system based on multimodal fusion according to claim 1, characterized in that, The test scenario construction module includes: The scene parameter extraction unit is used to extract scene parameters such as ambient light intensity, background noise decibel value, and operating temperature range from the interaction feature fusion data; A scene template matching unit is used to match a basic test template corresponding to the scene parameters from a preset scene template library; The scene detail filling unit is used to fill in the touch response latency, voice command recognition accuracy and interface switching smoothness parameters of the basic test template based on the interaction feature fusion data to obtain the standardized test scene.

8. The cockpit display and control interaction test system based on multimodal fusion according to claim 1, characterized in that, The fusion feature processing module includes: A visual feature extraction subunit is used to perform edge detection and color space conversion on the visual signal to obtain a visual feature vector; An auditory feature extraction unit is used to perform frequency analysis and time-domain waveform extraction on the auditory signal to obtain an auditory feature vector; The tactile feature extraction unit is used to identify the pressure distribution and count the contact duration of the tactile signal to obtain a tactile feature vector; A multimodal feature alignment unit is used to perform timestamp synchronization and dimension unification processing on the visual feature vector, auditory feature vector and tactile feature vector to obtain an aligned feature vector set; The feature fusion unit is used to perform weighted summation on the aligned feature vector set to obtain the interactive feature fusion data.

Citation Information

Patent Citations

  • Intelligent cabin man-machine interaction evaluation method, device, equipment and medium

    CN119917395A