Cabin display control interaction test system based on multi-modal fusion
By designing a cockpit display and control interaction test system based on multimodal fusion, the shortcomings of traditional test systems in multimodal signal processing, signal acquisition adaptability and test scenario construction are solved, and more efficient and accurate interactive performance testing is achieved.
Patent Information
- Application Number
- CN202511171292.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Traditional cockpit display and control interaction test systems have shortcomings in multimodal signal processing, signal acquisition adaptability, test scenario construction and test process automation. They cannot fully reflect the interactive performance of the equipment in actual use, resulting in inaccurate test results and low efficiency.
A cockpit display and control interaction test system based on multimodal fusion is designed, including an interactive data input module, a multimodal signal acquisition module, a fusion feature processing module, a test scenario construction module and an interactive performance verification module. By obtaining basic equipment parameters, collecting and processing visual, auditory and tactile signals, standardized test scenarios are generated, and interactive performance verification is performed.
It improves the quality and efficiency of signal acquisition, generates test scenarios that are closer to the actual environment, realizes the automation and intelligence of the test process, reduces manual operations and errors, and improves the accuracy and efficiency of the test.
Smart Images

Figure CN120653501A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cockpit display and control interaction testing, and in particular to a cockpit display and control interaction testing system based on multimodal fusion. Background Art
[0002] With the rapid development of aerospace, automotive, and other fields, cockpit display and control equipment, as a key interface for human-machine interaction, has a direct impact on the safety, convenience, and comfort of operation. However, traditional cockpit display and control interaction test systems have many shortcomings and cannot meet the testing requirements of modern complex cockpit environments.
[0003] Traditional test systems often collect and process signals from only a single modality, such as visual or auditory signals, while ignoring the synergistic effects of multimodal signals. This single-modality testing approach fails to fully reflect the interactive performance of cockpit display and control devices in actual use, resulting in significant limitations in test results. For example, during actual driving, drivers not only rely on vision to obtain information, but also rely on multiple senses such as hearing and touch to interact. Single-modality testing cannot simulate this complex interactive scenario.
[0004] Furthermore, traditional test systems lack flexibility and adaptability in signal acquisition. Different models of cockpit display and control devices have varying basic parameters, such as device model, screen size, touch sensitivity, and voice recognition threshold. Traditional systems struggle to automatically adjust signal acquisition strategies based on these parameters, resulting in poor signal quality and impacting the accuracy of test results. For example, a traditional system might use the same visual signal acquisition range for devices with varying screen sizes, failing to accurately capture all interactive information on the screen.
[0005] When it comes to test scenario construction, traditional systems typically use fixed test scenarios and are unable to dynamically adjust based on actual interaction characteristics and data fusion. These fixed test scenarios differ significantly from the actual cockpit environment, making the test results inaccurately reflect the device's performance in real-world use. For example, factors like lighting intensity, background noise decibels, and operating temperature in the actual cockpit environment constantly change, and traditional systems are unable to simulate the impact of these dynamic environmental factors on device interaction performance.
[0006] Traditional test systems lack effective algorithms and models for multimodal feature fusion processing, making it difficult to deeply fuse signals from multiple modalities such as vision, hearing, and touch, making it impossible to extract comprehensive and accurate interactive feature fusion data. This makes it impossible to extract comprehensive and accurate interactive feature fusion data. This makes the subsequent test scenario construction and interactive performance verification testing lack a reliable basis, resulting in low credibility of the test results. The interactive performance verification test process of traditional test systems lacks intelligence and automation, and requires a large amount of manual operation and data processing, which is not only inefficient but also prone to human errors. For example, after the test scenario is input into the device to be tested, manual observation and recording are required to verify the test results. This method is not only time-consuming and labor-intensive, but may also lead to inaccurate recorded results due to human factors.
[0007] Traditional cockpit display and control interaction test systems have obvious shortcomings in multimodal signal processing, signal acquisition adaptability, test scenario construction, feature fusion algorithm and test process automation. There is an urgent need for a cockpit display and control interaction test system based on multimodal fusion to solve these problems and improve the accuracy, comprehensiveness and efficiency of cockpit display and control equipment interaction performance testing. Summary of the Invention
[0008] The purpose of the present invention is to provide a cockpit display and control interaction test system based on multimodal fusion to solve the problems raised in the above background technology.
[0009] To achieve the above objectives, the present invention provides the following technical solution: a cockpit display and control interaction test system based on multimodal fusion, the system comprising: An interactive data input module is used to obtain basic parameters of the cockpit display and control device to be tested input by the user, including device model, screen size, touch sensitivity, and voice recognition threshold; A multimodal signal acquisition module, configured to acquire visual, auditory, and tactile signals in the cockpit environment based on the basic parameters of the cockpit display and control device to be tested; a fusion feature processing module, configured to perform multimodal feature fusion processing on the visual signal, the auditory signal, and the tactile signal to obtain interactive feature fusion data; A test scenario construction module, configured to generate a standardized test scenario based on the interaction feature fusion data; The interactive performance verification module is used to input the standardized test scenario into the cockpit display and control device to be tested to perform an interactive performance verification test, and transmit the verification test results to a display terminal for presentation.
[0010] Preferably, the multimodal signal acquisition module includes: An environmental perception unit, configured to extract configuration information of a visual signal acquisition device, an auditory signal acquisition device, and a tactile signal acquisition device from a sensor array; a signal type classification unit, configured to extract a function description of each signal acquisition device in the configuration information to obtain a set of signal type descriptions; a signal feature encoding unit, configured to perform feature encoding on each of the basic parameters of the cockpit display and control device to be tested and each signal type description in the set of signal type descriptions to obtain a set of device parameter encoding vectors and signal type encoding vectors; a multimodal association analysis unit, configured to perform a multimodal association analysis on the set of the device parameter coding vector and the signal type coding vector to obtain a device-signal association optimized coding vector; A signal acquisition control unit is used to generate an acquisition strategy for the visual signal, the auditory signal and the tactile signal based on the device-signal association optimization coding vector.
[0011] Preferably, the multimodal association analysis unit includes: a signal type coding feature aggregation subunit, configured to perform feature aggregation processing based on hierarchical mapping on the set of signal type coding vectors to obtain a guidance template; The cross-domain collaborative coding subunit is used to perform cross-domain collaborative coding on the set of the device parameter coding vector and the signal type coding vector based on the guiding template to obtain the device-signal association optimized coding vector.
[0012] Preferably, the signal type encoding feature aggregation subunit includes: a signal type coding hierarchical mapping secondary subunit, configured to perform hierarchical mapping processing on each signal type coding vector in the set of signal type coding vectors using a mapping matrix to obtain a set of hierarchically transformed signal type coding vectors; A signal type coding matrix arrangement secondary subunit is used to perform matrix arrangement on the set of signal type coding vectors after the hierarchical transformation to obtain a signal type coding feature matrix; The signal type coding matrix extreme value extraction secondary sub-unit is used to extract the extreme values of the signal type coding vectors after each level transformation in the signal type coding feature matrix to obtain the signal type coding feature matrix key vector as the guiding template.
[0013] Preferably, the signal type coding level mapping secondary sub-unit includes: The signal type coding vector is point-multiplied by the mapping matrix and then positionally added to the mapping bias vector to obtain the signal type coding vector after hierarchical transformation.
[0014] Preferably, the cross-domain collaborative encoding subunit includes: a device parameter hierarchical mapping secondary subunit, configured to perform hierarchical mapping processing on the device parameter encoding vector using a query matrix and a value matrix to obtain a device parameter query vector and a device parameter value vector; a template-guided heterogeneous conversion encoding secondary subunit, configured to input the device parameter query vector, the device parameter value vector, the transformed signal type coding vectors at each level in the signal type coding feature matrix, and the guiding template into a template-guided heterogeneous conversion structure to obtain a sequence of device-signal cross-domain collaborative coding vectors; The position mean calculation secondary subunit is used to calculate the position mean vector of the sequence of the device-signal cross-domain collaborative coding vector to obtain the device-signal association optimized coding vector.
[0015] Preferably, the template-guided isomerous conversion encoding secondary subunit includes: After calculating the product between the device parameter query vector and the transposed vector of the hierarchically transformed signal type coding vector, dividing the obtained device-signal association feature matrix by the modulus length of the key vector of the signal type coding feature matrix by position to obtain a device-signal association weight matrix; Inputting the device-signal association weight matrix into a normalization function for processing to obtain a device-signal association normalized weight matrix; After multiplying the device-signal association normalized weight matrix with the key vector of the signal type coding feature matrix, the obtained feature vector is multiplied by the device parameter value vector according to the position to obtain the device-signal cross-domain collaborative coding vector.
[0016] Preferably, the signal acquisition control unit includes: Inputting the device-signal association optimized coding vector into a signal acquisition strategy recommendation module based on a decision maker to obtain acquisition strategies for the visual signal, the auditory signal, and the tactile signal; Based on the acquisition strategy, the acquisition range of the visual signal, the auditory signal, and the tactile signal is determined.
[0017] Preferably, the test scenario construction module includes: A scene parameter extraction unit, configured to extract scene parameters of ambient light intensity, background noise decibel value, and operating temperature range from the interactive feature fusion data; A scene template matching unit, configured to match a basic test template corresponding to the scene parameters from a preset scene template library; A scene detail filling unit is used to fill in the details of the touch response delay, voice command recognition accuracy and interface switching smoothness parameters of the basic test template based on the interaction feature fusion data to obtain the standardized test scene.
[0018] Preferably, the fusion feature processing module includes: a visual feature extraction subunit, configured to perform edge detection and color space conversion on the visual signal to obtain a visual feature vector; an auditory feature extraction unit, configured to perform frequency analysis and time domain waveform extraction on the auditory signal to obtain an auditory feature vector; a tactile feature extraction unit, configured to perform pressure distribution recognition and contact duration statistics on the tactile signal to obtain a tactile feature vector; a multimodal feature alignment unit, configured to perform timestamp synchronization and dimension unification processing on the visual feature vector, the auditory feature vector, and the tactile feature vector to obtain an aligned feature vector set; A feature fusion unit is used to perform weighted sum processing on the aligned feature vector set to obtain the interactive feature fusion data.
[0019] Compared with the prior art, the present invention has the following beneficial effects: The system acquires the basic parameters of the cockpit display and control devices under test through an interactive data input module, providing an accurate basis for subsequent testing. Based on these basic parameters, the multimodal signal acquisition module extracts the configuration information of various signal acquisition devices from the sensor array, classifies and encodes the signal types, and uses multimodal correlation analysis to derive optimized device-signal correlation coding vectors. This in turn generates targeted acquisition strategies, ensuring more accurate and comprehensive visual, auditory, and tactile signals, improving both the quality and efficiency of signal acquisition.
[0020] The fusion feature processing module processes the collected multimodal signals, extracts the feature vectors of each modal through operations such as edge detection, frequency analysis, and pressure distribution identification, and then performs alignment processing such as timestamp synchronization and dimension unification. Finally, the weighted summation is performed to obtain the interactive feature fusion data, realizing the deep fusion of multimodal features and providing reliable data support for the subsequent test scenario construction.
[0021] The test scenario construction module extracts scene parameters such as ambient light intensity and background noise decibel value from the interactive feature fusion data, matches the basic test templates in the preset scene template library, and fills in details for parameters such as touch response delay to generate standardized test scenarios. This makes the test scenarios closer to the actual cockpit environment, improving the authenticity and effectiveness of the test.
[0022] The interactive performance verification module inputs standardized test scenarios into the device to be tested to perform verification tests, and transmits the results to the display terminal for presentation, thus realizing the automation and intelligence of the test process, reducing manual operations and human errors, and improving test efficiency and the accuracy of results.
[0023] This system utilizes multimodal fusion technology to comprehensively consider the synergistic effects of multiple modal signals, including visual, auditory, and tactile, enabling a more comprehensive reflection of the interactive performance of cockpit display and control equipment. Furthermore, the system automatically adjusts signal acquisition strategies and constructs test scenarios based on the device's basic parameters, enhancing its adaptability and flexibility. Furthermore, the system's various modules collaborate to form a complete testing process. From signal acquisition to feature fusion, scenario construction, and performance verification, each step has been meticulously designed and optimized, ensuring the efficiency and reliability of the entire testing process and providing strong technical support for the development and optimization of cockpit display and control equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a working principle diagram of the cockpit display and control interaction test system based on multimodal fusion according to the present invention; Figure 2 This is the workflow diagram of the multimodal signal acquisition module; Figure 3 A workflow diagram for cross-domain collaborative coding subunits; Figure 4 Detailed diagram encoding for template-guided heterogeneous transformation; Figure 5 This is the workflow diagram of the fusion feature processing module. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0026] See also Figure 1-Figure 5 The present invention relates to a cockpit display and control interactive test system based on multimodal fusion, which includes: an interactive data input module, a multimodal signal acquisition module, a fusion feature processing module, a test scenario construction module, and an interactive performance verification module. The specific implementation method is as follows: The interactive data input module obtains the basic parameters of the cockpit display and control device to be tested input by the user. The basic parameters include device model, screen size, touch sensitivity and voice recognition threshold. The multimodal signal acquisition module collects visual signals, auditory signals and tactile signals in the cockpit environment based on the basic parameters of the cockpit display and control device to be tested. The fusion feature processing module performs multimodal feature fusion processing on the visual signals, auditory signals and tactile signals to obtain interactive feature fusion data. The test scenario construction module generates standardized test scenarios based on the interactive feature fusion data. The interactive performance verification module inputs the standardized test scenarios into the cockpit display and control device to be tested to perform interactive performance verification tests, and transmits the verification test results to the display terminal for presentation.
[0027] Example 1: A multimodal signal acquisition module includes an environmental perception unit, a signal type classification unit, a signal feature encoding unit, a multimodal correlation analysis unit, and a signal acquisition control unit. The environmental perception unit operates by establishing a data exchange channel with the sensor array, parsing the raw configuration information data stream output by the sensor array through a preset protocol, and filtering out configuration parameters related to the visual signal acquisition device, the auditory signal acquisition device, and the tactile signal acquisition device. These parameters, such as physical properties such as camera resolution, microphone sampling rate, and pressure sensor array distribution, are then structured and stored to form a device configuration information database.
[0028] After obtaining the device configuration information output by the environmental perception unit, the signal type classification unit performs semantic parsing on the functional description fields of each signal acquisition device. Using keyword extraction techniques from natural language processing, it identifies core functional keywords such as "image capture," "audio recording," and "pressure sensing" from the functional description text. Devices with the same functional keywords are grouped into the same signal type, thereby constructing a signal type description set encompassing categories such as visual signal type, auditory signal type, and tactile signal type. Each element of the set corresponds to an abstract functional description of a specific device.
[0029] The signal feature encoding unit's processing is divided into two parts. The basic parameters of the cockpit display and control device to be tested, including device model, screen size, touch sensitivity, and voice recognition threshold, are first standardized. The device model is converted into a one-hot encoded vector. Numerical parameters such as screen size, touch sensitivity, and voice recognition threshold are normalized and mapped to the [0, 1] interval. Finally, a multidimensional device parameter encoding vector is combined. For each signal type description in the signal type description set, a word embedding model is used to convert the text description into a low-dimensional dense semantic vector. Model training captures the semantic associations between different signal types and generates the corresponding signal type encoding vector. All signal type encoding vectors constitute a set.
[0030] After the multimodal association analysis unit receives the device parameter coding vector and the signal type coding vector set, it first performs feature aggregation processing on the signal type coding vector set. Through hierarchical mapping, each signal type coding vector is decomposed at different semantic levels to capture its feature representations in different dimensions such as the basic function layer and the application scenario layer. These hierarchical features are then aggregated to form a guiding template that can characterize the core features of the entire signal type set. Afterwards, based on the guiding template, the device parameter coding vector and the signal type coding vector set are cross-domain collaboratively encoded. Through matrix operations and feature transformations, the device parameter features and the signal type features are associated and mapped in the same feature space to generate a device-signal association optimization coding vector that can reflect the matching relationship between the device and the signal. This vector contains the association weight information between the device characteristics and the signal acquisition requirements.
[0031] After obtaining the optimized device-signal association encoding vector, the signal acquisition control unit inputs it into the signal acquisition strategy recommendation module based on the decision maker. This module incorporates a pretrained decision model. Based on the device characteristics and signal association weights in the encoding vector, combined with a pre-set signal acquisition rule library, the model generates acquisition strategies for visual, auditory, and tactile signals through logical reasoning and weight calculation. These acquisition strategies include parameter settings such as signal acquisition frequency, sampling accuracy, and trigger conditions. Based on the generated acquisition strategy, the acquisition scope for each signal type is further determined, such as the viewing angle for visual signals, the pickup range for auditory signals, and the sensing area for tactile signals. Through coordinate mapping and parameter configuration, the abstract acquisition strategy is converted into specific device control instructions and sent to the corresponding signal acquisition device, achieving precise control of the signal acquisition process. Through the coordinated operation of various units, the entire multimodal signal acquisition module achieves intelligent and precise acquisition of multimodal signals based on basic device parameters, providing high-quality raw data for subsequent multimodal feature fusion and test scenario construction.
[0032] Example 2: The multimodal association analysis unit includes a signal type coding feature aggregation subunit and a cross-domain collaborative coding subunit. The signal type coding feature aggregation subunit performs a feature aggregation process based on hierarchical mapping on the set of signal type coding vectors to obtain a guidance template. The workflow of the signal type coding hierarchical mapping secondary subunit is as follows: for each signal type coding vector in the signal type coding vector set, , using the preset mapping matrix Perform hierarchical mapping processing. The specific operation is to multiply the signal type encoding vector with the mapping matrix and then multiply it with the mapping bias vector Perform positional addition processing to obtain the signal type coding vector after hierarchical transformation , its mathematical expression is: ; Where, Indicates the signal type encoding vectors, whose dimensions are , contains semantic feature information of signal type; Dimension The mapping matrix is used to map the original signal type encoding vector to the new feature space, where is the feature dimension after mapping; Dimension The mapping bias vector is used to adjust the feature offset after mapping; is the signal type encoding vector after hierarchical transformation, with dimension Through this transformation, the decomposition and reconstruction of signal type features at different levels are achieved.
[0033] The signal type coding matrix arrangement secondary sub-unit arranges the set of signal type coding vectors after the hierarchical transformation into a matrix. Assuming that the hierarchical transformation results in Signal type encoding vector , each vector dimension is , then these vectors are arranged in rows to form a Signal type encoding feature matrix ,The matrix stores the feature information of the signal type encoding vector after all levels of transformation in a structured form, which is convenient for subsequent feature extraction and analysis.
[0034] Signal type encoding matrix extreme value extraction Secondary sub-unit signal type encoding feature matrix Processing is performed to extract the extreme values of the signal type coding vector after each level transformation in the matrix. Specifically, for the matrix For each column (i.e. each feature dimension), find the maximum and minimum values in the column to form a two-dimensional extreme value vector, and then combine the extreme value vectors of all columns in order to obtain a dimension of Signal type encoding feature matrix key vector ,This vector serves as a guiding template, condenses the key feature information in the signal type encoding feature matrix, and can characterize the core feature distribution of the entire set of signal types.
[0035] The cross-domain collaborative coding subunit performs cross-domain collaborative coding on the set of device parameter coding vectors and signal type coding vectors based on the guidance template to obtain the device-signal association optimized coding vector. It contains the basic parameter characteristics of the cockpit display and control equipment to be tested, such as device model, screen size and other information, and forms a multi-dimensional feature vector after encoding processing. In the cross-domain collaborative encoding process, it is first necessary to associate and map the device parameter encoding vector with the signal type encoding feature matrix in the same feature space, and use the guiding template as a bridge to achieve cross-domain fusion of device parameter characteristics and signal type characteristics, so that the generated device-signal association optimization encoding vector can reflect the matching relationship between device characteristics and signal acquisition requirements, and provide key feature association basis for subsequent signal acquisition strategy generation. The entire multimodal association analysis unit realizes the effective fusion and association analysis of device parameter characteristics and signal type characteristics through the collaborative work of the signal type encoding feature aggregation subunit and the cross-domain collaborative encoding subunit, providing important feature processing support for the precise control of the multimodal signal acquisition module, ensuring that the acquisition strategy can better adapt to the characteristics of the cockpit display and control equipment to be tested, and improving the quality and efficiency of signal acquisition. During the processing process, through operations such as hierarchical mapping, matrix arrangement, and extreme value extraction, key features are gradually extracted from the signal type coding vector set, and cross-domain collaborative encoding is performed with the device parameter features to form an optimized coding vector containing device-signal association information, laying a solid feature foundation for the subsequent processing links of the entire test system.
[0036] Example 3: The cross-domain collaborative coding subunit includes a device parameter hierarchical mapping secondary subunit, a template-guided heterogeneous conversion coding secondary subunit, and a position mean calculation secondary subunit. The working principle of the device parameter hierarchical mapping secondary subunit is to use the query matrix and the value matrix to perform hierarchical mapping processing on the device parameter coding vector. It is a multidimensional vector obtained by standardizing and encoding the basic parameters of the cockpit display and control device to be tested (such as device model, screen size, touch sensitivity, voice recognition threshold, etc.). Its dimension is Query Matrix The dimension is , value matrix The dimension is ,in and are the dimensions of the mapped query vector and value vector respectively. Through matrix multiplication, the device parameter encoding vector With the query matrix Multiply to get the device parameter query vector , whose expression is ; Similarly, the device parameter encoding vector and value matrix Multiply to get the device parameter value vector , the expression is This process realizes the decomposition of device parameter features in different feature subspaces, providing a basis for subsequent cross-domain interaction with signal type features.
[0037] The template guides the heterogeneous conversion encoding secondary subunit to convert the device parameter query vector , device parameter value vector , signal type coding vector after each level transformation in the signal type coding feature matrix ( , is the number of signal types) and the guide template Input is a template-guided heterogeneous transformation structure. First, the device parameter query vector is calculated and the signal type encoding vector after level transformation The transposed vector of The product between them gives the device-signal correlation feature matrix , whose dimensions are ( is the dimension of the signal type encoding vector after hierarchical transformation), the expression is Next, the device-signal correlation feature matrix Divide the key vector of the feature matrix by the signal type encoding by position Length of mold , get the device-signal association weight matrix ,Right now ,in" " represents element-wise division. This step normalizes the associated features by guiding the module length of the template, enhancing the stability and comparability of the features.
[0038] Then, the device-signal association weight matrix The input is processed into a normalization function (such as the softmax function) to obtain the device-signal association normalized weight matrix The role of the softmax function is to map the elements in the weight matrix to The interval is such that the sum of the elements in the same row is 1, thereby obtaining the relative importance weight of each feature dimension. Specifically, for the matrix Each element in , whose normalized elements are .
[0039] Next, the device-signal association normalized weight matrix Key vector of feature matrix encoded with signal type Multiply them to get the eigenvector ,Right now The feature vector combines the associated weight information of the device parameter query vector and the signal type feature, and then combines it with the device parameter value vector Perform point multiplication by position to obtain the device-signal cross-domain collaborative coding vector , the expression is ,in" " represents element-wise multiplication. Through this series of operations, the cross-domain fusion of device parameter features and signal type features is achieved, so that the generated encoding vector can simultaneously reflect the semantic association between device characteristics and signal types.
[0040] Position mean calculation of the sequence of secondary subunits for device-signal cross-domain collaborative encoding vectors For each encoding vector , whose dimensions are (consistent with the dimension of the device parameter value vector), the corresponding dimensions of all encoding vectors are averaged by position. Specifically, for the dimensions ( ), calculate all The sum of the elements of the encoded vectors in this dimension, divided by , and get the mean vector No. Elements ,in For the Encoded vector No. elements. The final mean vector It is the device-signal association optimization encoding vector, whose dimension is ,This vector integrates the cross-domain collaborative encoding results of all signal types and ,device parameters, condenses the overall correlation features between the device and the ,signal, and provides a comprehensive and representative feature representation for subsequent ,signal acquisition and control.
[0041] The entire cross-domain collaborative encoding subunit decomposes the device parameter encoding vector into a query vector and a value vector through hierarchical mapping. Using a guiding template as a bridge, the device parameter features are correlated with the signal type features, weighted normalized, and fused within a heterogeneous transformation structure. Ultimately, the optimized encoding vector is obtained through positional mean calculation. This process fully considers the semantic association between device characteristics and signal types. Through matrix and vector operations, features from different modalities and sources are mapped into the same feature space for collaborative processing. This ensures that the generated encoding vector accurately reflects the matching relationship between the device and the signal, providing critical feature support for the generation of multimodal signal acquisition strategies. This allows the signal acquisition process to more accurately adapt to the parameter characteristics of the cockpit display and control device under test, improving the relevance and effectiveness of signal acquisition. Throughout the processing process, the mathematical operations in each step strictly adhere to the basic rules of linear algebra. Through reasonable matrix dimension design and operational logic, the rationality and accuracy of feature transformation and fusion are guaranteed, providing high-quality feature input for subsequent modules of the entire test system.
[0042] Example 4: The workflow of the signal acquisition control unit is to input the device-signal association optimization coding vector into the signal acquisition strategy recommendation module based on the decision maker. The module has a preset decision model built according to the association between device parameters and signal types. Taking a certain model of cockpit display and control equipment as an example, the screen size in its device parameter coding vector is 10.1 inches, the touch sensitivity parameter is 0.85 after normalization, the voice recognition threshold parameter is -35dB, the visual signal association weight component in the device-signal association optimization coding vector is 0.42, the auditory signal association weight component is 0.38, and the tactile signal association weight component is 0.2. Based on these weight components and the preset acquisition rules, the decision maker determines that the priority of visual signal acquisition is higher than that of auditory and tactile signals.
[0043] When generating a visual signal acquisition strategy, considering the large screen size and high touch sensitivity of the device, the decision model will recommend a higher image acquisition resolution, such as setting it to 1920×1080 pixels and setting the acquisition frame rate to 30 frames per second to ensure that the touch operation details on the screen can be clearly captured. At the same time, based on the device-signal association optimization coding vector, the spatial feature components of the visual signal are determined to be 120 degrees horizontally and 80 degrees vertically to ensure that the entire screen display area is covered. For auditory signal acquisition, since the device's voice recognition threshold is -35dB, which means that voice commands can be recognized even in a relatively low noise environment, the decision model will recommend that the microphone sampling rate be set to 44.1kHz, the pickup gain be adjusted to a medium level, and the pickup range be set to a 3-meter radius area centered on the device, which can effectively collect voice commands while reducing environmental noise interference.
[0044] The generation of the tactile signal acquisition strategy is based on the device's touch sensitivity parameter, which is normalized to 0.85, indicating that the device is relatively sensitive to touch pressure. Therefore, the decision model recommends setting the pressure sensor's sampling frequency to 100Hz, adjusting the sensing accuracy to a higher level, and ensuring that the sensing area covers the entire screen touch surface to accurately capture touch operations of varying pressure levels. After generating the above acquisition strategy, the signal acquisition control unit must further determine the acquisition range of each signal. Taking visual signals as an example, the acquisition viewing angle range is converted into specific camera installation positions and angle parameters through coordinate mapping. For example, the camera is installed 1.5 meters in front of the device, with a horizontal lens deflection angle of 0 degrees and a vertical lens deflection angle of 5 degrees to ensure that the captured image fully covers the screen.
[0045] For auditory signals, the microphones are positioned to meet a 3-meter radius based on the pickup range setting. Two microphones can be evenly distributed around the device, located on the upper left and upper right sides, 2.5 meters from the center, to achieve omnidirectional voice capture. The tactile signal acquisition range directly corresponds to the touch area of the screen. By configuring the scanning range of the pressure sensor array, every pixel on the screen is covered. For example, for a 10.1-inch screen, the pressure sensor array is configured as a 100×100 sensing point matrix, with each sensing point corresponding to a small area on the screen, thus capturing the pressure distribution across the entire touch surface.
[0046] In the test scene construction module, the scene parameter extraction unit extracts scene parameters from the interactive feature fusion data. Assume that the interactive feature fusion data contains information that the ambient light intensity is 500 lux, the background noise decibel value is 40 decibels, and the operating temperature range is 20-30 degrees Celsius. Based on these parameters, the scene template matching unit searches for a matching basic test template from the preset scene template library. For example, an ambient light intensity of 500 lux belongs to medium lighting conditions, a background noise of 40 decibels is a quiet environment, and an operating temperature of 20-30 degrees Celsius is a normal working temperature. Based on this, the basic test template numbered SC-003 is matched, which corresponds to a conventional functional test scene in a medium lighting and quiet environment.
[0047] The scene detail filling unit fills in the parameters of the basic test template based on the interactive feature fusion data. The interactive feature fusion data may contain information such as an average measurement value of 80 milliseconds for touch response delay, 95% for voice command recognition accuracy, and 60 frames per second for interface switching smoothness parameters. These parameters are filled in the corresponding fields of the basic test template, for example, 80 milliseconds for touch response delay parameters, 95% for voice command recognition accuracy, and 60 frames per second for interface switching smoothness, thereby forming a complete standardized test scenario. This scenario specifically includes testing the touch response delay of the device in an environment of 500 lux lighting and 40 decibel noise, requiring the response time to be recorded under different pressure touch operations; performing a voice command recognition test, using a command set containing different keywords to verify the recognition accuracy; and performing an interface switching test, monitoring frame rate changes when switching between different functional modules to ensure that the smoothness meets the 60 frames per second standard.
[0048] During the working process of the entire embodiment, the signal acquisition control unit generates a targeted acquisition strategy and range through the device-signal association optimization coding vector to ensure that the collected multimodal signal can accurately reflect the characteristics of the device; the test scenario construction module extracts the scenario parameters and fills in the details based on the interaction feature fusion data to generate a standardized test scenario that conforms to the actual application environment. Taking the specific device model and parameters as an example, from the generation of signal acquisition strategy, determination of acquisition range, to parameter extraction, template matching and detail filling of test scenarios, each link is closely centered around the actual characteristics and interaction data of the device, making the entire test process more targeted and effective, and providing reliable scenarios and data support for subsequent interaction performance verification tests. Through such a specific implementation method, a complete process from device parameters to signal acquisition control to test scenario construction is realized, ensuring that the test system can accurately evaluate the interactive performance of cockpit display and control equipment.
[0049] Example 5: The fusion feature processing module includes a visual feature extraction subunit, an auditory feature extraction subunit, a tactile feature extraction subunit, a multimodal feature alignment unit, and a feature fusion unit. Taking the test of a certain type of cockpit display and control equipment as an example, the visual feature extraction subunit processes the visual signal of the screen display image captured by the camera. Assuming that the original visual signal collected is an RGB image of 1920×1080 pixels, the edge detection algorithm is first used to identify the outline of the controls in the screen interface, such as the edge features of interactive elements such as buttons and sliders. The Canny operator is used to extract the edges of the image to highlight the boundary information of the interface elements. Then, a color space conversion is performed to convert the RGB image into the HSV color space, and the hue, saturation, and brightness components are separated to facilitate the extraction of the color features of the interface elements. For example, the hue value range of the red button is extracted to form the color feature component in the visual feature vector, and finally a multidimensional visual feature vector containing edge contours and color information is obtained.
[0050] The auditory feature extraction subunit processes the voice command auditory signal collected by the microphone. Assuming that the collected voice signal is audio data with a sampling rate of 44.1kHz, frequency analysis is first performed. The time domain signal is converted into a frequency domain signal through fast Fourier transform (FFT), and the characteristic frequency components in the voice command are extracted, such as the typical frequency range of initials and finals in Mandarin Chinese. At the same time, time domain waveform extraction is performed to identify the starting and ending points of the voice signal and calculate the duration, fundamental frequency and other parameters of the voice. For example, for the voice command "open navigation", the energy distribution characteristics of 200-800Hz in the frequency domain and the waveform characteristics of the command lasting 1.5 seconds in the time domain are extracted to form an auditory feature vector containing frequency and time domain information.
[0051] The tactile feature extraction subunit processes the tactile signals of touch operations collected by the pressure sensor. Assuming that the touch signal collected by the pressure sensor array is 100×100 pressure distribution data, the pressure distribution is first identified to determine the position and pressure of the touch point. For example, the pressure value generated by a touch operation at the screen coordinate (200,300) is 0.5N. After eliminating the noise through Gaussian smoothing, a clear pressure distribution matrix is obtained. At the same time, contact duration statistics are performed to record the press time and release time of the touch operation and calculate the contact duration. For example, the contact duration of the touch operation is 300 milliseconds. The pressure distribution and contact duration information are quantized and encoded to form a tactile feature vector, which contains characteristic components such as touch position, pressure value and contact duration.
[0052] The multimodal feature alignment unit synchronizes the timestamps and unifies the dimensions of the visual feature vectors, auditory feature vectors, and tactile feature vectors. For example, in a certain interactive operation, the acquisition timestamp of the visual feature vector is 10.234 seconds, the timestamp of the auditory feature vector is 10.236 seconds, and the timestamp of the tactile feature vector is 10.235 seconds. The timestamps of the auditory and tactile feature vectors are aligned to the timestamp of the visual feature vector 10.234 seconds through linear interpolation to ensure that the multimodal features are consistent in the time dimension. In terms of dimensional unification, assuming that the dimension of the visual feature vector is 128, the dimension of the auditory feature vector is 64, and the dimension of the tactile feature vector is 32, the dimensions of the auditory and tactile feature vectors are expanded to 128 by filling with zeros or linear transformation, so that the three have the same dimension, forming an aligned feature vector set.
[0053] The feature fusion unit performs weighted summation on the aligned feature vector set to obtain interactive feature fusion data. Assume that the aligned visual feature vector is , the auditory feature vector is , the tactile feature vector is According to the associated weights of device parameters and signal types, the weight of visual features is determined to be 0.5, the weight of auditory features is 0.3, and the weight of tactile features is 0.2. The three feature vectors are multiplied by the corresponding weights and then added together to obtain the interactive feature fusion data. The fused data integrates multimodal interaction features. For example, in the interactive operation of "touching the screen and issuing voice commands", the fused data simultaneously includes the visual features of the screen interface changes, the auditory features of the voice commands, and the tactile features of the touch pressure, comprehensively representing the multimodal information of the interactive operation.
[0054] Taking a specific interactive scenario as an example, when a tester performs the operation of "clicking the volume adjustment button on the screen and saying 'volume up'" on the cockpit display and control device, the visual feature extraction subunit captures the edge features and color features of the interface changes before and after the button is clicked, forming a visual feature vector; the auditory feature extraction subunit extracts the frequency and time domain features of the "volume up" voice command, forming an auditory feature vector; the tactile feature extraction subunit records the pressure distribution and contact duration when clicking the button, forming a tactile feature vector. The multimodal feature alignment unit synchronizes the timestamps of the three feature vectors to the same moment when the operation occurred and unifies the dimensions. The feature fusion unit performs a weighted summation of the three feature vectors according to preset weights to obtain interactive feature fusion data containing visual, auditory, and tactile features. This data comprehensively describes the multimodal information of this interactive operation and provides rich feature input for subsequent test scenario construction and interactive performance verification.
[0055] The entire fusion feature processing module, through the collaborative work of its subunits, achieves feature extraction, alignment, and fusion of multimodal signals, converting raw visual, auditory, and tactile signals into representative interactive feature fusion data. Using specific operational examples and starting from the processing flow of each subunit, this paper details the conversion process from raw signals to fused features, ensuring the accuracy and comprehensiveness of feature fusion. This provides high-quality feature data support for subsequent functional modules of the cockpit display and control interaction test system, enabling the system to more comprehensively analyze and evaluate the interactive performance of the device.
[0056] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0057] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. The cockpit display and control interaction test system based on multimodal fusion is characterized by: include: An interactive data input module is used to obtain basic parameters of the cockpit display and control device to be tested input by the user, including device model, screen size, touch sensitivity, and voice recognition threshold; A multimodal signal acquisition module, configured to acquire visual, auditory, and tactile signals in the cockpit environment based on the basic parameters of the cockpit display and control device to be tested; a fusion feature processing module, configured to perform multimodal feature fusion processing on the visual signal, the auditory signal, and the tactile signal to obtain interactive feature fusion data; A test scenario construction module, configured to generate a standardized test scenario based on the interaction feature fusion data; The interactive performance verification module is used to input the standardized test scenario into the cockpit display and control device to be tested to perform an interactive performance verification test, and transmit the verification test results to a display terminal for presentation.
2. The cockpit display and control interactive testing system based on multimodal fusion according to claim 1 is characterized in that: The multimodal signal acquisition module includes: An environmental perception unit, configured to extract configuration information of a visual signal acquisition device, an auditory signal acquisition device, and a tactile signal acquisition device from a sensor array; a signal type classification unit, configured to extract a function description of each signal acquisition device in the configuration information to obtain a set of signal type descriptions; a signal feature encoding unit, configured to perform feature encoding on each of the basic parameters of the cockpit display and control device to be tested and each signal type description in the set of signal type descriptions to obtain a set of device parameter encoding vectors and signal type encoding vectors; a multimodal association analysis unit, configured to perform a multimodal association analysis on the set of the device parameter coding vector and the signal type coding vector to obtain a device-signal association optimized coding vector; A signal acquisition control unit is used to generate an acquisition strategy for the visual signal, the auditory signal and the tactile signal based on the device-signal association optimization coding vector.
3. The cockpit display and control interactive testing system based on multimodal fusion according to claim 2 is characterized in that: The multimodal association analysis unit comprises: a signal type coding feature aggregation subunit, configured to perform feature aggregation processing based on hierarchical mapping on the set of signal type coding vectors to obtain a guidance template; The cross-domain collaborative coding subunit is used to perform cross-domain collaborative coding on the set of the device parameter coding vector and the signal type coding vector based on the guiding template to obtain the device-signal association optimized coding vector.
4. The cockpit display and control interactive testing system based on multimodal fusion according to claim 3 is characterized in that: The signal type encoding feature aggregation subunit includes: a signal type coding hierarchical mapping secondary subunit, configured to perform hierarchical mapping processing on each signal type coding vector in the set of signal type coding vectors using a mapping matrix to obtain a set of hierarchically transformed signal type coding vectors; A signal type coding matrix arrangement secondary subunit is used to perform matrix arrangement on the set of signal type coding vectors after the hierarchical transformation to obtain a signal type coding feature matrix; The signal type coding matrix extreme value extraction secondary sub-unit is used to extract the extreme values of the signal type coding vectors after each level transformation in the signal type coding feature matrix to obtain the signal type coding feature matrix key vector as the guiding template.
5. The cockpit display and control interactive testing system based on multimodal fusion according to claim 4 is characterized in that: The signal type coding level mapping secondary sub-unit includes: The signal type coding vector is point-multiplied by the mapping matrix and then positionally added to the mapping bias vector to obtain the signal type coding vector after hierarchical transformation.
6. The cockpit display and control interactive testing system based on multimodal fusion according to claim 5 is characterized in that: The cross-domain collaborative encoding subunit includes: a device parameter hierarchical mapping secondary subunit, configured to perform hierarchical mapping processing on the device parameter encoding vector using a query matrix and a value matrix to obtain a device parameter query vector and a device parameter value vector; a template-guided heterogeneous conversion encoding secondary subunit, configured to input the device parameter query vector, the device parameter value vector, the transformed signal type coding vectors at each level in the signal type coding feature matrix, and the guiding template into a template-guided heterogeneous conversion structure to obtain a sequence of device-signal cross-domain collaborative coding vectors; The position mean calculation secondary subunit is used to calculate the position mean vector of the sequence of the device-signal cross-domain collaborative coding vector to obtain the device-signal association optimized coding vector.
7. The cockpit display and control interactive testing system based on multimodal fusion according to claim 6 is characterized in that: The template-guided heterogeneous conversion encoding secondary subunit includes: After calculating the product between the device parameter query vector and the transposed vector of the hierarchically transformed signal type coding vector, dividing the obtained device-signal association feature matrix by the modulus length of the key vector of the signal type coding feature matrix by position to obtain a device-signal association weight matrix; Inputting the device-signal association weight matrix into a normalization function for processing to obtain a device-signal association normalized weight matrix; After multiplying the device-signal association normalized weight matrix with the key vector of the signal type coding feature matrix, the obtained feature vector is multiplied by the device parameter value vector according to the position to obtain the device-signal cross-domain collaborative coding vector.
8. The cockpit display and control interactive testing system based on multimodal fusion according to claim 7 is characterized in that: The signal acquisition control unit includes: Inputting the device-signal association optimized coding vector into a signal acquisition strategy recommendation module based on a decision maker to obtain acquisition strategies for the visual signal, the auditory signal, and the tactile signal; Based on the acquisition strategy, the acquisition range of the visual signal, the auditory signal, and the tactile signal is determined.
9. The cockpit display and control interactive testing system based on multimodal fusion according to claim 1 is characterized in that: The test scenario building module includes: A scene parameter extraction unit, configured to extract scene parameters of ambient light intensity, background noise decibel value, and operating temperature range from the interactive feature fusion data; A scene template matching unit, configured to match a basic test template corresponding to the scene parameters from a preset scene template library; A scene detail filling unit is used to fill in the details of the touch response delay, voice command recognition accuracy and interface switching smoothness parameters of the basic test template based on the interaction feature fusion data to obtain the standardized test scene.
10. The cockpit display and control interaction test system based on multimodal fusion according to claim 1, characterized in that: The fusion feature processing module includes: a visual feature extraction subunit, configured to perform edge detection and color space conversion on the visual signal to obtain a visual feature vector; an auditory feature extraction unit, configured to perform frequency analysis and time domain waveform extraction on the auditory signal to obtain an auditory feature vector; a tactile feature extraction unit, configured to perform pressure distribution recognition and contact duration statistics on the tactile signal to obtain a tactile feature vector; a multimodal feature alignment unit, configured to perform timestamp synchronization and dimension unification processing on the visual feature vector, the auditory feature vector, and the tactile feature vector to obtain an aligned feature vector set; A feature fusion unit is used to perform weighted sum processing on the aligned feature vector set to obtain the interactive feature fusion data.
Citation Information
Patent Citations
Intelligent scene identification method and equipment of intelligent interaction control unit
CN119475251A
Intelligent cabin man-machine interaction evaluation method, device, equipment and medium
CN119917395A
Cross-platform UI automatic testing method based on multi-modal large model
CN120429233A
Method and apparatus for processing of multi-modal data
WO2022228958A1
Cited By
Device for testing performance of central control screen of new energy automobile
CN120877626A