A dynamic interface interaction system based on multimodal gesture recognition

By using a multimodal gesture recognition system to dynamically adjust interaction strategies, the limitations of traditional interface interaction systems in complex lighting environments and their adaptability to various scenarios are solved, achieving stable gesture recognition performance and improved user experience.

CN120762574BActive Publication Date: 2025-11-14WUHAN VOCATIONAL COLLEGE OF SOFTWARE & ENG (WUHAN OPEN UNIV)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511261359.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-14
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Traditional interface interaction systems cannot automatically adjust their interaction strategies according to driving scenarios, and user experience and safety are limited in complex lighting environments.

Method used

A multimodal gesture recognition system is adopted to collect gesture information synchronously from multiple devices, generate multimodal gesture data, and dynamically allocate weights based on environmental parameters and reliability coefficients. Combined with scene type and risk assessment, a dynamic function mapping matrix is ​​generated to dynamically adjust the interactive interface.

Benefits of technology

It enhances the continuity and consistency of the user interaction experience, adapts to the safety requirements of different driving scenarios, ensures the intuitiveness and convenience of the interaction, and focuses on user safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762574B_ABST
    Figure CN120762574B_ABST
Patent Text Reader

Abstract

This invention relates to the field of dynamic interaction technology, specifically to a dynamic interface interaction system based on multimodal gesture recognition. The system comprises: a multimodal gesture perception module that synchronously collects gesture information to generate first multimodal gesture data; a multimodal gesture update module that obtains the reliability coefficients of each modality gesture based on the first multimodal gesture data and environmental parameters, assigns dynamic weights to each modality, and generates second multimodal gesture data; a gesture function mapping module that identifies a first scene type; a scene-gesture-function three-dimensional mapping matrix that dynamically determines the function mapping of the current gesture based on the second multimodal gesture data and the first scene type; a risk assessment module that generates an interaction safety coefficient based on the first scene type; and a dynamic interface interaction module that dynamically adjusts the interaction interface based on the function mapping matrix and the interaction safety coefficient. This invention effectively improves user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dynamic interaction technology, specifically to a dynamic interface interaction system based on multimodal gesture recognition. Background Technology

[0002] With the continuous development of intelligent driving, contactless gesture interaction, as a safe and convenient dynamic interface interaction method, has been gradually applied to intelligent driving.

[0003] However, in practical applications, traditional interface interaction systems generally adopt a fixed gesture-function mapping mechanism, which cannot automatically adjust the interaction strategy according to different driving scenarios. Furthermore, the limitations of traditional interface interaction systems in complex lighting environments can easily lead to limitations in user experience and safety.

[0004] To address this, a dynamic interface interaction system based on multimodal gesture recognition is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a dynamic interface interaction system based on multimodal gesture recognition. This invention relates to the field of dynamic interaction technology, specifically a dynamic interface interaction system based on multimodal gesture recognition, comprising: a multimodal gesture perception module synchronously collecting gesture information to generate first multimodal gesture data; a multimodal gesture update module obtaining the reliability coefficient of each modality gesture based on the first multimodal gesture data and environmental parameters, assigning dynamic weights to each modality, and generating second multimodal gesture data; a gesture function mapping module identifying a first scene type; obtaining a scene-gesture-function three-dimensional mapping matrix based on the second multimodal gesture data and the first scene type, dynamically determining the function mapping of the current gesture; a risk assessment module generating an interaction safety coefficient based on the first scene type; and a dynamic interface interaction module dynamically adjusting the interaction interface based on the function mapping matrix and the interaction safety coefficient. This invention can effectively improve the user experience.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A dynamic interface interaction system based on multimodal gesture recognition, comprising:

[0008] The multimodal gesture perception module collects gesture information synchronously from multiple devices to generate the first multimodal gesture data;

[0009] The multimodal gesture update module obtains the reliability coefficient of each modality gesture based on the first multimodal gesture data and environmental parameters, assigns dynamic weights to each modality, and generates the second multimodal gesture data.

[0010] The gesture function mapping module identifies the first scene type; based on the second multimodal gesture data and the first scene type, it obtains a three-dimensional mapping matrix of scene-gesture-function, and dynamically determines the function mapping of the current gesture;

[0011] The risk assessment module generates an interaction safety coefficient based on the first scenario type.

[0012] The dynamic interface interaction module dynamically adjusts the interactive interface based on the function mapping matrix and the interaction safety coefficient.

[0013] Preferably, the multimodal gesture perception module includes: capturing hand posture using a high-definition RGB camera; acquiring three-dimensional hand information using a structured light depth camera; detecting minute hand movements using millimeter-wave radar; and capturing finger touch and swipe movements using an inertial sensor.

[0014] Preferably, the multimodal gesture update module includes:

[0015] The environmental monitoring unit is used to monitor light intensity, noise level, degree of field of view obstruction, and degree of vibration interference in real time.

[0016] The quality assessment unit is used to calculate visual modal quality, depth modal quality, radar modal quality, and inertial modal quality based on environmental parameters.

[0017] The weight allocation unit is used to convert each modal quality into a reliability coefficient; and to allocate weights to different modal gestures based on each reliability coefficient.

[0018] The feature fusion unit performs weighted summation based on weights and gesture data from each modality to obtain the second multimodal gesture.

[0019] Preferably, the first scenario type includes four types: stationary mode, low-speed driving mode, high-speed driving mode, and emergency state.

[0020] Preferably, the gesture function mapping module includes:

[0021] The gesture recognition unit identifies gesture types based on second multimodal gesture data, including static gestures, dynamic gestures, and composite gestures.

[0022] The scene adaptation unit determines the interaction constraints in the current driving environment based on the first scene type;

[0023] The mapping matrix storage unit pre-stores a scene-gesture-function three-dimensional mapping matrix, which is constructed based on the principles of human factors engineering.

[0024] The mapping filtering unit filters out a subset of valid gesture-function mappings for the current scene from the three-dimensional mapping matrix based on the current scene type.

[0025] The context analysis unit receives the user's recent operation history data and provides contextual reference for the currently recognized gesture;

[0026] The function determination unit, in conjunction with the gesture type, current scene adaptation conditions, and context reference, determines the final function mapping result from the effective mapping subset;

[0027] The security priority unit performs a security check on the function mapping results to ensure that security-related functions are always available and respond with priority.

[0028] The mapping learning unit records user interaction feedback and continuously optimizes the three-dimensional mapping matrix based on usage frequency and acceptance.

[0029] Preferably, the risk assessment module includes:

[0030] The risk assessment module includes:

[0031] The vehicle status monitoring unit collects dynamic parameters in real time, including vehicle speed and vehicle acceleration.

[0032] The environmental perception unit collects environmental perception parameters in real time; these parameters include information on road curvature, vehicle distance, visibility, and road surface type collected through a forward-facing camera and vehicle-mounted radar.

[0033] The driver's physiological monitoring unit collects physiological monitoring parameters in real time, including pupil diameter change rate, blink frequency, head posture deviation angle, and steering wheel grip pressure.

[0034] The risk assessment network uses a deep neural network to map the above parameter data to a continuous risk spectrum.

[0035] The interactive safety factor generation unit converts the continuous risk spectrum into an interactive safety factor in the [0,1] interval through an S-shaped mapping function.

[0036] Preferably, the dynamic interface interaction module dynamically adjusts the size of the available gesture set, the complexity of interface elements, the interaction confirmation threshold, and the salience of visual feedback based on the interaction safety coefficient; when the interaction safety coefficient exceeds the preset threshold, only the key safety functions are retained.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0038] 1. This invention solves the limitations of traditional gesture interaction systems in complex lighting environments by using multimodal sensor fusion and intelligent quality assessment mechanisms; it enables the system to maintain stable gesture recognition performance; it ensures the continuity and consistency of the interactive experience; it can effectively improve the intuitiveness and convenience of the interaction; compared with the problem of frequent failures of traditional single-modal systems when the environment changes, this system provides a multimodal gesture interaction experience, effectively improving the user interaction experience.

[0039] 2. This invention introduces a dynamic interaction strategy adjustment mechanism based on driving scenarios. Traditional gesture interaction systems use fixed gesture-function mapping and interface presentation methods, which cannot adapt to the safety requirements of different driving scenarios. However, this system, through a risk assessment network and a scenario-gesture-function three-dimensional mapping matrix, can accurately perceive the current driving scenario and dynamically adjust the interaction strategy. It can fully consider the relationship between driving scenarios, gestures and functions, enrich the interface interaction mode, and thus effectively improve the user interaction experience.

[0040] 3. This invention uses a dynamic interface interaction module to dynamically adjust the size of the available gesture set, the complexity of interface elements, the interaction confirmation threshold, and the salience of visual feedback based on the interaction safety coefficient. When the interaction safety coefficient exceeds the preset threshold, only key safety functions are retained. This effectively improves the user interaction experience while also focusing on user safety. Attached Figure Description

[0041] Figure 1 A schematic diagram of the structure of a dynamic interface interaction system based on multimodal gesture recognition provided in an embodiment of the present invention;

[0042] Figure 2 This is a schematic diagram of the structure of a multimodal gesture update module provided in an embodiment of the present invention. Detailed Implementation

[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] Example 1

[0045] To improve user A's experience with the in-vehicle dynamic interface interaction system, a dynamic interface interaction system based on multimodal gesture recognition was applied, such as... Figure 1 A schematic diagram of a dynamic interface interaction system based on multimodal gesture recognition provided in an embodiment of the present invention includes:

[0046] The multimodal gesture perception module collects gesture information synchronously from multiple devices to generate the first multimodal gesture data;

[0047] Furthermore, the multimodal gesture perception module includes: capturing hand posture through a high-definition RGB camera; acquiring three-dimensional information of the hand through a structured light depth camera; detecting minute hand movements through millimeter-wave radar; and capturing finger touch and swipe movements through an inertial sensor; the sampling rates of the relevant devices are shown in Table 1.

[0048] Table 1 Data Collection Table

[0049]

[0050] The multimodal gesture update module obtains the reliability coefficient of each modality gesture data based on the first multimodal gesture data and environmental parameters, assigns dynamic weights to each modality, and generates the second multimodal gesture data. Figure 2 This is a schematic diagram of the structure of a multimodal gesture update module provided in an embodiment of the present invention;

[0051] The gesture data for each modality includes visual modal gesture data, depth modal gesture data, radar modal gesture data, and inertial modal gesture data;

[0052] Visual modal gesture data includes gesture motion trajectory and gesture outline; depth modal gesture data includes hand 3D coordinates and gesture geometry; radar modal gesture data includes hand movement speed and hand angle information; hand angle information represents the angle of the hand relative to the radar; inertial modal gesture data includes acceleration information (acceleration of the hand in 3D space) and angular velocity information (rotation speed of the hand).

[0053] Furthermore, the multimodal gesture update module includes:

[0054] The environmental monitoring unit is used to monitor light intensity, noise level, degree of field of view obstruction, and degree of vibration interference in real time.

[0055] Furthermore, some environmental data are shown in Table 2;

[0056] Table 2 Partial Environmental Data Table

[0057]

[0058] The quality assessment unit is used to calculate visual modal quality, depth modal quality, radar modal quality, and inertial modal quality based on environmental parameters.

[0059] This embodiment solves the limitations of traditional gesture interaction systems in complex lighting environments by using multimodal sensor fusion and intelligent quality assessment mechanisms; it enables the system to maintain stable gesture recognition performance; it ensures the continuity and consistency of the interactive experience; it can effectively improve the intuitiveness and convenience of the interaction; compared with the problem of frequent failures of traditional single-modal systems when the environment changes, this system provides a multimodal gesture interaction experience, effectively improving the user interaction experience.

[0060] The specific process for obtaining the visual modal quality, depth modal quality, radar modal quality, and inertial modal quality is as follows: Gesture recognition tests are performed on the four modalities (visual, depth, radar, and inertial) using environmental data from each group; the gesture recognition accuracy of each modality is determined, and the average value is taken to obtain the quality of each modality;

[0061] The weight allocation unit is used to convert each modal quality into a reliability coefficient; and to allocate weights to different modal gestures based on each reliability coefficient.

[0062] The reliability coefficient is:

[0063] ;

[0064] in, Indicates the first Reliability coefficient of each mode; Indicates the first Modal quality of each mode; Indicates the number of modes;

[0065] The feature fusion unit performs weighted summation based on weights and gesture data from each modality to obtain the second multimodal gesture data.

[0066] The gesture function mapping module identifies the first scene type; based on the second multimodal gesture data and the first scene type, it obtains a three-dimensional mapping matrix of scene-gesture-function, and dynamically determines the function mapping of the current gesture;

[0067] Furthermore, the first scenario type includes four scenario types: stationary mode, low-speed driving mode, high-speed driving mode, and emergency state. The first scenario type is determined by vehicle speed and vehicle status. If the vehicle speed is 0, it is determined to be stationary mode. If the vehicle speed is within a preset low-speed range, it is determined to be low-speed driving mode. If the vehicle speed is within a preset high-speed range, it is determined to be high-speed driving mode. If the vehicle exhibits abnormal behavior, such as sudden braking, sudden acceleration, or sudden U-turn, it is determined to be an emergency state.

[0068] Furthermore, the gesture function mapping module includes:

[0069] The gesture recognition unit identifies gesture types based on second multimodal gesture data, including static gestures, dynamic gestures, and composite gestures. Static gestures represent gestures that do not move, such as clenching a fist or extending fingers. Dynamic gestures represent gestures with a clear movement trajectory, such as sliding or rotating. Composite gestures are complex gestures composed of multiple static or dynamic gestures.

[0070] The scene adaptation unit determines the interaction constraints in the current driving environment based on the first scene type;

[0071] The scene adaptation unit is used to determine specific gesture type constraints according to the current scene type. For example, when driving at high speed, complex gestures are restricted; in low-speed scenes, complex gestures are not restricted.

[0072] The mapping matrix storage unit pre-stores a three-dimensional mapping matrix of scene-gesture-function, which is constructed based on the principles of human factors engineering. The principles of human factors engineering include the reasonable coordination and optimization of the driver's natural movements and habits with the vehicle functions according to the principles of human-computer interaction design. Through these designs, it is possible to ensure that the system provides the driver with an intuitive, convenient and efficient interactive experience in different driving states.

[0073] For example, in low-speed driving mode, the driver's operation is more stable, and there may be more time for fine-tuning. Therefore, the system can support slight gestures or fingertip touches to complete some routine settings, such as adjusting the volume, adjusting the seat, or navigation. In emergency situations, the driver's reaction needs to be quick and intuitive. The system should prioritize responding to rapid, drastic gestures to ensure the rapid execution of critical functions such as emergency braking and activation of automatic obstacle avoidance, thereby helping the driver deal with unexpected situations as quickly as possible.

[0074] The three-dimensional mapping matrix is:

[0075] ;

[0076] in, Represents a three-dimensional mapping matrix; Indicates the scene type; This represents the second multimodal gesture data; Indicates function mapping;

[0077] Where 0 and 1 represent non-compliance, respectively. Functions and compliance Function;

[0078] Functions include, for example, changing songs, confirming navigation, and answering phone calls;

[0079] The mapping filtering unit filters out a subset of valid gesture-function mappings for the current scene from the three-dimensional mapping matrix based on the current scene type.

[0080] The context analysis unit receives the user's recent operation history data and provides contextual reference for the currently recognized gesture;

[0081] Contextual reference specifically includes: the system can perform contextual reasoning based on the driver's operation history data; for example, if the driver frequently adjusts navigation settings in the past low-speed driving mode, the system determines that the driver has a need to adjust navigation settings in the same low-speed scenario;

[0082] The function determination unit, in conjunction with the recognized gesture type, current scene adaptation conditions, and context reference, determines the final function mapping result from the effective mapping subset;

[0083] The security priority unit performs a security check on the function mapping results to ensure that security-related functions are always available and respond with priority.

[0084] The mapping learning unit records user interaction feedback and continuously optimizes the three-dimensional mapping matrix based on usage frequency and acceptance.

[0085] Based on user feedback data, the mapping relationships in the 3D mapping matrix are adjusted to improve the accuracy of gesture recognition and the efficiency of function triggering; adjusting the mapping relationships in the 3D mapping matrix includes increasing or decreasing the functional range under fixed scene types and gesture data;

[0086] This embodiment introduces a dynamic interaction strategy adjustment mechanism based on driving scenarios. Traditional gesture interaction systems use fixed gesture-function mapping and interface presentation methods, which cannot adapt to the safety requirements of different driving scenarios. However, this system, through a risk assessment network and a scenario-gesture-function three-dimensional mapping matrix, can accurately perceive the current driving scenario and dynamically adjust the interaction strategy. It can fully consider the relationship between driving scenarios, gestures and functions, enrich the interface interaction mode, and thus effectively improve the user interaction experience.

[0087] The risk assessment module generates an interaction safety coefficient based on the first scenario type.

[0088] Furthermore, the risk assessment module includes:

[0089] The risk assessment module includes:

[0090] The vehicle status monitoring unit collects dynamic parameters in real time, including vehicle speed and vehicle acceleration.

[0091] The environmental perception unit collects environmental perception parameters in real time; these parameters include information on road curvature, vehicle distance, visibility, and road surface type collected through a forward-facing camera and vehicle-mounted radar.

[0092] The driver's physiological monitoring unit collects physiological monitoring parameters in real time, including pupil diameter change rate, blink frequency, head posture deviation angle, and steering wheel grip pressure.

[0093] The risk assessment network uses a deep neural network to map the aforementioned parameter data to a continuous risk spectrum. Specifically, the deep neural network compares the current parameter data with the corresponding historical reference data, calculating the similarity score between each current parameter data point and its corresponding historical reference data using the Pearson coefficient. This similarity score reflects the degree of similarity between the current driving state and the historical data. Based on this similarity, the deep neural network further infers the current risk level, thus obtaining the continuous risk spectrum.

[0094] The interactive safety factor generation unit converts the continuous risk spectrum into an interactive safety factor in the [0,1] interval through an S-shaped mapping function.

[0095] The interaction security factor is:

[0096] ;

[0097] in, Indicates the interaction safety factor; This represents the risk value in a continuous risk spectrum; Indicates the center point of the risk spectrum; The parameter indicating the adjustment speed of the mapping;

[0098] The dynamic interface interaction module dynamically adjusts the interactive interface based on the function mapping matrix and the interaction safety coefficient.

[0099] Furthermore, the dynamic interface interaction module dynamically adjusts the size of the available gesture set, the complexity of interface elements, the interaction confirmation threshold, and the salience of visual feedback based on the interaction safety coefficient; when the interaction safety coefficient exceeds the preset threshold, only the key safety functions are retained.

[0100] This embodiment uses a dynamic interface interaction module to dynamically adjust the size of the available gesture set, the complexity of interface elements, the interaction confirmation threshold, and the salience of visual feedback based on the interaction safety coefficient. When the interaction safety coefficient exceeds the preset threshold, only key safety functions are retained. This effectively improves the user interaction experience while also focusing on user safety.

[0101] Specifically, this is achieved by acquiring various modal gesture data of user A in multiple environments within the cockpit; these modal gesture data include visual modal gesture data, depth modal gesture data, radar modal gesture data, and inertial modal gesture data.

[0102] The method includes visual modal gesture data (gesture trajectory and contour), depth modal gesture data (3D hand coordinates and gesture geometry), radar modal gesture data (hand movement speed and angle information, with the hand angle information representing the angle of the hand relative to the radar), and inertial modal gesture data (acceleration information of the hand in 3D space and angular velocity information of the hand rotation speed). It calculates the recognition accuracy of each modality in specific environments and obtains the weights of the gesture data for each modality. This method overcomes the inadequacy of traditional gestures in complex environments, achieving accurate gesture interaction for multimodal acquisition in various environments, thereby improving customer experience.

[0103] Furthermore, the multimodal gesture data is weighted based on the weights of the gesture data of each modality to obtain the second multimodal gesture data;

[0104] Get the current first scenario type of user A; the first scenario type includes four scenario types: stationary mode, low-speed driving mode, high-speed driving mode, and emergency state;

[0105] Based on the second multimodal gesture data and the first scene type, a three-dimensional mapping matrix of scene-gesture-function is obtained, and the function mapping of the current gesture is dynamically determined. The three-dimensional mapping matrix is ​​constructed based on the principles of human-computer interaction. The principles of human-computer interaction include the reasonable coordination and optimization of the driver's natural movements and habits with the vehicle functions according to the principles of human-computer interaction design. Through these designs, it can be ensured that the system provides the driver and user A with an intuitive, convenient and efficient interactive experience in different driving states.

[0106] Traditional fixed gesture-function mapping cannot adapt to changes in driving scenarios; this system dynamically adjusts function mapping through a three-dimensional mapping matrix, which can enrich the driver's interaction experience with the interface, thereby effectively improving the user experience and safety.

[0107] Example 2

[0108] To enhance user experience for user B with the in-vehicle dynamic interface interaction system, a dynamic interface interaction system based on multimodal gesture recognition was applied, such as... Figure 1 A schematic diagram of a dynamic interface interaction system based on multimodal gesture recognition provided in an embodiment of the present invention includes:

[0109] The multimodal gesture perception module collects gesture information synchronously from multiple devices to generate the first multimodal gesture data;

[0110] Furthermore, the multimodal gesture perception module includes: capturing hand posture via a high-definition RGB camera; acquiring three-dimensional hand information via a structured light depth camera; detecting minute hand movements via millimeter-wave radar; and capturing finger touch and swipe movements via an inertial sensor.

[0111] The multimodal gesture update module obtains the reliability coefficient of each modality gesture data based on the first multimodal gesture data and environmental parameters, assigns dynamic weights to each modality, and generates the second multimodal gesture data.

[0112] The gesture data for each modality includes visual modal gesture data, depth modal gesture data, radar modal gesture data, and inertial modal gesture data;

[0113] Visual modal gesture data includes gesture motion trajectory and gesture outline; depth modal gesture data includes hand 3D coordinates and gesture geometry; radar modal gesture data includes hand movement speed and hand angle information; hand angle information represents the angle of the hand relative to the radar; inertial modal gesture data includes acceleration information (acceleration of the hand in 3D space) and angular velocity information (rotation speed of the hand).

[0114] Furthermore, the multimodal gesture update module includes:

[0115] The environmental monitoring unit is used to monitor light intensity, noise level, degree of field of view obstruction, and degree of vibration interference in real time.

[0116] Furthermore, some environmental data are shown in Table 3;

[0117] Table 3 Partial Environmental Data Table

[0118]

[0119] The quality assessment unit is used to calculate visual modal quality, depth modal quality, radar modal quality, and inertial modal quality based on environmental parameters.

[0120] The specific process for obtaining the visual modal quality, depth modal quality, radar modal quality, and inertial modal quality is as follows: Gesture recognition tests are performed on the four modalities (visual, depth, radar, and inertial) using environmental data from each group; the gesture recognition accuracy of each modality is determined, and the average value is taken to obtain the quality of each modality;

[0121] The weight allocation unit is used to convert each modal quality into a reliability coefficient; and to allocate weights to different modal gestures based on each reliability coefficient.

[0122] The reliability coefficient is:

[0123] ;

[0124] in, Indicates the first Reliability coefficient of each mode; Indicates the first Modal quality of each mode; Indicates the number of modes;

[0125] The feature fusion unit performs weighted summation based on weights and gesture data from each modality to obtain the second multimodal gesture data.

[0126] The gesture function mapping module identifies the first scene type; based on the second multimodal gesture data and the first scene type, it obtains a three-dimensional mapping matrix of scene-gesture-function, and dynamically determines the function mapping of the current gesture;

[0127] Furthermore, the first scenario type includes four scenario types: stationary mode, low-speed driving mode, high-speed driving mode, and emergency state. If the vehicle speed is 0, it is determined to be stationary mode; if the vehicle speed is within a preset low-speed range, it is determined to be low-speed driving mode; if the vehicle speed is within a preset high-speed range, it is determined to be high-speed driving mode; if the vehicle exhibits abnormal behavior, such as sudden braking, sudden acceleration, or sudden U-turn, it is determined to be an emergency state.

[0128] Furthermore, the gesture function mapping module includes:

[0129] The gesture recognition unit identifies gesture types based on second multimodal gesture data, including static gestures, dynamic gestures, and composite gestures. Static gestures represent gestures that do not move, such as clenching a fist or extending fingers. Dynamic gestures represent gestures with a clear movement trajectory, such as sliding or rotating. Composite gestures are complex gestures composed of multiple static or dynamic gestures.

[0130] The scene adaptation unit determines the interaction constraints in the current driving environment based on the first scene type;

[0131] The scene adaptation unit is used to determine specific gesture type constraints according to the current scene type. For example, when driving at high speed, complex gestures are restricted; in low-speed scenes, complex gestures are not restricted.

[0132] The mapping matrix storage unit pre-stores a three-dimensional mapping matrix of scene-gesture-function, which is constructed based on the principles of human-computer interaction. These principles include the reasonable coordination and optimization of the driver's natural movements and habits with the vehicle's functions according to the principles of human-computer interaction design. Through these designs, the system can ensure that it provides the driver with an intuitive, convenient and efficient interactive experience in different driving states.

[0133] For example, in low-speed driving mode, the driver's operation is more stable, and there may be more time for fine-tuning. Therefore, the system can support slight gestures or fingertip touches to complete some routine settings, such as adjusting the volume, adjusting the seat, or navigation. In emergency situations, the driver's reaction needs to be quick and intuitive. The system should prioritize responding to rapid, drastic gestures to ensure the rapid execution of critical functions such as emergency braking and activation of automatic obstacle avoidance, thereby helping the driver deal with unexpected situations as quickly as possible.

[0134] The three-dimensional mapping matrix is:

[0135] ;

[0136] in, Represents a three-dimensional mapping matrix; Indicates the scene type; This represents the second multimodal gesture data; Indicates function mapping;

[0137] Where 0 and 1 represent non-compliance, respectively. Functions and compliance Function;

[0138] Functions include, for example, changing songs, confirming navigation, and answering phone calls;

[0139] The mapping filtering unit filters out a subset of valid gesture-function mappings for the current scene from the three-dimensional mapping matrix based on the current scene type.

[0140] The context analysis unit receives the user's recent operation history data and provides contextual reference for the currently recognized gesture;

[0141] Contextual reference specifically includes: the system can perform contextual reasoning based on the driver's operation history data; for example, if the driver frequently adjusts navigation settings in the past low-speed driving mode, the system determines that the driver has a need to adjust navigation settings in the same low-speed scenario;

[0142] The function determination unit, in conjunction with the recognized gesture type, current scene adaptation conditions, and context reference, determines the final function mapping result from the effective mapping subset;

[0143] The security priority unit performs a security check on the function mapping results to ensure that security-related functions are always available and respond with priority.

[0144] The mapping learning unit records user interaction feedback and continuously optimizes the three-dimensional mapping matrix based on usage frequency and acceptance.

[0145] Based on user feedback data, the mapping relationships in the 3D mapping matrix are adjusted to improve the accuracy of gesture recognition and the efficiency of function triggering; adjusting the mapping relationships in the 3D mapping matrix includes increasing or decreasing the functional range under fixed scene types and gesture data;

[0146] The risk assessment module generates an interaction safety coefficient based on the first scenario type.

[0147] Furthermore, the risk assessment module includes:

[0148] The risk assessment module includes:

[0149] The vehicle status monitoring unit collects dynamic parameters in real time, including vehicle speed and vehicle acceleration.

[0150] The environmental perception unit collects environmental perception parameters in real time; these parameters include information on road curvature, vehicle distance, visibility, and road surface type collected through a forward-facing camera and vehicle-mounted radar.

[0151] The driver's physiological monitoring unit collects physiological monitoring parameters in real time, including pupil diameter change rate, blink frequency, head posture deviation angle, and steering wheel grip pressure.

[0152] The risk assessment network uses a deep neural network to map the aforementioned parameter data to a continuous risk spectrum. Specifically, the deep neural network compares the current parameter data with the corresponding historical reference data, calculating the similarity score between each current parameter data point and its corresponding historical reference data using the Pearson coefficient. This similarity score reflects the degree of similarity between the current driving state and the historical data. Based on this similarity, the deep neural network further infers the current risk level, thus obtaining the continuous risk spectrum.

[0153] The interactive safety factor generation unit converts the continuous risk spectrum into an interactive safety factor in the [0,1] interval through an S-shaped mapping function.

[0154] The interaction security factor is:

[0155] ;

[0156] in, Indicates the interaction safety factor; This represents the risk value in a continuous risk spectrum; Indicates the center point of the risk spectrum; The parameter indicating the adjustment speed of the mapping;

[0157] The dynamic interface interaction module dynamically adjusts the interactive interface based on the function mapping matrix and the interaction safety coefficient.

[0158] Furthermore, the dynamic interface interaction module dynamically adjusts the size of the available gesture set, the complexity of interface elements, the interaction confirmation threshold, and the salience of visual feedback based on the interaction safety coefficient; when the interaction safety coefficient exceeds the preset threshold, only the key safety functions are retained.

[0159] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A dynamic interface interaction system based on multimodal gesture recognition, characterized in that, include: The multimodal gesture perception module collects gesture information synchronously from multiple devices to generate the first multimodal gesture data; The multimodal gesture update module obtains the reliability coefficient of each modality gesture based on the first multimodal gesture data and environmental parameters, assigns dynamic weights to each modality, and generates the second multimodal gesture data. The gesture function mapping module identifies the first scene type; Based on the second multimodal gesture data and the first scene type, a three-dimensional mapping matrix of scene-gesture-function is obtained, and the function mapping of the current gesture is dynamically determined. The gesture function mapping module includes: The gesture recognition unit identifies gesture types based on second multimodal gesture data, including static gestures, dynamic gestures, and composite gestures. The scene adaptation unit determines the interaction constraints in the current driving environment based on the first scene type; The mapping matrix storage unit pre-stores a scene-gesture-function three-dimensional mapping matrix, which is constructed based on the principles of human factors engineering. The mapping filtering unit filters out a subset of valid gesture-function mappings for the current scene from the three-dimensional mapping matrix based on the current scene type. The context analysis unit receives the user's recent operation history data and provides contextual reference for the currently recognized gesture; The function determination unit, in conjunction with the gesture type, current scene adaptation conditions, and context reference, determines the final function mapping result from the effective mapping subset; The security priority unit performs a security check on the function mapping results to ensure that security-related functions are always available and respond with priority. The mapping learning unit records user interaction feedback and continuously optimizes the three-dimensional mapping matrix based on usage frequency and acceptance. The risk assessment module generates an interaction safety coefficient based on the first scenario type. The dynamic interface interaction module dynamically adjusts the interactive interface based on the function mapping matrix and the interaction safety coefficient.

2. The dynamic interface interaction system based on multimodal gesture recognition according to claim 1, characterized in that: The multimodal gesture perception module includes: capturing hand posture through a high-definition RGB camera; acquiring three-dimensional information of the hand through a structured light depth camera; detecting minute hand movements through millimeter-wave radar; and capturing finger touch and swipe movements through an inertial sensor.

3. The dynamic interface interaction system based on multimodal gesture recognition according to claim 1, characterized in that: The multimodal gesture update module includes: The environmental monitoring unit is used to monitor light intensity, noise level, degree of field of view obstruction, and degree of vibration interference in real time. The quality assessment unit is used to calculate visual modal quality, depth modal quality, radar modal quality, and inertial modal quality based on environmental parameters. The weight allocation unit is used to convert each modal quality into a reliability coefficient; and to allocate weights to different modal gestures based on each reliability coefficient. The feature fusion unit performs weighted summation based on weights and gesture data from each modality to obtain the second multimodal gesture.

4. The dynamic interface interaction system based on multimodal gesture recognition according to claim 1, characterized in that: The first scenario type includes four types: stationary mode, low-speed driving mode, high-speed driving mode, and emergency state.

5. A dynamic interface interaction system based on multimodal gesture recognition according to claim 1, characterized in that: The risk assessment module includes: The vehicle status monitoring unit collects dynamic parameters in real time, including vehicle speed and vehicle acceleration. The environmental perception unit collects environmental perception parameters in real time; these parameters include information on road curvature, vehicle distance, visibility, and road surface type collected through a forward-facing camera and vehicle-mounted radar. The driver's physiological monitoring unit collects physiological monitoring parameters in real time, including pupil diameter change rate, blink frequency, head posture deviation angle, and steering wheel grip pressure. The risk assessment network uses a deep neural network to map the above parameter data to a continuous risk spectrum. The interactive safety factor generation unit converts the continuous risk spectrum into an interactive safety factor in the [0,1] interval using an S-shaped mapping function.

6. A dynamic interface interaction system based on multimodal gesture recognition according to claim 1, characterized in that: The dynamic interface interaction module dynamically adjusts the size of the available gesture set, the complexity of interface elements, the interaction confirmation threshold, and the salience of visual feedback based on the interaction safety coefficient; when the interaction safety coefficient exceeds the preset threshold, only the key safety functions are retained.

Citation Information

Patent Citations

  • Human-computer interaction method and device based on gesture recognition, equipment and storage medium

    CN118131915A

  • Multi-mode interactive intelligent control system

    CN118226967A