Gesture instruction recognition method, electronic equipment and vehicle
By fusing multimodal sensor data and calculating modal contribution, the problems of low gesture recognition accuracy and high response latency in vehicle cockpit scenarios are solved, achieving highly robust gesture interaction control with a low false judgment rate and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, gesture recognition has not been optimized for vehicle cabin scenarios, resulting in low recognition accuracy, high response latency, poor user experience, limited functionality, and no involvement of gesture interaction.
By acquiring multimodal sensor data, including joint sensor data, limb sensor data, camera data, and millimeter-wave radar data, the system identifies initial gesture features, calculates initial gesture fusion features based on modal contribution, generates final gesture fusion features, and controls the vehicle to execute gesture commands.
It achieves highly robust and low-false-rate gesture interaction control in the vehicle cabin, improving the accuracy and response speed of gesture command recognition and enhancing the user experience.
Smart Images

Figure CN121884455A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cockpit gesture control technology, and in particular to a gesture command recognition method, electronic device, and vehicle. Background Technology
[0002] Among related technologies, the entire process of voice acquisition, recognition, parsing, classification, execution, feedback confirmation, and anomaly handling can be used to accurately interpret intent and verify syntax by combining system status and user preferences. This multimodal feedback enables efficient control of the intelligent cockpit, significantly improving recognition accuracy and user operation satisfaction. Alternatively, an original sequence dataset can be generated by constructing action frames based on sensor timestamps. After environmental compensation preprocessing, weights are allocated according to sensor contribution and multimodal hand motion data features are fused. This data is then input into a gesture recognition model via spatiotemporal coding, achieving effective fusion of multimodal sensor data and high-precision gesture recognition resistant to environmental noise.
[0003] However, the relevant technologies only focus on voice control and do not involve gesture interaction, resulting in limited functionality. Furthermore, multi-sensor gesture recognition is not optimized for vehicle cabin scenarios and does not clearly select and generate accurate execution actions based on fused features. This leads to low gesture command recognition accuracy, high response latency, and poor user experience, which urgently need improvement. Summary of the Invention
[0004] This application provides a gesture command recognition method, electronic device, and vehicle, aiming to improve the problems in related technologies that only focus on voice control and do not involve gesture interaction, resulting in limited functionality; and that multi-sensor gesture recognition is not optimized for vehicle cabin scenarios, does not clearly select and generate accurate execution actions based on fusion features, and suffers from low gesture command recognition accuracy, high response latency, and poor user experience.
[0005] A first aspect of this application provides a method for recognizing gesture commands, comprising the following steps: in response to a gesture command from at least one occupant, acquiring multimodal sensor data corresponding to the gesture command, wherein the multimodal sensor data includes at least joint sensor data, limb sensor data, camera data, and millimeter-wave radar data; identifying initial gesture feature data corresponding to each modal sensor data, and determining initial gesture fusion feature data between the modal sensor data based on the modal contribution of the modal sensor data and the initial gesture feature data; determining final gesture fusion feature data based on the initial gesture fusion feature data, and generating a gesture command execution action for the vehicle based on the final gesture fusion feature data, thereby controlling the vehicle to execute the gesture command execution action.
[0006] The above technical solution can respond to the gesture commands of drivers and passengers, acquire corresponding multimodal sensor data, identify the corresponding initial gesture feature data based on each modal sensor data, and determine the initial gesture fusion feature data by combining the modal contribution of the modal sensor data, and then determine the final gesture fusion feature data, thereby generating and executing the vehicle's gesture command execution action. Through multimodal sensor data fusion and feature calculation based on modal contribution, effective modal features are accurately selected and execution actions are generated, achieving highly robust and low-misjudgment-rate interactive control of gesture commands in the vehicle cabin.
[0007] Optionally, in one embodiment of this application, before determining the initial gesture fusion feature data between the modal sensor data, the method further includes: identifying the basic robustness coefficient of the modal sensor data and the scene factor corresponding to the modal sensor data, wherein the basic robustness coefficient characterizes an indicator used to measure the effectiveness, stability, and consistency of the modal sensor data, and the scene factor characterizes a parameter for the working environment of the sensor; obtaining the influence weight of the scene factor and the performance loss rate of the modal sensor data under the influence of the scene factor, wherein the performance loss rate is used to measure the proportion of performance degradation of the modal sensor data under the influence of the scene factor; and calculating the modal contribution based on the basic robustness coefficient, the scene factor, the influence weight, and the performance loss rate.
[0008] The above technical solution allows for the calculation of modal contribution based on the basic robustness coefficient of the identified modal sensor data, the scene factor corresponding to the modal sensor data, the influence weight of the scene factor, and the performance loss rate of the modal sensor data under the influence of the scene factor, before calculating the initial gesture fusion feature data. By comprehensively considering the basic robustness coefficient, scene factor, influence weight, and performance loss rate to calculate the modal contribution, the effectiveness of each modal sensor data in a specific scenario can be more accurately quantified, thereby improving the accuracy and reliability of gesture fusion feature data.
[0009] Optionally, in one embodiment of this application, the expression for the modal contribution can be, but is not limited to, the following: , in, For modal sensor data Modal contribution For modal sensor data The basic robustness coefficient, As a scene factor, Scene factors Influence weight, For modal sensor data In scene factors The performance loss rate is as follows. This represents the total number of modes.
[0010] The above technical solution clarifies the expression for modal contribution and, by integrating the basic robustness coefficient with the weighted performance loss of multiple scene factors (illuminance, speed, occlusion), achieves accurate quantitative evaluation of the contribution of sensor data under different environmental conditions, significantly improving the reliability and environmental adaptability of modal selection.
[0011] Optionally, in one embodiment of this application, determining the initial gesture fusion feature data between the modal sensor data based on the modal contribution of the modal sensor data and the initial gesture feature data includes: calculating the final timestamp and final coordinate value of the modal sensor data based on the modal contribution; calculating the predicted gesture fusion feature data between the modal sensor data based on the modal contribution; and obtaining the initial gesture fusion feature data based on the final timestamp, the final coordinate value, the predicted gesture fusion feature data, the modal contribution, and the initial gesture feature data.
[0012] The above technical solution can be used to calculate the final timestamp, final coordinate value and predicted gesture fusion feature data of modal sensor data based on modal contribution. Then, the final timestamp, final coordinate value, predicted gesture fusion feature data, modal contribution and initial gesture feature data are integrated to generate initial gesture fusion feature data. By dynamically fusing the final timestamp, final coordinate value and predicted gesture fusion feature data of multi-sensor data through modal contribution, the accuracy and anti-interference ability of gesture feature fusion are effectively improved, and the robustness of gesture recognition in complex scenarios is enhanced.
[0013] Optionally, in one embodiment of this application, the formula for calculating the final timestamp and the final coordinate value may be, but is not limited to, the following: , in, , For modal sensor data The corrected timestamp and 3D coordinates, i.e., the final timestamp and final coordinate values. , For the initial timestamp and initial coordinates, , For the system's reference time and reference coordinates, , The average timestamp and average coordinates for all modal sensor data. This is a time correction factor. This is the spatial correction factor. For modal sensor data Modal contribution For the vehicle's real-time acceleration, For modal sensor data The original time deviation, This refers to the system's time resolution. The expression for the predicted gesture fusion feature data is: , in, For prediction Predictive gesture fusion feature data at any given time. Weighted by historical features for Historical gesture feature data at any given moment This is the trend amplification factor. For modal sensor data Modal contribution , For modal sensor data exist The characteristic first derivative and characteristic second derivative at time t. Represents the predicted time interval. This represents the total number of modes.
[0014] By introducing key parameters such as modal contribution, real-time vehicle acceleration, and feature derivatives, and combining system benchmarks and historical feature data, the above technical solutions enable accurate correction of the spatiotemporal parameters of modal sensor data and reliable prediction of gesture fusion features, thereby improving the scene adaptability of data processing and the accuracy of prediction results.
[0015] Optionally, in one embodiment of this application, determining the final gesture fusion feature data based on the initial gesture fusion feature data includes: acquiring the current vehicle speed and the fatigue index of the at least one driver / passenger; calculating a gesture feature matching threshold for the initial gesture fusion feature data based on the current vehicle speed and the fatigue index; identifying the instruction mapping information of the initial gesture fusion feature data based on the current vehicle speed; determining the interaction mode of the initial gesture fusion feature data based on the fatigue index; and determining the final gesture fusion feature data based on the initial gesture fusion feature data, the gesture feature matching threshold, the instruction mapping information, and the interaction mode.
[0016] The above technical solution can calculate the gesture feature matching threshold, identify command mapping information, and determine the interaction mode of the initial gesture fusion feature data based on the vehicle's current speed and the fatigue index of the driver and passengers. Then, the final gesture fusion feature data can be determined. By dynamically combining vehicle speed and the fatigue index of the driver and passengers, the gesture feature matching threshold, command mapping and interaction mode can be adaptively adjusted to achieve accurate screening and personalized response of gesture commands in complex driving scenarios, which can significantly improve interaction safety and user experience.
[0017] Optionally, in one embodiment of this application, generating the gesture command execution action of the vehicle based on the final gesture fusion feature data includes: determining the initial gesture intent of the final gesture fusion feature data based on a preset gesture intent matching rule; calculating the intent confidence of the initial gesture intent based on the modal contribution; and generating the gesture command execution action based on the gesture intent whose intent confidence is greater than a preset threshold.
[0018] The above technical solution can determine the initial gesture intent based on the preset gesture intent matching rules, calculate the intent confidence based on the modal contribution, and then generate gesture command execution actions based on gesture intents with intent confidence greater than a preset threshold. The intent confidence is quantified by the preset gesture intent matching rules and modal contribution, and gesture command execution actions are generated based on the preset threshold, which effectively improves the accuracy of gesture intent recognition and the reliability of command execution, and avoids the risk of misoperation.
[0019] Optionally, in one embodiment of this application, generating the gesture command execution action of the vehicle based on the final gesture fusion feature data includes: in response to the interaction mode being a multi-person interaction mode, obtaining the gesture distance value of the corresponding driver / passenger based on the final gesture fusion feature data; if the gesture distance value meets a preset distance condition, obtaining the identity identifier of the corresponding driver / passenger; determining the priority weight value of the corresponding driver / passenger based on the identity identifier; and generating the gesture command execution action based on the gesture distance value and the priority weight value.
[0020] The above technical solution can respond to multi-person interaction mode by obtaining the gesture distance value of the driver and passenger based on the final gesture fusion feature data. When the gesture distance value meets the preset distance conditions, the identity identifier of the driver and passenger is extracted and the corresponding priority weight value is determined. Then, the gesture distance value and priority weight value are fused to generate gesture command execution action. In multi-person interaction mode, through the dual judgment logic of preset distance condition filtering and driver and passenger identity priority weighting, the accurate recognition and differentiated execution of gesture commands are realized, improving the accuracy of command response and interaction efficiency in multi-person interaction scenarios.
[0021] A second aspect of this application provides a gesture command recognition device, comprising: a first acquisition module, configured to acquire multimodal sensor data corresponding to a gesture command from at least one occupant, the multimodal sensor data including at least joint sensor data, limb sensor data, camera data, and millimeter-wave radar data; a determination module, configured to identify corresponding initial gesture feature data based on each modal sensor data, and determine initial gesture fusion feature data between the modal sensor data based on the modal contribution of the modal sensor data and the initial gesture feature data; and a control module, configured to determine final gesture fusion feature data based on the initial gesture fusion feature data, generate a gesture command execution action for the vehicle based on the final gesture fusion feature data, and control the vehicle to execute the gesture command execution action.
[0022] The above technical solution can respond to the gesture commands of drivers and passengers, acquire corresponding multimodal sensor data, determine the corresponding initial gesture feature data based on each modal sensor data, and determine the initial gesture fusion feature data by combining the modal contribution of the modal sensor data, and then determine the final gesture fusion feature data, thereby generating and executing the vehicle's gesture command execution action. Through multimodal sensor data fusion and feature calculation based on modal contribution, effective modal features are accurately selected and execution actions are generated, achieving highly robust and low-misjudgment-rate interactive control of gesture commands in the vehicle cabin.
[0023] Optionally, in one embodiment of this application, it further includes: an identification module, configured to identify a basic robustness coefficient of the modal sensor data and a scene factor corresponding to the modal sensor data before determining the initial gesture fusion feature data between the modal sensor data, wherein the basic robustness coefficient characterizes an index used to measure the effectiveness, stability, and consistency of the modal sensor data, and the scene factor characterizes a parameter for the working environment of the sensor; a second acquisition module, configured to acquire the influence weight of the scene factor and the performance loss rate of the modal sensor data under the influence of the scene factor, wherein the performance loss rate is used to measure the proportion of performance degradation of the modal sensor data under the influence of the scene factor; and a calculation module, configured to calculate the modal contribution based on the basic robustness coefficient, the scene factor, the influence weight, and the performance loss rate.
[0024] The above technical solution allows for the calculation of modal contribution based on the basic robustness coefficient of the identified modal sensor data, the scene factor corresponding to the modal sensor data, the influence weight of the scene factor, and the performance loss rate of the modal sensor data under the influence of the scene factor, before calculating the initial gesture fusion feature data. By comprehensively considering the basic robustness coefficient, scene factor, influence weight, and performance loss rate to calculate the modal contribution, the effectiveness of each modal sensor data in a specific scenario can be more accurately quantified, thereby improving the accuracy and reliability of gesture fusion feature data.
[0025] Optionally, in one embodiment of this application, the expression for the modal contribution can be, but is not limited to, the following: , in, For modal sensor data Modal contribution For modal sensor data The basic robustness coefficient, As a scene factor, Scene factors Influence weight, For modal sensor data In scene factors The performance loss rate is as follows. This represents the total number of modes.
[0026] The above technical solution clarifies the expression for modal contribution and, by integrating the basic robustness coefficient with the weighted performance loss of multiple scene factors (illuminance, speed, occlusion), achieves accurate quantitative evaluation of the contribution of sensor data under different environmental conditions, significantly improving the reliability and environmental adaptability of modal selection.
[0027] Optionally, in one embodiment of this application, the determining module includes: a first calculation unit, configured to calculate the final timestamp and final coordinate value of the modal sensor data based on the modal contribution; a second calculation unit, configured to calculate the predicted gesture fusion feature data between the modal sensor data based on the modal contribution; and a first generation unit, configured to obtain the initial gesture fusion feature data based on the final timestamp, the final coordinate value, the predicted gesture fusion feature data, the modal contribution, and the initial gesture feature data.
[0028] The above technical solution can be used to calculate the final timestamp, final coordinate value and predicted gesture fusion feature data of modal sensor data based on modal contribution. Then, the final timestamp, final coordinate value, predicted gesture fusion feature data, modal contribution and initial gesture feature data are integrated to generate initial gesture fusion feature data. By dynamically fusing the final timestamp, final coordinate value and predicted gesture fusion feature data of multi-sensor data through modal contribution, the accuracy and anti-interference ability of gesture feature fusion are effectively improved, and the robustness of gesture recognition in complex scenarios is enhanced.
[0029] Optionally, in one embodiment of this application, the formula for calculating the final timestamp and the final coordinate value may be, but is not limited to, the following: , in, , For modal sensor data The corrected timestamp and 3D coordinates, i.e., the final timestamp and final coordinate values. , For the initial timestamp and initial coordinates, , For the system's reference time and reference coordinates, , The average timestamp and average coordinates for all modal sensor data. This is a time correction factor. This is the spatial correction factor. For modal sensor data Modal contribution For the vehicle's real-time acceleration, For modal sensor data The original time deviation, This refers to the system's time resolution. The expression for the predicted gesture fusion feature data can be, but is not limited to, the following: , in, For prediction Predictive gesture fusion feature data at any given time. Weighted by historical features for Historical gesture feature data at any given moment This is the trend amplification factor. For modal sensor data Modal contribution , For modal sensor data exist The characteristic first derivative and characteristic second derivative at time t. Represents the predicted time interval. This represents the total number of modes.
[0030] By introducing key parameters such as modal contribution, real-time vehicle acceleration, and feature derivatives, and combining them with system benchmarks and historical feature data, the above technical solution achieves accurate correction of the spatiotemporal parameters of modal sensor data and reliable prediction of gesture fusion features. This improves the scene adaptability of data processing and the accuracy of prediction results. Optionally, in one embodiment of this application, the control module includes: a first acquisition unit, configured to acquire the current vehicle speed and the fatigue index of the at least one driver or passenger; a third calculation unit, configured to calculate a gesture feature matching threshold of the initial gesture fusion feature data based on the current vehicle speed and the fatigue index; a recognition unit, configured to recognize the instruction mapping information of the initial gesture fusion feature data based on the current vehicle speed; a first determination unit, configured to determine the interaction mode of the initial gesture fusion feature data based on the fatigue index; and a second determination unit, configured to determine the final gesture fusion feature data based on the initial gesture fusion feature data, the gesture feature matching threshold, the instruction mapping information, and the interaction mode.
[0031] The above technical solution can calculate the gesture feature matching threshold, identify command mapping information, and determine the interaction mode of the initial gesture fusion feature data based on the vehicle's current speed and the fatigue index of the driver and passengers. Then, the final gesture fusion feature data can be determined. By dynamically combining vehicle speed and the fatigue index of the driver and passengers, the gesture feature matching threshold, command mapping and interaction mode can be adaptively adjusted to achieve accurate screening and personalized response of gesture commands in complex driving scenarios, which can significantly improve interaction safety and user experience.
[0032] Optionally, in one embodiment of this application, the control module includes: a third determining unit, configured to determine the initial gesture intent of the final gesture fusion feature data based on a preset gesture intent matching rule; a fourth calculating unit, configured to calculate the intent confidence of the initial gesture intent based on the modal contribution; and a second generating unit, configured to generate the gesture instruction execution action based on the gesture intent whose intent confidence is greater than a preset threshold.
[0033] The above technical solution can determine the initial gesture intent based on the preset gesture intent matching rules, calculate the intent confidence based on the modal contribution, and then generate gesture command execution actions based on gesture intents with intent confidence greater than a preset threshold. The intent confidence is quantified by the preset gesture intent matching rules and modal contribution, and gesture command execution actions are generated based on the preset threshold, which effectively improves the accuracy of gesture intent recognition and the reliability of command execution, and avoids the risk of misoperation.
[0034] Optionally, in one embodiment of this application, the control module includes: a second acquisition unit, configured to acquire the gesture distance value of the corresponding driver / passenger based on the final gesture fusion feature data in response to the interaction mode being a multi-person interaction mode; a fourth determination unit, configured to determine the priority weight value of the corresponding driver / passenger based on the identity identifier; and a third generation unit, configured to generate the gesture command execution action based on the gesture distance value and the priority weight value.
[0035] The above technical solution can respond to multi-person interaction mode by obtaining the gesture distance value of the driver and passenger based on the final gesture fusion feature data. When the gesture distance value meets the preset distance conditions, the identity identifier of the driver and passenger is extracted and the corresponding priority weight value is determined. Then, the gesture distance value and priority weight value are fused to generate gesture command execution action. In multi-person interaction mode, through the dual judgment logic of preset distance condition filtering and driver and passenger identity priority weighting, the accurate recognition and differentiated execution of gesture commands are realized, improving the accuracy of command response and interaction efficiency in multi-person interaction scenarios.
[0036] A third aspect of this application provides an electronic device, including a processor and a memory, wherein the memory is used to store computer programs; and the processor is used to execute the programs stored in the memory to implement the gesture command recognition method as described in the above embodiments.
[0037] A fourth aspect of this application provides a vehicle that includes electronic equipment as described in the above embodiments.
[0038] A fifth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the gesture command recognition method described above. Attached Figure Description
[0039] Figure 1 This is a flowchart of a gesture command recognition method provided in an embodiment of this application; Figure 2 This is a flowchart illustrating the working principle of a gesture command recognition method provided in one embodiment of this application; Figure 3 This is a structural diagram of the gesture command recognition device provided in the embodiments of this application; Figure 4 This is a structural diagram of the electronic device provided in the embodiments of this application.
[0040] Figure label: Among them, 10-Gesture command recognition device; 100-First acquisition module, 200-Determination module, 300-Control module; 20-Electronic device; 410-Processor, 420-Memory. Detailed Implementation
[0041] To make the technical problems, technical solutions, and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0042] Example 1 This application provides a method for recognizing gesture commands. Please refer to [the relevant documentation]. Figure 1 This includes the following steps: In step S101, in response to a gesture command from at least one occupant, multimodal sensor data corresponding to the gesture command is acquired. The multimodal sensor data includes at least joint sensor data, limb sensor data, camera data, and millimeter-wave radar data.
[0043] It is understood that, in the embodiments of this application, multimodal sensor data may include, but is not limited to, joint sensor data, limb sensor data, RGB (Red Green Blue) camera data, camera data based on Taylor series distortion correction, optical flow camera data, depth camera data, infrared camera data, and millimeter-wave radar data, etc., and this application does not impose specific limitations.
[0044] Among them, the joint sensor data includes the angles and movement trajectories of the driver's and passengers' hand joints, which can accurately capture the basic movement details of gestures and provide accurate movement dimension data for subsequent feature extraction. This data can be obtained through the joint sensor.
[0045] Body sensor data includes the height of the driver's or passenger's arm as it is raised and the angle of its swing. This data can supplement the overall posture information of hand gestures and avoid the bias in action judgment caused by relying solely on hand data. This data can be obtained through body sensors.
[0046] RGB camera data includes color images of the driver's and passengers' hands and the surrounding environment. It is used to reproduce the color differences between the hands and the environment, which can help distinguish the hand area from background interference elements. This data can be obtained through an RGB camera.
[0047] Camera data based on Taylor series distortion correction is used to describe high dynamic range gesture images, which can be acquired by a camera based on Taylor series distortion correction.
[0048] Optical flow camera data includes the direction and speed of the driver's and passengers' hand gestures, which is used to clarify the movement trend of the gestures and provide dynamic evidence. This data can be obtained through optical flow cameras.
[0049] Depth camera data includes the relative distance between the driver's / passenger's hands and the cockpit equipment, as well as three-dimensional contour point clouds. This data can determine the spatial position of the hands within the smart cockpit, avoiding misjudgments of commands due to the hands being outside the effective interaction range. This data can be obtained through depth cameras.
[0050] Because of the ample lighting, the infrared camera can capture clear grayscale images of the hand, which can then be used as supplementary data to the color image, further improving the clarity of the hand image. This data can be obtained through the infrared camera.
[0051] Millimeter-wave radar data can detect unobstructed hands, and the radial velocity is consistent with the hand movement speed, eliminating invalid data caused by hand obstruction. It also verifies the authenticity of the hand movement speed. All collected raw data is stored synchronously, providing complete and reliable data source support for subsequent processing modules. This data can be obtained through millimeter-wave radar.
[0052] For example, in a high-speed driving scenario, when executing a driver's song-changing command, the embodiments of this application can collect joint sensor data such as the joint angle and movement trajectory of the driver's hand sliding horizontally to the left through joint point sensors; collect limb sensor data such as the slight lifting height and swing angle of the arm through limb sensors; collect RGB camera data such as color images of the hand and the surrounding environment through an RGB camera; collect optical flow camera data such as the gesture movement direction being horizontal to the left and the movement speed being stable through an optical flow camera; collect depth camera data such as the relative distance between the hand and the central control screen being 0.6 meters and the three-dimensional contour point cloud of the hand through a depth camera; collect infrared camera data such as clear grayscale images of the hand through an infrared camera as supplementary data for the color images; and detect that the hand is unobstructed and that the radial velocity is consistent with the gesture movement speed through millimeter-wave radar data, eliminating invalid data caused by the hand being obstructed, and verifying the authenticity of the gesture movement speed.
[0053] In addition, in urban parking scenarios, when a passenger adjusts the right rearview mirror, the embodiments of this application can collect joint sensor data such as the joint angle and movement trajectory of the passenger's hand as it first makes a counterclockwise circle and then points to the lower right; collect limb sensor data such as the appropriate arm raising height and swing angle; collect color images of the hand and the surrounding environment through an RGB camera to restore the color difference between the hand and the cabin interior; collect optical flow camera data such as the movement state and speed of the gesture as it first moves in a circle and then points in a fixed direction; collect depth camera data such as the relative distance between the hand and the central control screen being 0.4 meters and the three-dimensional contour point cloud of the hand; collect infrared camera data such as the grayscale image of the hand under normal lighting in the cabin as supplementary data for the color images; and detect that the hand is not obstructed by millimeter-wave radar data to exclude invalid data caused by the hand being obstructed by the passenger's own clothing or cabin items.
[0054] In some embodiments, this application can acquire multimodal sensor data corresponding to a gesture command when at least one occupant issues such a command. The occupant may include, but is not limited to, the driver and passengers; this application does not impose specific limitations.
[0055] Furthermore, in the embodiments of this application, gesture commands can be understood as a non-contact interaction method that conveys control intentions to the vehicle system through specific movements, postures, or movement trajectories of the driver's or passenger's hands or arms. These may include, but are not limited to, hand joint angles, movement trajectories, movement directions, movement speeds, the relative distance between the hand and the cockpit equipment, and the three-dimensional contour point cloud of the hand, etc. This application does not impose specific limitations.
[0056] Optionally, in one embodiment of this application, before determining the initial gesture fusion feature data between modal sensor data, the method further includes: identifying the basic robustness coefficient of the modal sensor data and the scene factor corresponding to the modal sensor data, wherein the basic robustness coefficient characterizes an index used to measure the effectiveness, stability, and consistency of the modal sensor data, and the scene factor characterizes parameters for the working environment of the sensor; obtaining the influence weight of the scene factor and the performance loss rate of the modal sensor data under the influence of the scene factor, wherein the performance loss rate is used to measure the proportion of performance degradation of the modal sensor data under the influence of the scene factor; and calculating the modal contribution based on the basic robustness coefficient, the scene factor, the influence weight, and the performance loss rate. The expression for the modal contribution is: , in, For modal sensor data Modal contribution For modal sensor data The basic robustness coefficient, As a scene factor, Scene factors Influence weight, For modal sensor data In scene factors The performance loss rate is as follows. This represents the total number of modes.
[0057] It is understood that, in the embodiments of this application, modal sensor data can be understood as data collected from different channels, such as joint sensor data, limb sensor data, RGB camera data, camera data based on Taylor series distortion correction, optical flow camera data, depth camera data, infrared camera data, millimeter-wave radar data, etc., and this application does not impose specific limitations.
[0058] The basic robustness coefficient is a core quantitative indicator that measures the validity, stability, and consistency of modal sensor data when faced with external interference, noise, environmental fluctuations, or minor equipment failures. It is usually a dimensionless value between 0 and 1: the closer the coefficient is to 1, the stronger the data's resistance to interference, the less affected it is by external factors, and the better its robustness; the closer the coefficient is to 0, the more susceptible the data is to interference, the poorer its stability, and the weaker its robustness.
[0059] Scene factors refer to key characteristic parameters that characterize the working environment of a sensor. These parameters directly or indirectly affect the sensor's acquisition accuracy, data validity, and the decision-making logic of subsequent algorithms. They may include, but are not limited to, scene factors such as illumination intensity, motion speed, and occlusion.
[0060] Among them, the illumination intensity scene factor is used to quantify the light intensity of the environment in which the sensor is located; the motion speed scene factor is used to quantify the relative motion rate between the sensor and the target being measured; the occlusion level scene factor is used to quantify the proportion or degree to which the sensor's detection path is blocked by obstacles, etc., and this application does not impose specific limitations.
[0061] Furthermore, in the embodiments of this application, the influence weight of the scene factor can be used to describe the relative importance of the scene factor to the performance of the modal sensor data. The weight can be assigned based on expert experience, such as setting the influence weight of the light intensity scene factor to 0.7. This application does not impose specific limitations. Alternatively, the influence weight can be learned through a data-driven model, such as a neural network. The specific settings can be made by those skilled in the art according to the actual situation. This application does not impose specific limitations. The performance loss rate can be understood as the percentage decrease in the data performance of a modal sensor, such as recognition accuracy and response speed, under the influence of scene factors.
[0062] In some embodiments, before determining the initial gesture fusion feature data between modal sensor data, this application can first identify the basic robustness coefficients of each modal sensor data and the scene factors corresponding to the modal sensor data, and obtain the influence weights of the scene factors and the performance loss rate of each modal sensor data under the influence of the corresponding scene factors. Based on the basic robustness coefficients, scene factors, influence weights, and performance loss rates, the modal contribution of each modal sensor data can be calculated. The expression for the modal contribution can be, but is not limited to, as follows: , in, For modal sensor data Modal contribution For modal sensor data The basic robustness coefficient, For scene factors, where, The scene factor representing light intensity. Represents the scene factor of motion speed. Scene factors representing the degree of occlusion. Scene factors Influence weight, For modal sensor data In scene factors The performance loss rate is as follows. This represents the total number of modes.
[0063] For example, in a high-speed driving scenario, when executing the driver's song-changing command, the optical flow camera is crucial for judging the dynamic characteristics of the song-changing gesture because it accurately captures the direction and speed of the gesture movement. Therefore, its modal contribution is relatively high. On the other hand, the infrared camera, due to sufficient lighting, has a weaker supplementary effect on the overall gesture recognition by acquiring grayscale images. Therefore, its modal contribution is relatively low.
[0064] In addition, in the urban parking scenario of this application embodiment, when the passenger executes the command to adjust the right rearview mirror, the joint point sensor accurately captures the step-by-step joint movements of the complex gesture, and the depth camera accurately obtains the spatial position and three-dimensional shape of the hand. Both are crucial to judging the complex gesture intention of "adjusting the right rearview mirror", and therefore have a high modal contribution. On the other hand, because the current lighting is normal, the grayscale image of the infrared camera has a low supplementary value to the overall gesture recognition, and therefore has a low modal contribution.
[0065] Before acquiring the initial gesture fusion feature data, this application introduces a basic robustness coefficient and a scene factor. By quantifying the loss rate and influence weight of the scene on sensor performance, the modal contribution of each modal sensor data is calculated, which serves as the weight basis for subsequent fusion processing. By constructing a modal contribution evaluation mechanism that considers the basic robustness of fusion and dynamic scene loss, the quality of each modal sensor data is refined and dynamically quantified, thereby improving the robustness and data effectiveness of gesture feature fusion in complex and ever-changing environments.
[0066] In step S102, the corresponding initial gesture feature data is identified based on each modal sensor data, and the initial gesture fusion feature data between the modal sensor data is determined based on the modal contribution of the modal sensor data and the initial gesture feature data.
[0067] It is understood that, in the embodiments of this application, the initial gesture feature data can be understood as the raw feature data extracted independently from each modal sensor data and not fused. It may include, but is not limited to, gesture feature data, such as the pixel coordinates or spatial coordinates of gesture joints (such as fingers, wrists, etc., which are not specifically limited in this application), the displacement sequence of the gesture center point in consecutive frames, etc., which are not specifically limited in this application; environmental feature data, such as light intensity, light direction, light color, background complexity, similarity between the background and the gesture, and the relative position of the gesture and surrounding objects, etc., which are not specifically limited in this application; and state feature data, such as the body posture of the driver and passengers, the current vehicle speed, etc., which are not specifically limited in this application.
[0068] Furthermore, in the embodiments of this application, the initial gesture fusion feature data can be understood as the comprehensive feature obtained by fusing the initial gesture feature data of each modality, which is used for subsequent gesture classification or regression tasks.
[0069] As one possible implementation, embodiments of this application can obtain the initial gesture feature data corresponding to each modal sensor data from each modal sensor data, and calculate the initial gesture fusion feature data between each modal sensor data based on the modal contribution of each modal sensor data and the obtained initial gesture feature data.
[0070] Optionally, in one embodiment of this application, initial gesture fusion feature data between modal sensor data is determined based on the modal contribution of the modal sensor data and the initial gesture feature data. This includes calculating the final timestamp and final coordinate value of the modal sensor data based on the modal contribution, calculating the predicted gesture fusion feature data between the modal sensor data based on the modal contribution, and obtaining the initial gesture fusion feature data based on the final timestamp, final coordinate value, predicted gesture fusion feature data, modal contribution, and initial gesture feature data. The formulas for calculating the final timestamp and final coordinate value can be, but are not limited to, as follows: , in, , For modal sensor data The corrected timestamp and 3D coordinates, i.e., the final timestamp and final coordinate values. , For the initial timestamp and initial coordinates, , For the system's reference time and reference coordinates, , The average timestamp and average coordinates for all modal sensor data. This is a time correction factor. This is the spatial correction factor. For modal sensor data Modal contribution For the vehicle's real-time acceleration, For modal sensor data The original time deviation, This refers to the system's time resolution. The expression for predicting gesture fusion feature data can be, but is not limited to, as follows: , in, For prediction Predictive gesture fusion feature data at any given time. Weighted by historical features for Historical gesture feature data at any given moment This is the trend amplification factor. For modal sensor data Modal contribution , For modal sensor data exist The characteristic first derivative and characteristic second derivative at time t. Represents the predicted time interval. This represents the total number of modes.
[0071] In some embodiments, the present application can calculate the final timestamp and final coordinate value of modal sensor data based on modal contribution.
[0072] The formulas for calculating the final timestamp and final coordinate values can be, but are not limited to, as follows: , in, , For modal sensor data The corrected timestamp and 3D coordinates, i.e., the final timestamp and final coordinate values. , For the initial timestamp and initial coordinates, , For the system's reference timestamp and reference coordinates, , The average timestamp and average coordinates for all modal sensor data. This is a time correction factor. This is the spatial correction factor. For modal sensor data Modal contribution For the vehicle's real-time acceleration, For modal sensor data The original time deviation, This is the system time resolution.
[0073] In this embodiment, the initial timestamp can be understood as the time point when each modal sensor data is first acquired, typically based on the system clock or the sensor's internal clock; the initial coordinates can be understood as the position coordinates of each modal sensor data in the original acquisition space; the reference timestamp, as a time synchronization reference standard, is typically selected from the system's global clock timestamp to eliminate clock deviations between acquisition devices; the reference coordinates, as a spatial alignment reference standard, are typically selected from predefined calibration points to address coordinate differences caused by different sensor viewing angles or installation positions; and the time correction coefficient is used to adjust the scaling factor of the modal sensor data timestamp to compensate for sensor sampling. Frequency inconsistency or clock drift; coordinate correction coefficients are transformation parameters used to convert the coordinate values of modal sensor data, which may include, but are not limited to, rotation, translation, scaling, etc., and this application does not impose specific limitations; the initial time deviation value can be understood as the fixed deviation between the initial timestamp of the modal sensor data and the reference timestamp, which may be caused by sensor start-up delay or transmission delay. By compensating for the initial deviation, the time synchronization accuracy is further refined; the average timestamp can be understood as the statistical average of the timestamps of multimodal sensor data, used to calculate the global time center; the average coordinate value can be understood as the statistical average of the coordinate values of multimodal sensor data, used to calculate the global spatial center.
[0074] In some embodiments, the present application can calculate predictive gesture fusion feature data between modal sensor data based on modal contribution.
[0075] The expression for predicting gesture fusion feature data can be, but is not limited to, expressed as: , in, For prediction Real-time fusion features Weighted by historical features for Historical gesture feature data at any given moment This is the trend amplification factor. For modal sensor data Modal contribution , For modal sensor data exist The characteristic first derivative and characteristic second derivative at time t. Represents the predicted time interval. This represents the total number of modes.
[0076] Among them, the first derivative of the feature is used to describe the instantaneous rate of change of the initial gesture feature data in the time dimension; the second derivative of the feature is used to describe the rate of change of the rate of change of the initial gesture feature data in the time dimension, reflecting the acceleration or curvature change of the gesture movement; the historical gesture fusion feature data can represent the multimodal fusion feature sequence calculated in the past time step, while the historical feature weight value can represent the importance coefficient of the historical gesture fusion feature data in predicting the current feature; the predicted gesture fusion feature data can represent the fusion feature at the current moment predicted based on the historical gesture fusion feature data, the historical feature weight value, the first derivative of the feature, and the second derivative of the feature.
[0077] In some embodiments, the present application embodiments can obtain the corresponding initial gesture fusion feature data based on the final timestamp, final coordinate value, predicted gesture fusion feature data, modal contribution, and initial gesture feature data.
[0078] It can be understood that the initial gesture fusion feature data in the embodiments of this application may include, but is not limited to, the final timestamp, the final coordinate value, the predicted gesture fusion feature data, the modal contribution, and the initial gesture feature data.
[0079] This application embodiment, based on modal contribution, combines real-time vehicle acceleration and historical feature derivatives to accurately correct the spatiotemporal coordinates of sensor data and predict future feature trends. The corrected spatiotemporal parameters, predicted features, and original features are then fused in multiple dimensions to generate initial gesture fusion feature data. Through the spatiotemporal adaptive correction guided by modal contribution and the derivative-based trend prediction mechanism, the sensor data deviation caused by vehicle motion and environmental interference is effectively compensated, significantly improving the spatiotemporal consistency, prediction accuracy, and anti-interference capability of gesture feature data.
[0080] In step S103, the final gesture fusion feature data is determined based on the initial gesture fusion feature data, and the gesture command execution action of the vehicle is generated according to the final gesture fusion feature data, thereby controlling the vehicle to execute the gesture command execution action.
[0081] In some embodiments, the present application can determine the final gesture fusion feature data based on the initial gesture fusion feature data, thereby generating a gesture command execution action corresponding to the vehicle, and controlling the vehicle to execute the corresponding gesture command execution action.
[0082] For example, in a high-speed driving scenario, when executing a driver's song-changing command, the "song-changing" intent can be parsed and converted into a gesture command that the intelligent cockpit entertainment control unit can recognize. The gesture command is then transmitted to the intelligent cockpit entertainment control unit, triggering the control unit to execute the song-changing operation. After the song-changing operation is completed, a "song changed" text prompt is displayed on the central control screen, and a short confirmation sound effect is played through the car audio system. This allows the driver to know in a timely manner that the song-changing command has been successfully executed, avoiding repeated triggering of the gesture. It also improves the intuitiveness and convenience of the driver's interaction with the intelligent cockpit, ensuring that the entire song-changing command execution process forms a closed loop.
[0083] Furthermore, in urban parking scenarios, when a passenger executes a command to adjust the right rearview mirror, the intent to "adjust the right rearview mirror angle downwards" can be parsed and converted into a gesture command that the rearview mirror control unit in the intelligent cockpit vehicle control unit can recognize. This ensures that the rearview mirror control unit can accurately receive and understand the command content and transmit the gesture command to the rearview mirror control unit, triggering the control unit to perform the downward adjustment of the right rearview mirror angle. After the rearview mirror control unit completes the angle adjustment, it displays the text "Right rearview mirror adjusted downwards" on the central control screen. Simultaneously, a visual icon of the current angle of the right rearview mirror is displayed on the vehicle control interface of the central control screen, allowing the passenger to intuitively know that the command has been successfully executed and the adjusted angle meets expectations. This avoids passengers repeatedly triggering complex gestures due to unclear feedback and ensures that the entire complex gesture command execution process forms a complete closed loop, improving the passenger's interactive experience with the intelligent cockpit.
[0084] Optionally, in one embodiment of this application, determining the final gesture fusion feature data based on the initial gesture fusion feature data includes: acquiring the current vehicle speed and the fatigue index of at least one driver or passenger; calculating the gesture feature matching threshold of the initial gesture fusion feature data based on the current vehicle speed and the fatigue index; identifying the instruction mapping information of the initial gesture fusion feature data based on the current vehicle speed; determining the interaction mode of the initial gesture fusion feature data based on the fatigue index; and determining the final gesture fusion feature data based on the initial gesture fusion feature data, the gesture feature matching threshold, the instruction mapping information, and the interaction mode.
[0085] In some embodiments, the present application can first obtain the current vehicle speed and the fatigue index of at least one driver or passenger, then calculate the gesture feature matching threshold corresponding to the initial gesture fusion feature data based on the current vehicle speed and the fatigue index, and identify the instruction mapping information of the initial gesture fusion feature data based on the current vehicle speed; at the same time, determine the interaction mode of the initial gesture fusion feature data based on the fatigue index, and thus determine the final gesture fusion feature data by combining the initial gesture fusion feature data, the gesture feature matching threshold, the instruction mapping information and the interaction mode.
[0086] The fatigue index can be understood as an indicator that quantifies the degree of fatigue among drivers and passengers. Its calculation formula can be, but is not limited to, expressed as: , in, It is the fatigue index of drivers and passengers, among which, 0 represents complete wakefulness, and 1 represents extreme fatigue. It is the visual feature fusion value. It is a fusion value of physiological characteristics. It is a fusion value of vehicle state features. These are the weights of each modal feature.
[0087] The gesture feature matching threshold is a dynamic threshold used to filter initial gesture fusion feature data. It is adjusted in real time based on the current vehicle speed and fatigue index to improve the robustness of gesture recognition. Its calculation formula can be, but is not limited to, expressed as: , in, Thresholds are matched for gesture features. As the baseline threshold, The coefficient representing the influence of vehicle speed. Current vehicle speed High-speed threshold, The fatigue effect coefficient is... It is the fatigue index of drivers and passengers.
[0088] For example, in this embodiment of the application, when the current vehicle speed is >80km / h or the fatigue index is ≥0.7, the gesture feature matching threshold is increased by 20%-30% compared to the benchmark value. When the current vehicle speed is ≤30km / h and the fatigue index is <0.3, the gesture feature matching threshold is decreased by 10%-15% compared to the benchmark value. The specific settings can be made by those skilled in the art according to the actual situation, and this application does not impose any specific limitations.
[0089] Command mapping information can be divided according to the vehicle's current speed. In high-speed scenarios, only the mapping between single-step gestures and basic commands is retained; in urban scenarios, the mapping between two-step combined gestures and complex commands is allowed; and in parking scenarios, the mapping between three-step and more complex gestures and refined control commands is enabled.
[0090] The interaction modes may include, but are not limited to, single-gesture interaction mode, dual interaction mode with gesture and vibration feedback confirmation, and multi-person interaction mode, etc., and this application does not impose specific limitations. For example, in the embodiments of this application, a single-gesture interaction mode can be used when the driver is not fatigued, and the interaction mode can be switched to dual interaction mode with gesture and vibration feedback confirmation when the driver is fatigued.
[0091] For example, in a high-speed driving scenario, when executing the driver's song-changing command, at a current vehicle speed of 100 km / h and a driver fatigue index of 0.2, the gesture feature matching threshold is increased by 25% compared to the baseline value.
[0092] Furthermore, to improve the accuracy of gesture recognition in high-speed scenarios and reduce misrecognition caused by vehicle bumps or slight hand tremors of the driver, this embodiment of the application uses a command mapping information that only retains the mapping between single-step gestures and basic commands. This simplifies the command triggering logic, allowing the driver to trigger the song-changing command without complex operations during high-speed driving, reducing operational complexity and minimizing distraction from driving attention. The interaction mode uses single-gesture interaction based on the driver's non-fatigue state, eliminating the need for additional confirmation steps and improving command triggering efficiency. Simultaneously, this embodiment of the application detects that only the driver's hand movements are currently active, with no other passengers' hands participating in the interaction, thus eliminating multi-person interaction gesture conflicts and eliminating the need for conflict handling. This allows for the determination of the gesture feature matching threshold, command mapping information, and interaction mode.
[0093] In addition, in the urban parking scenario, when the passenger's command to adjust the right rearview mirror is executed, the current vehicle speed is 20km / h and the passenger's fatigue index is 0.1, so the gesture feature matching threshold is reduced by 12% compared with the baseline value.
[0094] Furthermore, to improve the recognition success rate of complex gestures in parking scenarios and avoid recognition failures caused by subtle deviations in complex gesture movements, this embodiment of the application adjusts the command mapping information to allow mapping between complex gestures of three steps or more and refined control commands. This meets the needs of commands requiring precise control, such as "adjusting the rearview mirror angle," in parking scenarios, allowing passengers to perform refined operations through complex gestures. The interaction mode adopts single-gesture interaction based on the passenger's non-fatigue state, eliminating the need for additional confirmation steps and improving the triggering efficiency of complex gesture commands. Simultaneously, this embodiment of the application uses a depth camera combined with a seat pressure sensor to monitor the number and position of hands in the cabin. It detects that only the passenger is making hand movements, with no other passengers participating in the interaction, thus eliminating multi-passenger interaction gesture conflicts and eliminating the need for conflict handling. This allows for the determination of the gesture feature matching threshold, command mapping information, and interaction mode.
[0095] This application embodiment dynamically combines vehicle speed and driver / passenger fatigue index to adaptively adjust gesture feature matching threshold, command mapping, and interaction mode, thereby achieving accurate selection and personalized response of gesture commands in complex driving scenarios, significantly improving interaction safety and user experience.
[0096] Optionally, in one embodiment of this application, generating a gesture command execution action for a vehicle based on the final gesture fusion feature data includes: determining the initial gesture intent of the final gesture fusion feature data based on a preset gesture intent matching rule; calculating the intent confidence of the initial gesture intent based on modal contribution; and generating a gesture command execution action based on a gesture intent with an intent confidence greater than a preset threshold.
[0097] It is understood that, in the embodiments of this application, the preset gesture intent matching rules may include, but are not limited to, basic gesture rules, combined gesture rules, and complex gesture rules, etc., and can be set by those skilled in the art according to the actual situation. This application does not impose any specific restrictions.
[0098] The basic gesture rules may include, but are not limited to, swiping horizontally to the left to skip to the next song, swiping horizontally to the right to skip to the next song, making a fist to mute, raising the palm to increase the volume, pressing the palm down to decrease the volume, pointing the thumb up to confirm, pointing the thumb down to cancel, pointing to the left side of the screen to switch to the left menu, pointing to the right side of the screen to switch to the right menu, etc. This application does not impose specific restrictions.
[0099] The rules for combining gestures may include, but are not limited to, making a fist and swiping horizontally to the right to skip to the next track and increase the volume, opening your palm and pointing at the screen to confirm the current option, closing your palm and pointing at the screen to cancel the current option, swiping horizontally to the left and raising your palm to skip to the next track and decrease the volume, etc. This application does not impose specific restrictions.
[0100] Complex gesture rules may include, but are not limited to, drawing a circle clockwise and pointing to the lower left to adjust the left rearview mirror angle upwards, drawing a circle counterclockwise and pointing to the lower right to adjust the right rearview mirror angle downwards, horizontal back-and-forth swiping and opening the palm to turn on the air conditioning automatic mode, and vertical up-and-down swiping and clenching a fist to adjust the air conditioning temperature, etc. This application does not impose specific restrictions.
[0101] The formula for calculating the confidence level of intent can be, but is not limited to, as follows: , in, For the confidence level of the intention, The feature matching score between the current gesture and the template is given. Thresholds are matched for gesture features. For contextual influence coefficients, For contextual relevance, and , For modal sensor data Modal contribution This represents the total number of modes.
[0102] In some embodiments, this application can determine the initial gesture intent corresponding to the final gesture fusion feature data based on a preset gesture intent matching rule, and accurately calculate the intent confidence of the initial gesture intent based on the modal contribution of each modal sensor data. Then, gesture intents with intent confidence greater than a preset threshold are selected, thereby generating corresponding gesture command execution actions. The preset threshold can be set by those skilled in the art according to actual conditions, and this application does not impose specific limitations.
[0103] For example, in a high-speed driving scenario, when executing a driver's song-changing command, this embodiment of the application can perform semantic matching based on preset gesture intent matching rules to clearly identify the horizontal left swipe as a song-changing command, directly establishing the association between the current gesture and the preset command, providing a rule-based basis for intent judgment. Next, combined with time-space context awareness, the time sequence of the current gesture is analyzed to confirm that the gesture has no other potential intent, eliminating the possibility of the gesture being misjudged as another command. Subsequently, the intent confidence is calculated using the intent confidence calculation formula. Because the current gesture has a high feature matching score with the song-changing command template, strong contextual correlation, and stable modal contributions, the calculated intent confidence is greater than a preset threshold, ensuring that the determined "song-changing" intent is genuine and reliable, avoiding incorrect command output due to insufficient confidence, and outputting the intent parsing result of "song-changing," thereby converting the "song-changing" intent parsing result into the corresponding gesture command execution action.
[0104] Furthermore, in urban parking scenarios, when a passenger executes a command to adjust the right rearview mirror, this embodiment can perform semantic matching based on preset gesture intent matching rules. This clarifies that "drawing a circle clockwise and pointing to the lower left corresponds to adjusting the left rearview mirror angle upwards," and "drawing a circle counter-clockwise and pointing to the lower right corresponds to adjusting the right rearview mirror angle downwards." This directly establishes the association between the current combination of "drawing a circle counter-clockwise + pointing to the lower right" gesture and the command to "adjust the right rearview mirror angle downwards," providing a clear rule basis for judging complex gesture intents. Next, by combining time-space context awareness, the temporal sequence and spatial location of the current gesture are analyzed to confirm that the gesture has no other potential intent, eliminating the possibility of complex gestures being misinterpreted as other commands. Subsequently, the intent confidence score is calculated using the intent confidence calculation formula. Because the current complex gesture has a high feature matching score with the "adjust the right rearview mirror angle downward" instruction template, strong contextual relevance, and stable contribution of each key modality, the calculated intent confidence score is greater than the preset threshold. This ensures that the determined intent of "adjust the right rearview mirror angle downward" is real and reliable, avoids the erroneous output of complex instructions due to insufficient confidence, and outputs the intent parsing result of "adjust the right rearview mirror angle downward". This converts the intent parsing result of "adjust the right rearview mirror angle downward" into the corresponding gesture instruction execution action.
[0105] This application embodiment quantifies intent confidence by pre-setting gesture intent matching rules and modal contribution metrics, and generates gesture command execution actions based on preset thresholds, effectively improving the accuracy of gesture intent recognition and the reliability of command execution, and avoiding the risk of misoperation.
[0106] Optionally, in one embodiment of this application, generating a gesture command execution action for the vehicle based on the final gesture fusion feature data includes: in response to the interaction mode being a multi-person interaction mode, obtaining the gesture distance value of the corresponding driver / passenger based on the final gesture fusion feature data; if the gesture distance value meets a preset distance condition, obtaining the identity identifier of the corresponding driver / passenger; determining the priority weight value of the corresponding driver / passenger based on the identity identifier; and generating a gesture command execution action based on the gesture distance value and the priority weight value.
[0107] In some embodiments, when the interaction mode is a multi-person interaction mode, the present application can obtain the gesture distance value of the corresponding driver or passenger based on the final gesture fusion feature data, and when the gesture distance value meets the preset distance condition, obtain the identity identifier of the corresponding driver or passenger, such as whether it is a driver or a passenger, and then determine the corresponding priority weight value based on the determined identity identifier, and then generate the corresponding gesture command to execute the action. The preset distance condition can be set by those skilled in the art according to the actual situation, and the present application does not impose specific limitations.
[0108] The gesture distance value can be understood as the straight-line distance between the driver's hand and the central control screen, which can be obtained by combining a depth camera with a seat pressure sensor. This application does not impose specific restrictions.
[0109] For example, in this embodiment of the application, when the interaction mode is a multi-person interaction mode, the interaction can be monitored by a depth camera combined with a seat pressure sensor. When the gesture distance value is less than 0.8 meters, the hands of the corresponding driver and passenger are included in the candidate recognition range. That is, the closest hand is given priority as the recognition object, and the identity identifier of the corresponding driver and passenger is obtained. Then, the corresponding priority weight value is determined according to the identity identifier. For example, the priority weight value of the driver's hand is set to 3, and the priority weight value of the passenger's hand is set to 1. The product of the priority weight value of each hand and the gesture distance value is calculated, and the one with the largest product is selected as the recognition object. When the gesture distance value between the passenger's hand and the central control screen is less than 0.5 meters, the priority weight values of the driver's hand and the passenger's hand are both adjusted to 1, and the one with the smallest distance is selected as the recognition object, thereby generating the corresponding gesture command to execute the action.
[0110] In the multi-person interaction mode, this application embodiment filters valid commands by using gesture distance conditions, determines priority weights by combining driver and passenger identity identifiers, and generates the final execution action by integrating distance values and priority weights. It introduces a dual judgment mechanism of distance filtering and identity priority, which effectively solves the command conflict problem in multi-person interaction scenarios, realizes accurate identification and differentiated response to the intentions of different drivers and passengers, and significantly improves the safety of interaction and user experience.
[0111] The working principle of the gesture command recognition method will be introduced below with reference to several specific embodiments.
[0112] in, Figure 2 This is a flowchart illustrating the working principle of a gesture command recognition method provided in one embodiment of this application.
[0113] Step S201: Multimodal sensor data acquisition module.
[0114] The multimodal sensor data acquisition module integrates data collected by joint sensors, limb sensors, RGB cameras, optical flow cameras, depth cameras, infrared cameras, and millimeter-wave radar, and simultaneously collects and stores multimodal sensor data such as the gestures of drivers and passengers, cabin ambient lighting, vehicle surrounding space, and the original state data of drivers and passengers.
[0115] Step S202: Multimodal fusion low-latency processing module.
[0116] The multimodal fusion low-latency processing module is used to identify the corresponding initial gesture feature data based on each modal sensor data in the multimodal sensor data, calculate the real-time contribution of each modal sensor data through the expression of modal contribution, and fuse the multimodal sensor data based on the real-time contribution of each modality. This allows different modal sensor data to participate in the fusion process according to their importance, avoiding fusion errors caused by the deviation of single modal sensor data. The module also performs temporal and spatial synchronization correction on the modal sensor data, unifies the timestamps and three-dimensional coordinates of each modal sensor data, determines the final timestamp and final coordinate values, eliminates the spatiotemporal deviation caused by differences in acquisition speed and installation position of different sensors, ensures the consistency of fused features in time and space dimensions, and avoids misjudgment of gesture actions due to spatiotemporal asynchrony.
[0117] Furthermore, in this embodiment of the application, the predicted gesture fusion feature data can be calculated by the expression of the predicted gesture fusion feature data, capturing the gesture movement trend, predicting the subsequent movement direction of the driver's hand in advance, reducing the delay of command recognition, and responding to gesture intentions faster, thereby obtaining the initial gesture fusion feature data.
[0118] Step S203: Context-aware adaptation module.
[0119] The context-aware adaptation module receives the initial gesture fusion feature data and monitors the vehicle's current speed and the fatigue index of the driver and passengers in real time. This allows it to determine the gesture feature matching threshold, command mapping information, and interaction mode, providing the semantic parsing module with adaptation data that matches the current scenario and the driver's state, ensuring that the semantic parsing process can accurately match the gesture intent.
[0120] Step S204: Semantic parsing module.
[0121] The semantic parsing module receives the gesture feature matching threshold, instruction mapping information and interaction mode output by the context-aware adaptation module, then performs semantic matching based on the preset gesture intent matching rules and calculates the intent confidence.
[0122] Step S205: Instruction output module.
[0123] The instruction output module receives the intent parsing result from the semantic parsing module and converts the intent into a gesture command that the intelligent cockpit entertainment control unit can recognize, ensuring that the control unit can accurately receive and understand the instruction content and avoid execution failure due to incompatible instruction formats.
[0124] Combination Figure 2 As shown in the embodiment of this application, the driver's song-changing command is executed in a high-speed driving scenario.
[0125] The multimodal sensor data acquisition module integrates joint point sensors, limb sensors, RGB cameras, optical flow cameras, depth cameras, infrared cameras, and millimeter-wave radar. It simultaneously collects and stores multimodal sensor data on the driver's hand gestures, cabin ambient lighting, vehicle surrounding space, and driver status. Specifically, the joint point sensors capture the joint angle and trajectory of the driver's hand sliding horizontally to the left, accurately capturing the basic details of the gesture and providing accurate motion dimension data for subsequent feature extraction. The limb sensors capture the slight lifting height and swing angle of the arm, supplementing the overall posture information of the gesture and avoiding the biased judgment caused by relying solely on hand data. The RGB camera captures color images of the hand and the surrounding environment, restoring the color difference between the hand and the environment, helping to distinguish the hand area from background interference elements. The optical flow camera captures dynamic data of the gesture movement direction being horizontal to the left with a stable speed, clarifying the movement trend of the gesture and providing motion characteristics for judging whether the gesture conforms to the song-cutting command. Provides dynamic data; the depth camera captures the relative distance of the hand to the central control screen (0.6 meters) and the three-dimensional contour point cloud of the hand, determining the spatial position of the hand within the smart cockpit and avoiding misjudgments of commands due to the hand's position exceeding the effective interaction range; the infrared camera, due to sufficient current lighting, captures a clear grayscale image of the hand, serving as supplementary data to the color image and further improving the clarity of the hand image; the millimeter-wave radar detects that the hand is unobstructed and that its radial velocity matches the gesture movement speed, eliminating invalid data due to hand obstruction and verifying the authenticity of the gesture movement speed. All collected modal sensor data are stored synchronously, providing a complete and reliable data source for subsequent processing modules.
[0126] After receiving multimodal sensor data, the multimodal fusion low-latency processing module first extracts features from the modal sensor data collected by each sensor, and then identifies the corresponding initial gesture feature data. Subsequently, it calculates the real-time modal contribution of each modal sensor data using an expression for modal contribution. The optical flow camera, due to its accurate capture of the gesture's direction and speed, is crucial for determining the dynamic characteristics of the hand-cutting gesture; therefore, it has a high modal contribution. The infrared camera, due to sufficient current lighting, captures grayscale images that have a weaker supplementary effect on overall gesture recognition; therefore, its modal contribution is lower. This highlights the role of key modal sensor data and reduces the interference of redundant data on the fusion result. Based on the real-time modal contribution of each modality, multimodal features are dynamically fused, allowing data from different modal sensors to participate in the fusion process according to their importance. This avoids fusion errors caused by biases in single-modal sensor data. Next, temporal and spatial synchronization correction is performed to unify the timestamps and 3D coordinates of data from each modality. The final timestamps and coordinate values of each modality's data are calculated, eliminating spatiotemporal deviations caused by differences in acquisition speed and installation location between different sensors. This ensures consistency of fused features in both time and space, preventing misjudgments of gestures due to spatiotemporal asynchrony. Subsequently, the expression of the predicted gesture fusion feature data is used to capture the gesture movement trend and generate predicted gesture fusion feature data. This allows for early prediction of the driver's subsequent hand movement direction, reducing command recognition delays and enabling subsequent modules to respond to gesture intentions more quickly. Finally, initial gesture fusion feature data is output, providing clearly structured and accurate feature data support for the context-aware adaptation module.
[0127] After receiving the initial gesture fusion feature data, the context-aware adaptation module monitors and obtains the vehicle's current speed and the fatigue index of the driver and passengers in real time. For example, when the current speed is 100km / h and the driver's fatigue index is 0.2, the gesture feature matching threshold is increased by 25% compared to the baseline value.
[0128] Furthermore, to improve the accuracy of gesture recognition in high-speed scenarios and reduce misrecognition caused by vehicle bumps or slight hand tremors of the driver, this embodiment of the application uses a command mapping information that only retains the mapping between single-step gestures and basic commands. This simplifies the command triggering logic, allowing the driver to trigger the song-changing command without complex operations during high-speed driving, reducing operational complexity and minimizing distraction from driving attention. The interaction mode uses single-gesture interaction based on the driver's non-fatigue state, eliminating the need for additional confirmation steps and improving command triggering efficiency. Simultaneously, the context-aware adaptation module detects that only the driver's hand movements are currently active, with no other passengers' hands participating in the interaction, thus eliminating multi-person interaction gesture conflicts and eliminating the need for conflict handling. This allows the module to determine the gesture feature matching threshold, command mapping information, and interaction mode, providing the semantic parsing module with adaptation data that matches the current scenario and driver state, ensuring that the semantic parsing process accurately matches the gesture intent.
[0129] The semantic parsing module receives the gesture feature matching threshold, command mapping information, and interaction mode output by the context-aware adaptation module. It then performs semantic matching based on preset gesture intent matching rules, clearly defining the horizontal left swipe as a song-skipping command, directly establishing the association between the current gesture and the preset command, providing a rule-based basis for intent judgment. Next, it combines time-space context awareness to analyze the time series of the current gesture, confirming that the gesture has no other potential intent and eliminating the possibility of the gesture being misjudged as other commands. Subsequently, it calculates the intent confidence using the intent confidence calculation formula. Because the current gesture has a high feature matching score with the song-skipping command template, strong contextual correlation, and stable modal contributions, the calculated intent confidence is greater than the preset threshold, ensuring that the determined "song-skipping" intent is genuine and reliable, avoiding erroneous command output due to insufficient confidence, and outputting the intent parsing result of "song-skipping," providing clear intent information to the command output module.
[0130] The command output module receives the "song change" intent parsing result from the semantic parsing module and converts it into a gesture command execution action that the intelligent cockpit entertainment control unit can recognize. This ensures the control unit can accurately receive and understand the command content, avoiding execution failures due to incompatible command formats. The gesture command execution action is then transmitted to the intelligent cockpit entertainment control unit, triggering the control unit to execute the song change operation. After the control unit completes the song change operation, the command output module provides feedback on the command execution status in a multimodal manner. Specifically, it displays a "song changed" text prompt on the central control screen and plays a brief confirmation sound effect through the vehicle's audio system, allowing the driver to immediately know that the song change command has been successfully executed, avoiding repeated gesture triggers. This enhances the intuitiveness and convenience of the driver's interaction with the intelligent cockpit, ensuring a closed loop in the entire song change command execution process.
[0131] In summary, in the high-speed driving song-changing scenario, the various modules of the system operate collaboratively in this embodiment. The multimodal sensor data acquisition module integrates multiple types of sensors to comprehensively collect gesture, environmental, and status data, providing a complete data source for subsequent processing; the multimodal fusion low-latency processing module outputs initial gesture fusion feature data, ensuring data accuracy and timeliness; the context-aware adaptation module improves the gesture feature matching threshold by 25% using a gesture feature matching threshold algorithm, adapting to the needs of high-speed scenarios; the semantic parsing module accurately parses the song-changing intent based on preset gesture intent matching rules and intent confidence calculation formulas; and the instruction output module completes instruction conversion and feedback, achieving efficient and accurate gesture instruction execution in high-speed scenarios, matching the system's adaptability to highly dynamic scenarios.
[0132] Combination Figure 2 As shown in the embodiment of this application, in an urban parking scenario, the passenger's instruction to adjust the right rearview mirror is executed.
[0133] The multimodal sensor data acquisition module integrates joint point sensors, limb sensors, RGB cameras, optical flow cameras, depth cameras, infrared cameras, and millimeter-wave radar to simultaneously collect and store raw data on passenger hand gestures, cabin ambient lighting, vehicle surrounding space, and passenger status. Specifically, the joint point sensors capture the joint angle and movement trajectory of the passenger's hand as it first draws a counter-clockwise circle and then points to the lower right, accurately capturing the step-by-step details of complex gestures and providing core motion data for subsequently distinguishing the combination of "circling" and "pointing." The limb sensors collect data on the appropriate arm height and swing angle, supplementing the overall posture information of the gesture and avoiding errors in judging the continuity of complex gestures due to relying solely on hand data. The RGB camera captures color images of the hand and surrounding environment, restoring the color difference between the hand and the cabin interior, clearly delineating the hand area from background interference, and ensuring that the gesture area is not obscured by environmental elements for recognition. The optical flow camera captures the motion state and speed of the gesture, first in a circular motion and then pointing in a fixed direction, clarifying the movement trend of each stage of the complex gesture, and providing data for judging "counter-clockwise" and "pointing." The combination of "circling the clock hand + pointing to the lower right" provides dynamic evidence; the depth camera collects the relative distance of the hand and the central control screen at 0.4 meters and the three-dimensional contour point cloud of the hand to determine that the hand is within the effective interaction range of the smart cockpit. At the same time, the three-dimensional point cloud restores the three-dimensional shape of the hand to avoid misjudgment of gestures caused by planar images; the infrared camera collects grayscale images of the hand under normal lighting in the cockpit as supplementary data for the color images, further improving the clarity of the edge details of the hand and ensuring that the subtle movements of complex gestures are not missed; the millimeter-wave radar detects that the hand is unobstructed, eliminating invalid data caused by the hand being obstructed by the passenger's own clothing or cabin items. All collected modal sensor data are stored synchronously to provide complete and coherent data source support for complex gestures for subsequent processing modules.
[0134] After receiving multimodal sensor data, the multimodal fusion low-latency processing module first extracts features from the modal sensor data collected by each sensor, and then identifies the corresponding initial gesture feature data, such as the step-by-step gesture feature data of "drawing a circle counterclockwise" and "pointing to the lower right", cabin ambient lighting feature data, and passenger hand movement status feature data. Subsequently, the real-time modal contribution of each modal sensor data is calculated through the expression of modal contribution. Among them, the joint point sensor is crucial for judging the complex gesture intention of "adjusting the right rearview mirror" because it accurately captures the step-by-step joint movements of complex gestures, and the depth camera is crucial for judging the complex gesture intention of "adjusting the right rearview mirror" because it accurately acquires the spatial position and three-dimensional shape of the hand. Therefore, the modal contribution of the infrared camera is relatively low because the current lighting is normal and its grayscale image has less supplementary value for the overall gesture recognition. This algorithm can highlight the key modal sensor data for complex gesture recognition and reduce the interference of redundant data on the fusion result. Based on the real-time modal contribution of each modality, multimodal features are dynamically fused, allowing data from different modal sensors to participate in the fusion process according to their importance. This avoids errors in the judgment of complex gesture combination logic caused by the bias of data from a single modal sensor. Next, temporal and spatial synchronization correction is performed to unify the timestamps and 3D coordinates of data from each modal sensor, calculating the final timestamps and coordinate values of each modal sensor data. This eliminates spatiotemporal deviations caused by differences in acquisition speed and installation location between different sensors, ensuring that the "circling" and "pointing" action stages are temporally continuous and spatially consistent, avoiding the fragmentation of complex gesture judgments due to spatiotemporal asynchrony. Then, the gesture movement trend is captured by predicting the expression of the gesture fusion feature data, generating predictive gesture fusion feature data. This allows for the early judgment of the motion connection logic of the passenger's hand transitioning from "circling" to "pointing," reducing the delay in complex gesture recognition and enabling subsequent modules to respond more quickly to the complete intent of the combined gesture. This outputs initial gesture fusion feature data, providing the context-aware adaptation module with clearly structured and logically coherent complex gesture feature data support.
[0135] After receiving the initial gesture fusion feature data, the context-aware adaptation module monitors and obtains the vehicle's current speed and the fatigue index of the driver and passengers in real time. For example, if the vehicle is currently in an urban parking scenario with a current speed of 20 km / h and a passenger fatigue index of 0.1, it can accurately grasp the parking scenario that requires fine-grained control and the passenger's non-fatigue state, providing scenario and state basis for subsequent parameter adjustments and reducing the gesture feature matching threshold by 12% compared to the baseline value.
[0136] Furthermore, to improve the recognition success rate of complex gestures in parking scenarios and avoid recognition failures caused by subtle deviations in complex gesture movements, this embodiment of the application adjusts the instruction mapping information to allow mapping between complex gestures of three steps or more and refined control instructions. This meets the needs of precise control instructions such as "adjusting the rearview mirror angle" in parking scenarios, allowing passengers to perform refined operations through complex gestures. The interaction mode adopts single-gesture interaction based on the passenger's non-fatigue state, eliminating the need for additional confirmation steps and improving the triggering efficiency of complex gesture instructions. Simultaneously, the context-aware adaptation module monitors the number and position of hands in the cabin using a depth camera combined with seat pressure sensors. It detects that only the passenger has hand movements, with no other passengers' hands participating in the interaction, thus eliminating the need for conflict handling due to conflicts between multiple passengers' gestures. This determines the gesture feature matching threshold, instruction mapping information, and interaction mode, providing the semantic parsing module with adaptation data that conforms to the current refined control scenario and passenger state, ensuring that the semantic parsing process can accurately match the intent of complex gestures.
[0137] The semantic parsing module receives the gesture feature matching threshold, command mapping information, and interaction mode output by the context-aware adaptation module. Then, based on preset gesture intent matching rules, it performs semantic matching to clarify that "drawing a circle clockwise and pointing to the lower left corresponds to adjusting the left rearview mirror angle upwards," and "drawing a circle counter-clockwise and pointing to the lower right corresponds to adjusting the right rearview mirror angle downwards." This directly establishes the association between the current combination gesture of "drawing a circle counter-clockwise + pointing to the lower right" and the command "adjust the right rearview mirror angle downwards," providing a clear rule basis for judging complex gesture intents. Next, combined with time-space context awareness, it analyzes the time sequence and spatial location of the current gesture to confirm that the gesture has no other potential intent, eliminating the possibility of complex gestures being misinterpreted as other commands. Subsequently, the intent confidence score is calculated using the intent confidence calculation formula. Because the current complex gesture has a high feature matching score with the "adjust the right rearview mirror angle downward" instruction template, strong contextual correlation, and stable contribution of each key modality, the calculated intent confidence score is greater than the preset threshold. This ensures that the determined intent of "adjust the right rearview mirror angle downward" is real and reliable, avoids the erroneous output of complex instructions due to insufficient confidence, and outputs the intent parsing result of "adjust the right rearview mirror angle downward", providing clear and accurate complex gesture intent information to the instruction output module.
[0138] The command output module receives the "adjust right rearview mirror angle downwards" intent from the semantic parsing module and converts it into a gesture command that the rearview mirror control unit in the intelligent cockpit vehicle control unit can recognize. This ensures the rearview mirror control unit can accurately receive and understand the command content, avoiding adjustment operation failures due to incompatible command formats or mismatched command parameters. The gesture command is then transmitted to the rearview mirror control unit, triggering it to adjust the right rearview mirror angle downwards. After the rearview mirror control unit completes the angle adjustment, the command output module provides multimodal feedback on the command execution status. Specifically, it displays a text prompt "Right rearview mirror adjusted downwards" on the central control screen, and simultaneously displays a visual icon of the current right rearview mirror angle on the vehicle control interface of the central control screen. This allows passengers to intuitively know that the command has been successfully executed and the adjusted angle meets expectations, preventing passengers from repeatedly triggering complex gestures due to ambiguous feedback. It also ensures that the entire complex gesture command execution process forms a complete closed loop, improving the passenger's interactive experience with the intelligent cockpit.
[0139] In summary, this application's embodiment adapts to refined requirements in the scenario of adjusting the right rearview mirror while parking in an urban area. The multimodal sensor data acquisition module captures complex gestures and environmental data, laying the foundation for complex command recognition; the multimodal fusion low-latency processing module processes the data through a series of algorithms, outputting initial gesture fusion feature data; the context-aware adaptation module reduces the threshold by 12% using a gesture feature matching threshold algorithm, opening up complex gesture mapping; the semantic parsing module clarifies the adjustment intent based on preset gesture intent matching rules and intent confidence calculation formulas; and the command output module converts commands and provides visual feedback, meeting the refined control requirements in parking scenarios and demonstrating the system's adaptability to complex gestures and specific scenarios.
[0140] The gesture command recognition method proposed in this application can respond to the gesture commands of drivers and passengers, acquire corresponding multimodal sensor data, determine the corresponding initial gesture feature data based on each modal sensor data, and determine the initial gesture fusion feature data by combining the modal contribution of the modal sensor data, and then determine the final gesture fusion feature data, thereby generating and executing the vehicle's gesture command execution action. Through multimodal sensor data fusion and feature calculation based on modal contribution, effective modal features are accurately selected and execution actions are generated, achieving highly robust and low-false-rate interactive control of gesture commands in the vehicle cabin. This solves the problems in related technologies that focus only on voice control, do not involve gesture interaction, have limited functionality, and whose multi-sensor gesture recognition is not optimized for the vehicle cabin scenario, does not clearly select and generate accurate execution actions based on fusion features, resulting in low gesture command recognition accuracy, high response latency, and poor user experience.
[0141] Example 2 This application provides a gesture command recognition device. Please refer to... Figure 3 The gesture command recognition device 10 includes a first acquisition module 100, a determination module 200, and a control module 300.
[0142] The first acquisition module 100 is used to acquire multimodal sensor data corresponding to the gesture command in response to a gesture command from at least one driver or passenger. The multimodal sensor data includes at least joint sensor data, limb sensor data, camera data, and millimeter-wave radar data.
[0143] The determination module 200 is used to identify the corresponding initial gesture feature data based on each modal sensor data, and to determine the initial gesture fusion feature data between modal sensor data based on the modal contribution of the modal sensor data and the initial gesture feature data.
[0144] The control module 300 is used to determine the final gesture fusion feature data based on the initial gesture fusion feature data, and generate the gesture command execution action of the vehicle according to the final gesture fusion feature data, and control the vehicle to execute the gesture command execution action.
[0145] Optionally, in one embodiment of this application, it further includes: an identification module, a second acquisition module, and a calculation module.
[0146] The recognition module is used to identify the basic robustness coefficient of the modal sensor data and the scene factor corresponding to the modal sensor data before determining the initial gesture fusion feature data between the modal sensor data. The basic robustness coefficient is a metric used to measure the effectiveness, stability and consistency of the modal sensor data, and the scene factor is a parameter used for the working environment of the sensor.
[0147] The second acquisition module is used to acquire the influence weight of scene factors and the performance loss rate of modal sensor data under the influence of scene factors. The performance loss rate is used to measure the proportion of performance degradation of modal sensor data under the influence of scene factors.
[0148] The calculation module is used to calculate the modal contribution based on the basic robustness coefficient, scenario factor, influence weight, and performance loss rate.
[0149] Optionally, in one embodiment of this application, the expression for modal contribution may be, but is not limited to, as: , in, For modal sensor data Modal contribution For modal sensor data The basic robustness coefficient, As a scene factor, Scene factors Influence weight, For modal sensor data In scene factors The performance loss rate is as follows. This represents the total number of modes.
[0150] Optionally, in one embodiment of this application, the determining module 200 includes: a first calculation unit, a second calculation unit, and a first generation unit.
[0151] The first calculation unit is used to calculate the final timestamp and final coordinate value of the modal sensor data based on the modal contribution.
[0152] The second computing unit is used to calculate the predicted gesture fusion feature data between modal sensor data based on modal contribution.
[0153] The first generation unit is used to obtain initial gesture fusion feature data based on the final timestamp, final coordinate value, predicted gesture fusion feature data, modal contribution, and initial gesture feature data.
[0154] Optionally, in one embodiment of this application, the formula for calculating the final timestamp and the final coordinate value may be, but is not limited to, the following: , in, , For modal sensor data The corrected timestamp and 3D coordinates, i.e., the final timestamp and final coordinate values. , For the initial timestamp and initial coordinates, , For the system's reference time and reference coordinates, , The average timestamp and average coordinates for all modal sensor data. This is a time correction factor. This is the spatial correction factor. For modal sensor data Modal contribution For the vehicle's real-time acceleration, For modal sensor data The original time deviation, This refers to the system's time resolution. The expression for predicting gesture fusion feature data can be, but is not limited to, as follows: , in, For prediction Predictive gesture fusion feature data at any given time. Weighted by historical features for Historical gesture feature data at any given moment This is the trend amplification factor. For modal sensor data Modal contribution , For modal sensor data exist The characteristic first derivative and characteristic second derivative at time t. Represents the predicted time interval. This represents the total number of modes.
[0155] Optionally, in one embodiment of this application, the control module 300 includes: a first acquisition unit, a third calculation unit, an identification unit, a first determination unit, and a second determination unit.
[0156] The first acquisition unit is used to acquire the vehicle's current speed and the fatigue index of at least one driver or passenger.
[0157] The third calculation unit is used to calculate the gesture feature matching threshold of the initial gesture fusion feature data based on the current vehicle speed and fatigue index.
[0158] The recognition unit is used to recognize the instruction mapping information of the initial gesture fusion feature data based on the current vehicle speed.
[0159] The first determining unit is used to determine the interaction mode of the initial gesture fusion feature data based on the fatigue index.
[0160] The second determining unit is used to determine the final gesture fusion feature data based on the initial gesture fusion feature data, gesture feature matching threshold, instruction mapping information and interaction mode.
[0161] Optionally, in one embodiment of this application, the control module 300 includes: a third determining unit, a fourth calculating unit, and a second generating unit.
[0162] The third determining unit is used to determine the initial gesture intent of the final gesture fusion feature data based on the preset gesture intent matching rules.
[0163] The fourth calculation unit is used to calculate the intent confidence of the initial gesture intent based on the modal contribution.
[0164] The second generation unit is used to generate gesture instructions to perform actions based on gesture intentions with an intent confidence level greater than a preset threshold.
[0165] Optionally, in one embodiment of this application, the control module 300 includes: a second acquisition unit, a third acquisition unit, a fourth determination unit, and a third generation unit.
[0166] The second acquisition unit is used to acquire the gesture distance value of the corresponding driver and passenger based on the final gesture fusion feature data in response to the interaction mode being a multi-person interaction mode.
[0167] The third acquisition unit is used to acquire the identity identifier of the corresponding driver or passenger when the gesture distance value meets the preset distance conditions.
[0168] The fourth determining unit is used to determine the priority weight value of the corresponding driver and passenger based on the identity identifier.
[0169] The third generation unit is used to generate gesture commands to execute actions based on gesture distance values and priority weight values.
[0170] The gesture command recognition device proposed in this application can respond to the gesture commands of drivers and passengers, acquire corresponding multimodal sensor data, determine the corresponding initial gesture feature data based on each modal sensor data, and determine the initial gesture fusion feature data by combining the modal contribution of the modal sensor data, and then determine the final gesture fusion feature data, thereby generating and executing the vehicle's gesture command execution action. Through multimodal sensor data fusion and feature calculation based on modal contribution, effective modal features are accurately selected and execution actions are generated, achieving highly robust and low-false-rate interactive control of gesture commands in the vehicle cabin. This solves the problems in related technologies that focus only on voice control, do not involve gesture interaction, have limited functionality, and whose multi-sensor gesture recognition is not optimized for the vehicle cabin scenario, does not clearly select and generate accurate execution actions based on fusion features, resulting in low gesture command recognition accuracy, high response latency, and poor user experience.
[0171] This application also provides an electronic device 20, please refer to... Figure 4 It includes a processor 410 and a memory 420, wherein the memory 410 is used to store computer programs; the processor 420 is used to execute the programs stored in the memory 410 to implement the gesture command recognition method described in any embodiment of this application.
[0172] This application also provides a vehicle that includes the electronic equipment described in any embodiment of this application.
[0173] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the gesture command recognition method described in any embodiment of this application.
[0174] In this application, "multiple" refers to two or more.
[0175] In this application, unless otherwise expressly defined, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0176] The terms “first,” “second,” “third,” “fourth,” etc., in this application (if present) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0177] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0178] Unless otherwise specified, all steps in this application may be performed sequentially or randomly. For example, if a method includes steps A and B, it means that the method may include steps A and B performed sequentially, or it may include steps B and A performed sequentially. For example, if a method may also include step C, it means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or it may include steps A, C, and B, or it may include steps C, A, and B, etc.
[0179] The above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for recognizing gesture commands, characterized in that, Includes the following steps: In response to a gesture command from at least one driver or passenger, multimodal sensor data corresponding to the gesture command is acquired, wherein the multimodal sensor data includes at least joint sensor data, limb sensor data, camera data, and millimeter-wave radar data; The initial gesture feature data corresponding to each modal sensor data is identified, and the initial gesture fusion feature data between the modal sensor data is determined based on the modal contribution of the modal sensor data and the initial gesture feature data. Based on the initial gesture fusion feature data, the final gesture fusion feature data is determined, and the gesture command execution action of the vehicle is generated according to the final gesture fusion feature data, thereby controlling the vehicle to execute the gesture command execution action.
2. The method according to claim 1, characterized in that, Before determining the initial gesture fusion feature data between the modal sensor data, the method further includes: Identify the basic robustness coefficient of the modal sensor data and the scene factor corresponding to the modal sensor data, wherein the basic robustness coefficient is a metric used to measure the effectiveness, stability and consistency of the modal sensor data, and the scene factor is a parameter used for the working environment of the sensor; Obtain the influence weight of the scene factor and the performance loss rate of the modal sensor data under the influence of the scene factor, wherein the performance loss rate is used to measure the proportion of performance degradation of the modal sensor data under the influence of the scene factor; The modal contribution is calculated based on the basic robustness coefficient, the scenario factor, the influence weight, and the performance loss rate.
3. The method according to claim 1, characterized in that, The expression for the modal contribution is: , in, For modal sensor data Modal contribution For modal sensor data The basic robustness coefficient, As a scene factor, Scene factors Influence weight, For modal sensor data In scene factors The performance loss rate is as follows. This represents the total number of modes.
4. The method according to claim 1, characterized in that, The step of determining the initial gesture fusion feature data between the modal sensor data based on the modal contribution of the modal sensor data and the initial gesture feature data includes: Based on the modal contribution, calculate the final timestamp and final coordinate value of the modal sensor data; Based on the modal contribution, predictive gesture fusion feature data between the modal sensor data are calculated; The initial gesture fusion feature data is obtained based on the final timestamp, the final coordinate value, the predicted gesture fusion feature data, the modal contribution, and the initial gesture feature data.
5. The method according to claim 4, characterized in that, in, The formulas for calculating the final timestamp and the final coordinate value are as follows: , in, 、 For modal sensor data The corrected timestamp and 3D coordinates, i.e., the final timestamp and final coordinate values. , For the initial timestamp and initial coordinates, , For the system's reference time and reference coordinates, , The average timestamp and average coordinates for all modal sensor data. This is a time correction factor. This is the spatial correction factor. For modal sensor data Modal contribution For the vehicle's real-time acceleration, For modal sensor data The original time deviation, This refers to the system's time resolution. The expression for the predicted gesture fusion feature data is: , in, For prediction Predictive gesture fusion feature data at any given time. Weighted by historical features for Historical gesture feature data at any given moment This is the trend amplification factor. For modal sensor data Modal contribution , For modal sensor data exist The characteristic first derivative and characteristic second derivative at time t. Represents the predicted time interval. This represents the total number of modes.
6. The method according to claim 1, characterized in that, The step of determining the final gesture fusion feature data based on the initial gesture fusion feature data includes: Obtain the current speed of the vehicle and the fatigue index of the at least one driver and passenger; Based on the current vehicle speed and the fatigue index, calculate the gesture feature matching threshold of the initial gesture fusion feature data; Based on the current vehicle speed, the instruction mapping information of the initial gesture fusion feature data is identified; Based on the fatigue index, the interaction mode of the initial gesture fusion feature data is determined; The final gesture fusion feature data is determined based on the initial gesture fusion feature data, the gesture feature matching threshold, the instruction mapping information, and the interaction mode.
7. The method according to claim 6, characterized in that, The step of generating the gesture command execution action for the vehicle based on the final gesture fusion feature data includes: In response to the interaction mode being a multi-person interaction mode, the gesture distance value of the corresponding driver and passenger is obtained based on the final gesture fusion feature data; If the gesture distance value meets the preset distance condition, the identity identifier of the corresponding driver or passenger is obtained; Based on the identity identifier, the priority weight value of the corresponding driver and passenger is determined; Based on the gesture distance value and the priority weight value, the gesture command is generated to execute the action.
8. The method according to claim 1, characterized in that, The step of generating the gesture command execution action for the vehicle based on the final gesture fusion feature data includes: Based on preset gesture intent matching rules, the initial gesture intent of the final gesture fusion feature data is determined; Based on the modal contribution, calculate the intent confidence of the initial gesture intent; The gesture instruction is generated to perform an action based on the gesture intent whose confidence level is greater than a preset threshold.
9. An electronic device, characterized in that, Including processor and memory, among which, Memory, used to store computer programs; A processor for executing a program stored in memory to implement the gesture command recognition method according to any one of claims 1-8.
10. A vehicle, characterized in that, It includes the electronic device as described in claim 9.
Citation Information
Cited By
A vehicle control method and a vehicle
CN122300209B