Simulated driving training interaction method and system based on multi-channel gesture fusion
Through multi-channel gesture fusion technology and neural network recognition, the existing simulated driving training system has been solved, and a high-precision and flexible simulated driving training interactive method is realized.
Patent Information
- Application Number
- CN202510663703.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
AI Technical Summary
The existing simulated driving training system has limited usage scenarios, poor flexibility and lacks realism; the volleyball natural gesture interaction method has poor stability and recognition effect, making it difficult to customize gestures for different models and scenarios.
Using a simulated driving training interaction method based on multi-channel gesture fusion, two RGB-d depth cameras are used to collect multi-angle data of the user's hands, and through neural networks and reinforcement learning technology, multi-channel data is fused to identify user gestures and generate control instructions.
It improves the authenticity and flexibility of simulated driving training, enhances the accuracy and robustness of gesture recognition, and enables customized gesture recognition and control for different models and scenarios.
Smart Images

Figure CN120179080A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human-computer interaction, and particularly relates to a simulation driving training interaction method and system based on multi-channel gesture fusion, which is applicable to driving assistance training based on three-dimensional simulation scenarios such as large screens, AR, and VR. Background Art
[0002] In recent years, the development of new energy vehicles has promoted the popularization of electric vehicles, posing new challenges to automobile driving training. Simulation driving training helps beginners experience the driving environment and get familiar with vehicle operations in advance, which is a necessary link before scholars use real cars for training. The existing simulation driving training systems have the following problems:
[0003] Professional driving training systems generally rely on large vehicle simulation devices and are difficult to use at home. Most simulation training software depends on vehicle simulation input peripherals, which cannot adapt to the complex and diverse innovative designs such as irregular steering wheels, column shifters, and intelligent instrument panels in the automobile market. In addition, using traditional computer interaction input devices, such as keyboards and mice, for simulation driving training does not conform to the operation habits of users in real driving environments, and users lack a sense of presence and immersion during the training process.
[0004] Air natural gesture interaction is a natural and intuitive interaction method that enables users to get rid of physical interaction input devices and use hand postures to operate computer systems. The method of combining air natural gesture interaction technology with automobile simulation driving allows users to obtain a relatively real simulation of vehicle driving and operation experience of in-vehicle devices with only ordinary computers and gesture sensor devices at a lower cost. The existing air natural gesture interaction methods have the following problems in actual driving training simulation applications:
[0005] The input data is single-channel, and the anti-noise and anti-interference performance is poor. The system robustness is insufficient, and it is easy to lose the tracking target when the sensor is blocked by an obstacle. The sensor cannot solve the self-occlusion problem of air gesture interaction, and the error of hand movement tracking is large, making it impossible to recognize refined gesture operations. The system mainly supports general gestures and it is difficult to customize different gestures for different vehicle models and scenarios. Summary of the Invention
[0006] The purpose of the present invention is to solve the deficiencies of the existing simulation driving training systems, such as limited usage scenarios and poor flexibility; lack of realism in simulation training software; poor stability and recognition effect of air natural gesture interaction methods; and difficulty in customizing different gestures for different vehicle models and scenarios. The present invention proposes a simulation driving training interaction method and system based on multi-channel gesture fusion to solve the problems existing in the above background.
[0007] The basic hardware required for the technical solution of the present invention includes two RGB-d depth cameras, a display device, and a computer host. The two RGB-d depth cameras are placed at a 90-degree angle in the user's hand operation area to collect data directly in front of and directly below the user's hand respectively.
[0008] Preferably, the above RGB-d depth camera consists of a number of infrared fill lights, a wide-angle infrared camera, and a wide-angle color camera, and is used for image acquisition and position tracking of the user's hands. The RGB-d depth camera can be replaced by a dedicated gesture sensor device, such as LeapMotion.
[0009] An interactive method for simulated driving training based on multi-channel gesture fusion is as follows:
[0010] Step S1: The user's hands are within the common field of view of the two RGB-d depth cameras, and the depth cameras are used to obtain a sequence of dual-camera hand color images and a sequence of spatial position data of dual-camera hand key points.
[0011] Step S2: Preprocess the sequence of dual-camera hand color images obtained in S1 to obtain a high-quality contour of the user's hand movement image. Fuse the contour of the user's hand movement image with the pre-rendered two-dimensional static image of the simulated cockpit of the corresponding vehicle model to generate multi-channel hand movement scene time-series data.
[0012] Step S3: Use a neural network to process the multi-channel hand movement scene time-series data obtained in S2 to obtain the matching preset action types Ml(a), Mr(a) and their confidence levels Cl(a), Cr(a); the multi-channel gesture scene action intention Wo, and its confidence level Co.
[0013] Step S4: Preprocess the spatial position data of the dual-camera user hand key points obtained in S1. Subsequently, perform a spatial transformation to make them in the same coordinate system and merge them into the time-series data of the fused spatial coordinates of the hand key points.
[0014] Step S5: Use a neural network to process the time-series data of the fused spatial coordinates of the hand key points obtained in S4 to obtain the matching multi-channel key point gesture action signals Ml(b), Mr(b) and their confidence levels Cl(b), Cr(b); the gesture key point action intention Wm, and its confidence level Cm.
[0015] Step S6: According to Ml(a), Mr(a), Cl(a), Cr(a), Wo, Co obtained in S3 and Ml(b), Mr(b), Cl(b), Cr(b), Wm, Cm obtained in S5, input them into the dynamic decision-making unit based on reinforcement learning and vector database to obtain the final control instruction and dynamic control coefficient, and at the same time calculate the real-time in-vehicle device hot area list to control the simulated vehicle.
[0016] Through the above steps, the driving states of vehicles with different layouts and models in the simulation system can be controlled according to the gesture pose information, and the operations of in-vehicle devices can be completed, such as controlling the vehicle to accelerate, decelerate, change the direction of the vehicle, shift gears, control interior functions, multimedia, intelligent driving systems, etc.
[0017] Further, step S1 includes:
[0018] S101: Use two RGB-d depth cameras to capture the spatial position data sequence of the user's both hands and the dual-camera hand color image sequence.
[0019] Further, the two RGB-d cameras or gesture sensors are respectively located directly in front of and directly below the user's hand positions, with a 90-degree placement angle, and collect the user's hand data from front to back and from bottom to top respectively. The user's hand operation area is located at the intersection of the acquisition ranges of the two depth cameras. Each RGB-d camera simultaneously captures the hand spatial position data and hand color images of both hands, obtaining the dual-camera hand spatial position data sequence and the dual-camera hand color image sequence.
[0020] S102: Extract the spatial three-dimensional coordinates of the key bone points of each hand from the spatial position data sequence of both hands.
[0021] Further, in a group of spatial position data sequences, the spatial three-dimensional coordinates of each hand include the palm center and a total of h key points such as the base of each finger, the first joint, the second joint, and the fingertip. Each hand obtains two sets of spatial three-dimensional coordinates of bone points from different cameras.
[0022] S103: Record the confidence of the key point coordinates to provide a basis for subsequent data filtering and fusion, ensuring that the output data has sufficient accuracy. If the depth camera can obtain the confidence of each key point coordinate, record the confidence of each key point coordinate, otherwise record the coordinate confidence of the overall data of each channel respectively.
[0023] S104: Divide the key points into two groups according to the left and right hands and encapsulate them into a multi-channel hand key point spatial position data sequence. Among them, the multi-channel hand key point spatial position data sequence contains two-channel data. In channel one, the 2h key point coordinates and confidences of the two cameras of the left hand are stored; in channel two, the 2h key point coordinates and confidences of the two cameras of the right hand are stored.
[0024] Further, step S2 includes:
[0025] S201: Perform image preprocessing on the dual-camera hand color image sequence captured in S1, including removing blurred images and image noise reduction, to provide clear images for subsequent recognition.
[0026] Further, extract the region of interest (ROI) from the image and remove the background, respectively extract and enhance the contour features of the left and right hands, and obtain a sequence of background-free hand image data.
[0027] S202: Perform temporal synchronization processing on the sequence of background-free hand image data to ensure that the image data of the two cameras match in time.
[0028] Further, use a global clock synchronization mechanism to align the data, and obtain a preprocessed sequence of hand action images. If there is time drift, use linear interpolation to adjust the sampling time to ensure accurate frame matching. Use the time window method to construct a continuous sequence of hand action images, providing high-quality input data for subsequent gesture recognition with scenes.
[0029] S203: Identify and separate the left and right hands from the sequence of background-free hand image data, and respectively generate multi-channel hand action image sequences for both hands.
[0030] S204: Obtain the two-dimensional static image of the simulated cockpit corresponding to the selected vehicle model pre-rendered in the computer system. The two-dimensional static image of the simulated cockpit includes two different perspective pictures, namely the front view and the top view of the interior of the cockpit of the vehicle model selected by the user.
[0031] S205: Deform and scale the multi-channel hand action image sequence according to the pre-calibrated conversion parameters, and superimpose and fuse it with the two-dimensional cockpit images from two different perspectives. Obtain two sets of hand-scene superimposed image sequences from different perspectives.
[0032] S206: Generate multi-channel hand action scene temporal data.
[0033] Further, the multi-channel hand action scene temporal data contains three-channel data. In channel one, it contains the multi-channel hand action image sequence of the left hand after preprocessing in step S203; in channel two, it contains the multi-channel hand action image sequence of the right hand after preprocessing in step S203; in channel three, it contains the two sets of hand-scene superimposed image sequences from different perspectives in step S204.
[0034] Further, step S3 includes:
[0035] S301: Import the multi-channel hand action scene temporal data obtained in step S2 into a neural network for recognition. Among them, the hand action image sequences of channels one and two are input into the first set of neural network models; the hand action image sequence of channel three is input into the second set of neural network models.
[0036] Furthermore, the first group of neural network models are pre-trained gesture feature models, which are used to extract feature data of mid-air gestures trained based on dual-camera images; the first group of neural network models are pre-trained instruction feature models, which are used to extract feature data describing the interaction relationship between hands and equipment in the cockpit and operation instructions.
[0037] Preferably, the first group of neural network models uses CNN (convolutional neural network) to extract local features of the hand image.
[0038] Preferably, the second group of neural network models combines LSTM (Long Short-Term Memory Network) to process temporal information to capture dynamic changes of gestures.
[0039] S302: The first group of neural network models uses the data from channels one and two to identify and obtain the preset action types Ml(a), Mr(a) and their confidences Cl(a), Cr(a) for the left and right hands; the second group of neural network models uses the data from channel three to identify and obtain the multi-channel gesture scene action intention Wo and its confidence Co.
[0040] Further, step S4 includes:
[0041] S401: Determine their respective local coordinate systems based on the intrinsic and extrinsic parameters of the two RGB-D depth cameras. According to the pre-calibrated rotation matrix R and translation vector T, convert the dual-camera hand key point data in the multi-channel hand key point spatial position data sequence obtained in step S1 to the same global coordinate system to ensure the spatial consistency of the data.
[0042] S402: Perform time synchronization processing on the converted key point data to ensure that the key point data of the two cameras match in time and avoid position deviation caused by delay, so as to obtain two sets of corrected user hand key point spatial position data.
[0043] Furthermore, the timestamps collected by the two depth cameras are read and the data are aligned using a global clock synchronization mechanism.
[0044] Furthermore, Kalman filtering is used to predict and update the spatial position data of the key points of the user's hand to improve the smoothness and continuity of the coordinates and avoid jumps or jitters.
[0045] S403: The two sets of corrected user hand key point spatial position data obtained in the previous step are coordinate fused according to a confidence algorithm to form a set of user hand key point fused spatial coordinate time series data.
[0046] Suppose the coordinates of the same hand key points collected by two RGB-D depth cameras are corrected as follows: , , the confidence levels of a certain key point for the two cameras are respectively , .
[0047] The confidence level is weighted as: , .
[0048] The coordinates after weighting are: .
[0049] If the confidence level of the key point of a certain camera is low due to occlusion, and the confidence level difference of the coordinates of a certain key point obtained by the two cameras is greater than 0.3, then the coordinate data with low confidence is directly discarded, and the coordinates of the other camera are adopted.
[0050] The time window method is used for division, and the weighted coordinates are merged into the time series data of the spatial coordinates of the hand key point fusion. In the merged time series data of the spatial coordinates of the hand key point fusion, the spatial coordinate data of h key points of the left hand and h key points of the right hand after weighted fusion are saved.
[0051] Further, step S5 includes:
[0052] S501: Process the time series data of the spatial coordinates of the hand key point fusion obtained in S4, and calculate the acceleration and spatial vector of each key point in the sequence.
[0053] S502: Input the spatial coordinates, acceleration and spatial vectors of the hand key point fusion of the left and right hands into a pre-trained gesture recognition network to classify gesture information. The gesture recognition network is a pre-trained model based on Transformer, and a number of aerial gesture feature data based on spatial coordinate information and motion data are saved in the model. A large number of aerial gesture feature data based on spatial coordinate information and motion data are saved in the model.
[0054] Further, the results of gesture information classification include the gesture action types Ml(b), Mr(b) of the left and right hands and their confidence levels Cl(b), Cr(b). And the motion intention Wm and its confidence level Cm.
[0055] Further, step S6 includes:
[0056] S601: Extract the position information of h key points of each hand from the time series data of the spatial coordinates of the hand key point fusion obtained in S4, and calculate the overall spatial position of the hand through summation and averaging. The overall spatial position of the hand can stably represent the overall position of a single hand and reduce the influence of the jitter of individual fingers on the judgment.
[0057] S602: Use the spatial Euclidean distance calculation formula to calculate the distance from the overall spatial position of each hand of the user to the in-vehicle device preset in the simulated driving environment, and output a real-time in-vehicle device hot area list.
[0058] S603: Obtain the gesture action types Ml(a), Mr(a) and their confidence levels Cl(a), Cr(a) obtained in step S3; the gesture action types Ml(b), Mr(b) and their confidence levels Cl(b), Cr(b) obtained in step S5. Obtain the gesture intention Wo and its confidence level Co obtained in step S3; the motion intention Wm and its confidence level Cm obtained in step S5.
[0059] Further, use a non-linear enhancement function to transform Cl(a), Cr(a), Cl(b), Cr(b), Co, Cm, and obtain confidence level weight coefficients S1, S2, S 3…… S6.
[0060] S604: Perform pseudo-natural language encoding on Ml(a), Mr(a), Ml(b), Mr(b), Wo, Wm, S1, S2, S 3…… S6.
[0061] S605: Use BERT to process the pseudo-natural language generated in S604 as a single representation of the entire operation intention, and finally obtain a vector of length d , as the simulated driving operation intention feature.
[0062] Preferably, use BERT or Word2Vec embedding network to generate the intention feature vector. Use the JAVA language to develop and construct a simulated driving vector database network for vector retrieval.
[0063] S606: Use the simulated driving vector data unit to retrieve the simulated driving operation intention feature to obtain the control instruction Ctl, the weight factor k, and calculate the dynamic control coefficient. A large number of pre-trained vectors corresponding control instructions and weight factors have been stored in the simulated driving vector data unit.
[0064] Further, the method for adding new vectors to the simulated driving vector data network is as follows: Insert the feature vectors generated by BERT or BART into the storage structure and save them together with the corresponding control instructions and weight factors. The simulated driving vector data network stores vectors and their corresponding control instructions and weight factors in JSON format.
[0065] Further, the retrieval process described in step S606 is to input a query vector, find the most similar vector by calculating the similarity, and return the corresponding control instruction and weight factor. Use the cosine similarity algorithm to measure the similarity in the direction of two vectors.
[0066] Further, the specific process of calculating the dynamic control coefficient U is to calculate the dynamic control coefficient U of each control instruction based on the weight factor k, and the calculation method is as follows.
[0067]
[0068]
[0069]
[0070] In the above formula, is the ratio of the sum of the movement paths of the fused spatial coordinates of the hand key points to the sum of the movement paths of the corresponding coordinates in the three-dimensional scene within the same time interval. is the ratio of the sum of the movement speeds of the fused spatial coordinates of the hand key points collected within the same time interval to the sum of the movement speeds of the corresponding coordinates in the three-dimensional scene. is the adjustment value pre-stored in the system for calibration. is the weighted movement speed of the action execution. is the speed threshold pre-stored in the system. Further, the weighted movement speed of the action execution The calculation formula is as follows:
[0071]
[0072] In the above formula represents the movement speeds of represents the th key point's weight; represents the total number of hand key points participating in the operation.
[0073] Further, the movement speed of each key point The calculation formula is as follows:
[0074]
[0075] In the above formula, is the coordinate of the th key point at time t, is the time interval.
[0076] S607: Based on the control instruction C, the dynamic control coefficient U, and the list of in-vehicle devices in the hot area, the computer control system controls the simulated vehicle on the screen, and renders and feeds back the operation results using the computer system.
[0077] By dynamically controlling the coefficient U through real-time calculation, the intensity of the user's operation intention when performing hand movements can be analyzed, and then various weight values can be dynamically adjusted in combination with other parameters to achieve dynamic semi-closed-loop control.
[0078] A simulation driving training interaction system based on multi-channel gesture fusion is used to implement the described simulation driving training interaction method based on multi-channel gesture fusion, including a hand movement acquisition module, a gesture movement processing module, a simulation driving interaction control module, a user interface and menu module. Between each module, each module is connected by wired or wireless means to ensure the smooth operation of the simulation driving scenario and the precise control of the vehicle. Among them:
[0079] Hand movement acquisition module: Use a depth camera to capture the spatial position of the user's hand and the real-time hand image from multiple angles, and generate multi-channel gesture tracking data, including the time-series data of the key points of the spatial position of the hand with two cameras and the time-series data of the hand movement images with two cameras.
[0080] Gesture movement processing module: Use a computer system to process the multi-channel gesture tracking data, generate high-precision user hand position and posture data, identify the user's gesture interaction actions and intentions, and parse these data into specific system control instructions.
[0081] Simulation driving interaction control module: Execute corresponding simulation driving operations according to the system control instructions, such as controlling the steering angle of the vehicle in the simulation scenario according to the steering gesture, and switching the gear and speed of the vehicle according to the gear-shifting gesture, etc. The system ensures the real-time nature of the operation, ensures the coherence of the driving process, and synchronizes the changes in the vehicle state in the simulation scenario.
[0082] User interface and menu module: The user can call out the menu and control interface through specific gesture actions to perform operations such as vehicle selection, driving settings, and scene switching. The menu can be displayed at the user-defined position, and the menu has multi-functional switching, such as switching to simulate driving under different weather scenarios, entering the multi-person cooperative driving mode, or resetting the driving scene, etc.
[0083] The beneficial effects of the present invention are as follows:
[0084] The present invention provides a low-cost and real simulation of the operation method and system of a vehicle in reality for the driving training process. Among them, the natural gesture interaction method conforms to the user's intuition, and the training method is closer to reality; the hardware required for the system to achieve natural gesture interaction is only an ordinary computer, a display device, and two depth cameras, with strong flexibility and easy to deploy and promote.
[0085] The present invention combines multiple RGB-d depth cameras to obtain multi-channel fusion time-series data of gestures, effectively reducing data loss caused by occlusion, and improving the accuracy and robustness of gesture recognition.
[0086] The method of the present invention for superimposing the hand contour on the vehicle cockpit screen for recognition can recognize the fine movements of operating in-vehicle devices during simulated driving, and reduce the adaptation and customization costs for different vehicle models and scenarios.
[0087] The system of the present invention performs gesture recognition at different levels (hand images and spatial coordinates) respectively, and uses a semantic vector model for comprehensive decision-making, which can more accurately perceive the user's operation intention, correct errors, and provide a better user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] Figure 1 is the hardware layout of the present invention;
[0089] Figure 2 is the principle framework diagram of the present invention;
[0090] Figure 3 is the step description of the present invention;
[0091] Figure 4 is a schematic diagram of a simulated driving interface according to an embodiment of the present invention;
[0092] Figure 5 is a schematic diagram of a control instruction according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0093] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only partial embodiments of the present invention, rather than all embodiments. The accompanying drawings are only for illustration and are not used to limit the present invention in any way. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0094] The present invention relates to a simulated driving training interaction method based on multi-channel gesture fusion, belonging to the technical field of human-computer interaction. This system is mainly applicable to driving assistance training based on three-dimensional simulation scenarios such as large screens, AR, or VR. Its core technology lies in fusing multi-channel hand data from different angles and different types, and parsing the user's gestures into control instructions through deep learning algorithms, and making comprehensive decisions by the computer system to achieve precise control of the vehicle in the simulated driving scenario.
[0095] This embodiment achieves the above object through the following technical solutions. Its system mainly includes two RGB-d depth cameras, a display device, and a computer host. Each module works in coordination to ensure the high real-time performance and smooth interaction of the system.
[0096] Embodiment 1:
[0097] System hardware composition: 1. Two RGB-d cameras are respectively located directly in front of and directly below the user's hand position, with a 90-degree placement angle, and collect user hand data from front to back and from bottom to top respectively. 2. The computer host preprocesses the collected image data and depth data, recognizes gesture actions and intentions, makes dynamic intelligent decisions, and generates final system control instructions. 3. The display device can be a large screen, AR or VR display terminal, which is used to display the simulated driving scene, vehicle dynamic state and interactive feedback information in real time. 4. Other auxiliary devices include a transmission interface, which is used to ensure the stability and synchronization of data transmission between modules.
[0098] The placement method of the two RGB-d depth cameras is as Figure 1 shown.
[0099] As Figure 2 shown, the main modules of the system include a hand movement acquisition module, a gesture action processing module, a simulated driving interaction control module, and a user interface and menu system module. These modules can be divided into two parts according to their functions: user interaction and intelligent decision-making. Among them, the user interaction part: is used to directly interact with the user, and its function is to receive user gesture operation data input using a depth camera, obtain control instructions, and output a simulated driving screen. The intelligent decision-making part: processes the user gesture signal, maps it to action instructions and intentions, makes dynamic decisions based on the action instructions, intentions and weight parameters, and outputs the control instructions to the user interaction module.
[0100] As Figure 2 shown, the user operates through natural in-air gestures. The system collects the action data of the user's hand through sensors or cameras, and performs multi-channel data fusion to extract key features. Subsequently, the system processes this data to generate corresponding control instructions. These instructions are sent to different functional modules. Finally, the system provides feedback output according to the execution results, and dynamically adjusts the parameters of each model and decision-making system according to the operation feedback to achieve semi-closed-loop control.
[0101] The following is an expansion of the implementation process of each step.
[0102] S1 Hand data acquisition: When the user operates, their hands are respectively placed within the intersection of the field of view of the two RGB-D depth cameras. Each depth camera simultaneously captures real-time hand images and depth data, obtaining a dual-camera hand spatial position data sequence and a dual-camera hand color image sequence.
[0103] Among them, the depth data includes the three-dimensional coordinates (X, Y, Z) and confidence (C) of each hand key point. The sampling frequency is generally set to 120 frames per second to ensure the continuity and real-time nature of the data.
[0104] The collected hand key point data includes the palm center and the key joint points of each finger (such as the finger root, the first joint, the second joint, and the fingertip), a total of 21 key points. The data format is for example: Key point 1(X1, Y1, Z1, C1); Key point 2(X2, Y2, Z2, C2); …; Key point 21(X 21 , Y 21 , Z 21 , C 21 ).
[0105] The above key point data is divided into two groups according to the left and right hands and encapsulated into a multi-channel hand key point spatial position data sequence. Among them, the multi-channel hand key point spatial position data sequence contains two-channel data. In channel one, the coordinates and confidence levels of 42 key points of two camera positions of the left hand are stored; in channel two, the coordinates and confidence levels of 42 key points of two camera positions of the right hand are stored.
[0106] S2 Generate multi-channel hand action scene time series data: The computer host performs clarity detection on the image data collected by two depth cameras, and eliminates blurred or low-quality data; at the same time, performs time series synchronization processing on the two groups of data to ensure the consistency of multi-angle data in time.
[0107] Furthermore, extract the region of interest (ROI) of the image and remove the background, respectively extract and enhance the contour features of the left and right hands, and obtain a background-free hand image data sequence.
[0108] The linear interpolation method is used to adjust the sampling time. Specifically, the system first detects the magnitude of the time drift. For example, if the image frame timestamps of camera A are t1, t2, t3, …, and the timestamps of camera B are t1+Δt, t2+Δt, t3+Δt, …, then the time drift Δt can be calculated by comparing the timestamp sequences of the two cameras. Subsequently, perform linear interpolation processing on the image frames of camera B to generate image frames that match the timestamps of camera A, so as to ensure the accuracy of frame matching.
[0109] In this embodiment, the time window method is used to construct a continuous hand action image sequence. Specifically, the system arranges the aligned image frames in chronological order and performs a sliding window process with a fixed time window length (for example, 1 second) and step size (for example, 0.5 second). For example, for a 10-second hand action image sequence, the system will sequentially extract the following windows:
[0110] Window 1: 0.0 second to 1.0 second, Window 2: 0.5 second to 1.5 second, Window 3: 1.0 second to 2.0 second. And so on, until the entire 10-second image sequence is covered. The image frames within each window are combined into a continuous image sequence for subsequent gesture recognition tasks.
[0111] Identify and separate the left and right hands from a sequence of hand image data without background. The system extracts simulated cockpit images of the vehicle model being operated by the user from the front and top views, and fuses the hand action images with the simulated cockpit images from the front and top views respectively.
[0112] Use the bilinear interpolation method to scale the hand action image to a size that matches the simulated cockpit image. First, the offset parameters are pre-calibrated in the system, and the hand action image is placed at the specified position of the simulated cockpit image. Subsequently, the hand action image is rotated according to the pre-calibrated rotation angle parameter in the system. Finally, three-dimensional deformation correction is performed on the hand action image according to the pre-calibrated distortion deformation parameter in the system.
[0113] Overlay and fuse the processed hand action images with the cockpit two-dimensional images from two different views to obtain two sets of hand-scene overlay image sequences from different views, encapsulate and generate multi-channel hand action scene time series data.
[0114] Furthermore, the multi-channel hand action scene time series data contains three-channel data. In channel one and channel two, there are respectively multi-channel hand action image sequences of the pre-processed left and right hands; in channel three, there are two sets of hand-scene overlay image sequences from different views.
[0115] S3 Generate scene-based multi-channel gesture action signals: Import the multi-channel hand action scene time series data obtained in step S2 into a neural network for recognition. Among them, the hand action image sequences of channel one and two are input into the first set of neural network models; the hand action image sequences of channel three are input into the second set of neural network models.
[0116] In this embodiment, a CNN (Convolutional Neural Network) pre-stored with a standard feature model is used to extract local features of the hand image. An LSTM (Long Short-Term Memory Network) pre-stored with a standard feature model is used to process time series information to capture the dynamic changes of gestures.
[0117] For example, when the user makes an action of pushing the shift lever, consecutive frames show the user's hand pinching and moving, and the hand image after being superimposed with the three-dimensional scene shows the user moving the shift lever to the 5th grid. Use the data of channel one and two to identify and obtain the gesture action signals of the left and right hands, and use the data of channel three to identify and obtain the scene-based gesture operation signals.
[0118] Furthermore, the gesture action signals of the left and right hands include the action types Ml(a), Mr(a) and their confidence levels Cl(a), Cr(a), and the scene-based gesture operation signals include the gesture intention Wo and its confidence level Co.
[0119] S4 Generate fused spatial coordinate time series data:
[0120] According to the internal and external parameters of two RGB - d depth cameras, their respective local coordinate systems are determined separately. The computer system has a pre - calibrated rotation matrix R and translation vector T built - in, and the dual - camera hand key - point data in the multi - channel hand key - point spatial position data sequence obtained in step S1 is converted into the same global coordinate system to ensure the spatial consistency of the data.
[0121] Furthermore, the 42 key - point coordinates of the left hand stored in channel one come from different cameras with different coordinate systems. After calibration, they are located in the same spatial coordinate system. The same calibration process is performed on the 42 key - point coordinates of the right hand stored in channel two.
[0122] For example: The original coordinates of the left - hand wrist key - point captured by camera A are Pleft_A=(100, 200, 130), and the original coordinates of the left - hand wrist key - point captured by camera B are Pleft_B=(136, 216, 102).
[0123] After calibration, Pleft_A becomes Pleft_A_corrected=(108, 222, 310), and after calibration, Pleft_B becomes Pleft_B_corrected=(109, 219, 308).
[0124] Perform time synchronization processing on the converted key - point data to ensure that the key - point data of the two cameras match in time, avoiding position deviation caused by delay, and obtaining two sets of corrected user hand key - point spatial position data.
[0125] Furthermore, read the timestamps collected by the two depth cameras and use a global clock synchronization mechanism to align the data. Calculate the time deviation ΔT. If the deviation is within the threshold range, directly align it; otherwise, perform interpolation processing.
[0126] Furthermore, use Kalman filtering to predict and update the user hand key - point spatial position data, improving the smoothness and continuity of the coordinates and avoiding jumps or jitters.
[0127] Fuse the coordinates of the two sets of corrected user hand key - point spatial position data obtained in the previous step according to the confidence algorithm, and merge them into a set of user hand key - point fused spatial coordinate time - series data.
[0128] If the confidence difference between two points ΔC < 0.3, fuse the two sets of hand key - point coordinates with weights. The process is as follows:
[0129] Suppose the coordinates of the same hand key - point collected by two RGB - d depth cameras are, after calibration, respectively , , and the confidences of the two key - points are respectively , 。
[0130] The confidence weighting is as follows: , 。
[0131] The coordinates after weighting are: 。
[0132] For example, for a certain key point: Pleft_A_corrected = (108, 222, 310), = 0.89, Pleft_B_corrected = (109, 219, 308), = 0.93. After using weighted average (considering their respective confidence levels), the fused data obtained is (108.51, 220.47, 308.98), thereby improving the accuracy of the overall data.
[0133] X(fuse) = ((108 × 89) + (109 × 93)) / (89 + 93) ≈ 108.51
[0134] Y(fuse) = ((222 × 89) + (219 × 93)) / (89 + 93) ≈ 220.47
[0135] Z(fuse) = ((109 × 89) + (308 × 93)) / (89 + 93) ≈ 308.98
[0136] If the confidence level of a key point of a certain camera is low due to occlusion, and the confidence level difference ΔC of the coordinates of a certain key point obtained by two cameras is > 0.3, then the group of data with the lower confidence level is directly discarded, and the coordinates of the other camera are used.
[0137] Using the time window method, the weighted coordinates are merged into the time series data of the fused spatial coordinates of the hand key points. The merged time series data of the fused spatial coordinates of the hand key points stores the spatial coordinate data of 21 key points on the left hand and 21 key points on the right hand after weighted fusion.
[0138] S5: Generate multi-channel gesture action signals based on key points.
[0139] Process the time series data of the fused spatial coordinates of the hand key points obtained in S4, and calculate the acceleration and spatial vectors of each key point in the sequence.
[0140] Among them, the accelerations of 21 key points on the left hand are respectively al(1), al(2), al(3) …… al(21), and the spatial vectors are respectively vl(1), vl(2), vl(3) …… vl(21); the accelerations of 21 key points on the right hand are respectively ar(1), ar(2), ar(3) …… ar(21), and the spatial vectors are respectively vr(1), vr(2), vr(3) …… vr(21). The fused spatial coordinates, accelerations and spatial vectors of the hand key points on both hands are input into a pre-trained gesture recognition network to classify gesture information.
[0141] When the user makes a forward and backward pushing gesture simulating gear shifting, the model analyzes the motion patterns of the key points in consecutive frames and identifies it as a gear shifting action. The system generates the gesture action types Ml(b), Mr(b) of the left and right hands and their confidence levels Cl(b), Cr(b). There is a standard feature model pre-stored in the gesture recognition network, and a large amount of air gesture feature data based on spatial coordinate information and motion data is stored in this model.
[0142] Preferably, a gesture recognition network based on Transformer is used for gesture classification.
[0143] S6: Control the simulated vehicle based on a dynamic decision-making method.
[0144] Calculate the spatial distance between the virtual hand position of the user in the simulated driving scenario and the device in the simulated vehicle. The distance here is used to cooperate with the recognition of hand key points to determine whether the user's hand is in the control area of the simulated in-vehicle device.
[0145] Extract the position information of 21 key points of each hand from the fused spatial coordinate time series data of the hand key points obtained in S4, and calculate the overall spatial position of the hand. The overall spatial position of the hand can stably represent the overall position of a single hand and reduce the influence of the jitter of individual fingers on the judgment.
[0146] The formula is as follows: , , . Among them, n = 20 is the maximum number of key points, numbered from zero, with a total of n + 1 key points; ( , , ) is the spatial position of the th hand.
[0147] Calculate the distance from the user's single hand to the in-vehicle device preset in the simulated driving environment, and output the list of in-vehicle devices currently in the hot area.
[0148] Furthermore, the calculation formula is as follows: , where the spatial position of the user's single hand is ( , , ), the position of the in-vehicle device in the simulated driving environment is ([ , , ).
[0149] Obtain the gesture action types Ml(a), Mr(a) and their confidence levels Cl(a), Cr(a) obtained in step S3; the gesture action types Ml(b), Mr(b) and their confidence levels Cl(b), Cr(b) obtained in step S5; obtain the gesture intention Wo and its confidence level Co obtained in step S3; the motion intention Wm and its confidence level Cm obtained in step S5.
[0150] Use a non-linear enhancement function to transform Cl(a), Cr(a), Cl(b), Cr(b), Co, Cm, and obtain the confidence weight coefficients S1, S2, S 3…… S6, the formula is , where is a pre-calibrated parameter.
[0151] Furthermore, perform pseudo-natural language encoding on Ml(a), Mr(a), Ml(b), Mr(b), Wo, Wm, S1, S2, S 3…… S6.
[0152] Furthermore, use BERT to process the pseudo-natural language of the grouped and transformed data, convert the original data into a vector form, and generate a simulated driving operation intention feature vector.
[0153] In the output of the BERT model, take the output vector f at the first position of the sequence (i.e., the position corresponding to the [CLS] token) as the length representation of the entire input, with a dimension of d. Finally, obtain a vector with a length of d , as the simulated driving operation intention feature.
[0154] In this embodiment, assume d = 768, then the output feature vector may be: f = [0.12, -0.34, 0.56, 0.01,..., 0.89]. Where f is a vector with a length of 768, containing various features extracted from the "pseudo-natural language" text.
[0155] Use the JAVA language to develop and construct a simulated driving vector database network for vector retrieval.
[0156] Furthermore, the method for adding new vectors to the simulated driving vector data network is as follows: The feature vectors generated by BERT or BART are inserted into the storage structure and saved together with the corresponding control instructions and weight factors. The simulated driving vector data network stores vectors and their corresponding control instructions and weight factors in JSON format.
[0157] Furthermore, an input query vector is used to find the most similar vector by calculating the similarity, and the corresponding control instruction and weight factor are returned. The cosine similarity algorithm is used to measure the similarity in the direction of two vectors.
[0158] The simulated driving vector data network is used to retrieve the control instruction C and weight factor k for the simulated driving operation intention feature vector f, as Figure 5 shown.
[0159] For example, when the input feature vector f = [0.12, -0.34, 0.56, 0.01,..., 0.89], the nearest matching result is retrieved.
[0160] Furthermore, based on the weight factor k, the dynamic control coefficient U of each control instruction is calculated. The formula is:
[0161]
[0162] In the above formula, represents the movement path of the fused spatial coordinates of hand key points, is the ratio of the sum of the movement paths of the corresponding coordinates in the three-dimensional scene, is the movement speed of the fused spatial coordinates of hand key points collected within the same time interval, is the movement speed of the corresponding coordinates in the three-dimensional scene, is the adjustment value saved in advance in the system for calibration. is the weighted speed of action execution, is the speed threshold saved in advance in the system.
[0163] Based on the control instruction Ctl, the dynamic control coefficient U, and the list of in-vehicle devices in the hot zone, the computer system generates the final control signal to control the simulated vehicle on the screen, and the operation result is rendered and fed back using the computer system.
[0164] In this embodiment, for example, when the user makes a gesture of "pushing the gear lever to the 5th gear" with the right hand at a very fast speed, the signal that the image-based model may output is "holding the gear lever and switching to the 5th gear, confidence 0.95"; the signal that the model based on hand key points may output is "quickly pushing the gear lever forward, confidence 0.93"; the distance information may output a signal: right hand [(gear lever, 8), (center console, 11)]. After the system performs feature vector retrieval, it is found that this situation conforms to the preset situation of "executing the gear shifting action too quickly", and the control result generated based on the dynamic control coefficient may be "gear shifting successful, but the vehicle shakes"
[0165] If a traditional multi-channel gesture recognition scheme and a simulation driving system are used, the possible recognition result of the above action may be "successful hand operation for gear shifting", and the control result executed by the system may be "shifting to the fifth gear and the vehicle accelerating". Compared with the traditional scheme, the present invention can detect situations where the user's gear shifting process is not standard or in place, and at the same time can dynamically control the strength of the interaction feedback of the simulation driving system according to the user's actual operation amplitude. A series of control experiments show that this system has a more refined gesture operation recognition ability, a driving feedback that is more in line with the actual situation, and a better user experience.
[0166] As Figure 4 shown, it is a schematic diagram of the control instructions preset by the present invention, including a) release and hold: for releasing and holding in-vehicle operating devices; b) steering: for controlling the left-right direction offset and steering of the simulated vehicle; c) acceleration and deceleration: for operating the throttle and brakes of the simulated vehicle; d) positive and negative feedback: for feedback on the current training experience and adjusting the training difficulty; e) finger flick: for controlling the vehicle gear and lighting operations; f) click: for selection operations on the user interface; g) confirm: for confirmation operations on the user interface.
[0167] Gesture actions include but are not limited to Figure 4 the control instruction actions preset in
[0168] In one embodiment, as Figure 3 shown, it simulates the perspective of the driver inside the vehicle, and the picture includes in-vehicle devices and the road environment outside the vehicle. When the user's hands are within the detection range of the sensor or camera, two dots on the screen represent the position of the user's hands in the simulated driving scenario. When the user's hands approach the position of the steering wheel in the simulated driving scenario, the indicating dots change color, indicating that the steering wheel device is activated, and at this time the user can operate the steering wheel through gestures.
[0169] As a further specific optimization of the technical solution of the present invention: gesture actions include but are not limited to steering wheel turning gestures, gear switching gestures, light on and off gestures, and menu calling gestures. The system can flexibly set personalized gestures for different simulated vehicles and scenes.
[0170] The specific implementation methods described above have described in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
[0171] In addition, the present invention also provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can also be executed by one or more processors to implement the steps of each embodiment of the above-mentioned mid-air gesture interaction method.
[0172] The specific implementation of the computer-readable storage medium of the present invention is basically the same as the above-mentioned embodiments of the mid-air gesture interaction method, and will not be repeated here.
[0173] A simulated driving training interactive system based on multi-channel gesture fusion is used to implement the simulated driving training interactive method based on multi-channel gesture fusion, including a hand motion acquisition module, a gesture motion processing module, a simulated driving interactive control module, a user interface and a menu module. The modules are connected by wire or wireless means to ensure the smooth operation of the simulated driving scene and the precise control of the vehicle, wherein:
[0174] Hand motion acquisition module: Use a depth camera to capture the user's hand spatial position and real-time hand images from multiple angles, and generate multi-channel gesture tracking data, including dual-camera hand spatial position key point timing data, and dual-camera hand motion image timing data.
[0175] Gesture processing module: Use the computer system to process multi-channel gesture tracking data, generate high-precision user hand position and posture data, identify the user's gesture interaction actions and intentions, and parse these data into specific system control instructions.
[0176] Simulated driving interactive control module: performs corresponding simulated driving operations according to system control instructions, such as controlling the steering angle of the vehicle in the simulated scene according to steering gestures, switching the vehicle gear and speed according to gear shift gestures, etc. The system ensures the real-time operation, ensures the continuity of the driving process, and synchronizes the changes in the vehicle status in the simulated scene.
[0177] User Interface and Menu Module: The user can call out the menu and control interface through specific gesture actions to perform operations such as vehicle selection, driving settings, and scene switching. The menu can be displayed at user-defined positions, and the menu has multi-functional switching, such as switching to simulated driving in different weather scenarios, entering the multi-player cooperative driving mode, or resetting the driving scene, etc.
[0178] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
[0179] The present invention proposes a method for natural gesture interaction based on multi-channel fusion, which uses two RGB-D depth cameras in the front and below to respectively collect images and key point data of the user's hand. Through methods such as coordinate transformation, time series synchronization, and weighted fusion, multi-channel data is integrated into high-precision hand key point coordinate time series data and multi-angle hand image time series data. Compared with existing methods, this mechanism significantly improves the accuracy and robustness of gesture recognition, especially in scenarios with partial occlusion or changing lighting conditions.
[0180] The present invention adopts a multi-modal dual gesture recognition mechanism, which superimposes the hand movement image time series data on the interaction scene, and performs neural network recognition on the image and key points respectively. It can simultaneously utilize the information quantity advantage of the image data and the accuracy advantage of the key point data. In addition to gestures and movement trajectories, it can recognize the fine interaction between the user and the simulated driving space. This also enables the present invention to be directly compatible with various simulated driving scene data such as three-dimensional data and two-dimensional videos, greatly simplifying the workload of scene development and adaptation.
[0181] The present invention adaptively weights multi-modal information, uses a text embedding model to construct a high-dimensional semantic intention vector database, and uses the vector database for in-depth matching of control instructions. Compared with traditional gesture interaction systems, the present invention can achieve in-depth understanding of the user's operation intention and dynamic adaptive adjustment of the user's operation feedback, which helps to truly simulate and restore the operation feedback of real driving training and improve the effect of simulated driving training.
Claims
1. A simulation driving training interaction method based on multi-channel gesture fusion, characterized in that, It includes the following steps: Step S1: Use a depth camera to obtain a sequence of dual-camera hand color images and a sequence of spatial position data of hand key points for dual cameras; Step S2: Preprocess the sequence of dual-camera hand color images to obtain the contour features of the user's hand action images, and then generate multi-channel time-series data of the hand action scene; Step S3: Use a neural network to process the multi-channel time-series data of the hand action scene to obtain the matching preset action types and their confidence levels, multi-channel gesture scene action intentions and their confidence levels; Step S4: Preprocess the spatial position data of the dual-camera user's hand key points, and then perform spatial transformation and merge them into a time-series data of the fused spatial coordinates of the hand key points; Step S5: Use a neural network to process the time-series data of the fused spatial coordinates of the hand key points to obtain the matching multi-channel key point gesture action signals and their confidence levels, gesture key point action intentions and their confidence levels; Step S6: According to the data obtained in Step S3 and Step S5, obtain control instructions and dynamic control coefficients, and at the same time calculate the real-time in-vehicle device hot area list to control the simulated vehicle.
2. The simulation driving training interaction method based on multi-channel gesture fusion according to claim 1, characterized in that, The specific content of Step S1 includes: S101: Two depth cameras capture a sequence of spatial position data of both hands and a sequence of hand color images; S102: Extract the three-dimensional spatial coordinates of the key bone points of each hand from the spatial position data of both hands; In a sequence of spatial position data, the three-dimensional spatial coordinates of each hand include h key points; each hand obtains two sets of three-dimensional spatial coordinates of bone points from different cameras; S103: If the depth camera can obtain the confidence level of each key point coordinate, record the confidence level of each key point coordinate, otherwise record the coordinate confidence level of the overall data of each channel; S104: Divide the key points into two groups according to the left and right hands and encapsulate them into a multi-channel time-series data of hand key point spatial positions. The multi-channel time-series data of hand key point spatial positions contains two-channel data; in Channel 1, store the 2h key point coordinates and confidence levels of the two cameras of the left hand; in Channel 2, store the 2h key point coordinates and confidence levels of the two cameras of the right hand.
3. The simulation driving training interaction method based on multi-channel gesture fusion according to claim 2, characterized in that, The specific implementation process of Step S2 is as follows: S201: Perform image preprocessing on the sequence of dual-camera hand color images, including removing blurred images, image denoising, and extracting and removing the background of the region of interest (ROI) of the image, respectively extract and enhance the contour features of the left and right hands to obtain a sequence of background-free hand image data; S202: Perform time-series synchronization processing on the sequence of background-free hand image data to make the image data of the two cameras match in time; S203: Identify and separate the left and right hands from the sequence of background-free hand image data, and generate multi-channel hand action image sequences for both hands respectively; S204: Obtain a pre-rendered two-dimensional static image of the simulated cockpit of the corresponding vehicle model. The two-dimensional static image of the simulated cockpit includes two different perspective pictures, namely the front view and the top view of the interior of the cockpit of the vehicle model selected by the user; S205: deforming and scaling the multi-channel hand motion image sequence according to the calibrated conversion parameters, and superimposing them with the two-dimensional cockpit images from two different perspectives to obtain two sets of hand-scene superimposed image sequences from different perspectives; S206: Generate multi-channel hand motion scene time series data, including three channel data, channel one includes a multi-channel hand motion image sequence of the left hand; channel two includes a multi-channel hand motion image sequence of the right hand; channel three includes two groups of hand-scene superimposed image sequences from different perspectives.
4. The simulation driving training interaction method based on multi-channel gesture fusion according to claim 3, characterized in that, The specific implementation process of step S3 is as follows: S301: Importing multi-channel hand action scene time series data into a neural network for recognition, wherein the hand action image sequences of channels one and two are input into a first set of neural network models; and the hand action image sequence of channel three is input into a second set of neural network models; The first set of neural network models is a pre-trained gesture feature model, which is used to extract the feature data of mid-air gestures trained based on dual-camera images; the first set of neural network models is a pre-trained command feature model, which is used to extract feature data describing the interaction relationship between the hand and the equipment in the cockpit and the operation command; S302: The first group of neural network models uses the data from channels one and two to identify and obtain the preset action types Ml(a), Mr(a) and their confidences Cl(a), Cr(a) for the left and right hands; the first group of neural network models uses the data from channel three to identify and obtain the multi-channel gesture scene action intention Wo and its confidence Co.
5. The simulation driving training interaction method based on multi-channel gesture fusion according to claim 4, characterized in that, The specific implementation process of step S4 is as follows: S401: Determine the local coordinate systems of the two depth cameras according to the intrinsic parameters and extrinsic parameters of the two depth cameras; transform the dual-camera hand key point data in the multi-channel hand key point spatial position data sequence into the same global coordinate system according to the pre-calibrated rotation matrix R and translation vector T; S402: Performing time synchronization processing on the converted key point data to obtain two sets of corrected user hand key point spatial position data; S403: The two sets of corrected user hand key point spatial position data are fused according to the confidence algorithm to form a set of user hand key point fused spatial coordinate time series data. The specific implementation is as follows: Suppose the coordinates of the same hand key point collected by two depth cameras are and respectively after calibration, and the confidence levels of the two cameras for a certain key point are and ; The confidence weighting is: , ; The coordinates after weighting are: ; If the confidence difference between the coordinates of a key point obtained by the two cameras is greater than 0.3, the coordinate data with low confidence is directly discarded and the coordinates of the other camera are used; The weighted coordinates are merged into the hand key point fusion spatial coordinate time series data. In the merged hand key point fusion spatial coordinate time series data, the spatial coordinate data of h key points on the left hand and h key points on the right hand after weighted fusion are saved.
6. The simulated driving training interaction method based on multi-channel gesture fusion according to claim 5, characterized in that The specific implementation process of step S5 is as follows: S501: Processing the fusion spatial coordinate time series data of the key points of the hand to calculate the acceleration and spatial vector of each key point in the sequence; S502: Fusing the spatial coordinates, acceleration and spatial vector of the key points of the left and right hands, and inputting them into a pre-trained gesture recognition network to classify the gesture information; the gesture recognition network is a pre-trained model based on Transformer, and the model stores a number of mid-air gesture feature data related to spatial coordinate information and motion data; The results of gesture information classification include the gesture action types Ml(b), Mr(b) of the left and right hands and their confidences Cl(b), Cr(b), as well as the movement intention Wm and its confidence Cm.
7. The simulated driving training interaction method based on multi-channel gesture fusion according to claim 6, characterized in that The specific implementation process of step S6 is as follows: S601: extracting the position information of h key points of each hand from the hand key point fusion spatial coordinate time series data, and calculating the overall spatial position of the hand by summing and averaging; S602: using a spatial Euclidean distance calculation formula, calculating the distance from the overall spatial position of each hand of the user to the pre-set vehicle-mounted equipment in the simulated driving environment, and outputting a real-time vehicle-mounted equipment hot zone list; S603: Obtain the gesture action types Ml(a), Mr(a) and their confidence levels Cl(a), Cr(a) obtained in step S3; the gesture action types Ml(b), Mr(b) and their confidence levels Cl(b), Cr(b) obtained in step S5; obtain the gesture intention Wo and its confidence level Co obtained in step S3; the motion intention Wm and its confidence level Cm obtained in step S5; use a non-linear enhancement function to transform Cl(a), Cr(a), Cl(b), Cr(b), Co, Cm to obtain confidence weight coefficients S1, S2, …… S6; S604: Take Ml(a), Mr(a), Ml(b), Mr(b), Wo, Wm, S1, S2, …… S6, perform pseudo-natural language encoding; S605: Process the pseudo-natural language generated in S604 using BERT as a single representation of the entire operation intention, and finally obtain a vector of length d , as the feature of the simulated driving operation intention; S606: using the simulated driving vector data unit to retrieve the simulated driving operation intention feature to obtain the control instruction and the weight factor k, and calculate the dynamic control coefficient; the simulated driving vector data unit has stored the control instruction and the weight factor corresponding to the pre-trained vector; S607: Based on the control instructions, dynamic control coefficients, and the list of on-board equipment in the hot area, the simulated vehicle on the screen is controlled, and the operation results are rendered and fed back using a computer system.
8. The simulated driving training interaction method based on multi-channel gesture fusion according to claim 7, characterized in that The retrieval process in step S606 is to input a query vector, that is, a simulated driving operation intention feature, find the most similar vector by calculating the similarity, and return the corresponding control instruction and weight factor.
9. The simulated driving training interaction method based on multi-channel gesture fusion according to claim 8, characterized in that The specific process of calculating the dynamic control coefficient U is as follows: The dynamic control coefficient U of each control instruction is calculated based on the weight factor k, and the calculation method is as follows: ;; ; ; In the above formula, is the ratio of the sum of the movement paths of the fused spatial coordinates of the hand key points to the sum of the movement paths of the corresponding coordinates in the three-dimensional scene within the same time interval, is the ratio of the sum of the movement speeds of the fused spatial coordinates of the hand key points collected within the same time interval to the sum of the movement speeds of the corresponding coordinates in the three-dimensional scene, is the corrected adjustment value; is the weighted speed of action execution, is the speed threshold.
10. A simulated driving training interaction system based on multi-channel gesture fusion, for implementing the simulated driving training interaction method according to any one of claims 1 to 9, characterized in that It includes hand motion acquisition module, gesture motion processing module, simulated driving interactive control module, user interface and menu module; The hand motion acquisition module uses a depth camera to capture the user's hand spatial position and real-time hand image at multiple angles, and generates multi-channel gesture tracking data, including dual-camera hand spatial position key point time series data and dual-camera hand motion image time series data; The gesture action processing module processes the multi-channel gesture tracking data, generates the user's hand position and posture data, identifies the user's gesture interaction actions and intentions, and interprets these data into system control instructions; The simulated driving interactive control module performs corresponding simulated driving operations according to the system control instructions to ensure the continuity of the driving process and synchronize the changes in the vehicle status in the simulated scene; The user interface and menu module allows the user to call out the menu and control interface through gestures. The menu is displayed in a user-defined location and has multi-function switching.
Citation Information
Patent Citations
Method and device for identifying abnormal equipment in wireless network
CN111757365A
Gesture interaction method and system of vehicle-mounted AR-HUD
CN112241204A
Gesture recognition method, device and system and vehicle
CN113646736A
Gesture recognition method and system, computer equipment and readable storage medium
CN115223239A
Automobile cabin gesture recognition system and method based on DVS
CN117218716A
Cited By
Dynamic interface interaction system based on multi-modal gesture recognition
CN120762574A