Self-powered intelligent rail driver gesture recognition sensor based on triboelectric nano generator (TENG)
Through self-powered gesture recognition sensors and deep learning technology based on triboelectric nanogenerators, the lighting dependence and wearable comfort problems of existing gesture recognition technologies are solved, and high-precision and real-time driver gesture recognition are achieved, which improves the safety and interaction efficiency of the intelligent transportation system.
Patent Information
- Application Number
- CN202510598227.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-22
AI Technical Summary
The existing driver gesture recognition technology has strong dependence on lighting conditions, poor anti-occlusion ability, wearable comfort and signal susceptible to mechanical vibration interference, and a single signal analysis dimension, making it difficult to achieve efficient digital interaction, which limits the improvement of the safety level of intelligent driving.
A self-powered intelligent rail driver gesture recognition sensor based on triboelectric nanogenerator (TENG) is used to collect hand motion electrical signals by combining Ecoflex resin electronegative layer and copper electrode layer, and multi-dimensional feature extraction is performed through GRU/LSTM/CNN neural network, and real-time gesture recognition is achieved by combining driver model bone reconstruction.
It realizes high-precision and real-time driver gesture recognition, improves human-computer interaction efficiency and safety, and provides an effective interaction method for intelligent transportation systems, suitable for future road, ocean and air traffic fields.
Smart Images

Figure CN120523322A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart transportation technology, and specifically relates to a self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG). Background Art
[0002] With the rapid development of smart transportation, the driver's gestures will become key information during driving. Existing gesture recognition technology has the following limitations:
[0003] Visual methods: Depend on lighting conditions, have poor anti-occlusion capabilities, and are difficult to adapt to complex working conditions;
[0004] Inertial sensor: poor wearing comfort, and the signal is easily disturbed by mechanical vibration;
[0005] Traditional wearable devices: Signal analysis is limited in dimension and difficult to capture complex gesture features;
[0006] The lack of an efficient digital interaction system for driver gesture commands in intelligent rail systems hinders improvements in intelligent driving safety. Given this situation, it is necessary to develop a self-powered, high-precision, and robust gesture recognition system that can map driver gesture commands to digital commands in real time, thereby improving the efficiency and safety of human-computer interaction in intelligent transportation. Summary of the Invention
[0007] In response to the problem in the prior art of how to provide a self-powered, high-precision, and highly robust gesture recognition system, the present invention provides a self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG).
[0008] The technical solution adopted in the present invention is as follows:
[0009] A self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG). The sensor is designed based on the principle of a triboelectric nanogenerator and has a structure comprising:
[0010] The housing is used to form the basic frame of the sensor and has a flexible stress relief structure;
[0011] The electronegative layer of Ecoflex resin is used to collect electrical signals from changes in hand movements;
[0012] The copper electrode layer is arranged between the shell and the electronegative layer of the Ecoflex resin, and is used to receive and transmit the hand movement change electrical signal collected by the electronegative layer of the Ecoflex resin, and the copper electrode is electrically connected to the ground wire.
[0013] Preferably, the Ecoflex resin electronegative layer comprises a substrate, and a plurality of bionic suction cups are provided on one side of the substrate.
[0014] Preferably, the sensor is a ring-shaped structure.
[0015] A method for using a self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG) comprises the following steps:
[0016] S1: Collect the driver's hand movement electrical signal through the sensor;
[0017] S2: By integrating the GRU / LSTM / CNN neural network, it extracts the time-frequency-time-frequency multi-dimensional features of the hand motion signals collected by S1, thereby recognizing various driving gestures.
[0018] Preferably, the method further includes S3, wherein the specific contents of S3 include: reconstructing the posture scene based on the driving gestures recognized in S2 and the skeleton of the driver model, specifically:
[0019] S301: The recognition result of S2 is transmitted to the Unity3D engine in real time via the Socket protocol;
[0020] S302: The Unity3D engine drives the 3D driver model skeleton binding system to reconstruct the driving scene in the cloud based on a predefined motion mapping database.
[0021] Preferably, the specific steps of collecting the driver's hand motion electrical signal by the sensor in S1 are as follows:
[0022] S101: The electronegative layer of the Ecoflex resin comes into contact with the driver's hand. When the driver makes a gesture, the electronegative layer of the Ecoflex resin deforms, causing charge transfer between the driver's hand skin and the electronegative layer of the Ecoflex resin.
[0023] S102: The transfer of charge creates a potential difference between the copper electrode layer and the ground wire, causing current to flow from the ground wire to the copper electrode; when the skin moves away from the electronegative Ecoflex resin layer, the induced charge increases, but because there is no direct contact, the potential difference between the skin and the electronegative Ecoflex resin layer does not disappear, so the current continues to flow from the ground wire to the copper electrode layer; when the skin approaches the electronegative Ecoflex resin layer again, the induced charge decreases, and the positive charge on the copper electrode layer begins to flow back to the ground wire; finally, when the skin contacts the electronegative Ecoflex resin layer again, a complete hand movement electrical signal is formed.
[0024] Preferably, the specific steps of S2 include:
[0025] S201: Digitize the hand movement electrical signal collected in S1 to obtain digital signal data;
[0026] S202: intercepting a time domain segment of the digitized signal data and superimposing a Gaussian noise enhanced data set;
[0027] S203: Use the fusion GRU / LSTM / CNN neural network to extract time series and time-frequency features in parallel.
[0028] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0029] The present invention uses a portable triboelectric sensor ring (TENG) to capture the driver's gesture signals. By comparing the use of recurrent neural networks (RNN) and convolutional neural networks (CNN) in deep learning technology, real-time and accurate recognition (98.9%) of the driver's gestures and the generation of virtual actions are achieved. Finally, Unity3D is used to generate taxi scenes in real time, and real-time virtual interaction between the train driver and various transport units is achieved. The application of these technologies not only improves the accuracy of gesture recognition, but also provides an effective technical method for human-computer interaction and human-machine interaction in driver-based intelligent transportation systems. It lays the foundation for the application of intelligent transportation digital twin technology in the field of metaverse. In addition, its technical principles and methods are also applicable to key areas such as future road, marine and air transportation, where human gesture recognition is crucial to safety and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a schematic diagram of the structure of the self-powered gesture recognition sensor;
[0031] Figure 2 This is a schematic diagram of the structure of the components of the self-powered gesture recognition sensor;
[0032] Figure 3 is a structural schematic diagram of the shell;
[0033] Figure 4 The following are conceptual diagrams of the present invention, wherein a is an application scenario diagram of the present invention in a train cab; FIG. b is a usage and structural diagram of the self-powered gesture recognition sensor;
[0034] Figure 5 Figures 1 and 2 illustrate the working principle and electrical performance of the self-powered gesture recognition sensor. (a) shows the principle of finger tendon interaction, (b) the principle of electrical signal generation in the self-powered gesture recognition sensor, (c) a voltage simulation diagram showing the contact between the finger and the electronegative layer of Ecoflex resin, and (d) an experimental analysis of the electrical performance of the self-powered gesture recognition sensor.
[0035] Figure 6Figure 1 is a physical picture of the self-powered gesture recognition sensor, where a is a physical picture of each component of the self-powered gesture recognition sensor, b is a physical picture of each component and assembly of the self-powered gesture recognition sensor, and c is a physical picture of the vibration signal transmission experiment of the self-powered gesture recognition sensor;
[0036] Figure 7 Figure 2 is a diagram of driver command detection based on deep learning, where a is the waveform of sensors at different positions corresponding to gesture commands, b is a demonstration diagram of the gesture command recognition process based on deep learning, c is the confusion matrix of GRU, d is the confusion matrix of LSTM, and e is the confusion matrix of Conv1D.
[0037] Figure 8 A diagram showing the virtual scene modeling of the digital twin of the train driver's gesture commands. (a) is a flow chart of the DC-SS system, and (b) is a diagram showing the real-time recognition of the train driver's gesture commands and the virtual scene modeling based on deep learning (only four gestures are shown).
[0038] Figure 9 Diagrams for constructing multidisciplinary test scenarios and digital twin platforms, where a is the multidisciplinary test scenario, b is the digital twin platform demonstration diagram, and c is the test scenario diagram in the real CRH high-speed train cab and CRH high-speed train simulation;
[0039] Among them, 1-shell, 101-base, 102-installation groove, 103-wire hole, 2-Ecoflex resin electronegative layer, 201-base plate, 202-bionic suction cup, 3-copper electrode layer. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.
[0041] In the description of the embodiments of the present application, it should be noted that the terms "upper", "lower", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings, or are the orientation or positional relationship in which the product of the invention is usually placed when in use. They are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, they cannot be understood as limiting the present application. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0042] Example 1
[0043] A self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG). The sensor is designed based on the principle of a triboelectric nanogenerator and has a structure comprising:
[0044] Shell 1, such as Figure 3 As shown, it is used to constitute the basic frame of the sensor and has a flexible stress release structure. Specifically, in this embodiment, the shell 1 is made of polylactic acid material and has a flexible stress release structure that can adapt to the flexion and extension movement of the fingers; the shell 1 includes a base 101, and the base 101 is an annular structure. The inner side of the annular structure is provided with a mounting groove 102 arranged circumferentially around the base 101. The mounting groove 102 passes through the base 101 in a circular manner, and a contact port is provided on the side of the mounting groove 102 facing the inner side of the base 101. The Ecoflex resin electronegative layer 2 and the copper electrode layer 3 are both provided in the mounting groove 102, and the Ecoflex resin electronegative layer 2 contacts the driver's hand skin through the contact port. The base 101 is also provided with a wire hole 103 connected to the mounting groove 102, and the wire hole 103 is used to pass a wire connected to the copper electrode layer 3 and the ground wire;
[0045] Ecoflex resin electronegative layer 2, used to collect electrical signals of hand movement changes;
[0046] The copper electrode layer 3 is provided between the housing 1 and the electronegative Ecoflex resin layer 2 and is used for receiving and transmitting the hand movement change electrical signal collected by the electronegative Ecoflex resin layer 2 , and the copper electrode is electrically connected to the ground wire.
[0047] To prepare the electronegative Ecoflex resin layer 2, the first step is reverse molding using a 3D-printed PLA mold: Ecoflex A and B raw materials are mixed in a 1:1 ratio and stirred for 5 minutes. After stirring, a vacuum treatment is applied to release air bubbles from the mixture. The material is then kept at a constant temperature of 60°C for 3 hours before demolding and cutting. For the outer shell 1, the first step is to improve comfort by rounding corners in areas that may come into contact with the human body, while also providing a guide for the wires. Finally, the various components of the sensor are assembled using 3D printing.
[0048] In one embodiment, the Ecoflex resin electronegative layer 2 includes a substrate 201, and a plurality of biomimetic suction cups 202 are provided on one side of the substrate 201. Figure 1 and Figure 2 As shown, the bionic suction cup 202 is an "8"-shaped structure.
[0049] In one embodiment, the sensor is a ring-shaped structure.
[0050] In one embodiment, the ring-shaped structure is an open ring, and the opening angle is 60°.
[0051] Example 2
[0052] The method of use includes the following steps:
[0053] S1: If Figure 4-9 As shown, the above-mentioned sensor is used to collect the driver's hand movement electrical signal; the specific steps are as follows:
[0054] S101: If Figure 5 As shown, when the wearer makes a hand movement, Figure 5 As shown in Figure a, the tendons in the relevant area contract or relax, causing the soft tissue in the finger area to expand or contract. This muscle contraction increases the cross-sectional area and circumference of the finger, creating a squeezing effect between the finger and the sensor. This exerts pressure on the sensor's contact surface, increasing the contact area between the electronegative Ecoflex resin layer 2 and the skin. Simultaneously, as the finger's skin penetrates further, the gap between the skin and the electronegative Ecoflex resin layer 2 further decreases. Figure 5b shows a single cycle of the signal generated by the self-powered gesture recognition sensor. The electronegative Ecoflex resin layer 2 undergoes four phases during a single motion cycle: bending initiation, bending completion, stretching initiation, and stretching completion. As the skin approaches the electronegative Ecoflex resin layer 2, the skin acts as the anode, while the layer 2 acts as the cathode. Charge transfer occurs between the two, creating a potential difference between the copper electrode layer 3 connected to the layer 2 and the ground, causing current to flow from the ground to the copper electrode layer 3.
[0055] S102: When the skin moves away from the electronegative Ecoflex resin layer 2, the induced charge increases. However, due to the lack of direct contact, the potential difference between the skin and the electronegative Ecoflex resin layer 2 persists, causing current to continue flowing from the ground to the copper electrode layer 3. When the skin approaches the electronegative Ecoflex resin layer 2 again, the induced charge decreases, and the positive charge on the copper electrode layer 3 begins to flow back to the ground. Finally, when the skin re-contacts the electronegative Ecoflex resin layer 2, a complete electrical signal representing the hand movement is generated. If the movement remains in any of the aforementioned states, no current is generated. The voltage amplitude is positively correlated with the bending angle (30° → 45mV, 90° → 145mV). For example, the train driver's "pass" gesture: the driver's directing hand extends the index and middle fingers, which are not bent, while the other fingers are fully bent. Sensors worn on bent fingers produce different charge transfer than those worn on straight fingers, resulting in different electrical signals. Therefore, it is worth noting that the sensor primarily captures the signal generated by the finger's extension and retraction during the gesture, rather than directional or rotational movement. When the user wears the sensor, although some relative motion will inevitably occur between the sensor and the hand, due to the contact separation principle of the sensor, these tiny rotations and longitudinal displacements will not be collected as the main signal. The voltage signal generated by the sensor is input into the OpenBCIADC sampling interface through a soft wire. Then, it is low-pass filtered by the built-in digital filter to eliminate the high-frequency noise in the input voltage. Finally, the data is transmitted to the computer software of the host computer via WI-FI for real-time data display and subsequent data processing. We used COMSOL Multiphysics to perform triboelectric simulation of a single sensor, and the simulation results are consistent with the actual situation ( Figure 6 c).
[0056] Considering the frequency of human finger use and the gesture commands issued by train drivers, a total of 7 sensors were worn on the user's finger area, of which 2 sensors were worn on the front of the index finger and middle finger respectively to obtain more gesture features. Subsequently, the basic actions of 8 gesture commands commonly used by train drivers using sensors were obtained, such as Figure 7As shown in a. The first is when the train passes: extend the index finger and middle finger of your right hand together, and then punch to the left. The second is when the train stops or slows down, clench your fist. The third is to stop or pass the siding: clench your right hand, extend your little finger and thumb; the fourth is to stop on the main line: extend your thumb and clamp the other fingers tightly; the fifth is departure time 1: clench your fist and extend your little finger; the sixth is departure time 2: clench your fist, extend your little finger and ring finger; the seventh is departure time 3: clench your fist, extend your little finger, ring finger and middle finger; the eighth is a relaxed state. From Figure 7 As can be seen, each action has significant differences in the output of each channel. The most obvious difference between them is that when performing actions 5, 6, and 7, there are almost no motion signals from the thumb and index finger when these thumb and index finger are almost motionless, which is consistent with the actual action situation.
[0057] Sensor response was measured using a horizontal vibration meter and OpenBCI. The test frequency was chosen based on the fact that human activity frequencies typically range from 0 to 20 Hz, with 98% of human activity occurring at frequencies below 10 Hz. Therefore, the vibration meter's test frequency range was set to 1 to 5 Hz. OpenBCI's minimum sampling frequency is 250 Hz, which satisfies the Nyquist sampling theorem for human activity signal acquisition. Figure 5 d shows the electrical performance test of the sensor under external excitation at lower frequency and amplitude. Figure 5 d II and Figure 5 Figure d III shows the sensor's output voltage waveform under external vibration (amplitude 2mm) from 1.5Hz to 2.5Hz, with intervals of 0.1Hz. As shown in the figure, the output voltage amplitude gradually decreases as the vibration frequency increases. As the vibration frequency increases from 1.5Hz to 2.5Hz, the maximum output voltage drops from approximately 75mV to approximately 40mV. It is noteworthy that the internal stress of the DC-SS ring sensor is primarily determined by displacement. Since the displacement set by the vibration table is constant, the maximum friction voltage decreases accordingly with higher vibration frequency and shorter contact time. As the vibration frequency increases, the frequency of the output voltage remains consistent with the external excitation frequency. Figure 5 Figure dV shows the sensor's output voltage waveform at 1Hz intervals when the external excitation and the amplitude (2mm) are 1Hz-5Hz. Similarly, as the vibration frequency increases, the input voltage amplitude gradually decreases, and the maximum output voltage drops from approximately 90mV to approximately 20mV. At the same time, the output voltage still effectively represents the frequency of the external excitation.
[0058] Due to the randomness of different wearer sizes and wearing styles, the tightness and angle of the sensor when worn are inconsistent, and larger wearers tend to produce larger electrical signal outputs. Therefore, we compared the electrical signal output of sensors of corresponding sizes at different frequencies under different body shapes. (e.g. Figure 5 As shown in the figure (dVI), the output voltage of the sensor worn by the larger wearer is significantly higher, but it still well reflects the frequency of external vibration and the characteristics of the motion process (for example, the motion process of size 3). Therefore, differences in sensor output voltage between people of different body types are inevitable. This problem cannot be avoided through early design, but it can be effectively addressed through back-end deep learning. For wearers of different body types, the sensor can obtain the wearer's dataset through the same data transmission method for local relearning, and then perform gesture command recognition under the new environmental conditions.
[0059] Figure 5 Figure d IV shows the output voltage waveform of the sensor when the finger is bent at different angles while worn on a user's hand. As shown in the figure, the greater the degree of finger bending, the greater the peak output voltage. At a 30° bend, the peak output voltage is approximately 45mV, at a 60° bend, approximately 95mV, and at a 90° bend, approximately 145mV. The peak voltage is nearly proportional to the bending angle. Therefore, the sensor can record the relevant characteristics of different gesture commands, facilitating analysis and processing. Experimental results show that within the typical frequency range of driver gesture commands (1-5Hz), the sensor's output voltage can significantly distinguish external stimuli of different frequencies. Furthermore, durability testing was conducted to infer performance under long-term use. The sensor was subjected to 14,400 repeated cycles (simulating 2-3 years of use) while monitoring signal amplitude decay. The results show that the signal output maintains above 83% of its initial amplitude, and no physical damage was observed in the Ecoflex or PLA components.
[0060] S2: By integrating the GRU / LSTM / CNN neural network, we extract the time-frequency-time-frequency multi-dimensional features of the hand motion signals collected by S1, thereby identifying various driving gestures. The specific steps include:
[0061] The dataset was collected from 8 participants who performed 8 standardized train driver gestures ( Figure 7 a). The sensor transmits the electrical signal containing the driver's gesture command to OpenBCI through wires, and processes the data through various methods such as ADC sampling, digital filtering, FFT, etc. The data is then wirelessly transmitted to the host computer via Wi-Fi to form a data visualization platform ( Figure 7c). The raw data was initially sampled at 250 Hz using an analog-to-digital converter (ADC) and then bandpass filtered using a Butterworth filter with a cutoff frequency range of 0.5–30 Hz to attenuate high- and low-frequency noise components. Subsequently, a fast Fourier transform (FFT) was applied to calculate a real-time spectral representation of the signal and visualized on a graphical interface on the host computer. Simultaneously, the filtered signal was segmented using a sliding window for data augmentation. The window size was set to 500 (corresponding to a gesture duration of approximately 2 seconds), and the sliding window step size was set to 5. The dataset was split using a training:test ratio of 8:2 to ensure proportional representation of all gesture categories. A GRU, LSTM, and a one-dimensional convolutional neural network (Conv1D) were selected for training. Both the GRU and LSTM had 128 hidden layers. The number of filters in the first layer of the Conv1D was set to 128, and the number of filters in the second layer was set to 256. The convolution kernel size was set to 3, and the pooling layer window size was set to 2 to improve the model's expressiveness and generalization capabilities. Figure 7 Figures c, 7d, and 7e show the confusion matrices obtained for the three models. The results demonstrate that the model is able to effectively recognize most driver gesture commands. Comparing the training results of various neural networks, the GRU achieved a prediction accuracy of 93.5%, the LSTM achieved a prediction accuracy of 95.4%, and the CNN achieved the highest prediction accuracy of 98.9%.
[0062] Regarding deep learning model parameter settings, we compared GRU and LSTM models with different numbers of hidden layers. Comparative experiments were conducted with 64, 128, and 256 hidden layers. For the GRU model, the accuracies for models with 64, 128, and 256 hidden layers were 88.9%, 93.5%, and 93.4%, respectively. For the LSTM model, the corresponding accuracies for models with 64, 128, and 256 hidden layers were 89.2%, 95.4%, and 94.4%, respectively. These results confirm that 128 neurons provide the best balance between capacity and generalization for our dataset. The LSTM is more capable of performing this task than the GRU. Note that the accuracies of the GRU and LSTM models with 128 and 256 layers are very similar. This slight difference is primarily due to the marginal effect of increasing the number of model units. Increasing model size does not significantly improve recognition accuracy and may actually increase the risk of overfitting. Furthermore, increasing the number of hidden layers significantly increases model parameters, which is detrimental to model deployment costs and impairs real-time recognition performance. To further analyze the advantages and disadvantages of GRU and LSTM with 128 and 256 hidden layers, we conducted further experiments and compared their F1 scores. The results show that when the number of hidden layers is 128 and 256, the GRU achieves an F1 score of 0.929 and 0.908, respectively; the LSTM achieves an F1 score of 0.899 and 0.895, respectively, which are very close. Regarding model parameters, the parameters of the GRU architecture increase from 53,640 with a 128-unit hidden layer to 205,576 with a 256-unit hidden layer. Similarly, the parameters of the LSTM increase from 70,664 (128 units) to 272,392 (256 units). This significant parameter expansion (approximately 3.8 times for both architectures) clearly demonstrates that expanding to 256 hidden units significantly impacts model lightweightness. Therefore, the configuration of 128 units becomes a better choice. For Conv1D, we conducted a series of comparative experiments by changing the number of filters, kernel size and pooling layer configuration. Figure 7 As shown in the confusion matrices shown in 7c, 7d, and 7e, the GRU and LSTM models have lower recognition performance for categories 2, 3, and 4. This may be because gestures 2, 3, and 4 involve more similar curved finger patterns, which challenges the sequential modeling of RNNs, while CNNs are good at capturing spatial differences in sensor signals.
[0063] To verify the recognition efficiency of the model, we conducted a comparative test on the recognition time of the trained models. The results showed that the inference time (mean ± standard deviation) for each model was: GRU: 36.08 ± 3.19 ms; LSTM: 41.04 ± 3.32 ms; Conv1D: 28.75 ± 8.80 ms. Therefore, Conv1D achieved the fastest inference time (28.75 ms) and the highest recognition accuracy, effectively addressing the latency issue in real-time recognition and making it the optimal choice for this task.
[0064] In addition, to avoid overfitting problems, we adopted the following specific regularization methods:
[0065] In the architectural design of the GRU and LSTM models, a dual L2 regularization strategy was implemented to enhance generalization. Specifically, L2 regularization constraints were applied simultaneously to the kernel weights (input-to-hidden layer connections) and the recurrent weights (hidden-to-hidden layer recurrent connections), with the regularization coefficient set to 0.001. This mechanism introduces a squared penalty term for the weight parameter in the loss function, effectively suppressing the excessive growth of the network parameter magnitude and reducing the risk of overfitting. Furthermore, a dropout layer with a rate of 0.1 was added after the GRU / LSTM layer, randomly shutting down 10% of the neurons during training to reduce the risk of overfitting.
[0066] In the CNN architecture, a multi-level regularization framework is systematically integrated into its various hierarchical structures. A hybrid L1-L2 regularization scheme is applied to the kernel weights of the two convolutional layers, with sparsity-inducing L1 (λ=0.01) and magnitude-constrained L2 (λ=0.01) penalties operating simultaneously. This composite regularization balances feature selection through parameter sparsity control and stable weight distribution control. The parameter size is set based on empirical practices found in existing literature. Spatial dropout layers (rate=0.3) are interleaved after each convolutional feature extraction block to disrupt local activation patterns. Simultaneously, a terminal dropout layer with enhanced suppression (rate=0.5) is inserted before the classification head to strategically reduce inter-neuronal dependencies in the high-dimensional feature representation.
[0067] Each gesture corresponds to a different electrical signal in each channel (such as Figure 7 (as shown in a), which brings features that are very conducive to machine learning feature extraction. The maximum fluctuation amplitude of the signal (peak to peak) is close to 1V. When comparing large and small movements, the difference in signal fluctuation amplitude is also very obvious. Sliding windows are used for data enhancement. After obtaining the dataset, GRU, LSTM and CNN are used for model training. The specific model parameters are as follows Figure 7 As shown in b.
[0068] In one embodiment, the method further includes S3, wherein the specific content of S3 includes: reconstructing the posture scene based on the driving gesture recognized in S2 and the skeleton of the driver model, specifically:
[0069] After completing gesture recognition on the Python side, we built a digital twin interaction platform for driver training based on Unity3D. Figure 8 As shown in Figure 2, Unity3D is a powerful cross-platform game engine widely used for creating and developing 3D content. In this study, we used Unity3D to build an interactive 3D animation environment that can respond in real time to driver gesture commands used to train deep learning models for recognition.
[0070] First, we created a library of 3D models containing eight gesture recognition actions that a trained driver might perform. These models ensure natural and highly realistic animations through precise skeletal rigging and animation keyframes. Next, we developed a real-time communication interface that allows Unity3D to receive gesture recognition results from a Python deep learning model. To implement this functionality, we employed the Socket protocol, a networking technology for full-duplex communication between client and server. Through Socket, the Python backend can send recognition results to the Unity3D frontend in real time, and Unity3D can dynamically adjust the character's movements in the 3D animation based on the received data. Within the Unity3D environment, we implemented a gesture mapping system that associates gesture commands output by the deep learning model with corresponding actions on the 3D character. This allows the driver to perform a gesture, the deep learning model recognizes and sends the command, and the 3D character in Unity3D immediately displays the corresponding gesture. Finally, we conducted comprehensive testing of the entire system to verify the synchronization and accuracy between the Unity3D animation and the deep learning model. Test results demonstrated that the system accurately recognized the driver's gestures and displayed the corresponding 3D animation in Unity3D in real time, meeting the research requirements.
[0071] Through this work, we have provided a new platform for the application of gesture recognition technology for train drivers and other drivers. Unity3D, as a basic platform for real-time 3D animation display, lays the foundation for the development of future intelligent transportation systems. In future research, we can explore the integration of deep learning models with more types of sensor data (such as sound, facial expressions, etc.) to achieve more comprehensive driver status monitoring and interaction. In addition, combined with AR technology, virtual information can be superimposed on the real world, providing drivers with a more intuitive interactive control experience. At the same time, using real-time gesture recognition technology, it is possible to develop an intelligent driving safety monitoring system to monitor driver fatigue or abnormal behavior and issue timely warnings.
[0072] To validate the applicability of DC-SS to a wide range of users, we conducted motion acquisition and recognition experiments with eight subjects, including two women and six men, ranging in age from 23 to 25. In terms of weight, two men were thin, two were of medium height, and two were tall. Of the two women, one was thin and one was of medium height. Figure 9 a shows a photo of the test site with 8 subjects. As for the test results, after data acquisition and model regularization, the trained deep learning model performed well in recognizing the subjects' instructions.
[0073] In addition, we have built a digital twin platform display website. The website interface can intuitively display the driver's operations and operation recognition results to promote interaction among operators of various intelligent transportation vehicles. Figure 9 b shows the display interface of the digital twin platform.
[0074] DC-SS tests were conducted in a real CRH high-speed train cab and a CRH simulation cab (e.g. Figure 9 c). The results show that the signal acquisition module, data transmission module, deep learning model, and cloud data (web platform) can operate effectively, completing the complete combination of wearable devices and digital twins, laying a good foundation for the development of intelligent transportation based on the metaverse.
[0075] The above-described embodiments merely represent specific implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of protection of the present application. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the technical concept of the present application, and all such variations and improvements fall within the scope of protection of the present application.
Claims
1. A self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG), characterized by: Its structure includes: A housing (1) is used to form a basic frame of the sensor and has a flexible stress release structure; The electronegative layer (2) of Ecoflex resin is used to collect the electrical signal of the hand movement change; The copper electrode layer (3) is arranged between the housing (1) and the electronegative Ecoflex resin layer (2) and is used for receiving and transmitting the hand movement change electrical signal collected by the electronegative Ecoflex resin layer (2), and the copper electrode layer is electrically connected to the ground wire.
2. The self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG) according to claim 1 is characterized by: The Ecoflex resin electronegative layer (2) comprises a substrate (201), and a plurality of bionic suction cups (202) are provided on one side of the substrate (201).
3. The self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG) according to claim 1 is characterized by: The sensor is a ring-shaped structure.
4. The self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG) according to claim 2 is characterized by: The ring-shaped structure is an open ring, and the opening angle is 60°.
5. A method for using a self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG), characterized by: The following steps are involved: S1: Wear the sensor on the driver's hand and collect the driver's hand movement electrical signal through the sensor; S2: By integrating the GRU / LSTM / CNN neural network, it extracts the time-domain-frequency-time-frequency multi-dimensional features of the hand motion signals collected by S1, thereby recognizing various driving gestures.
6. The method for using the self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG) according to claim 5 is characterized by: S3 is also included. The specific content of S3 includes: reconstructing the posture scene based on the driving gestures recognized by S2 and the skeleton of the driver model, specifically: S301: The recognition result of S2 is transmitted to the Unity3D engine in real time via the Socket protocol; S302: The Unity3D engine drives the 3D driver model skeleton binding system to reconstruct the driving scene in the cloud based on a predefined motion mapping database.
7. The method for using the self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG) according to claim 5 is characterized by: The specific steps of collecting the driver's hand movement electrical signal through the sensor in S1 are as follows: S101: The Ecoflex resin electronegative layer (2) contacts the driver's hand. When the driver makes a gesture, the Ecoflex resin electronegative layer (2) deforms, causing charge transfer between the driver's hand skin and the Ecoflex resin electronegative layer (2). S102: The transfer of charge creates a potential difference between the copper electrode layer (3) and the ground wire, causing current to flow from the ground wire to the copper electrode. When the skin moves away from the electronegative Ecoflex resin layer (2), the induced charge increases. However, due to the lack of direct contact, the potential difference between the skin and the electronegative Ecoflex resin layer (2) does not disappear, so the current continues to flow from the ground wire to the copper electrode layer (3). When the skin approaches the electronegative Ecoflex resin layer (2) again, the induced charge decreases, and the positive charge on the copper electrode layer (3) begins to flow back to the ground wire. Finally, when the skin contacts the electronegative layer of Ecoflex resin (2) again, a complete electrical signal of hand movement is formed.
8. The method for using the self-powered smart rail driver gesture recognition sensor based on a triboelectric nanogenerator (TENG) according to claim 5 is characterized by: The specific steps of S2 include: S201: Digitize the hand movement electrical signal collected in S1 to obtain digital signal data; S202: intercepting a time domain segment of the digitized signal data and superimposing a Gaussian noise enhanced data set; S203: Use the fusion GRU / LSTM / CNN neural network to extract time series and time-frequency features in parallel; S204: After fusing the time series and time-frequency features, a Softmax classifier is used to identify various driving gestures.