Gesture recognition method of optical fiber composite flexible sensor and LSTM neural network model based on SNTNS structure
By combining SNTNS structure fiber optic composite flexible sensor and LSTM neural network model with deep learning, the illumination and privacy issues of machine vision gesture recognition are solved, achieving high-precision gesture recognition and anti-electromagnetic interference capability, which is suitable for wearable devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-27
AI Technical Summary
Existing machine vision gesture recognition methods have limitations in terms of lighting, privacy, and line-of-sight occlusion, while traditional fiber optic sensors are not stretchable and cannot meet the sensitivity and electromagnetic interference resistance requirements of wearable devices.
We designed a fiber optic composite flexible sensor based on the SNTNS structure, combined with FBG and LSTM neural network models, to recognize hand gestures by changing optical power. We used PDMS encapsulation to improve curvature sensitivity and combined deep learning for gesture classification.
It achieves high-precision gesture recognition, solves the lighting and privacy issues of traditional methods, improves the sensitivity and anti-electromagnetic interference capabilities of the sensor, and is suitable for wearable devices.
Smart Images

Figure CN121743971A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a composite flexible sensor based on a single-mode coreless-three-coreless-single-mode fiber (SMF-NCF-TCF-NCF-SMF, SNTNS) structure coated with polydimethylsiloxane (PDMS) and a long short-term memory (LSTM) neural network model for human gesture recognition, belonging to the field of fiber optic sensing and gesture recognition technology. Background Technology
[0002] Fiber optic sensing technology is a sensing method that utilizes the transmission characteristics of light waves in optical fibers to detect and measure physical, chemical, or other variables. Since its initial proposal, fiber optic sensors have attracted widespread attention due to their advantages such as high sensitivity, resistance to electromagnetic interference, miniaturization, lightweight design, and stable operation in harsh environments. With the continuous advancement of fiber optic manufacturing and photoelectric detection technologies, fiber optic sensing has evolved from initial laboratory research into a widely used technology, playing a vital role in various fields such as industrial monitoring, environmental monitoring, medical diagnosis, and civil engineering.
[0003] Currently, gesture recognition methods primarily utilize machine vision, employing image processing techniques to identify gesture postures by marking feature points and combining these with recognition algorithms. While continuous innovation in image processing technology has enabled machine vision-based gesture recognition methods to achieve good real-time performance and high accuracy, they also exhibit significant limitations, as follows:
[0004] (1) The lighting in the scene when the image is acquired affects the recognition result. Poor imaging effect will lead to a decrease in image recognition effect.
[0005] (2) When the camera collects images, it may collect too much information, which may cause privacy issues. This method cannot be used for collection in special circumstances.
[0006] (3) Image acquisition is affected by the fixed installation position of the camera, and the acquisition of human posture is limited to a certain area, and cannot recognize human posture in all directions.
[0007] The problems mentioned above are unavoidable in the process of recognizing gestures and postures based on machine vision. For example, environmental occlusion and real-time monitoring at night may occur during health monitoring.
[0008] In recent years, distributed fiber optic sensor systems based on quartz glass have been widely used in systems requiring minimal or no stretching, such as strain, pressure, vibration, acceleration, temperature, and humidity sensing. To overcome the instretchability of quartz fiber, researchers have done considerable work and developed many types of stretchable sensors. However, researchers have focused more on optical waveguide materials capable of withstanding large strain and repeated deformation. Traditional silica fiber has attracted considerable attention due to its advantages such as small size, ease of large-scale integration, low cost, and a mature market for supporting infrastructure. Therefore, its use as a flexible sensing element shows great promise. Furthermore, flexible sensing elements exhibit advantages over machine vision in human-computer interaction fields such as gesture recognition, overcoming limitations imposed by machine vision such as scene lighting, privacy, and line-of-sight occlusion.
[0009] This invention designs a human gesture recognition method based on a composite flexible sensor with a SNTNS structure and an LSTM neural network model. The invention aims to design a fiber-optic flexible composite structure with higher sensing sensitivity, apply it to a distributed sensing system, and then combine this system with a deep learning algorithm for gesture recognition. Summary of the Invention
[0010] The purpose of this invention is to provide a gesture recognition method based on a novel fiber optic composite flexible sensor structure and an LSTM neural network model. The core principle of this invention is to utilize a PDMS-coated SNTNS composite fiber structure to improve the curvature sensitivity of the fiber optic sensor. The bending angle is quantified by detecting changes in the optical power of the fiber optic composite sensor. This structure is then cascaded with a fiber Bragg grating (FBG) to identify the location of the bending. Combined with a deep learning model, gesture classification is performed, thereby achieving high-precision gesture recognition.
[0011] To achieve the above objectives, the present invention is implemented through the following technical solution:
[0012] The gesture recognition system upon which the gesture recognition method based on optical fiber composite flexible sensor and deep learning relies includes an optical fiber composite flexible sensor, a terminal processing and display device, an optical fiber spectral demodulation device, an optical fiber coupler, and a digital signal acquisition card.
[0013] The fiber optic composite flexible sensor, also known as a flexible sensor, comprises multiple SNTNS composite fiber structures and FBG. The SNTNS composite fiber structure is encapsulated by a PDMS film, and the radius of curvature of the bending is 1-2 cm.
[0014] The fiber optic spectral demodulation device has a built-in broadband light source and a photodetector for detecting the wavelength and intensity of the light reflected from the fiber grating.
[0015] The connection relationships of the components in the gesture recognition system are as follows:
[0016] The fiber optic spectral demodulation device is connected to the fiber optic coupler and the digital signal acquisition card respectively; the digital signal acquisition card is connected to the terminal processing and display device; the fiber optic coupler is connected to the fiber optic spectral demodulation device and the fiber optic composite flexible sensor respectively; the fiber optic composite flexible sensor is connected to the fiber optic coupler; one end of the fiber optic coupler is connected to multiple fiber optic composite flexible sensors, and the other end is connected to the fiber optic grating demodulator.
[0017] The functions of each component in the posture recognition system are as follows:
[0018] Fiber optic couplers couple multiple beams into a single beam, enabling optical path multiplexing; digital signal acquisition cards digitize the output signal of the demodulator and transmit it to terminal processing and display devices; fiber Bragg gratings are used for positioning, specifically by determining the position based on the one-to-one correspondence between the specific reflection center wavelength of the grating and its position; and fiber optic composite flexible sensors are used to transmit the angle at which bending occurs.
[0019] The gesture recognition method includes the following steps:
[0020] Step 1: The layout and cascade fixation of SNTNS structures. In this invention, medical-grade latex gloves are used as flexible substrates. Each SNTNS structure is laid on the back of the finger joints of the gloves and fixed by dot matrix bonding with UV curing adhesive. Single-mode optical fiber (SMF) is used for cascading in the middle, and FBG is cascaded at each fingertip of the gloves.
[0021] The number of SNTNS structures is the same as the number of finger joints that need to be detected.
[0022] Step 2: Record the relationship between the joint positions to be detected in Step 1 and the corresponding FBG center reflection wavelength;
[0023] Step 3: The broadband light emitted by the light source in the fiber optic spectral demodulation device enters different fiber optic flexible sensor branches through fiber optic couplers.
[0024] Step 4: The light entering the branch passes through the fiber optic composite flexible sensor. If the sensor bends at a certain angle, it will modulate the intensity of the input light. The optical power will be reduced due to macro-bending loss, and the output light will be the light after the first light intensity attenuation.
[0025] Step 5: After the light intensity is attenuated for the first time after step 4, the light passes through the FBG. The FBG backscatters the light whose wavelength meets the Bragg condition and transmits the rest of the light.
[0026] Step 6: The backscattered light undergoes a second light intensity attenuation through the fiber optic composite flexible sensor. The light that has undergone the second attenuation is combined with the light reflected from other branches through the fiber optic coupler and sent to the fiber optic spectral demodulation device for demodulation, outputting an analog signal carrying light intensity and wavelength information.
[0027] Step 7: The fiber optic spectral demodulation device connects the analog signal demodulated in Step 6 to the digital acquisition card, which then converts the corresponding analog signal into a digital signal.
[0028] Step 8: The digital acquisition card connects the digital signal from Step 7 to the terminal processing and display device, and the wavelength and intensity information of the light are observed through software on the terminal processing and display device.
[0029] Step 9: The terminal processing and display device decouples and separates the wavelength and intensity information of the light to obtain the intensity information that fluctuates over time.
[0030] Step 10: The terminal processing and display device performs low-pass filtering on the light intensity information that fluctuates over time to filter out high-frequency jitter and outputs the filtered light intensity information.
[0031] Step 11: Based on the filtered light intensity information output in Step 10, calculate the magnitude of the corresponding finger joint bending angle.
[0032] Step 12: Collect data on the bending angle of the finger joints. Users need to perform ten different hand gestures from 1 to 10 in accordance with the standard hand gesture posture. The amount of data collected for each gesture should be large enough to cover the gesture changes of different individuals in order to improve the generalization ability of the model.
[0033] Step 13: Neural Network Model Design and Training. A Long Short-Term Memory (LSTM) neural network model is used for training to avoid long-term dependency issues and the potential gradient vanishing or exploding problems during error backpropagation. To train and evaluate the neural network model, the dataset is divided into two subsets: a training set and a test set. The training set is used to train the neural network model and perform cross-validation to fine-tune hyperparameters such as the learning rate, while the test set is reserved for evaluating the neural network. After model parameter tuning and optimization, the trained neural network model is saved and can then be used for high-precision gesture recognition.
[0034] Beneficial effects
[0035] This invention provides a gesture recognition method based on a fiber optic composite flexible sensor with a SNTNS structure and an LSTM neural network model, which has the following advantages compared to existing technologies:
[0036] 1. The method provides a novel light composite flexible sensor structure with higher curvature sensitivity, which can be better applied to wearable devices;
[0037] 2. The method provides a gesture recognition method based on fiber optic sensing, which solves the problems of inconvenience in wearing due to sensor materials and susceptibility to electromagnetic interference in complex electronic environments when wearable devices measure gesture posture in the prior art.
[0038] 3. The method can be applied to systems that require human-computer interaction;
[0039] 4. The method uses a neural network model for training, which has high accuracy in human-computer interaction and can be widely applied to wearable devices when combined with other devices. Attached Figure Description
[0040] Figure 1 This is a structural diagram of a gesture recognition method based on an optical fiber composite flexible sensor and an LSTM neural network model according to the present invention.
[0041] Figure 2 This is a schematic diagram of gesture model reconstruction for a gesture recognition method based on an optical fiber composite flexible sensor with a SNTNS structure and an LSTM neural network model, according to the present invention.
[0042] Figure 3 This is a schematic diagram of the optical transmission processing proposed in the gesture recognition method based on the SNTNS structure optical fiber composite flexible sensor and LSTM neural network model of the present invention.
[0043] Figure 4 This is a schematic diagram of the SNTNS fiber optic structure of the gesture recognition method based on the SNTNS structure and the LSTM neural network model of the present invention.
[0044] Figure 5 This is a schematic diagram of a deep learning network structure provided by the present invention for a gesture recognition method based on an optical fiber composite flexible sensor with SNTNS structure and an LSTM neural network model; Detailed Implementation
[0045] The present invention will now be described in conjunction with the accompanying drawings and specific embodiments. The following drawings and embodiments are for illustrative purposes only and should not be construed as limiting the scope of this patent.
[0046] To better illustrate this embodiment, some parts of the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions;
[0047] It is understandable to those skilled in the art that some well-known details may be omitted from the accompanying drawings.
[0048] Example 1
[0049] This embodiment proposes a gesture recognition wearable device based on a PDMS-coated SNTNS fiber optic structure for sensing. Figure 1 This includes: wearable hand unit, terminal processing and display device, fiber optic spectral demodulation device, fiber optic coupler, and digital signal acquisition card.
[0050] Figure 2 This is a schematic diagram illustrating the application of the gesture recognition wearable device proposed in this embodiment of the invention.
[0051] See Figure 2 Multiple fiber optic composite flexible sensors are cascaded on different surfaces of the hand-wearing body. In this embodiment, the hand-wearing body is a medical-grade latex glove. Figure 2 (The main body of the hand to be worn is not shown); Taking the little finger of the glove as an example, three fiber optic composite flexible sensors 1 are fixed at the three joints of the little finger, and three FBGs (2) with different central reflection wavelengths are fixed at the fingertip. They are connected in series with the three fiber optic flexible composite sensors respectively. The three fiber optic composite flexible sensors are connected to the fiber optic coupler through single-mode optical fiber respectively.
[0052] The fiber optic spectral demodulation device in this embodiment is a fiber optic grating demodulator, which has a built-in broadband light source with an average light intensity of 2.4uW, a wavelength range of 1530-1560nm, and a 3dB bandwidth of 0.1nm for the reflected light from the FBG.
[0053] See Figure 3 Several wavelengths (λ1, λ2, λ3, ...) of light emitted from the light source are fed into the sensing optical fibers of the five fingers via fiber optic couplers. The correspondence between the reflected wavelength of the FBG and the sensor position is recorded. After modulation of the specific wavelength of light in each fiber by the sensing part, it is reflected by the corresponding FBG and modulated a second time at the sensing part (to improve sensitivity). The light is then coupled into the fiber optic demodulator via the fiber optic coupler. The fiber optic demodulator separates the light of several wavelengths (λ1, λ2, λ3, ...) and connects the signal to the digital acquisition card. The digital acquisition card converts the corresponding analog signal into a digital signal, which enters the terminal processing and display device. The terminal processing and display device decouples and separates the wavelength and intensity information of the light to obtain the intensity information that fluctuates over time. This information is then low-pass filtered to remove high-frequency jitter, and the filtered intensity information is output. The bending angle of the corresponding finger joint is calculated from this intensity information. The data on the relationship between the position of the finger joint and the wavelength are combined to obtain the feature data of the gesture to be recognized for gesture recognition.
[0054] See Figure 4The SNTNS structure is formed by splicing five optical fibers one after another. Each coreless optical fiber is 1mm long, and the three-core optical fiber in the center is 15mm long. The two ends of the structure are connected to other structures with single-mode optical fibers.
[0055] Coreless optical fibers are not particularly sensitive to temperature changes, but they are quite sensitive to external strain and bending (such as changes in hand gestures), effectively avoiding the influence of the measured hand temperature environment on the fiber optic sensor. Meanwhile, three-core optical fibers have strong anti-interference capabilities, are less affected by environmental noise, and can avoid the influence of slight hand tremors on the signal; furthermore, three-core optical fibers also have temperature self-compensation characteristics, resulting in good stability. The SNTNS fiber optic composite sensing unit proposed in this embodiment is relatively simple to manufacture and inexpensive, making it suitable for mass production.
[0056] Specifically, this embodiment uses a single-mode fiber with a core diameter of approximately 9 μm to guide the light emitted from the light source into the coreless-three-core-coreless fiber sensing unit. The light enters the first segment of the coreless fiber in the sensing unit through the single-mode fiber at the incident end. The light propagates perpendicularly to the fiber core, minimizing scattering. Since the external refractive index is lower than that of the coreless fiber, and they are very close, the coreless fiber can be approximated as a weakly guided multimode fiber. The energy of the single-mode fiber is fully coupled into the higher-order modes of the cladding of the coreless fiber. The higher-order modes of the cladding are fully coupled due to multimode interference. The three-core fiber in the sensing unit consists of three independent fiber cores, each surrounded by a cladding with a lower refractive index. The optical signal propagates within the fiber cores and is held within the cores by total internal reflection.
[0057] Since the core refractive index of a three-core optical fiber is 1.457, and the average refractive index of PDMS after curing is 1.4, the lower refractive index of PDMS can simulate the function of the fiber cladding when encapsulating optical fiber, effectively guiding light propagation within the fiber core. In addition, the mechanical properties of PDMS, such as its low Young's modulus and high Poisson's ratio, make it an ideal choice for encapsulating the sensitive area of optical fiber sensors, ensuring that the sensor maintains high sensitivity and stability when subjected to external forces.
[0058] Three-core optical fiber contains three independent cores, which allows multiple signals to be transmitted simultaneously in the same fiber. A large amount of high-order energy is coupled into the core of the three-core optical fiber at the fusion splice between the coreless fiber and the three-core fiber. At the same time, under multipath propagation, the signal transmission efficiency is improved and the anti-interference ability is stronger. As can be seen from the fiber interference theory, the core-cladding interference peak is sensitive to the refractive index, and this characteristic can be used for curvature measurement.
[0059] Example 2
[0060] This embodiment proposes a gesture recognition method based on an optical fiber composite flexible sensor with an SNTNS structure and an LSTM neural network model. It also includes a host computer, which receives the feature values and their corresponding feature vectors extracted by the information processing module and uses a deep learning neural network to learn the feature vectors.
[0061] Specifically, this embodiment uses a Long Short-Term Memory (LSTM) neural network to learn the feature vector. A schematic diagram of the LSTM structure is shown below. Figure 5 As shown.
[0062] In this embodiment, the user needs to perform ten different hand gestures from 1 to 10 according to standard gestures. The amount of data collected for each gesture should be large enough to cover the gesture variations of different individuals, thereby improving the model's generalization ability. All preprocessed (low-pass filtered) light intensity and knuckle position information are stored in an Excel spreadsheet for subsequent deep learning grid training. Through the above data preprocessing steps, this invention ensures the quality of the input data, improves the robustness and generalization ability of the model, and lays the foundation for high-precision gesture recognition.
[0063] The LSTM neural network model has a layered structure. In this embodiment, the classification task is relatively simple, so a single-layer LSTM is used for higher computational efficiency. The input to the LSTM is a time-series data sequence, and the input at each time step is a preprocessed light intensity information vector. Its output is the standard hand gesture for the digits 1 to 10 in the classification task. The LSTM neural network uses a gating structure to delete or add information in the neurons. The gating structure is a combination of a sigmoid layer and a point operation. This activation function controls how much information can pass through each part. LSTM computation uses three gates: the forget gate, the input gate, and the output gate. The forget gate determines what information to forget from the cell state, reading the output of the previous loop and the current input. The input gate consists of a sigmoid layer and a tanh layer. The sigmoid layer determines what information needs to be updated in the cell state, while the tanh layer updates the cell state with the necessary information. After updating the cell state, the output of the current loop is obtained based on the updated cell state, the output of the previous loop, and the input information. The output gate uses a sigmoid layer to determine which part of the cell state needs to be output, then uses a tanh layer to process the cell state, multiplying the result point-by-point with the output of the sigmoid layer to obtain the hidden state h. t The data is then processed through a Softmax layer to obtain the final output. For the loss function, the cross-entropy loss function is used for optimization to minimize the error between the true and predicted classes.
[0064] Furthermore, the host computer uses 80% of the received dataset as the training set and 20% as the test set. A threshold for prediction accuracy is set when the absolute value of the difference between the predicted and true values is less than 0.3, and the learning rate is set to 0.005. The deep learning neural network model is trained using the training set and validated using the validation set. When learning feature vectors, the deep learning neural network model matches these vectors with changes in external hand gestures to be recognized, i.e., it learns the features of external hand gestures. The model's performance is repeatedly evaluated and iterated using the test set to further improve its generalization ability. The deep learning neural network model capable of recognizing hand gestures is then transmitted to the information processing module for gesture recognition, enabling subsequent offline applications of the deep learning neural network model in the information processing module.
[0065] In summary, the above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A gesture recognition method based on a fiber optic composite flexible sensor with a SNTNS structure and an LSTM neural network model, comprising a single-mode-coreless-three-core-coreless-single-mode fiber (SMF-NCF-TCF-NCF-SMF, SNTNS) composite flexible sensor structure coated with polydimethylsiloxane (PDMS) and cascaded into a fiber optic distributed sensing system, using this distributed sensing system in conjunction with a Long Short-Term Memory (LSTM) neural network for gesture recognition, characterized in that... Includes the following steps: Step 1: Fabricate a flexible fiber optic composite sensor with a SNTNS structure and cascade it with a fiber Bragg grating to form a sensing element; Step 2: Distribute the various fiber optic composite flexible sensors at the finger joints of the wearable hand device (glove) and fix them with UV adhesive; Step 3: Record the finger joint position to be detected in Step 1 and the corresponding center reflection wavelength of the fiber Bragg grating; Step 4: The broadband light (1530-1560nm) emitted by the fiber optic spectral demodulation equipment enters each sensing branch through the optical coupler. In the sensing branch, due to the change in human hand gestures, the curvature of the sensor changes, and the broadband light is subjected to intensity modulation. The optical power is reduced due to macro-bending loss. After modulation, the light of the corresponding wavelength is reflected by the fiber Bragg grating and passes through the sensing element again for a second intensity modulation, and the optical power is further reduced. Step 5: The light that has been modulated twice enters the fiber optic spectral demodulation device through the optical coupler, and the analog signal with light intensity and wavelength information is demodulated. Step 6: The analog signal is converted into a digital signal by the digital acquisition card and input into the terminal processing and display device; Step 7: The terminal processing and display device decouples and separates the digital signal, and performs low-pass filtering to obtain the final light intensity information; Step 8: Calculate the angle of the corresponding finger joint bend based on the light intensity information; Step 9: Collect multiple sets of standard gesture data of numbers 1 to 10 as a dataset, construct a deep learning model of long short-term memory neural network, input preprocessed light intensity information, output gesture classification results, and train the model through cross-entropy loss function; Step 10: Input the real-time collected light intensity information into the trained deep learning model and output the corresponding gesture category.
2. The gesture recognition method based on a fiber optic composite flexible sensor with a SNTNS structure and an LSTM neural network model according to claim 1, characterized in that: A fiber composite structure of single-mode-coreless-three-core-coreless-single-mode is adopted, and PDMS encapsulation is used as the fiber composite flexible sensor.
3. The gesture recognition method based on a fiber optic composite flexible sensor with a SNTNS structure and an LSTM neural network model according to claim 1, characterized in that: In step 1, each fiber composite flexible sensor is connected to a fiber coupler at one end with a single-mode fiber, and the other end is fixed to a fiber Bragg grating at the fingertip with a single-mode fiber in series. The fiber coupler has a total of 15 fiber sensing branches, and each finger needs to be fixed with 3 sensing branches, located at its 3 finger joints.
4. The gesture recognition method based on a fiber optic composite flexible sensor with a SNTNS structure and an LSTM neural network model according to claim 1, characterized in that: In step 4, the change in reflected light power is measured by recording the center reflection wavelength of the fiber Bragg grating corresponding to the finger joint, thereby obtaining the change in the degree of bending of the corresponding finger joint.
5. The gesture recognition method based on a fiber optic composite flexible sensor with a SNTNS structure and an LSTM neural network model according to claim 1, characterized in that: In step 8, the light intensity information of each branch reflected back by each fiber Bragg grating is obtained through a distributed sensing system. The angle of the composite flexible sensor is proportional to the optical power loss caused by macro bending within a certain range. The bending angle of each finger joint can be deduced based on the demodulated light intensity information. On this basis, a human hand is modeled. The model is divided into 16 parts, including the palm and 15 finger joints. Given the angle of each finger joint, the hand gesture can be reconstructed.
6. The gesture recognition method based on a fiber optic composite flexible sensor with a SNTNS structure and an LSTM neural network model according to claim 1, characterized in that: In step 9, a deep learning model of a long short-term memory neural network is constructed. The input is a time-series data sequence, and the input at each time step is the light intensity vector x at each finger joint. t The output is the final gesture category (1-10). Its hidden layer first uses a single-layer LSTM. Single-layer LSTM is suitable for relatively simple tasks and has high computational efficiency. If the subsequent training effect is not good, a double-layer LSTM can be used. The double-layer LSTM is a stacked structure, and the output of the previous layer will be used as the input of the next layer to enhance its feature abstraction ability. Each LSTM layer uses sigmoid and tanh functions to achieve gating. The sigmoid function controls whether information is retained or discarded in each layer, while the tanh function maps the input to a symmetric interval, accelerating model convergence. This invention's LSTM neural network model uses cross-entropy loss as the loss function to measure the difference between the predicted probability distribution and the true distribution, making it suitable for this multi-gesture classification task. This model learns from preprocessed light intensity information to obtain a deep learning network model capable of recognizing gestures. This neural network model is then transmitted to a terminal processing and display device for gesture recognition.