Gesture recognition method based on optical fiber sensing technology and deep learning

Through the combination of multimode fiber sensing technology and deep learning, the speckle image changes caused by finger bends are solved, and the accuracy problem of existing gesture recognition technology under light and background complexity is achieved, achieving high-precision and robust gesture recognition.

CN120496128APending Publication Date: 2025-08-15TIANJIN POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510585080.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-01
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing gesture recognition technology is highly sensitive to light conditions, background complexity, occlusion and other factors. Inertial sensors are difficult to distinguish static gestures. The wearable electromyography signal method is low and susceptible to physiological changes. It is difficult for traditional methods to achieve high-precision gesture recognition.

Method used

Multimode fiber sensing technology is used to combine deep learning, and by fixing multimode fiber on the glove, using speckle image changes caused by finger bending, combining convolutional neural networks and self-attention mechanisms for gesture classification, designing a high-resolution CCD camera to collect real-time data preprocessing and enhancement, and building a high-precision gesture recognition model.

Benefits of technology

High-precision gesture recognition under different lighting conditions and complex backgrounds is realized, which improves the robustness and generalization ability of the model, and adapts to gesture changes of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496128A_ABST
    Figure CN120496128A_ABST
Patent Text Reader

Abstract

The invention provides a gesture recognition method based on an optical fiber sensing technology and deep learning, and relates to the technical field of optical fiber sensing. According to the gesture recognition method, gesture information is captured through speckle image changes of the multimode optical fiber fixed on a glove under different gestures by utilizing the deformation characteristic of the multimode optical fiber, and high-precision gesture classification is carried out in combination with a convolutional neural network (CNN) and a self-attention mechanism. The specific implementation of the method comprises the steps of layout and fixation of multimode optical fibers, speckle image acquisition, data preprocessing, neural network model design and training and the like. Wherein a 635nm semiconductor laser is used as a light source, a CCD camera is used for collecting speckle images in real time, a Python script is used for carrying out optimization processing and data enhancement on the images, a final deep learning model is composed of a feature extraction layer, a global feature modeling layer and a classification layer, and high-precision recognition of gestures is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of optical fiber sensing technology, and in particular to a gesture recognition method based on optical fiber sensing technology and deep learning. Background Art

[0002] Fiber optic sensing technology utilizes the propagation properties of light waves through optical fibers to detect and measure physical, chemical, or other variables. Since its initial introduction, fiber optic sensors have garnered widespread attention due to their high sensitivity, immunity to electromagnetic interference, miniaturization, lightweight design, and stable operation in harsh environments. With the continuous advancement of fiber optic manufacturing and photoelectric detection technologies, fiber optic sensing has evolved from its initial laboratory research into a widely used technology, playing a vital role in various fields, including industrial monitoring, environmental monitoring, medical diagnostics, and civil engineering.

[0003] Current gesture recognition technologies are mainly divided into the following categories: computer vision methods (such as RGB cameras, depth cameras): rely on hand images or depth information, combined with deep learning (CNN) for gesture classification. However, this method is highly sensitive to factors such as lighting conditions, background complexity, and occlusion; inertial sensor methods (such as accelerometers, gyroscopes): identify by measuring the motion acceleration and angular velocity of fingers or palms. This method is suitable for dynamic gestures, but it is difficult to distinguish static gestures, and sensor drift will affect long-term stability; electromyographic signal method: use surface electromyographic sensors to collect hand muscle activity information for gesture recognition. Although it has high biological specificity, it is easily interfered by factors such as physiological changes and skin contact resistance, has low wearing comfort, and requires calibration for long-term use.

[0004] Against this backdrop, the use of fiber optic sensing technology for gesture recognition has become a novel and promising approach. Multimode fiber (MMF) is a type of optical fiber that supports multiple modes of propagation. When laser light is incident on an MMF, multipath interference and inter-mode coupling within the fiber form a unique speckle pattern at the output. When the MMF is attached to the fingers of a glove, bending or stretching the finger causes the fiber to deform, altering the internal mode propagation paths and causing specific changes in the speckle pattern at the output. These changes contain rich gesture information, and due to the high-dimensional complexity of the speckle pattern, it can provide higher resolution and sensitivity than traditional sensors. By using deep learning to train speckle patterns for different gestures, a mapping model between gestures and speckle changes can be established, enabling high-precision gesture recognition. Summary of the Invention

[0005] The purpose of the present invention is to provide a gesture recognition method based on fiber optic sensing technology and deep learning. The core principle of the present invention is to utilize the multimode interference characteristics of multimode optical fiber, through the changes in speckle images caused by the deformation of multimode optical fiber under different gestures, combined with a deep learning model to perform gesture classification, thereby achieving high-precision gesture recognition.

[0006] The present invention is implemented by adopting the following technical solutions: a gesture recognition method based on fiber optic sensing technology and deep learning, including the layout and fixation of multimode optical fibers: fixing the multimode optical fibers on the gloves to obtain speckle images under different gestures; speckle image acquisition: using a CCD camera to capture speckle images at the output end of the multimode optical fibers in real time; data preprocessing: performing preprocessing operations on the collected speckle images; neural network model design and training: designing a suitable deep learning neural network model and using the collected data set for network training.

[0007] Furthermore, for the layout and fixation of multimode optical fibers, the present invention uses medical-grade latex gloves as a flexible substrate, and layouts multimode optical fibers on the finger surfaces, arranging them in a serpentine pattern along the sides of each finger. UV-curing glue is used for dot-matrix bonding and fixation to ensure that the deformation characteristics of the optical fibers can accurately reflect changes in gestures.

[0008] Furthermore, for speckle image acquisition, the present invention uses a semiconductor laser with a wavelength of 635nm as the laser light source. The laser passes through a section of single-mode optical fiber and is input into the input end of a multimode optical fiber placed on a medical-grade latex glove. After passing through the multimode optical fiber, a dynamic speckle image is formed at the output end. At the output end of the multimode optical fiber, the present invention uses a high-resolution CCD camera for real-time acquisition of speckle images, with the camera acquisition frame rate set to 10 frames per second. During the acquisition process, when the finger bends and causes the optical fiber to deform, the modal interference within the optical fiber changes, and the speckle pattern at the output end exhibits unique spatial distribution characteristics. The user is required to perform ten gestures from 1 to 10 in sequence according to standard gesture postures. The amount of data collected for each gesture should be large enough to cover the gesture variations of different individuals and improve the generalization ability of the model.

[0009] Furthermore, for data preprocessing, the present invention uses a CCD camera to capture speckle images under different gestures and stores them in JPG format. Since the original captured images contain a large amount of irrelevant black background in addition to the circular speckle area, directly inputting them into the neural network training may affect the learning effect of the model. Therefore, the present invention uses the Pycharm integrated development environment and Python to write a data preprocessing program to optimize the original images. Specifically, the circular speckle area is first accurately extracted through an image processing algorithm, and useless background information around it is removed, thereby retaining key information and improving data quality. In order to improve the generalization ability of the model in recognizing gestures of different users, the present invention further introduces a data enhancement strategy, performing random rotation, mirror flipping, contrast adjustment and other operations on the processed speckle images to simulate possible changes under different gesture states, thereby expanding the data distribution and enhancing the model's adaptability to complex scenes. Finally, all preprocessed and data-enhanced speckle images are classified according to gesture categories (1 to 10) and stored in a designated folder for subsequent deep learning network training. Through the above data preprocessing steps, the present invention ensures the quality of input data, improves the robustness and generalization ability of the model, and lays the foundation for high-precision gesture recognition.

[0010] Furthermore, the neural network model design and training utilizes a convolutional neural network (CNN) architecture combined with a self-attention mechanism to fully extract both local and global features of the speckle image. The CNN component learns local texture features of the speckle, such as local variations in the pattern distribution, while the self-attention mechanism focuses on the global characteristics of the entire speckle pattern, improving classification accuracy. To train and evaluate the neural network model, the dataset is divided into three subsets: training, validation, and test. The training and validation subsets are used for training and cross-validation to fine-tune hyperparameters such as the learning rate and mini-batch size, while the test subset is reserved for evaluation purposes. During model training, the present invention optimizes the model using a cross-entropy loss function and incorporates the Adam optimizer for parameter adjustment. The initial learning rate is set to 0.001, and a learning rate decay strategy is employed to enhance training stability. The network model used in this invention is a classification network model, with preprocessed speckle images as input and predicted gesture information as output. After model parameter tuning, the trained network model is saved and can be used for high-precision gesture recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a flowchart of a gesture recognition method based on optical fiber sensing technology and deep learning provided by an embodiment of the present invention;

[0012] Figure 2 This is a schematic diagram of a data acquisition system for optical fiber gesture recognition provided by an embodiment of the present invention;

[0013] Figure 3 This is a schematic diagram of a deep learning network model structure provided by an embodiment of the present invention;

[0014] Reference numerals: 1 is a 635 nm semiconductor laser source; 2 is a single-mode optical fiber; 3 is a multi-mode optical fiber; 4 is a medical-grade latex glove; 5 is a CCD camera; 6 is a network cable; and 7 is a laptop computer. DETAILED DESCRIPTION

[0015] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The steps in the following embodiments are only for the convenience of explanation and are not limited to the order of the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0016] See Figure 1 An embodiment of the present invention provides a flow chart of a gesture recognition method based on fiber optic sensing technology and deep learning, which includes several key steps. These steps include the deployment and fixing of multimode optical fibers, which are placed on medical-grade latex gloves as gesture sensing units. Speckle image acquisition involves using a laser to input optical signals into the optical fibers, and using a high-resolution CCD camera to capture speckle images of the optical fibers under different gestures. Different gestures cause different deformations in the optical fibers, resulting in changes in the speckle pattern. These speckle images therefore contain key information about the gestures. Data preprocessing involves preprocessing the collected speckle images to remove non-speckle areas and employing data augmentation strategies to expand the dataset. Finally, neural network model design and training involves constructing a deep neural network model. This invention utilizes a model that combines a deep convolutional neural network with a self-attention mechanism to fully extract local and global features of the speckle images. The deep learning model is trained using the collected training data, enabling it to learn the mapping relationship between speckle images and gestures. After training, the system can analyze speckle images at the optical fiber end in real time and utilize the trained neural network model for gesture recognition.

[0017] See Figure 2 An embodiment of the present invention provides a data acquisition device for a gesture recognition system using optical fiber sensing technology, comprising a 635nm semiconductor laser source 1, a single-mode optical fiber 2, a multimode optical fiber 3, a medical-grade latex glove 4, a CCD camera 5, a network cable 6, and a laptop computer 7. The 635nm semiconductor laser source 1 is connected to one end of the multimode optical fiber 3 via the single-mode optical fiber 2, and the other end of the multimode optical fiber 3 is connected to the CCD camera 5. The multimode optical fiber 3 is fixed to the side of each finger of the medical-grade latex glove 4 using UV curable adhesive in a serpentine routing manner. The CCD camera 5 is connected to the network cable interface of the laptop computer 7 via the network cable 6. The laptop computer 7 is used to observe and save the speckle image collected by the CCD camera 5.

[0018] Specifically, during the fiber optic layout and fixation process, the present invention uses UV-curing adhesive to firmly fix the multimode optical fiber to the surface of the medical-grade latex glove. Specifically, the multimode optical fiber is laid out in a serpentine pattern along the side of the finger, and UV-curing adhesive is used to bond the knuckles and the bends of the multimode optical fiber to ensure its stable layout, ensure that the multimode optical fiber produces a stable deformation effect when the finger is bent, and ensure that a dynamic speckle image is generated at the output end when the gesture changes.

[0019] Specifically, a fiber fusion splicer is used to fuse a multimode fiber to a single-mode fiber. The single-mode fiber is then connected to a 635nm semiconductor laser source. A high-resolution CCD camera is used at the output end of the multimode fiber to capture speckle images in real time, forming a complete speckle image acquisition device. When a user, wearing gloves, performs standard hand gestures from 1 to 10, the speckle image captured by the high-resolution CCD camera changes accordingly. This means that the desired gesture information is hidden in the captured speckle image. Based on this principle, the user repeatedly performs standard hand gestures from 1 to 10, while intentionally introducing random jitter in the joints to simulate natural movement deviations. The speckle image information corresponding to the gesture is then captured and saved.

[0020] Specifically, the acquisition frame rate of the high-resolution CCD camera is set to 10 frames per second. When the user performs the corresponding gesture, the finger joints are randomly jittered, and the speckle image at the output end shows dynamic changes. The speckle images under the corresponding gesture are collected. 6,000 speckle images are collected for each gesture, and a total of 60,000 speckle images (jpg format) are collected for 10 gestures to ensure the adequacy and robustness of the data.

[0021] Furthermore, during data preprocessing, the speckle images captured by the CCD camera contain the complete fiber output pattern, but also contain a large proportion of black background areas. These background areas do not provide effective feature information and may affect the training effect of the model. Therefore, during the data preprocessing stage, the original images need to be cropped and optimized to ensure that the data input to the neural network has higher validity and recognizability. The present invention uses Python to write image processing scripts and batch processes the original speckle images in the Pycharm integrated development environment. The system automatically detects and extracts the effective speckle areas in the image through image recognition algorithms, while removing irrelevant black background, ensuring that the size and content of all input images are consistent. To enhance the generalization ability of the model, the present invention introduces a series of data augmentation strategies. Specifically, the cropped speckle images are randomly rotated, mirrored, flipped, and the contrast and brightness are adjusted. In addition, slight noise is added to some data to improve the model's robustness to data deviations. All data-augmented images are classified according to gesture categories (1 to 10) and stored in a designated training dataset folder, providing high-quality data input for subsequent deep learning network training.

[0022] See Figure 3 An embodiment of the present invention provides a schematic diagram of the deep learning network model structure. This network extracts effective features from optical fiber speckle images and achieves high-precision gesture classification by learning the patterns of speckle changes under different gestures. This invention utilizes a deep learning architecture that combines a convolutional neural network (CNN) with a self-attention mechanism (Transformer) to fully extract local and global features of speckle images and improve adaptability to complex gestures.

[0023] Specifically, the deep learning network of the present invention consists of the following parts: input layer, feature extraction layer (convolutional layer), global feature modeling layer (Transformer), classification layer, and is equipped with an appropriate regularization mechanism to enhance the generalization ability of the model.

[0024] Specifically, the input layer is responsible for receiving the preprocessed fiber speckle image. The size of the input image is fixed at 128×128 pixels. Since the input data is a single-channel grayscale image, it is converted into a standardized tensor format and normalized at input to improve the stability of network training.

[0025] Specifically, in the feature extraction layer, the present invention employs multiple convolutional layers to extract local features. The first layer typically uses a larger convolution kernel (7×7) to capture the global pattern distribution of the speckle image, while subsequent layers use smaller convolution kernels (3×3) to learn more detailed speckle structural features. Each convolution layer is equipped with a ReLU activation function to increase the network's nonlinear expression capabilities, while batch normalization is used to accelerate convergence and improve generalization. To reduce computational complexity and enhance robustness to displacement transformations, a maximum pooling operation is added after the convolution layer to reduce data dimensionality while retaining the most representative features.

[0026] Specifically, at the global feature modeling layer, the present invention, based on the Transformer architecture, converts the feature maps extracted by the convolutional layer into a series of feature vectors and calculates the correlation between different regions through a self-attention mechanism. This approach enhances the model's ability to perceive global pattern changes and improves its adaptability to different user gestures.

[0027] Specifically, at the classification layer, the present invention uses the Softmax activation function to output the final gesture category (1-10). Specifically, the last layer of the network is a fully connected layer with the same number of nodes as the number of gesture categories, that is, 10 output units, each corresponding to a gesture category. The Softmax function converts the network output into a probability distribution and ultimately selects the category with the highest probability as the prediction result. As for the loss function, the cross-entropy loss function is used for optimization to minimize the error between the predicted category and the true category.

[0028] During network training, the present invention uses the Adam optimizer with an initial learning rate set to 0.001. A learning rate decay strategy is used to gradually decrease the learning rate over the training process to improve model stability. The training dataset is split into a training set and a validation set in an 8:2 ratio. Each epoch in training is evaluated on the validation set to monitor model performance and prevent overfitting. Furthermore, to enhance the model's generalization capabilities, the present invention incorporates data augmentation during training, including random rotation, mirror flipping, and contrast adjustment, to expand the data distribution and improve the model's robustness to different gesture transformations.

[0029] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A gesture recognition method based on fiber optic sensing technology and deep learning, characterized in that: The following steps are involved: (1) Layout and fixation of multimode optical fiber: The multimode optical fiber is laid out in a serpentine pattern along the side of the finger of a medical-grade latex glove and fixed in a dot matrix manner using UV curing glue; (2) Speckle image acquisition: Laser is input into the multimode optical fiber through a 635nm semiconductor laser source, and a CCD camera is used to capture dynamic speckle images at the output end of the multimode optical fiber under different gestures in real time; (3) Data preprocessing: The background of the collected speckle image is cropped to remove the black background area, and data enhancement is performed through random rotation, mirror flipping, contrast adjustment, and noise addition operations; (4) Neural network model training: A deep learning model that integrates convolutional neural network (CNN) and self-attention mechanism is constructed, and the preprocessed speckle image is input to output the gesture classification result. The model is trained using the cross entropy loss function and the Adam optimizer; (5) Gesture recognition: The real-time collected speckle image is input into the trained deep learning model, and the corresponding gesture category is output.

2. The method according to claim 1, characterized in that The multimode optical fiber is laid out in a serpentine path along the side of each finger of a medical latex glove and fixed at the bend of the knuckles with UV-curable glue to ensure that the optical fiber deformation is synchronized with the bending movement of the finger.

3. The method according to claim 1, characterized in that The acquisition frame rate of the CCD camera is 10 frames per second, and the acquired speckle images cover ten standard gestures, with no less than 6,000 images collected for each gesture.

4. The method according to claim 1, wherein The data preprocessing includes: automatically extracting speckle circular areas using an image processing algorithm and uniformly adjusting the image size to 128×128 pixels; data enhancement strategies include random rotation, mirror flipping, and adding slight noise to expand the data set and improve the model's adaptability to changes in different user gestures.

5. The method according to claim 1, wherein The structure of the deep learning model includes: an input layer that receives a 128×128 pixel single-channel grayscale speckle image; a feature extraction layer that contains multiple convolutional layers, uses 7×7 and 3×3 convolution kernels to extract local texture features, and is connected to a maximum pooling layer; a global feature modeling layer that calculates the global pattern correlation of the speckle image through a self-attention mechanism; and a classification layer that uses a fully connected layer and a softmax function to output gesture category probabilities from 1 to 10.

6. The method according to claim 1, characterized in that The deep learning model was optimized using the cross-entropy loss function and the Adam optimizer with an initial learning rate set to 0.001 and a learning rate decay strategy to improve training stability.

7. The method according to claim 1, characterized in that The training dataset is divided into a training set and a validation set in a ratio of 8:2, and the test set is retained independently to evaluate the generalization performance of the model.

8. The method according to claim 1, characterized in that The gesture recognition system includes: a 635nm laser source, a single-mode optical fiber, a multimode optical fiber arranged in a glove, a CCD camera, and a computer equipped with a deep learning model, wherein the output end of the multimode optical fiber is aligned with the optical path of the CCD camera.