Dynamic Sign Language Recognition Method, System and Device Based on Afterimage Diagram

Dynamic sign language recognition is converted into two-dimensional image classification through optical flow method and afterimage method, which solves the problems of large amount of calculation and poor real-time performance in the prior art, and achieves efficient and real-time sign language recognition.

CN115719512BActive Publication Date: 2025-07-29SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211425260.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-07-29
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

The prior art has a large amount of calculation in dynamic sign language recognition, poor real-time performance, and requires wearable devices to be used for recognition, which is insufficient practicality.

Method used

The optical flow method is used to fusion of the afterimage image method to convert three-dimensional video classification problems into two-dimensional image classification problems, and sign language recognition is performed through frame processing, dynamic skin segmentation, afterimage image synthesis and deep learning methods.

Benefits of technology

It reduces the amount of computing in sign language recognition, ensures real-time, and improves the practicality of recognition without wearing a device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719512B_ABST
    Figure CN115719512B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of image processing technology, and provides a dynamic sign language recognition method, system and device based on a residual image. The method includes the following steps: obtaining a video to be recognized, and performing frame division processing to obtain pictures to be recognized; using an optical flow method to perform dynamic skin segmentation processing on the pictures to be recognized, extracting the dynamic regions in the video frames, and extracting the skin-color regions in the dynamic regions of the video frames to obtain images of the hand regions performing sign language actions; synthesizing the images of the hand regions performing sign language actions into a residual image in the chronological order of the picture frames and according to the principle of decreasing transparency; using a deep learning method to perform image classification on the residual image to obtain the text information corresponding to the video to be recognized. The dynamic sign language recognition that combines the optical flow method and the residual image method is proposed, which converts the three-dimensional video classification problem into a two-dimensional image classification problem, thereby reducing the computational amount required for sign language recognition at the dimension level and ensuring the real-time performance of sign language recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of image processing. Specifically, it relates to a dynamic sign language recognition method, system, and device based on an afterimage map. Background Technique

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] Dynamic gesture recognition realizes the conversion from sign language video to text, including four aspects: gesture detection and segmentation, gesture tracking, feature extraction, and gesture classification. The present disclosure analyzes the existing technical characteristics of these four aspects, integrates the existing models, and proposes a dynamic sign language recognition method based on an afterimage map, aiming to make up for the defects of using only one technology for sign language recognition through a multi-process afterimage map method, thereby achieving better results.

[0004] Gesture detection and segmentation is to separate the hand from the background, and the gesture segmentation effect is directly related to the accuracy of gesture recognition. Deep convolutional networks have achieved great success in image hand segmentation. However, this CNN-based method mainly focuses on per-frame inference, which is inefficient for hand segmentation in videos because of the continuity and redundancy in consecutive video frames. Li et al. proposed a method to extend the convolutional neural network-based hand segmentation method from static images to video images, which consists of two main branches: flow-guided feature propagation and lightweight occlusion-aware detail enhancement. The flow-guided feature propagation branch results in a large accuracy drop due to distortion and occlusion problems in the distortion.

[0005] Gesture tracking essentially analyzes the sign language video frame by frame to locate the position of the target hand in each frame of the video image. Li Ming et al. proposed an algorithm for detecting moving targets in complex backgrounds based on the pyramid LK optical flow method combined with DBSCAN clustering to effectively eliminate the interference of complex backgrounds and achieved good results in detecting moving targets. However, this algorithm still has false judgments in the case of shadows and occlusions of moving targets.

[0006] Hand feature extraction is a key step in gesture recognition. Good features can not only improve the recognition accuracy but also reduce unnecessary computational complexity while fully exploring data information. Gesture features mainly include global features (color, texture, shape, etc.) and local features (corner-point local features and region local features).

[0007] The gesture classification process classifies the extracted spatio-temporal features of gestures, which is the last step in gesture recognition. Yang Yanfang et al. proposed an acceleration gesture recognition algorithm based on the combination of convolutional neural network and long short-term memory network. By collecting three-axis acceleration gesture data through Wiimote, in the test of 2,400 data of 8 acceleration gestures of 10 testers, an accuracy rate of 96.2% was achieved. ElBadawy et al. input the image frames extracted from the sign language video stream into a 3D convolutional network for the recognition of Arabic dynamic sign language, and the accuracy rate reached 90%. The current acceleration gesture recognition algorithms need to be recognized by wearing devices, and the practicability is poor. Summary of the Invention

[0008] To solve the above problems, the present disclosure proposes a dynamic sign language recognition method, system and device based on afterimage graphs, and proposes a dynamic sign language recognition that combines the optical flow method and the afterimage graph method, which converts the three-dimensional video classification problem into a two-dimensional image classification problem, thereby reducing the computational complexity required for sign language recognition at the dimension level and ensuring the real-time performance of sign language recognition. The sign language recognition implemented by the afterimage graph method only needs to input a video, and the practicability is stronger.

[0009] To achieve the above object, the present disclosure adopts the following technical solutions:

[0010] One or more embodiments provide a dynamic sign language recognition method based on afterimage graphs, including the following steps:

[0011] Obtain the video to be recognized, and perform frame division processing to obtain the pictures to be recognized;

[0012] Adopt the optical flow method to perform dynamic skin segmentation processing on the pictures to be recognized, extract the dynamic regions in the video frames, and extract the skin regions in the dynamic regions of the video frames to obtain the images of the hand regions performing sign language actions;

[0013] Synthesize the images of the hand regions performing sign language actions into afterimage graphs in the order of time of the picture frames and according to the principle of decreasing transparency;

[0014] Use the deep learning method to perform image classification on the afterimage graphs to obtain the text information corresponding to the video to be recognized.

[0015] One or more embodiments provide a dynamic sign language recognition system based on afterimage graphs, including:

[0016] Frame division module: configured to obtain the video to be recognized and perform frame division processing to obtain the pictures to be recognized;

[0017] Image segmentation module: configured to perform dynamic skin segmentation processing on the picture to be recognized by using the optical flow method, extract the dynamic region in the video frame, and extract the skin color region in the dynamic region of the video frame to obtain an image of the hand region for sign language actions;

[0018] Afterimage synthesis module: configured to synthesize the images of the hand regions for sign language actions into an afterimage according to the time sequence of the picture frames and the principle of decreasing transparency;

[0019] Image classification module: configured to perform image classification on the afterimage by using deep learning methods to obtain the text information corresponding to the video to be recognized.

[0020] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps described in the above method are completed.

[0021] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps described in the above method are completed.

[0022] Compared with the prior art, the beneficial effects of the present disclosure are as follows:

[0023] In the present disclosure, based on the optical flow method, targeted adjustments and innovative improvements are made for sign language recognition, and a dynamic sign language recognition method using the afterimage method is proposed. According to the position of each frame of the sign language video in the entire time sequence, the afterimage method synthesizes each frame image of the video from far to near into an afterimage according to the principle of decreasing transparency, thereby converting the three-dimensional video classification problem into a two-dimensional image classification problem, reducing the amount of calculation required for sign language recognition at the dimension level, and ensuring the real-time performance of sign language recognition.

[0024] The advantages of the present disclosure and the advantages of additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The specification drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute a limitation to the present disclosure.

[0026] Figure 1 It is a schematic diagram of the recognition process of the recognition method in Embodiment 1 of the present disclosure;

[0027] Figure 2 It is a flow chart of the recognition method in Embodiment 1 of the present disclosure;

[0028] Figure 3 It is a schematic diagram of the RAFT model structure in Embodiment 1 of the present disclosure;

[0029] Figure 4 It is the RAFT optical flow segmentation effect diagram of Embodiment 1 of the present disclosure;

[0030] Figure 5 They are the video frame images before and after dynamic skin segmentation in Embodiment 1 of the present disclosure;

[0031] Figure 6 It is in Embodiment 1 of the present disclosure that Figure 5 The synthesized afterimage diagram of the frame images;

[0032] Figure 7 It is the schematic diagram of the EfficientNet-B7 network structure in Embodiment 1 of the present disclosure;

[0033] Figure 8 It is the effect comparison diagram of classification using the EfficientNet-B7 network in Embodiment 1 of the present disclosure. Specific implementation manners

[0034] The present disclosure will be further described below in conjunction with the drawings and embodiments.

[0035] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0036] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features in the present disclosure can be combined with each other. The embodiments will be described in detail below in conjunction with the drawings.

[0037] Embodiment 1

[0038] In the technical solutions disclosed in one or more embodiments, as Figures 1 - 8 shown, the dynamic sign language recognition method based on the afterimage diagram includes the following steps:

[0039] Step 1, obtain the video to be recognized, and perform frame splitting processing to obtain the pictures to be recognized;

[0040] Step 2: Use the optical flow method to perform dynamic skin segmentation on the image to be recognized, extract the dynamic regions in the video frame, and extract the skin-color regions in the dynamic regions of the video frame to obtain the image of the hand region for sign language actions;

[0041] Step 3: Combine the images of the hand regions for sign language actions into a residual image in the order of frame time and according to the principle of decreasing transparency;

[0042] Step 4: Use the deep learning method to perform image classification on the residual image to obtain the text information corresponding to the video to be recognized.

[0043] In this embodiment, based on the optical flow method, targeted adjustments and innovative improvements are made for sign language recognition, and a dynamic sign language recognition method based on the residual image method is proposed. The residual image method synthesizes each frame image of the sign language video from far to near into a residual image according to the position of each frame image in the entire time series and the principle of decreasing transparency, thereby converting the three-dimensional video classification problem into a two-dimensional image classification problem, reducing the computational complexity required for sign language recognition at the dimension level, and ensuring the real-time performance of sign language recognition.

[0044] In Step 1, the video to be recognized is processed. First, the entire video is segmented using a sliding window to obtain multiple segments of videos, and any segment of video is frame-processed to obtain a set of pictures. The pictures are grouped, and Steps 2 to 4 are performed on each group of pictures once to obtain the recognition results, and a recognition result is obtained for each segment of video respectively.

[0045] In Step 2, for image segmentation processing, the optical flow method is used to perform preliminary processing on this set of pictures, thereby extracting the dynamic regions in the video frame. Further, edge detection, image segmentation, etc. are performed on the pictures to extract the skin-color regions in the dynamic regions of the video frame, thereby obtaining the "dynamic skin-color region" in the video frame, that is, the hand region for sign language actions.

[0046] The main goal of dynamic skin segmentation is to segment the moving parts in the skin-color regions of the video stream.

[0047] To address this problem, in this embodiment, the existing optical flow algorithm is improved by introducing the information of the YUV color space of the image for optimization to achieve dynamic skin segmentation. The method for performing dynamic skin segmentation on the image to be recognized using the optical flow method includes the following steps:

[0048] Step 21: Process the image to be recognized obtained in Step 1 using the optical flow method to obtain an optical flow image;

[0049] Specifically, in this embodiment, the RAFT optical flow method is used to process the image to be recognized.

[0050] The RAFT model of the RAFT optical flow method is as Figure 3 shown, including an encoder, four-dimensional correlation volumes (4D Correlation Volumes), and an optical flow field updater. The encoder includes a feature encoder and a context encoder. The feature encoder receives two adjacent frames of images as inputs simultaneously, extracts feature vectors, and generates four-dimensional correlation volumes through inner product operations; the context encoder receives the first frame of image and extracts feature vectors. Finally, the optical flow field updater uses the feature vectors of the first frame of image and combines the information of the four-dimensional correlation volumes to iteratively update the optical flow process. The effect of performing RAFT optical flow segmentation on a set of test images is as Figure 4 shown.

[0051] Step 22: Perform threshold segmentation on the optical flow image, convert it into a binary image, multiply the binary image with the original image, and extract the moving regions;

[0052] Among them, the binary image is specifically a 01 binary image.

[0053] Step 23: Further, convert the obtained moving region image to the YUV color space, and use the chrominance information of the image to segment the regions in the moving regions that are on the skin, that is, dynamic skin segmentation is achieved.

[0054] In step 3, it is a method for synthesizing a residual image. The images obtained after dynamic skin segmentation are synthesized into a residual image in groups of 60 in the order from far to near in time according to the principle of decreasing transparency.

[0055] The frame images of the video before and after dynamic skin segmentation are as Figure 5 shown. The images of each frame of the video after dynamic skin segmentation are synthesized into a single residual image according to the principle of decreasing transparency. In this way, a sign language video is converted into a single picture, that is, the 3D video recognition is converted into 2D image recognition. The Figure 5 residual image synthesized from the frame images of Figure 6 is shown.

[0056] Step 4: Use deep learning methods to classify the residual image to obtain the text information corresponding to the video to be recognized.

[0057] In step 3, the residual image method has converted the sign language video into a residual image, so the problem has correspondingly been converted into image classification.

[0058] Optionally, the deep learning method in this embodiment can use transfer learning to classify the obtained residual image. Specifically, the EfficientNet-B7 network can be used to classify the residual image.

[0059] The network structure of the EfficientNet-B7 network is as Figure 7 shown. The input of the network is the residual image, and the output is the text information corresponding to the gesture corresponding to the image classification. As Figure 1 shown, the information expressed by sign language is tea, and the finally output text information is "tea".

[0060] The training process of the EfficientNet-B7 network is as follows:

[0061] Step S1: Obtain segmented sign language video clips for recording as the sign language context training data set;

[0062] Step S2: Process the collected data according to Steps 1 to 3 of the residual image method, and convert each sign language segment into a residual image;

[0063] Step S3: Divide the training set, validation set, and test set according to the ratio of 7:2:1, set learning-rate = 0.001, and batch-size = 2. After the EfficientNet-B7 network model iterates about 40 epochs, the accuracy rates on the training set and the validation set both exceed 91%.

[0064] In this embodiment, the EfficientNet-B7 network is used to perform image classification on the residual image, which can greatly improve the accuracy of gesture recognition. The performance comparison between the EfficientNet itself and other networks is as Figure 8 shown and as shown in Table 1. Considering performance and accuracy comprehensively, EfficientNet-B7 is the optimal model for the image classification process of the residual image method.

[0065] Table 1

[0066] Model Top-1 Acc. Top-5 Acc. #Params EfficientNet-B0 77.1% 93.3% 5.3M EfficientNet-B1 79.1% 94.4% 7.8M EfficientNet-B2 80.1% 94.9% 9.2M EfficientNet-B3 81.6% 95.7% 12M EfficientNet-B4 82.9% 96.4% 19M EfficientNet-B5 83.6% 96.7% 30M EfficientNet-B6 84.0% 96.8% 43M EfficientNet-B7 84.3% 97.0% 66M

[0067] Optionally, in some other embodiments, the deep learning method in Step 4 can also adopt a 3D convolutional network, or an algorithm for gesture recognition based on the combination of a convolutional neural network and a long short-term memory network can also be adopted.

[0068] The gesture recognition network based on the combination of a convolutional neural network and a long short-term memory network includes two cascaded convolutional neural networks and a long short-term memory network.

[0069] 3D convolutional networks are mainly used in fields such as video classification and action recognition. It is modified based on 2D convolutional networks. In 2D convolutional networks, convolution is applied to 2D feature maps, and features are calculated only from the spatial dimension. When using video data to analyze problems, it is expected to capture the motion information encoded in multiple consecutive frames. For this purpose, 3D convolution is proposed to be performed in the convolution of CNN to calculate features in the spatial and temporal dimensions. 3D convolution stacks multiple consecutive frames to form a cube, and then a 3D convolution kernel is applied in the cube. Through this structure, the feature maps in the convolutional layer are all connected to multiple adjacent frames in the previous layer, thus capturing motion information.

[0070] Embodiment 2

[0071] Based on Embodiment 1, this embodiment provides a dynamic sign language recognition system based on afterimage maps, including:

[0072] Frame segmentation module: configured to obtain the video to be recognized and perform frame segmentation to obtain the pictures to be recognized;

[0073] Image segmentation module: configured to perform dynamic skin segmentation processing on the pictures to be recognized using the optical flow method, extract the dynamic regions in the video frames, and extract the skin-colored regions in the dynamic regions of the video frames to obtain the images of the hand regions performing sign language actions;

[0074] Afterimage map synthesis module: configured to synthesize the images of the hand regions performing sign language actions into afterimage maps in the chronological order of the picture frames and according to the principle of decreasing transparency;

[0075] Image classification module: configured to perform image classification on the afterimage maps using deep learning methods to obtain the text information corresponding to the video to be recognized.

[0076] It should be noted here that each module in this embodiment corresponds one by one to each step in Embodiment 1, and its specific implementation process is the same, so it will not be repeated here.

[0077] Embodiment 3

[0078] This embodiment provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps described in the method of Embodiment 1 are completed.

[0079] Embodiment 4

[0080] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps described in the method of Embodiment 1 are completed.

[0081] The foregoing are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

[0082] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, they do not limit the protection scope of the present disclosure. Those skilled in the art should understand that based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.

Claims

1. A dynamic sign language recognition method based on an afterimage map, characterized in that It includes the following steps: Obtain the video to be recognized, and perform frame division processing to obtain the pictures to be recognized; Use the optical flow method to perform dynamic skin segmentation processing on the pictures to be recognized, extract the dynamic regions in the video frames, and extract the skin regions in the dynamic regions of the video frames to obtain the images of the hand regions performing sign language actions; Synthesize the images of the hand regions performing sign language actions into a ghost image in the chronological order of the picture frames and according to the principle of decreasing transparency; Use the deep learning method to perform image classification on the ghost image to obtain the text information corresponding to the video to be recognized.

2. The dynamic sign language recognition method based on the afterimage diagram according to claim 1, characterized in that: Process the video to be recognized. First, use a sliding window to segment the entire video to obtain multiple segments of video, and perform frame division processing on any segment of video to obtain a set of pictures.

3. The dynamic sign language recognition method based on ghost images according to claim 1, wherein: The method of using the optical flow method to perform dynamic skin segmentation processing on the pictures to be recognized includes the following steps: Process the image to be recognized using the optical flow method to obtain an optical flow image; Perform threshold segmentation on the optical flow image, convert it into a binary image, multiply the binary image by the original image, and extract the moving regions; Convert the obtained moving region image to the YUV color space, and use the chrominance information of the image to segment the regions in the moving regions that are skin regions to achieve dynamic skin segmentation.

4. The dynamic sign language recognition method based on the afterimage map according to claim 3, characterized in that: Use the RAFT optical flow method to process the image to be recognized to obtain an optical flow image.

5. The dynamic sign language recognition method based on an afterimage graph according to claim 1, characterized in that: The RAFT model adopted by the RAFT optical flow method includes three parts: an encoder, a four-dimensional correlation body, and an optical flow field updater. The encoder includes a feature encoder and a context encoder; The feature encoder simultaneously receives two adjacent frames of images as inputs, extracts feature vectors, and generates a four-dimensional correlation body through inner product operations; The context encoder receives the first frame of image and extracts feature vectors; The optical flow field updater uses the feature vectors of the first frame of image and combines the information of the four-dimensional correlation body to iteratively update the optical flow process.

6. The dynamic sign language recognition method based on an afterimage diagram according to claim 1, wherein: Use the deep learning method to perform image classification on the ghost image. Specifically, use the method of transfer learning to classify the obtained ghost image; Alternatively, the deep learning method is to use a 3D convolutional network; Alternatively, the deep learning method is to use a gesture recognition algorithm based on the combination of a convolutional neural network and a long short-term memory network.

7. The dynamic sign language recognition method based on an afterimage map according to claim 1, characterized in that: The method of transfer learning specifically uses the EfficientNet-B7 network to perform image classification on the ghost image.

8. The dynamic sign language recognition system based on the afterimage map is characterized in that It includes: Frame division module: configured to obtain the video to be recognized and perform frame division processing to obtain the pictures to be recognized; Image segmentation module: configured to use the optical flow method to perform dynamic skin segmentation processing on the pictures to be recognized, extract the dynamic regions in the video frames, and extract the skin regions in the dynamic regions of the video frames to obtain the images of the hand regions performing sign language actions; Ghost image synthesis module: configured to synthesize the images of the hand regions performing sign language actions into a ghost image in the chronological order of the picture frames and according to the principle of decreasing transparency; Image classification module: configured to use the deep learning method to perform image classification on the ghost image to obtain the text information corresponding to the video to be recognized.

9. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored in the memory and running on the processor. When the computer instructions are run by the processor, the steps described in the method of any one of claims 1-7 are completed.

10. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the steps described in the method of any one of claims 1-7 are completed.

Citation Information

Patent Citations

  • Optical flow-based gesture motion direction recognition method

    CN104331151A

  • Small sample deep learning multi-modal sign language recognition method based on key frame sampling

    CN111666845A