A sign language recognition method

By acquiring color and depth images through a depth camera, combining pseudo-colorization and color gamut segmentation methods, and using a dual-channel feature fusion network for gesture segmentation and feature extraction, the problems of low background segmentation accuracy and model training efficiency in existing sign language recognition methods are solved, and efficient sign language recognition is achieved.

CN117953585BActive Publication Date: 2025-10-14WUXI ELECTRICAL & HIGHER VOCATIONAL SCHOOLS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410184774.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-10-14
Estimated Expiration
2044-02-19

AI Technical Summary

Technical Problem

Existing sign language recognition methods have shortcomings in network feature extraction effect, gesture background segmentation accuracy and network model training efficiency. The use of sensor equipment brings inconvenience to users, the model parameters are huge, the training amount is high, and the actual application effect is limited.

Method used

A depth camera is used to acquire color and depth images, and gesture segmentation is performed through pseudo colorization and color gamut segmentation. A dual-channel feature fusion network is combined to extract features, and the attention module is used to improve feature utilization. A dual-channel feature fusion network is used for gesture recognition.

Benefits of technology

The accuracy of sign language gesture segmentation is improved, the number of network model parameters and training overhead are reduced, the computing efficiency is improved, and the gesture background separation and recognition in complex backgrounds are realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117953585B_ABST
    Figure CN117953585B_ABST
Patent Text Reader

Abstract

The application provides a sign language recognition method, aiming at providing a sign language recognition solution for the hearing impaired population. In view of the image background interference, gesture deformation and similar features between classes, the 16-bit depth information is converted into 8-bit depth information, the depth pixel is pseudo-colorized, the sign language gesture is segmented from the background through the color gamut segmentation method, the color fusion is carried out on the gestures with large span through the pixel fusion method, the affine transformation is carried out on the depth image and the color image, the pixel alignment is realized, finally the double-channel feature fusion attention network (DFANet) is trained, the feature fusion is carried out in the later stage of the network, and the classification is carried out through the softmax classifier. Compared with the prior art, the network feature extraction effect of the application is better, the sign language gesture segmentation precision is higher, the network model parameter amount is smaller, the training cost is low, and the operation efficiency is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a sign language recognition method. Background Art

[0002] Sign language is a visual language and a primary means of communication for the hearing-impaired. Because sign language has unique grammar and vocabulary, specialized sign language recognition systems are needed to convert it into text or speech for better communication with non-hearing people.

[0003] The solution of the existing technology is to achieve gesture background separation through image morphological processing (such as color filtering, threshold segmentation) or gesture modeling to extract image identification features, and send them to the convolutional neural network for feature learning to complete the sign language recognition task. However, this method requires manual setting of image processing thresholds or detailed modeling of each hand shape. If the threshold setting is unreasonable or the modeling is wrong, it will affect the accuracy of sign language recognition. In addition, due to the existence of background interference, movement deformation, gesture occlusion and other problems, it becomes difficult to extract effective features from sign language images. To solve this problem, some researchers use sensors to obtain gesture motion information to assist in completing the sign language recognition task, but this method requires the subjects to wear motion sensor equipment, which causes certain troubles to the sign language recognition work.

[0004] Furthermore, the aforementioned schemes have deficiencies in the way they pre-process sign language data, resulting in a large amount of invalid and duplicate data in the dataset, which not only increases the network training workload but also affects the performance of gesture feature extraction. To improve feature utilization, some methods utilize multiple sensors, such as visible light RGB cameras, depth cameras, millimeter-wave radars, or calculate additional channels, such as optical flow, to improve performance. This results in large model parameters, high network training workload, and limited improvement in practical application scenarios. Gestures are temporally correlated and spatially continuous. This dependence on space and time indicates that the utilization of gesture spatial and temporal features is particularly important, but existing methods have shortcomings in this regard. Making full use of inter-frame temporal correlation information and the hand position, shape, and orientation encoding information in each frame is particularly important for improving gesture feature utilization and model recognition results.

[0005] In summary, how to optimize sign language recognition methods in terms of network feature extraction effect, gesture background segmentation accuracy, and network model training efficiency and cost has become a pain point in this field. Summary of the Invention

[0006] In response to the shortcomings of the above-mentioned prior art, the present application provides a sign language recognition method, which has the advantages of better network feature extraction effect, high sign language gesture segmentation accuracy and low error, small number of network model parameters, low training overhead and high computational efficiency.

[0007] The technical solutions adopted in the present invention are as follows:

[0008] A sign language recognition method comprises the following steps:

[0009] S1: Original sign language image acquisition, using a depth camera to obtain color and depth images of sign language, and normalize the size of the acquired images;

[0010] S2: Convert the normalized depth image into a grayscale image and perform pseudo-colorization based on the distance information;

[0011] S3: Segment the gestures of the image obtained in step S2 according to the color gamut segmentation method, identify the gestures after segmentation and perform color fusion;

[0012] S4: The generated sign language segmentation gestures are grayscaled and binarized, and the processed results are pixel-inverted to concentrate the image pixel information. After the processing is completed, the image is sent to the dual-channel feature fusion network for feature extraction;

[0013] S5: Connect the output features to the fully connected layer and finally output 24 classification probabilities through the softmax classifier. After the network is trained for a set number of rounds, save the trained model and load it for testing. Perform N cross-validations on the public dataset to verify the generalization of the model and save the best trained model.

[0014] S6: After the model training in step S5 is completed, load the model and perform sign language recognition.

[0015] Furthermore, in step S1, an image pyramid is generated based on the original sign language RGB-D image, and a selective search algorithm is used to obtain several regions of interest where targets may exist from the image pyramid. The regions of interest are scaled to a size of 227*227 to complete the acquisition and normalization of the sign language image.

[0016] Furthermore, in step S2, the original sign language gesture depth acquired by the depth camera is 16 bits, and the pixel range is 0 to 65535. The 16-bit depth pixel matrix is ​​normalized by the following formula:

[0017]

[0018] Where H, W, and C represent the image width, height, and number of channels, respectively. is the original image pixel matrix, is the normalized image pixel matrix.

[0019] The formula for determining the region of interest is:

[0020] Dst=Deal(Src)×Scale+Shift

[0021] in is the original image, is the output array after linear transformation, Scale is the scale factor, Shift is the offset, and the pixel inverse transformation matrix Deal is derived from the following linear transformation formula:

[0022]

[0023]

[0024]

[0025] Where R(x,y), G(x,y), and B(x,y) represent the color values ​​of the point in the red, green, and blue channels. f(x,y) and f represent the grayscale value of a specific point in the grayscale image and the grayscale value of the selected grayscale image, respectively. After the image is input, the array is scaled according to the scale factor Scale and the elements are offset. The offset is Shift. After scaling, the image depth information and pixel information change accordingly, thereby changing the color. The scale factor is determined according to the distance of the hand from the camera. The scale factor is determined by the following formula:

[0026] D×Scale=255

[0027] Where D is the distance from the region of interest to the camera.

[0028] Furthermore, in step S3, the image is converted from the RGB image space to the HSV color space, where the HSV color space includes hue (H), saturation (S), and value (V). The value of H is modified to determine the color to be segmented, and the values ​​of S and V are dynamically adjusted to determine the color range to be segmented.

[0029] Furthermore, in step S4, the grayscale of the color image is converted using a weighted method, with the ratio of R, G, and B being 3:6:1. The color of one spot in the image is defined as RGB (red: R, green: G, blue: B), and the calculation formula is:

[0030] Gray=R×0.3+G×0.59+B×0.11

[0031] Among them, R, G, and B are the three primary colors of the image, representing red, green, and blue respectively. Gray is the grayscale value of the image. The coefficient is the value obtained after weighted conversion. Finally, the image is thresholded and reversed. First, the image is binarized by the local threshold binarization method, and then the thresholded result is reversed by the following formula:

[0032] Reverse=255-binary

[0033] Reverse is the flipped image, and binary is a single-channel binary image.

[0034] Furthermore, in step S5, feature information is extracted through a dual-channel feature fusion network, which includes two convolutional batch activation layers, two reverse-shift bottleneck layers, two fused-shift convolutional layers, and four attention modules. A 96×96 image is used as input to extract sign language gesture texture features. The last convolution layer generates a feature map with a spatial resolution of 3×3. The last convolution layer is connected to the global average pooling, the fully connected layer, and the Softmax layer.

[0035] Furthermore, the attention module includes a channel attention module, a spatial attention module, a DPAM and a pixel attention module.

[0036] Furthermore, in order to speed up calculation and prevent gradient diffusion in the training model, binary cross entropy is used as the loss function:

[0037]

[0038] Where: N is the minimum batch size, K is the number of categories, y ij is the predicted probability that the i-th sample belongs to the j-th category, y ij Represents the sample label. A sparse-induced penalty term is added to the loss function. The loss function is:

[0039] L=L b +λΣ r∈Γ |γ|

[0040] Where γ is the scaling factor of batch normalization, Γ is the set of scaling factors in the network, and λ is the weight between the binary cross entropy loss and the sparsity-induced penalty. Finally, stochastic gradient descent with momentum is used as the optimizer.

[0041] Furthermore, in step S3, the value of the pixel of interest is subjected to color fusion by the following determination method:

[0042] Method 1: If (or ) is less than Eff_low, the current sign language segmentation image is designated as (or ).

[0043] Method 2: If (or ) is greater than Eff_high, the current sign language segmentation image is designated as (or ).

[0044] Method 3: If (or ) is greater than Eff_low and less than Eff_high, the image matrix is ​​fused using the following formula:

[0045]

[0046] in, is the image pixel matrix after image fusion.

[0047] Furthermore, the image segmentation method in step S3 is as follows: a mask is generated according to the original image size, a mask operation is performed on the HSV image pixels, the image pixel values ​​within the mask space range are changed to white, and the remaining image pixel values ​​are changed to black, and finally the original image is ANDed with the image processed according to the mask, the black is removed, the white is retained, and the mask position area of ​​the original image is obtained.

[0048] The beneficial effects of the present invention are as follows:

[0049] Compared with the existing technology, the present invention provides an image segmentation and feature extraction algorithm, which can help achieve gesture background separation and sign language recognition in complex backgrounds.

[0050] Compared with the existing technology, the present invention extracts image pixels and depth features respectively through dual channels, and improves the network feature extraction effect through the feature attention module.

[0051] Compared with the existing technology, the present invention realizes gesture background separation through background segmentation algorithm, avoids the influence of environmental factors such as skin color and light, and has high sign language gesture segmentation accuracy and low error.

[0052] Compared with the existing technology, the network model proposed in the present invention has fewer parameters, lower training overhead and higher computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flow chart of the sign language recognition method of the present invention;

[0054] Figure 2 Schematic diagram of the image segmentation process in the present invention;

[0055] Figure 3 This is the DFANet network architecture diagram in the present invention;

[0056] Figure 4 This is a diagram of the architecture of the SAM and CAM in the present invention;

[0057] Figure 5 This is the architecture diagram of the DPAM in the present invention;

[0058] Figure 6 This is a visualization diagram of feature extraction in the present invention;

[0059] Figure 7 Schematic diagram of image segmentation in the present invention. DETAILED DESCRIPTION

[0060] The specific embodiments of the present invention will be described below with reference to the accompanying drawings.

[0061] The present invention provides a sign language recognition method.

[0062] Step 1: Sign language information collection: obtain sign language color and depth images through the depth camera and normalize the size of the collected images.

[0063] Step 2: Convert the normalized depth image into a grayscale image and perform pseudo-colorization based on the distance information to achieve depth image visualization.

[0064] Step 3: Obtain the visualized depth map and segment the gestures based on the color space information. At the same time, judge the segmented gesture results according to the gesture fusion rules and perform color fusion.

[0065] Step 4: The generated sign language segmentation gestures are grayscaled and binarized, and the processed results are pixel-inverted to concentrate the image pixel information. After processing, the image is sent to the dual-channel feature fusion network for feature extraction.

[0066] Step 5: During training, the segmented depth map is normalized and fed into a dual-channel feature fusion network for feature extraction. The output features are then concatenated with a fully connected layer, and finally a softmax function is used to output 24 classification probabilities. After the network is trained for a set number of rounds, the trained model is saved and loaded for testing. N cross-validations are performed on a public dataset (N is the number of dataset targets) to verify model generalization and save the best trained model.

[0067] In some embodiments, in step 1, an image pyramid is generated based on the original sign language RGB-D image. A selective search algorithm is used to obtain several regions of interest (ROIs) from the image pyramid where objects may be present. The ROIs are scaled to a size of 227 x 227 to complete the acquisition and normalization of the sign language image.

[0068] In some embodiments, in step 2, the original sign language gesture depth acquired by the depth camera is 16 bits, and the pixel range is 0 to 65535. To facilitate subsequent processing, the 16-bit depth pixel matrix needs to be normalized using the following formula, where H, W, and C represent the image width, height, and number of channels, respectively:

[0069]

[0070] is the original image pixel matrix, is the normalized image pixel matrix. The region of interest is determined using equation (2).

[0071] Dst = Deal (Src) x Scale + Shift (2)

[0072] wherein is the original image, is the output array after linear transformation, Scale is the scale factor, Shift is the offset, and the pixel inverse transformation matrix Deal is derived from the following linear transformation equation:

[0073]

[0074]

[0075]

[0076] wherein R(x, y), G(x, y), B(x, y) represent the color values of the point in the red, green and blue channels, f(x, y) and f represent the gray value of the specific point of the gray image and the gray value of the selected gray image, respectively. After inputting the image, the array is scaled according to the scale factor Scale and the elements are offset by Shift. After scaling, the image depth information and pixel information change accordingly, so the color changes. The scale factor is determined according to the distance of the hand from the camera. The scale factor is determined by the following equation:

[0077] D x Scale = 255 (6)

[0078] wherein D is the distance of the region of interest from the camera. Setting different regions of interest can highlight different regions of color.

[0079] In some embodiments, step three, for the convenience of modifying the image pixels, the image needs to be converted from the RGB image space to the HSV color space. HSV is composed of three components, namely hue (H), saturation (S) and lightness (V). First, modify the value of H to determine the color to be segmented, and second, dynamically adjust the values of S and V to determine the color range to be segmented.

[0080] Table 1 Color space range table

[0081]

[0082]

[0083] If the color of interest is green, determine the green color space range from (35, 43, 46) to (77, 255, 255) based on color space table 1. Generate a mask based on the original image size. Perform a mask operation on the HSV image pixels, changing the image pixel values ​​within the mask space to 255 (white) and the remaining image pixel values ​​to 0 (black). Finally, perform an AND operation on the original image and the masked image to remove black and retain white. This results in the masked region of the original image, and the image segmentation is complete.

[0084] Some sign language gestures have a large span before and after, so the image after pseudo color linear transformation needs to be stored in two image matrices through color gamut segmentation. Fusion is performed based on the size of the pixel value of interest in the matrix using the following discrimination rules:

[0085] Rule ①: If (or ) is less than Eff_low, the current sign language segmentation image is designated as (or ).

[0086] Rule ②: If (or ) is greater than Eff_high, the current sign language segmentation image is designated as (or ).

[0087] Rule 3: If (or ) is greater than Eff_low and less than Eff_high, the image matrix is ​​fused using formula (3.7). Eff_low and Eff_high are set empirically and used to make fusion judgments on image pixels.

[0088]

[0089] In the above formula (7) is the image pixel matrix after image fusion. Through this formula, the two-color images can be fused together to display a complete sign language gesture diagram.

[0090] Furthermore, in step 4, the grayscale of the color image is converted using a weighted method, with the ratio of R, G, and B being 3:6:1. Assuming the color of a point is RGB (red: R, green: G, blue: B), the following calculation formula is used: Gray = R×0.3+G×0.59+B×0.11

[0091] Among them, R, G, and B are the three primary colors of the image, representing red, green, and blue respectively. Gray is the grayscale value of the image. The coefficient is the value obtained after weighted conversion. In order to segment the gesture target from the image with uneven brightness, it is not possible to use a unified threshold to filter the global target. Consider starting from the local pixel point and gradually calculating the threshold with the current pixel point as the center. In order to highlight the pixel features of the image, it is necessary to perform threshold inversion on the image. First, the image is binarized using the local threshold binarization method, and then the threshold result is inverted using the following formula:

[0092] Reverse=255-binary (8)

[0093] Reverse is the flipped image, and binary is a single-channel binary image. This method not only reduces the number of image channels but also makes the image features more concentrated.

[0094] In some embodiments, in step five, in order to effectively extract the feature information of sign language, a dual-channel feature fusion network is proposed. The network consists of two convolution batch activation layers (Conv BatchNormalization Active layers, CBNAct), two inverted mobile bottleneck layers (Inverted Mobile Bottlenecks, MBConv), two fused mobile convolution layers (Fused Mobile Inverted Bottleneck Conv, FMBConv) and four attention modules. The 96×96 size image is used as input to extract the texture features of sign language gestures. The last convolution layer generates a feature map with a spatial resolution of 3×3. The feature map can effectively preserve spatial information before being input into the pooling layer. The last convolution layer is connected to the global average pooling, the fully connected layer and the softmax layer.

[0095] To refine RGB and depth feature maps simultaneously, three attention modules emphasize representative features while suppressing meaningless features. These modules are the Channel Attention Module (CAM), the Spatial Attention Module (SAM), and the DPAM.

[0096] SAM focuses on selecting salient regions, while CAM exploits inter-channel relationships. SAM extracts the salient regions from an intermediate feature map with height H, width W, and number of channels C. Compute 2D spatial attention map Its attention process is given as follows:

[0097]

[0098] in Represents the vector product, and the 2D spatial attention feature map is calculated by formula (10):

[0099]

[0100] Where σ is the Sigmoid function, f 7×7 represents a convolution with 7×7 filters, AP C and MP C Represents the average pooling and maximum pooling along the channel respectively. CAM calculates the one-dimensional channel attention map The attention process is given by formula (11):

[0101]

[0102] The intermediate feature map M in the above formula C (F) is calculated as follows:

[0103] M C (F)=σ(W1δ(W0(AP S (S)))+W1δ(W0(MP S (F)))) (12)

[0104] Where δ and σ represent the calibrated linear unit and reduction rate, respectively. and is the shared weight, AP S and MP S Represents average pooling and maximum pooling along the spatial axis, respectively. SAM is placed after the ConvBNAct layer to infer the spatial relationship of the intermediate feature map under the high-dimensional feature map. In addition, CAM is placed after the MBConv layer to infer the inter-channel relationship of the intermediate feature map containing two channels. On the other hand, in the deep path, a SAM is placed after ConvBNAct, and the DPAM deep pixel attention mechanism is provided in the subsequent path. First, global average pooling is used to use the global deep spatial information at the channel level as a channel descriptor:

[0105]

[0106] where X c (i, j) represents the cth channel X at position (i, j) c The value of H p is the global pooling function, F c is the input feature map. The shape of the feature map changes from C×H×W to C×1×1. To obtain the weights of different channels, the features are processed through two layers of convolution and Sigmoid and ReLU activation functions.

[0107] CA c=σ(Conv(δ(Conv(g c )))) (14)

[0108] Where σ is the Sigmoid function and δ is the ReLU function. Finally, the input depth and pixel feature map and channel CA c The weights of the intermediate feature map are multiplied element by element to obtain

[0109]

[0110]

[0111] Considering that different hand shapes have similar joint pointing information in depth images, a pixel attention (PA) module is proposed to make the network pay more attention to information-rich features. The output of CA is fed into two convolutional layers with ReLU and Sigmoid activation functions. The shape changes from C×H×W to 1×H×W.

[0112] PA=σ(Conv(δ(Conv(F * )))) (17)

[0113] Finally, using element-wise multiplication, we can input and PA, and They are depth pixel-aware attention gain and RGB pixel-aware attention gain respectively.

[0114]

[0115]

[0116] At the same time, in order to speed up the calculation and prevent problems such as gradient diffusion, binary cross entropy is used as the loss function:

[0117]

[0118] In the above formula, N is the minimum batch size, K is the number of categories, and y ij is the predicted probability that the i-th sample belongs to the j-th category, y ij represents the sample label. In addition, the present invention adds a sparsity-induced penalty term to the loss function to improve generalization ability. This penalty forces the scaling factor in the batch normalization layer to become sparse. Therefore, the loss function with the sparsity-induced penalty term is given by

[0119] L=L b +λΣ r∈Γ|γ| (21)

[0120] Where γ is the scaling factor of batch normalization, Γ is the set of scaling factors in the network, and λ is the weight between the binary cross entropy loss and the sparsity-induced penalty. Finally, stochastic gradient descent with momentum is used as the optimizer.

[0121] During model training, the segmented depth map is normalized and fed into a convolutional neural network for feature extraction. The output features are then connected to a fully connected layer, and finally a softmax function is used to output 24 classification probabilities. After a set number of network training rounds, the trained model is saved and loaded for testing. N cross-validation runs are performed on public datasets (N represents the number of dataset targets) to verify model generalization and save the best trained model.

[0122] After model training is complete, the model is loaded and sign language recognition is performed. First, the camera captures a depth image of the person signing and segments it using the depth image segmentation algorithm proposed in this invention. The segmented depth image is then fed into the network for prediction, and the prediction results are output as text. Finally, the prediction results are recorded and compared with the true labels to verify the model's sign language recognition performance.

[0123] In summary, the sign language recognition method proposed in this paper addresses issues such as image background interference, gesture deformation, and inter-class feature similarity. By normalizing the RGB-D color-depth image acquired by the Microsoft Kinect depth camera, the 16-bit depth information is converted to 8-bit depth information. Depth pixels are pseudo-colored using a pseudo-color linear transformation and a color transformation strategy. A color gamut segmentation method is then used to accurately segment sign language gestures from the background. A pixel fusion method is then used to color-fuse gestures with large spans, completing the segmented sign language gestures. Affine transformations are then performed on the depth and color images to achieve pixel alignment. A proposed dual-channel feature fusion attention network (DFANet) is used to learn pixel and depth features related to the fine-grained gestures. Feature fusion is then performed in the later stages of the network, and classification is performed using a softmax classifier. After training, the network model is saved and used for sign language recognition.

[0124] Comparative experiment: To verify the effectiveness of the present invention, the following comparative experiment was conducted

[0125] The information of the verification data set of the present invention is as follows.

[0126] ASL-A: Data was recorded using a Microsoft Kinect depth camera. Ten people recorded 24 sign language gestures (excluding the non-static gestures J and Z). Volunteers recorded image data from multiple angles while changing their gestures, capturing each gesture 500 times.

[0127] (2) NTU-D: The NTU-D digit dataset was acquired using Kinect. It contains 10 hand gestures for digits from 0 to 9. Ten volunteers performed these gestures, and 10 samples were recorded for each digit gesture. The RGB-D image resolution is 640 × 480 pixels. The hand region in each image was not cropped or annotated, and the background was not removed.

[0128] (3) ASL-AP: Self-built dataset. A total of 24 sign language gestures (excluding non-static letters J and Z) were recorded from 8 volunteers. Each gesture has more than 400 samples. The subjects made gestures facing the Kinect device and then slowly moved their hands in different backgrounds and perspectives to collect data.

[0129] The initial learning rate used in this invention is 0.01, and the weight decay is 10 -4 The learning rate is reduced by half every 10 epochs. The hyperparameter λ is empirically set to 10 -4 , a momentum factor of 0.9, and a training mini-batch size of 32. All experiments were implemented using the PyTorch deep learning platform, and the model was trained using an NVIDIA GeForce GTX 2060 GPU. The proposed DFANet was evaluated on the ASL spelling dataset using leave-one-out cross-validation (LOOCV). Common evaluation metrics, namely accuracy, precision, recall, and F-score, were used to evaluate the proposed DFANet.

[0130] In order to compare the segmentation effects, the PSNR peak signal-to-noise ratio and FSIM feature similarity, commonly used in the field of image segmentation, were used to evaluate the segmentation effects. PSNR is expressed in dB, and FSIM is a unitless similarity index with a value between 0 and 1, where 1 indicates a perfect match and 0 indicates a complete mismatch. SD-Segment is the image segmentation method proposed in this paper. The comparison results are shown in the following table:

[0131] Table 2 Segmentation performance comparison results

[0132] Methods PSNR (dB) FSIM Depth Segment 25.1087 0.8914 SEM 26.0578 0.9037 SD-Segment 26.9689 0.9224

[0133] This paper verifies the impact of segmented images on sign language recognition accuracy. The results are shown in Table 3, where R_R, S_R, and S_RD represent the original image, segmented image, and segmented color and depth images, respectively. FLOPs represents the number of network floating-point operations, Param represents the number of network calculation parameters, and Avg represents the average recognition accuracy.

[0134] Table 3 Comparison results of segmentation image experiments

[0135] Methods FLOPs (B) Param (M) Avg (%) R_R 9.6 24.04 90.07 S_R 9.6 24.04 95.89 S_RD 9.6 48.08 98.72

[0136] The present invention conducts ablation experiments on the combination of different attention modules, and the results are shown in Table 4, where Avg represents the average recognition accuracy.

[0137] Table 4 Results of different attention module combinations

[0138]

[0139]

[0140] The present invention compares the DPAM arrangements of various convolutional blocks, and Table 5 lists the results of these ablation experiments.

[0141] Table 5 Results of placing DPAM after each module in the network

[0142] DPAM Placement FLOPs (B) Param (M) Avg (%) DNet 10.8 48.00 95.78 ConvBNAct 10.8 48.08 96.89 MBConv 10.8 48.08 97.34 ConvBNAct + MBConv 10.8 48.08 98.72

[0143] We conducted comparative experiments on both publicly available ASL datasets and self-constructed datasets, comparing the proposed DFANet method with existing methods. Table 6 lists the average classification accuracy of five subjects obtained using leave-one-out cross-validation (P, R, and F represent precision, recall, and F-score, respectively; AvgA and AvgN values ​​represent the accuracy on the ASL alphabetic and NTU digit datasets, respectively).

[0144] Table 6 Comparison with existing methods on public datasets

[0145] Methods P(%) R(%) F(%) AvgA (%) AvgN (%) PCANet

[72] 90.90 89.65 89.40 88.70 93.40 DES

[82] 93.50 93.05 92.71 92.70 98.90 RBM

[70] 95.33 94.67 94.37 94.04 99.15 D-Fuse

[73] 95.10 94.48 94.26 93.53 96.10 DFANet 98.16 97.96 97.65 98.72 99.35

[0146] The present invention uses the proposed algorithm to preprocess the data set and conducts comparative experiments. The experimental results are as follows:

[0147] Table 7 Segmentation effect comparison experiment

[0148]

[0149]

[0150] The above description is an explanation of the present invention, not a limitation of the present invention. The scope of the present invention is defined in the claims. Any modifications may be made within the scope of protection of the present invention.

Claims

1. A sign language recognition method, characterized in that: The steps include: S1: Original sign language image acquisition, using a depth camera to obtain color and depth images of sign language, and normalize the size of the acquired images; S2: Convert the normalized depth image into a grayscale image and perform pseudo-colorization based on the distance information; S3: Segment the gestures of the image obtained in step S2 according to the color gamut segmentation method, identify the gestures after segmentation and perform color fusion; S4: The generated sign language segmentation gestures are grayscaled and binarized, and the processed results are pixel-inverted to concentrate the image pixel information. After the processing is completed, the image is sent to the dual-channel feature fusion network for feature extraction; S5: Connect the output features to the fully connected layer and finally output 24 classification probabilities through the softmax classifier. After the network is trained for a set number of rounds, save the trained model and load it for testing. Perform N cross-validations on the public dataset to verify the generalization of the model and save the best trained model. S6: After the model training in step S5 is completed, load the model and perform sign language recognition; In step S5, feature information is extracted through a dual-channel feature fusion network, which includes two convolutional batch activation layers, two reverse-shift bottleneck layers, two fused mobile convolutional layers, and four attention modules. The image of size is used as input to extract the texture features of sign language gestures, and the final convolution layer generates a spatial resolution of The feature map of the last convolution layer is connected with the global average pooling, the fully connected layer and the Softmax layer; The attention module includes a channel attention module, a spatial attention module, a DPAM and a pixel attention module; The pixel attention module makes the network pay more attention to the information-rich features. It is the same as CA and converts the input 、 Feed with and Two convolutional layers with activation functions; the shape is given by becomes , the PA formula is: ; Finally, using element-wise multiplication, we can input 、 and ,in, and They are depth pixel-aware attention gain and RGB pixel-aware attention gain respectively; ; 。 2. A sign language recognition method according to claim 1, characterized in that: In step S1, an image pyramid is generated based on the original sign language RGB-D image. A selective search algorithm is used to obtain several regions of interest (ROIs) where objects may be present from the image pyramid. The ROIs are scaled to a size of 227*227 to complete the acquisition and normalization of the sign language image.

3. A sign language recognition method according to claim 2, characterized in that: In step S2, the original sign language gesture depth acquired by the depth camera is 16 bits, and the pixel range is 0 to 65535. The 16-bit depth pixel matrix is ​​normalized using the following formula: ; Where H, W, and C represent the image width, height, and number of channels, respectively. is the original image pixel matrix, is the normalized image pixel matrix; The formula for determining the region of interest is: ; in is the original image, is the output array after linear transformation, is the scale factor, is the offset, where the pixel inverse transformation matrix It is derived from the following linear transformation formula: ; ; ; in 、 、 Indicates the color value of the point in the red, green and blue channels. and Respectively represent the grayscale value of a specific point in the grayscale image and the grayscale value of the selected grayscale image. After inputting the image, the scale factor Scales the array and offsets the elements by After scaling, the image depth information and pixel information change accordingly, resulting in a color change. The scaling factor is determined based on the distance between the hand and the camera. The scaling factor is determined by the following formula: ; in is the distance from the region of interest to the camera.

4. The sign language recognition method according to claim 1, wherein: In step S3, the image is converted from the RGB image space to the HSV color space. The HSV color space includes hue (H), saturation (S), and value (V). The value of H is modified to determine the color to be segmented, and the values ​​of S and V are dynamically adjusted to determine the color range to be segmented.

5. The sign language recognition method according to claim 1, wherein: In step S4, the grayscale of the color image is converted using a weighted method, with the ratio of R, G, and B being 3:6:

1. The color of one spot in the image is defined as RGB (red: R, green: G, blue: B), and the calculation formula is: ; Among them, R, G, and B are the three primary colors of the image, representing red, green, and blue respectively. Gray is the grayscale value of the image. The coefficient is the value obtained after weighted conversion. Finally, the image is thresholded and reversed. First, the image is binarized by the local threshold binarization method, and then the thresholded result is reversed by the following formula: ; Reverse is the flipped image, and binary is a single-channel binary image.

6. The sign language recognition method according to claim 1, wherein: In order to speed up calculation and prevent gradient diffusion in the training model, binary cross entropy is used as the loss function: ; in: is the minimum batch size, is the number of categories, It is The samples belong to The predicted probability of each category, Represents the sample label. A sparsity-induced penalty term is added to the loss function. The loss function is: ; in is the scaling factor of batch normalization, is the set of scaling factors in the network, is the weight between controlling the binary cross entropy loss and the sparsity-inducing penalty.

7. A sign language recognition method according to claim 3, characterized in that: In step S3, the value of the pixel of interest is determined by the following method to perform color fusion: Method 1: If (or ) has a pixel value of interest less than , designate the current sign language segmentation image as (or ); Method 2: If (or ) has a pixel value greater than , designate the current sign language segmentation image as (or ); Method 3: If (or ) has a pixel value greater than and less than , the image matrix is ​​fused using the following formula: ; in, is the image pixel matrix after image fusion.

8. The sign language recognition method according to claim 1, wherein: The image segmentation method in step S3 is as follows: a mask is generated according to the original image size, a mask operation is performed on the HSV image pixels, the image pixel values ​​within the mask space range are changed to white, and the remaining image pixel values ​​are changed to black, and finally the original image is ANDed with the image processed according to the mask, the black is removed and the white is retained, and the mask position area of ​​the original image is obtained.

Citation Information

Patent Citations

  • Wearable electrocardiogram real-time diagnosis system based on deep neural network

    CN113749668A

  • Sign language alphabet spelling recognition method based on convolutional neural network

    CN115359562A

  • Target detection method and device based on sparse federal training, and electronic equipment

    CN117315388A