End-to-end rapid gesture segmentation method, device and equipment

By introducing a deformed spatial attention mechanism and a lightweight MobileNet structure into the gesture segmentation network, the problem of poor segmentation effect and large calculations of traditional gesture segmentation methods in complex backgrounds is solved, and high-precision and high-efficiency gesture segmentation are achieved.

CN120088851APending Publication Date: 2025-06-03WUHAN TIANYU INFORMATION IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510055793.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The traditional gesture segmentation method is not effective in segmentation under complex backgrounds and diverse gesture patterns, it is difficult to capture the detailed characteristics of gestures, and it is large in calculations, making it difficult to meet the needs of real-time applications.

Method used

The end-to-end gesture segmentation network model based on the deformable spatial attention mechanism is adopted, and gesture features of different scales and morphology are captured through the deformable convolutional layer and attention module, and the MobileNet structure is used in the backbone network to reduce the computational amount.

Benefits of technology

It significantly improves the accuracy and real-timeness of gesture segmentation, and can accurately segment gesture targets in complex contexts, enhancing the adaptability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088851A_ABST
    Figure CN120088851A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of gesture recognition, and discloses an end-to-end quick gesture segmentation method, device and equipment. The method comprises the steps of collecting gesture image data and performing labeling processing to obtain a training set and a test set; constructing an end-to-end gesture segmentation network model based on a deformation space attention mechanism; training and testing an end-to-end gesture segmentation network model based on the deformation space attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model; and inputting a to-be-predicted picture into the trained gesture segmentation network model, and outputting a binary gesture segmentation mask graph. The network model based on the deformation space attention mechanism is provided, the attention mechanism can capture gesture features of different scales and different forms, the modeling ability for complex gesture shapes and details is improved, the gesture segmentation precision is remarkably improved, the space attention distribution strategy of deformable convolution is utilized, and the gesture segmentation accuracy is improved. And the adaptability and robustness in a complex scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gesture recognition, and particularly to an end-to-end fast gesture segmentation method, device and equipment. Background Art

[0002] With the rapid development of human-computer interaction technology in the field of artificial intelligence, gesture recognition has gradually become one of the core technologies in augmented reality, virtual reality and intelligent human-computer interaction systems. In a vision-based gesture recognition system, gesture segmentation, as an important link in data preprocessing, its accuracy and real-time performance directly affect the overall recognition effect.

[0003] However, traditional gesture segmentation methods mainly rely on multiple preprocessing steps and complex feature engineering, such as color space transformation, background modeling, motion detection, skin color detection and image threshold segmentation, etc. These methods have poor adaptability to complex scenes and are easily affected by background noise, illumination changes and occlusion problems, resulting in unstable segmentation effects. Currently, deep learning methods based on convolutional neural networks have made significant progress in gesture segmentation. Using deep neural networks to automatically extract features can effectively improve the accuracy of segmentation, but it is still difficult to capture the detailed features of gestures in the face of complex backgrounds. At the same time, the computational complexity of the network model is large and it is difficult to be popularized in real-time applications. Summary of the Invention

[0004] The main object of the present invention is to provide an end-to-end fast gesture segmentation method, device and equipment, aiming to solve at least one of the above technical problems.

[0005] To achieve the above object, the present invention provides an end-to-end fast gesture segmentation method, including:

[0006] Collecting gesture image data and performing annotation processing to obtain a training set and a test set;

[0007] Constructing an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism;

[0008] Training and testing the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model;

[0009] Inputting a picture to be predicted into the trained gesture segmentation network model and outputting a binary gesture segmentation mask picture.

[0010] In some embodiments, the collecting gesture image data and performing annotation processing to obtain a training set and a test set includes:

[0011] Collecting gesture image data and extracting sample pictures according to the gesture image data;

[0012] Use a data annotation tool to perform polygon annotation on the hand region in the sample image, and fill the region enclosed by the polygon to generate a mask label;

[0013] Construct a data sample for model training based on the sample image and the annotated mask label;

[0014] Divide the data sample into a training set and a test set based on a preset ratio.

[0015] In some embodiments, constructing an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism includes:

[0016] Obtain a fully convolutional network model;

[0017] Modify the backbone network of the fully convolutional network model to a MobileNet structure;

[0018] Add a deformable spatial attention module after the backbone network to obtain an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism.

[0019] In some embodiments, the end-to-end gesture segmentation network model based on a deformable spatial attention mechanism includes: a MobileNet structure, a deformable spatial attention module, a transposed convolution layer, and a bilinear sampling layer.

[0020] In some embodiments, the deformable spatial attention module includes: a deformable convolutional layer, an output feature layer, a 1*1 convolutional layer, an activation layer, and an element-wise multiplication layer.

[0021] In some embodiments, training and testing the end-to-end gesture segmentation network model based on a deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model includes:

[0022] Train the end-to-end gesture segmentation network model based on a deformable spatial attention mechanism according to the training set;

[0023] Use the stochastic gradient descent method to optimize the parameter update of the end-to-end gesture segmentation network model based on a deformable spatial attention mechanism, and use cross-entropy loss as the loss function during model training;

[0024] After each round of training of the end-to-end gesture segmentation network model based on a deformable spatial attention mechanism, evaluate the accuracy of the network model on the test set according to the test set;

[0025] Select the network weights of the round of training with the highest accuracy as the model parameters to obtain a trained gesture segmentation network model.

[0026] In some embodiments, the accuracy is the mean intersection over union (MIoU) metric.

[0027] In some embodiments, inputting the picture to be predicted into the trained gesture segmentation network model and outputting a binary gesture segmentation mask map includes:

[0028] Obtaining an original image or video frame, and extracting the picture to be predicted according to the original image or video frame;

[0029] Performing normalization processing on the picture to be predicted to convert the picture to be predicted into an image input format required by the model;

[0030] Inputting the normalized picture to be predicted into the trained gesture segmentation network model to obtain a gesture segmentation result;

[0031] Performing bilinear interpolation on the gesture segmentation result to obtain a binary gesture segmentation mask map.

[0032] In addition, to achieve the above object, the present invention also provides an end-to-end fast gesture segmentation device, including:

[0033] An image annotation module, configured to collect gesture image data and perform annotation processing to obtain a training set and a test set;

[0034] A model construction module, configured to construct an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism;

[0035] A model training module, configured to train and test the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model;

[0036] A model prediction module, configured to input the picture to be predicted into the trained gesture segmentation network model and output a binary gesture segmentation mask map.

[0037] In addition, to achieve the above object, the present invention also provides an electronic device, the electronic device includes: a memory, a processor, and an end-to-end fast gesture segmentation program stored on the memory and executable on the processor, the end-to-end fast gesture segmentation program is configured to implement the end-to-end fast gesture segmentation method as described above.

[0038] The present invention provides an end-to-end fast gesture segmentation method, including: collecting gesture image data and performing annotation processing to obtain a training set and a test set; constructing an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism; training and testing the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model; inputting a picture to be predicted into the trained gesture segmentation network model and outputting a binary gesture segmentation mask map. In the present invention, a network model based on a deformable spatial attention mechanism is proposed and applied to the end-to-end gesture segmentation task. The attention mechanism can capture gesture features of different scales and forms, improve the modeling ability of complex gesture shapes and details, thereby significantly improving the accuracy of gesture segmentation. Moreover, by using the spatial attention allocation strategy of deformable convolution, the model can directly input the original image for fast segmentation without relying on complex preprocessing under complex background conditions, and finally can accurately segment the gesture target, improving the adaptability and robustness in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic structural diagram of an electronic device in the hardware operating environment related to the solution of the embodiment of the present invention;

[0040] Figure 2 It is a schematic flowchart of an embodiment of the end-to-end fast gesture segmentation method of the present invention;

[0041] Figure 3 It is a flowchart of the end-to-end fast gesture segmentation method based on the deformable spatial attention mechanism related to the solution of the embodiment of the present invention;

[0042] Figure 4 It is a schematic network structure diagram of the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism related to the solution of the embodiment of the present invention;

[0043] Figure 5 It is a schematic network structure diagram of the deformable spatial attention module related to the solution of the embodiment of the present invention;

[0044] Figure 6 It is a schematic flowchart of the end-to-end fast gesture segmentation example related to the solution of the embodiment of the present invention;

[0045] Figure 7 It is a schematic block diagram of an embodiment of the end-to-end fast gesture segmentation device of the present invention.

[0046] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0048] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0049] In addition, the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0050] Refer to Figure 1 , Figure 1 It is a schematic structural diagram of an electronic device for the hardware operating environment involved in the solution of the embodiment of the present invention.

[0051] Such as Figure 1As shown in the figure, the electronic device may include: a processor 1001, such as a Central Processing Unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard. Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed Random Access Memory (RAM memory) or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0052] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the electronic device, and it may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0053] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating system, a network communication module, a user interface module, and an end-to-end fast gesture segmentation program.

[0054] In Figure 1 the electronic device shown, the network interface 1004 is mainly used for data communication with a network server; the user interface 1003 is mainly used for data interaction with a user; the processor 1001 and the memory 1005 in the electronic device of the present invention may be arranged in the electronic device. The electronic device calls the end-to-end fast gesture segmentation program stored in the memory 1005 through the processor 1001 and executes the end-to-end fast gesture segmentation method provided by the embodiments of the present invention.

[0055] The present invention provides an end-to-end fast gesture segmentation method, device, and device, which solve the following problems existing in traditional gesture segmentation methods:

[0056] Improving segmentation accuracy: Traditional gesture segmentation methods have poor segmentation effects in complex backgrounds and diverse gesture forms, easily resulting in blurred boundaries or missed detection of target areas. The present invention aims to improve the focusing ability of the model on the gesture area and enhance the capture of detailed features by introducing the spatial attention mechanism of deformable convolution, thereby improving the segmentation accuracy.

[0057] Improve the segmentation speed: Due to the high model complexity and large computational volume of traditional deep convolutional neural networks, it is difficult to meet the real-time requirements. The present invention aims to design a lightweight attention mechanism module and optimize the number of model parameters and computational volume, so that the model can achieve fast response on resource-constrained devices.

[0058] Enhance the model robustness: Utilize the spatial attention allocation strategy of deformable convolution, so that the model can directly input the original image for fast segmentation without relying on complex preprocessing under complex background conditions, and finally can accurately segment the gesture target, achieving the goal of improving the applicability and stability of the model.

[0059] An embodiment of the present invention provides an end-to-end fast gesture segmentation method, referring to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the end-to-end fast gesture segmentation method of the present invention.

[0060] As Figure 2 shown, the end-to-end fast gesture segmentation method includes:

[0061] Step S100: Collect gesture image data and perform annotation processing to obtain a training set and a test set;

[0062] Step S200: Construct an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism;

[0063] Step S300: Train and test the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model;

[0064] Step S400: Input the picture to be predicted into the trained gesture segmentation network model and output a binary gesture segmentation mask map.

[0065] It should be noted that the execution subject in this embodiment can be an electronic device, which can be a computer device with data processing functions, or other devices that can implement the same or similar functions. This embodiment does not make any restrictions. In this embodiment, a computer device is used as an example for illustration.

[0066] It can be understood that this embodiment takes the spatial attention mechanism module based on deformable convolution as an example for illustration and applies it to an end-to-end fast gesture segmentation scheme. The spatial attention mechanism module based on deformable convolution can significantly improve the focusing ability on gesture targets in the end-to-end network structure, realize efficient and accurate gesture segmentation, improve the robustness of the gesture segmentation method in complex scenarios and improve the real-time performance.

[0067] In one embodiment, gesture image data is collected and annotated to obtain a training set and a test set, including: collecting gesture image data, and extracting sample pictures according to the gesture image data; performing polygon annotation on the hand regions in the sample pictures through a data annotation tool, and filling the regions enclosed by the polygons to generate mask labels; constructing data samples for model training according to the sample pictures and the annotated mask labels; dividing the data samples into a training set and a test set based on a preset ratio.

[0068] Specifically, referring to Figure 3 , data processing: constructing data samples required for network model training. Gesture image data is collected to obtain sample pictures in various scenarios for training the model. A data annotation tool (such as labelme) is used to perform polygon annotation on the hand regions in the sample pictures; filling is performed according to the regions enclosed by the polygons to complete the generation of mask labels, and finally samples and labels required for model training are constructed. The annotated data samples are divided into a training set and a test set according to a preset ratio (such as 9:1). Among them, the training set can be used for optimizing the training of the network model, and the test set can be used for testing the accuracy of the model.

[0069] It should be noted that the constructed training samples (sample pictures) and their mask labels can also be subjected to data augmentation processing. The augmentation processing methods include but are not limited to rotation, inversion, color transformation, and normalization, etc. The samples processed in this way are used for model training, and the model obtained after training will be applicable to various scenarios, greatly increasing the robustness of the model.

[0070] In one embodiment, an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism is constructed, including: obtaining a fully convolutional network model; modifying the backbone network of the fully convolutional network model into a MobileNet structure; adding a deformable spatial attention module after the backbone network to obtain an end-to-end gesture segmentation network model based on the deformable spatial attention mechanism.

[0071] In one embodiment, the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism includes: a MobileNet structure, a deformable spatial attention module, a transposed convolution layer, and a bilinear sampling layer.

[0072] It can be understood that the model designed in this embodiment needs to predict and output a binary segmentation map of the gesture area based on the input original image or video frame. The segmentation network used in this embodiment can be the open-source Fully Convolutional Networks (FCN). FCN is an algorithm framework for image semantic segmentation. Its input is a three-channel picture tensor after normalization, and the output is a binary mask with the same size as the input. The value at each position represents whether the corresponding pixel at that position belongs to the gesture area. FCN has the characteristics of being adaptable to any input size, low computational overhead, and relatively high segmentation accuracy. At the same time, the deconvolution layer is used in FCN to increase the image size and can output a more refined segmentation result.

[0073] In this embodiment, in order to better segment the gesture area, the network structure of FCN is improved. The MobileNet structure is used as the backbone network to improve the real-time performance of model prediction. At the same time, a novel deformable spatial attention module is added after its backbone network to enhance the model's segmentation ability for the gesture area in the picture. Among them, the MobileNet structure includes but is not limited to MobileNetv3, etc. It should be noted that other lightweight networks can also be used as the backbone network, and this embodiment does not limit this.

[0074] Specifically, referring to Figure 3 , for the network structure construction: modify the FCN network structure. The backbone network in the FCN network structure can be changed to use MobileNetv3, and the novel deformable spatial attention module proposed in this embodiment is added after the backbone network. The improved FCN network (i.e., the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism) structure is as Figure 4 shown.

[0075] In one embodiment, the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism is trained and tested according to the training set and the test set to obtain the trained gesture segmentation network model, including: training the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set; using the stochastic gradient descent method to optimize the parameter update of the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism, and using the cross-entropy loss as the loss function during model training; after each round of training of the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism, evaluate the accuracy of the network model on the test set according to the test set; select the network weights of the round of training with the highest accuracy as the model parameters to obtain the trained gesture segmentation network model.

[0076] In one embodiment, the accuracy is the mean intersection over union (MIoU) metric.

[0077] Specifically, referring toFigure 3 , Model Training: The data used for training the model includes images collected in various scenarios, such as 5000 images, and the test data set (test set) includes 2000 images. In this embodiment, the stochastic gradient descent method is used to optimize the parameter update of the network model. The loss function used during training can be the cross-entropy loss. The model is trained for 500 epochs, and the accuracy of the network model on the test set is evaluated in each epoch. Finally, the network weights of the epoch with the highest accuracy are selected as the parameters of the network model trained this time for application in the subsequent stage.

[0078] In one embodiment, inputting the picture to be predicted into the trained gesture segmentation network model to output a binary gesture segmentation mask map, including: obtaining the original image or video frame, and extracting the picture to be predicted according to the original image or video frame; performing normalization processing on the picture to be predicted to convert the picture to be predicted into the image input format required by the model; inputting the normalized picture to be predicted into the trained gesture segmentation network model to obtain a gesture segmentation result; performing bilinear interpolation on the gesture segmentation result to obtain a binary gesture segmentation mask map.

[0079] Specifically, referring to Figure 3 , Gesture Segmentation Prediction: The improved FCN network in this embodiment loads the model weights trained in the previous step to obtain a trained gesture segmentation network model. Inputting the original image or video frame containing the gesture into the trained gesture segmentation network model to output a gesture segmentation result.

[0080] It should be noted that the prediction result output by the trained gesture segmentation network model is one-fourth the size of the original input image. The prediction result is further bilinearly interpolated to improve its resolution and finally generate a binary gesture segmentation mask map with the same size as the input image. This binary gesture segmentation mask map can be further used for gesture recognition and gesture interaction tasks in action capture scenarios.

[0081] In one embodiment, the deformable spatial attention module includes: a deformable convolutional layer, an output feature layer, a 1*1 convolutional layer, an activation layer, and an element-wise multiplication layer.

[0082] Specifically, the activation layer can be a softmax activation layer. Referring to Figure 5 the schematic diagram of the network structure of the deformable spatial attention module shown, in the actual process, the input feature is input into the deformable convolutional layer of the deformable spatial attention module. The output feature of the deformable convolutional layer is output to the 1*1 convolutional layer through the output feature layer, and then passes through the softmax activation layer to output a single-channel attention map. The output feature of the deformable convolutional layer and the single-channel attention map pass through the element-wise multiplication layer, and finally output the deformable spatial attention feature.

[0083] It should be noted that the following combines Figure 6 and specific examples to elaborate in detail on the end-to-end fast gesture segmentation method based on the deformable spatial attention mechanism in this embodiment.

[0084] Specifically, Step 1: Obtain the pictures taken in each scenario containing the gesture area; Step 2: Use the annotation tool to annotate the obtained sample pictures to obtain the mask labels corresponding to the images; Step 3: Modify the FCN network structure, with its backbone network using MobileNetv3, and add the proposed new deformable spatial attention module after the backbone network (the improved FCN network structure is as Figure 4 shown); Step 4: Use the dataset composed of the pictures and labels constructed in Step 2 to train the gesture segmentation network model constructed in Step 3, and select the optimal model parameters according to the evaluation accuracy of the model on the test set; Step 5: Use the optimal network model parameters trained in Step 4 to deploy and apply the gesture segmentation network model to obtain the trained gesture segmentation network model, which can accept gesture image inputs of any size and output the binary segmentation mask map of the gesture area.

[0085] Exemplarily, as Figure 6 shown, Step 1: Collect image data: Organize the pictures taken in each scenario containing the gesture area that have been pre-collected.

[0086] Step 2: Annotate the image data: Use the annotation tool to annotate the obtained picture data to obtain the mask labels corresponding to the images. Then divide all the annotated data into a training set and a test set in a ratio of 9:1. The training set is used for optimizing the training of the network model, and the test set is used for testing the accuracy of the model. It should be noted that the mask label of the image data in Step 2 is the area containing the gesture boundary information in the image.

[0087] Step 3: Write code based on the pytorch framework of python to implement the training code of the improved FCN network model. Among them, the writing of the training code includes the following parts: 1) FCN network structure; 2) Random gradient descent optimizer configuration; 3) Cross-entropy loss function; 4) Data loader for loading training data.

[0088] Step 4: Construct the image enhancement preprocessing operations for the improved FCN network model before training. These image enhancement operations include translation, rotation, inversion, scaling, color transformation, and normalization. Write the code for implementing the image enhancement function in the data loader code of the training data written in the previous step (Step 3).

[0089] Step 5: Input all the training data obtained in Step 2 into the FCN network model for model training. In an example, an optimal model (the trained gesture segmentation network model) with an accuracy of 97.65% can be finally obtained.

[0090] It can be understood that the accuracy rate in Step 5 is the MIoU metric, that is, the Mean Intersection over Union, and its calculation formula is as follows:

[0091] IoU = Intersection area of a certain category / Union area of a certain category

[0092] MIoU = Sum of IoUs of all categories / Total number of categories.

[0093] Step 6: Obtain the picture to be predicted and perform corresponding processing on this picture. Among them, the picture to be predicted in Step 6 refers to the input picture that needs gesture segmentation, and the processing performed on the input picture is the normalization processing performed on the input pictures of the test set during the model training process, and the purpose is to convert the picture to be predicted into the image input format required by the model.

[0094] Step 7: Use the model obtained in Step 5 to predict the input picture and output the prediction result. Among them, the prediction result output by the model in Step 7 refers to a binary mask map that divides the gesture and the background outside the gesture area into 2 categories.

[0095] In this embodiment, an improved FCN network structure of an end-to-end fast gesture segmentation method based on a deformable spatial attention mechanism is proposed. To solve the problems of low robustness and poor real-time performance of traditional gesture segmentation methods in complex scenarios, the network architecture is improved based on the open-source FCN network. The backbone network is changed to the lightweight MobileNetv3 network, and at the same time, a novel deformable spatial attention module is added after its backbone network to enhance the model's segmentation ability for the gesture area in the picture.

[0096] It should be noted that in this embodiment, a trained gesture segmentation network model is obtained through processing and applied to the end-to-end gesture segmentation task. The innovations of this module mainly include the following aspects: Introduction of deformable convolutional kernels: Convolutional kernels with adaptively variable sizes are used for feature extraction, enabling the attention mechanism to capture gesture features of different scales and forms, and improving the modeling ability for complex gesture shapes and details. Lightweight network structure: While maintaining high-precision segmentation results, the number of parameters and computational complexity of the attention module are optimized, enabling this technology to run in real time on mobile devices and embedded systems. End-to-end training framework: By designing an adaptive attention learning strategy, the model can optimize both the segmentation accuracy and attention allocation parameters during end-to-end training, thereby improving the overall segmentation performance. This embodiment not only achieves significant improvements in the accuracy and real-time performance of gesture segmentation, but also demonstrates strong advantages in adaptability and robustness in complex scenarios, and has broad application prospects and market value.

[0097] This embodiment provides an end-to-end fast gesture segmentation method, including: collecting gesture image data and performing annotation processing to obtain a training set and a test set; constructing an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism; training and testing the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model; inputting a picture to be predicted into the trained gesture segmentation network model and outputting a binary gesture segmentation mask map. In this embodiment, a network model based on a deformable spatial attention mechanism is proposed and applied to the end-to-end gesture segmentation task. The attention mechanism can capture gesture features of different scales and forms, improve the modeling ability for complex gesture shapes and details, thereby significantly improving the accuracy and real-time performance of gesture segmentation. Moreover, using the spatial attention allocation strategy of deformable convolution enables the model to directly input the original image for fast segmentation without relying on complex preprocessing under complex background conditions, and finally can accurately segment the gesture target, improving the adaptability and robustness in complex scenarios.

[0098] In addition, an embodiment of the present invention also proposes a storage medium, on which an end-to-end fast gesture segmentation program is stored. When the end-to-end fast gesture segmentation program is executed by a processor, the steps of the end-to-end fast gesture segmentation method as described above are implemented.

[0099] Refer to Figure 7 , Figure 7 which is a structural block diagram of an embodiment of the end-to-end fast gesture segmentation device of the present invention.

[0100] As Figure 7 shown, the end-to-end fast gesture segmentation device includes:

[0101] An image annotation module 10, configured to collect gesture image data and perform annotation processing to obtain a training set and a test set;

[0102] A model construction module 20, configured to construct an end-to-end gesture segmentation network model based on a deformable spatial attention mechanism;

[0103] A model training module 30, configured to train and test the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model;

[0104] A model prediction module 40, configured to input a picture to be predicted into the trained gesture segmentation network model and output a binary gesture segmentation mask map.

[0105] Specifically, the end-to-end fast gesture segmentation device is used for fine segmentation and target enhancement of a gesture area. The device includes: an image annotation module 10, a model construction module 20, a model training module 30, and a model prediction module 40. Among them, the image annotation module 10 is used for data processing, performing gesture mask annotation on the input image data to obtain a mask label corresponding to a sample; the model construction module 20 is used for network structure construction, improving the FCN model to construct an end-to-end lightweight gesture segmentation network model with a deformable spatial attention mechanism; the model training module 30 is used for model training, inputting the processed training data and labels into the improved network model for training; the model prediction module 40 is used for gesture segmentation prediction, extracting a binary gesture segmentation mask map according to the output result of the model prediction.

[0106] In this embodiment, the device processes to obtain a trained gesture segmentation network model and applies it to an end-to-end gesture segmentation task. The innovations of this module mainly include the following aspects: Introduction of a deformable convolution kernel: A convolution kernel with an adaptively variable size is used for feature extraction, enabling the attention mechanism to capture gesture features of different scales and forms, and improving the modeling ability for complex gesture shapes and details. Lightweight network structure: While maintaining a high-precision segmentation effect, the number of parameters and the amount of computation of the attention module are optimized, enabling this technology to run in real time on mobile devices and embedded systems. End-to-end training framework: By designing an adaptive attention learning strategy, the model can simultaneously optimize the segmentation accuracy and attention allocation parameters in end-to-end training, thereby improving the overall segmentation performance. It not only achieves a significant improvement in the accuracy and real-time performance of gesture segmentation, but also shows strong advantages in adaptability and robustness in complex scenarios, and has broad application prospects and market value.

[0107] This embodiment provides an end-to-end fast gesture segmentation device, including an image annotation module 10, a model construction module 20, a model training module 30, and a model prediction module 40. In this embodiment, a network model based on a deformable spatial attention mechanism is proposed and applied to the end-to-end gesture segmentation task. The attention mechanism can capture gesture features of different scales and forms, improve the modeling ability for complex gesture shapes and details, thereby significantly improving the accuracy and real-time performance of gesture segmentation. Moreover, by using the spatial attention allocation strategy of deformable convolution, the model can directly input the original image for fast segmentation without relying on complex preprocessing under complex background conditions, and finally can accurately segment the gesture target, improving the adaptability and robustness in complex scenarios.

[0108] It should be noted that for the technical details not described in detail in this embodiment of the end-to-end fast gesture segmentation device, reference can be made to the application of the end-to-end fast gesture segmentation method provided in any embodiment of the present invention as described above, and details will not be repeated here.

[0109] It should be understood that the above is only an example and does not constitute any limitation to the technical solution of the present invention. In specific applications, those skilled in the art can set according to needs, and the present invention does not limit this.

[0110] It should be noted that the above-described work process is only illustrative and does not limit the protection scope of the present invention. In actual applications, those skilled in the art can select some or all of them according to actual needs to achieve the purpose of the solution of this embodiment, and details will not be described here.

[0111] In addition, it should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0112] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as Read Only Memory (ROM) / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0114] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. An end-to-end fast gesture segmentation method, characterized in that: include: Collect gesture image data and perform annotation processing to obtain training sets and test sets; Construct an end-to-end gesture segmentation network model based on deformable spatial attention mechanism; Training and testing the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model; The image to be predicted is input into the trained gesture segmentation network model, and a binary gesture segmentation mask image is output.

2. The method according to claim 1, characterized in that The collecting of gesture image data and labeling to obtain a training set and a test set include: Collecting gesture image data, and extracting sample images according to the gesture image data; Performing polygon annotation on the hand area in the sample image by using a data annotation tool, and filling the area enclosed by the polygon to generate a mask label; Constructing a data sample for model training according to the sample image and the annotated mask label; The data samples are divided into a training set and a test set based on a preset ratio.

3. The method according to claim 1, characterized in that The end-to-end gesture segmentation network model based on the deformable spatial attention mechanism is constructed, including: Get the full convolutional network model; Modify the backbone network of the fully convolutional network model into a MobileNet structure; A deformation space attention module is added after the backbone network to obtain an end-to-end gesture segmentation network model based on the deformation space attention mechanism.

4. The method according to claim 3, characterized in that The end-to-end gesture segmentation network model based on the deformable space attention mechanism includes: a MobileNet structure, a deformable space attention module, a deconvolution layer, and a bilinear sampling layer.

5. The method according to claim 3, characterized in that The deformable spatial attention module includes: a deformable convolution layer, an output feature layer, a 1*1 convolution layer, an activation layer and an element-level multiplication layer.

6. The method according to claim 1, characterized in that The step of training and testing the end-to-end gesture segmentation network model based on the deformable space attention mechanism according to the training set and the test set to obtain the trained gesture segmentation network model includes: Training the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set; The stochastic gradient descent method is used to optimize the parameter update of the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism, and the cross entropy loss is used as the loss function during model training; After each round of training of the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism, evaluating the accuracy of the network model on the test set according to the test set; The network weights of the round of training with the highest accuracy are selected as model parameters to obtain the trained gesture segmentation network model.

7. The method according to claim 6, characterized in that The accuracy is the mean intersection over union (MIoU) indicator.

8. The method according to any one of claims 1 to 7, characterized in that The step of inputting the image to be predicted into the trained gesture segmentation network model and outputting a binary gesture segmentation mask map comprises: Acquire an original image or video frame, and extract a picture to be predicted based on the original image or video frame; Normalizing the image to be predicted to convert it into an image input format required by the model; Inputting the normalized image to be predicted into the trained gesture segmentation network model to obtain a gesture segmentation result; Bilinear interpolation is performed on the gesture segmentation result to obtain a binary gesture segmentation mask image.

9. An end-to-end fast gesture segmentation device, characterized in that: include: The image annotation module is used to collect gesture image data and perform annotation processing to obtain training sets and test sets; Model building module, used to build an end-to-end gesture segmentation network model based on deformable spatial attention mechanism; A model training module, used for training and testing the end-to-end gesture segmentation network model based on the deformable spatial attention mechanism according to the training set and the test set to obtain a trained gesture segmentation network model; The model prediction module is used to input the image to be predicted into the trained gesture segmentation network model and output a binary gesture segmentation mask image.

10. An electronic device, characterized in that: The electronic device comprises: a memory, a processor, and an end-to-end fast gesture segmentation program stored in the memory and executable on the processor, wherein the end-to-end fast gesture segmentation program is configured to implement the end-to-end fast gesture segmentation method according to any one of claims 1 to 8.