A five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning

By using soft masking region representation and multi-task learning, the five-finger dexterity hand grasping detection is decomposed into multiple sub-tasks. By utilizing a grasping-oriented attention mechanism and an adaptive weighted loss function, the robustness and real-time performance issues of grasping detection in multi-object scenes in existing technologies are solved, achieving efficient and accurate grasping detection.

CN122299732APending Publication Date: 2026-06-30SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG UNIV OF SCI & TECH
Filing Date
2026-05-28
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing five-finger dexterous hand grasping and detection technologies suffer from insufficient robustness in multi-object scenarios, difficulty in achieving both real-time performance and continuous labeling, and challenges in balancing multi-task optimization. In particular, they struggle to achieve efficient and accurate grasping and detection in complex environments.

Method used

We employ a soft mask region representation and multi-task learning approach to decompose high-dimensional grasping parameters into three sub-tasks: grasping quality prediction, grasping width regression, and grasping gesture classification. We also construct a lightweight multi-task detection network through a grasping-oriented channel-space-geometric attention mechanism and dynamically balance the multi-task learning process by combining an adaptive weighted loss function with effective region constraints.

Benefits of technology

It improves the accuracy, real-time performance, and robustness of five-finger dexterity hand grasping detection, reduces model complexity, enhances grasping detection capabilities in complex backgrounds, and is suitable for multi-object scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122299732A_ABST
    Figure CN122299732A_ABST
Patent Text Reader

Abstract

This invention discloses a five-finger dexterity hand grasping detection method based on soft-mask region representation and multi-task learning. This method decomposes high-dimensional continuous grasping parameters into three sub-tasks: grasping quality prediction, grasping width regression, and grasping gesture classification. It constructs a soft-mask multi-color grasping region representation to generate pixel-level grasping quality, width, and gesture labels. A grasping-oriented channel-space-geometric attention mechanism is designed to construct a lightweight multi-task generative grasping detection network. Taking RGB-D images as input, it outputs grasping quality maps, width maps, and gesture maps in parallel. An adaptive weighted loss function based on effective region constraints is used for training to suppress background interference and dynamically balance multi-task learning. During inference, the grasping center is located by searching for peaks in the quality map, and the corresponding width and gesture are read to achieve single-target or multi-target grasping detection. This invention improves the accuracy, real-time performance, and robustness of grasping detection in multi-object scenes while reducing model complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot vision perception and intelligent grasping technology, specifically involving a five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning. Background Technology

[0002] Five-fingered dexterous hands possess significant advantages such as high degrees of freedom and manipulative dexterity, enabling them to perform precise grasping tasks in complex environments and demonstrating broad application prospects in fields such as intelligent manufacturing, service robots, and embodied intelligence. However, compared to traditional two-fingered grippers, the grasping process of five-fingered dexterous hands involves a high-dimensional parameter space, including hand pose, joint angles, and contact relationships, making grasping modeling and detection tasks more complex. Especially in multi-object scenarios, proximity interference between targets and background noise further exacerbate the difficulty of grasping perception.

[0003] Most current mainstream technologies for five-finger dexterity hand grasping detection employ an end-to-end approach to directly regress high-dimensional grasping parameters from visual input. While these methods offer strong expressive power, they often rely on large-scale high-dimensional labeled data and complex model structures. Training is unstable, data acquisition costs are high, and inference computation is expensive, making it difficult to meet real-time requirements. Furthermore, in multi-object scenes, they are susceptible to interference from neighboring objects and background noise, exhibiting insufficient robustness. In addition, some existing technologies transform the grasping problem into a classification or low-dimensional prediction problem through dimensionality reduction modeling. While this reduces the difficulty of grasping modeling, limitations remain in label representation, multi-task optimization balancing, and adaptability to multi-object scenes, including discontinuous label representation, difficulty in coordinating multi-task losses, and the challenge of simultaneously achieving both real-time performance and detection accuracy.

[0004] Therefore, there is an urgent need to provide a five-finger dexterous hand grasping detection method for multi-object scenarios, so as to achieve lightweight deployment and real-time detection while ensuring the semantic expression capability of grasping, and improve the accuracy, real-time performance and robustness of five-finger dexterous hand grasping detection in multi-object scenarios. Summary of the Invention

[0005] To address the aforementioned problems in existing technologies, this invention proposes a five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning. This method decomposes high-dimensional continuous grasping parameters into three sub-tasks: grasping quality prediction, grasping width regression, and grasping gesture classification. By constructing a soft mask multi-color grasping region representation, continuous modeling of the grasping region is achieved. Simultaneously, a grasping-oriented channel-space-geometric attention mechanism is designed to construct a lightweight multi-task grasping detection network. Combined with an adaptive weighted loss function based on effective region constraints, this method reduces model complexity while improving the accuracy, real-time performance, and robustness of grasping detection in multi-object scenarios.

[0006] A five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning includes the following steps: Step 1: Acquire aligned RGB and depth images of the target scene to form multimodal input data; based on the contact patterns and number of participating fingers of the five-finger dexterous hand, the dexterous hand grasping posture is discretized into four typical grasping gestures: three-finger grasping, i.e., the thumb grasps against the index and middle fingers; five-finger grasping, i.e., the thumb grasps against the other four fingers; three-finger wrapping grasping, i.e., the thumb, index, and middle fingers form a wrapping contact grasping; and five-finger full palm grasping, i.e., all fingers and the palm participate in grasping. Step 2: Construct a soft mask multi-color grasping region representation based on the original grasping rectangle annotation, perform continuous quality modeling on the grasping region, and simultaneously generate grasping quality labels, grasping width labels, and grasping gesture labels; Step 3: Design a grasping-oriented channel-space-geometric attention mechanism and construct a lightweight multi-task generative grasping detection network PAG-GraspNet. The multimodal data obtained in Step 1 is used as input, and the grasping quality map, grasping width map and grasping gesture map are jointly output through shared feature extraction and multi-scale feature fusion. Step 4: Construct an effective grasping region based on the grasping quality label, and train the lightweight multi-task generative grasping detection network using a multi-task adaptive weighted loss function based on the effective region constraint. Supervise each task only within the effective grasping region to suppress background interference, and design an adaptive weighting strategy based on task uncertainty to dynamically balance the multi-task joint learning process. Step 5: In the inference phase, based on the grasp quality map, a multi-target grasp joint parsing method is used to obtain the candidate grasp center position. At the corresponding positions in the grasp width map and grasp gesture map, the grasp width prediction value and grasp gesture category prediction value are read to obtain the final grasp detection result.

[0007] Furthermore, the specific method for constructing the soft mask multi-color grasping region representation in step 2 is as follows: Based on the original rotated rectangle, keeping the center and short side of the rectangle unchanged, the long side is shortened to half its original length to obtain a new rectangle. Then, multiple rectangles of the same target are merged into a whole region as the final grasping region, which is represented as follows: ; in, This represents a multi-color capture region using a soft mask. To capture the area, For width, This indicates the category of the grab gesture corresponding to the color code; The method for constructing the captured quality labels is as follows: A continuous quality map is generated within the captured area using an adaptive Gaussian distribution, and for any pixel... Its tag definition is: ; in, To capture the center pixel of the region, the standard deviation , For adjustment coefficients, Indicates the area to be captured; The method for constructing the grab width label is as follows: the initial grab width is set to 1.5 times the shorter side of the smallest bounding rectangle of the grab area, and it is mapped to the physically executable range [20mm, 120mm] and normalized for any pixel. Its tag definition is: ; in, To capture the short side of the smallest bounding rectangle of the region; The grasping gesture tags are constructed as follows: a five-channel grasping gesture map is constructed using One-Hot encoding. The first four channels correspond to four types of grasping gestures, and the fifth channel represents the background category for any pixel. Its tag definition is: ; in, To capture the gesture image in the first On each channel, pixel position The value at that location, This indicates the area to be crawled. The corresponding gesture category number, Indicates the channel index.

[0008] Furthermore, the lightweight multi-task generative grasping detection network PAG-GraspNet in step 3 adopts an encoder-decoder architecture based on a hybrid deep convolutional structure and a grasping-oriented channel-spatial-geometric attention mechanism, and finally outputs a grasping quality map through the joint output of multi-task output heads. 1. Capture the width image and the image of the grabbing gesture ; The encoder part constructs a lightweight coordinate-aware feature extraction unit using depthwise separable convolution and coordinate attention mechanisms. Through the design of multiple layers of lightweight coordinate-aware feature extraction units, shallow semantic features are obtained sequentially. Mid-level semantic features and deep semantic features In deep semantic features The subsequent design incorporates a mixed-direction depthwise convolution structure, MixDepConv, the specific process of which is as follows: ; ; ; ; in, This represents the features extracted by depthwise convolution in the horizontal direction. This represents the features extracted by depthwise convolution in the vertical direction. This indicates the result of directional feature fusion. This represents the output features obtained after passing through a single layer of mixed-direction depthwise convolutional structure, and the output features are obtained after passing through multiple layers of mixed-direction depthwise convolutional structures. ; The encoder output end is designed with a gripping-guided channel-space-geometric attention module (GCSGA). The calculation process of the channel attention branch is as follows: ; ; ; ; in, and These represent global average pooling and global max pooling, respectively. and These are the global channel descriptions after global average pooling and global max pooling, respectively. Represents a 1×1 convolution. and for and Channel-level intermediate features obtained by lightweight MLP This represents the Sigmoid activation function. This is the output channel attention map; Meanwhile, the calculation process for the spatial attention branch is as follows: ; ; in, and These represent the mean and maximum values ​​along the channel dimension, respectively. and These are the mean and maximum feature maps along the channel dimension, respectively. The output is a spatial attention map; The calculation process for directional geometric enhancement branch is as follows: ; in, For directional geometric enhancement features; Apply channel attention and spatial attention This yields the channel-space joint enhancement features, which are then added to the directional geometric enhancement features to obtain the complete attention features. Finally, these features are combined with... The deepest output features of the encoder are obtained through residual connection fusion. The calculation process is as follows: ; ; in, This indicates channel-by-channel multiplication. For channel-space joint enhancement features; The decoder section employs a top-down multi-scale feature fusion structure, which... , , , Cross-scale fusion is performed between levels. Each decoding unit consists of channel splicing, CSPBlock feature fusion, and bilinear interpolation upsampling. The calculation process is as follows: ; ; ; in, This indicates a bilinear interpolation upsampling operation. This refers to the CSPBlock module. This represents feature concatenation along the channel dimension; These are the shared features in the final output of the decoder; The multi-task output constraint module contains three parallel task output heads: the Quality Head is used to predict the capture quality map. Width Head is used to predict the crawl width map. Gesture Head is used to predict grasping gestures. Each output header contains two 3×3 convolutional layers. The first convolutional layer is followed by batch normalization and the SiLU activation function, and Dropout is introduced. The second convolutional layer is used to generate the final prediction result. Finally, only the quality image needs to be captured. Find the position with the largest response as the grab center, and then in the grab width map and grab gesture diagram The system reads the grab width and gesture category corresponding to the position to obtain the grab parameters at the same spatial position.

[0009] Furthermore, the total loss of the multi-task adaptive weighted loss function ERAMLoss based on effective region constraints in step 4 is... Loss of grasping quality Crawling width loss Grab gesture classification loss and location-assisted loss Composition; First, based on the captured quality label Construct an effective capture region mask : ; in, It is a threshold. It is an indicator function. Represents pixels If it is a valid capture area, otherwise it is considered a background area. To effectively capture the number of pixels in the region; The losses for each subtask are calculated within the effective region: ; in, , and Both use Huber loss. Employ multi-class cross-entropy loss; The location-assisted loss A two-dimensional Gaussian heatmap is constructed based on the center of the captured region. And during the training phase, the positional responses within the network are supervised: ; in, Auxiliary position response map predicted during the training phase; The total loss is calculated using an adaptive weighted strategy based on task uncertainty. ; in, Indicates the first The learnable uncertainty parameter corresponding to the loss of each task.

[0010] Furthermore, step 5 employs a multi-target grasping joint analysis method to generate multi-target grasping detection results, specifically including: Grab gesture diagram After performing probability normalization, pixel-level probability distribution maps of each grabbing gesture category are obtained. ; For the quality of the captured image Perform local peak detection to extract a set of candidate grab center points that meet the quality threshold and minimum peak spacing. ; For the candidate crawl center point set Each candidate grab center point According to the capture width map In position The predicted width value at that location is used as the crawl width parameter. According to the probability distribution diagram of the grasping gesture Read in position The probability vector at a given point is used to determine the gesture type, with the gesture category corresponding to the highest probability being selected. And record this probability value as the gesture confidence level. According to the capture quality map In position The quality response value at the location is used as the confidence level of the crawl center. ; Combined with the confidence level of the crawl center Gesture confidence and the capture width parameter The candidate grab center points are jointly screened, and the multi-target grab detection results that meet the constraints are output.

[0011] The beneficial technical effects of this invention are as follows: (1) This invention reconstructs the high-dimensional continuous grasping representation of five-finger dexterity hand into three low-dimensional sub-tasks: grasping quality prediction, grasping width regression, and grasping gesture classification, which reduces the difficulty of high-dimensional grasping modeling and improves the learnability and reasoning efficiency of detection problems. (2) This invention proposes a soft mask multicolor grasping region representation method, which replaces the traditional hard boundary region supervision with continuous quality distribution, effectively improving the continuity of grasping label expression, the stability of the training process and the robustness in complex backgrounds; (3) This invention constructs a lightweight multi-task generative grasping and detection network PAG-GraspNet. By designing a hybrid directional depthwise convolution and a grasping-oriented channel-space-geometric attention mechanism, while maintaining an extremely low number of parameters (0.91M), the network is endowed with strong directional awareness and global noise reduction capabilities, improving the real-time performance and robustness of grasping and detection in multi-object scenes without sacrificing accuracy; (4) The present invention designs a multi-task adaptive weighted loss function based on effective region constraints, which is supervised only within the effective grasping region and dynamically balances the multi-task loss weights to reduce background noise interference and improve the convergence stability and grasping center positioning accuracy of multi-task joint learning. (5) This invention achieves stable detection of multiple grasping centers in multi-object scenes through a multi-target grasping joint analysis algorithm based on local peak search, and can provide reliable priors for target region extraction, point cloud processing and grasping posture generation, and has good engineering application value. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating the overall method of the present invention; Figure 2 This is a schematic diagram of the soft mask multi-color capture region representation method in this invention; Among them, (a) is a schematic diagram of TPP capture representation; (b) is a schematic diagram of POG capture representation; (c) is a schematic diagram of TEG capture representation; and (d) is a schematic diagram of FPG capture representation. Figure 3 This is a visual diagram of the task training labels in this invention; Wherein, (a) is the grasping area; (b) is the quality label; (c) is the width label; and (d) is the gesture label. Figure 4 This is a schematic diagram of the channel-space-direction attention mechanism network structure for grasping guidance in this invention; Figure 5 This is a schematic diagram of the lightweight multi-task generative grasping and detection network structure in this invention; Figure 6 This is a schematic diagram illustrating the grasping and detection results of the present invention in a single-object scene; Among them, (a) is the grasping quality image; (b) is the grasping width image; (c) is the grasping gesture image; and (d) is the detection result image. Figure 7 This is a schematic diagram of the parallel grasping and detection results of the present invention in a multi-object scene; Among them, (a) is the grasping quality image; (b) is the grasping width image; (c) is the grasping gesture image; and (d) is the detection result image. Figure 8 This is a schematic diagram of the real-time serial grasping and detection results of the present invention in a multi-object scene; Figure 9 This is a schematic diagram illustrating the application effect of the target area extraction based on the grabbing prompts of the present invention; Figure 10 This is a schematic diagram comparing the real-time grasping and detection performance of the algorithm in multi-object scenes according to the present invention; Among them, (a) is a comparison chart of average accuracy; (b) is a comparison chart of average processing time. Detailed Implementation

[0013] The specific embodiments of the present invention will be further described below with reference to specific examples: A five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning, such as Figure 1As shown, the process is mainly divided into three stages: input and grasping representation construction, network training and grasping detection, and result output and downstream application. First, in the input and grasping representation construction stage, aligned RGB images and depth images of the target scene are obtained. Combined with the semantic discretization of the five-finger dexterity grasping type, the grasping semantic representation is constructed. On this basis, a soft-mask multi-color region representation (S-MCRR) is constructed, and grasping quality labels, grasping width labels, and grasping gesture labels are generated simultaneously. Secondly, in the network training and grasping detection phase, the aforementioned multimodal input data is fed into the lightweight multi-task generative grasping detection network (Position and Gesture GraspNet, PAG-GraspNet), and trained using an Effective Region Constrained Adaptive Multi-Task Loss (ERAM Loss) function. This allows the network to output a grasping quality map, a grasping width map, and a grasping gesture map, respectively. Subsequently, based on the grasping quality map, a single-object global peak search or a multi-object local peak search is performed to determine the candidate grasping center positions, and the width and gesture prediction values ​​at the corresponding positions are read to obtain the final grasping detection results. Finally, in the result output and downstream application phase, the grasping detection results are used as foreground cues to drive a lightweight instance segmentation model to generate a target object mask corresponding to the current grasping intent, providing a reliable visual perception prior for subsequent target point cloud extraction and dexterous hand grasping posture generation.

[0014] The specific steps of this invention are as follows: Step 1: Obtain aligned RGB and depth images of the target scene to form multimodal input data; based on the contact pattern of the five-finger dexterity hand and the number of participating fingers, the dexterity hand grasping posture is discretized into four typical grasping gestures, and a grasping semantic labeling system is constructed. To avoid directly regressing the high-dimensional continuous grasping posture parameters of the five-finger dexterity hand, this invention first performs semantic discretization on the grasping forms of the five-finger dexterity hand. Based on the involvement of fingers and palm in the grasping process, as well as the shape and size characteristics of the object, the grasping gestures are divided into four typical forms: 1. Three-finger grasp (TPP): The thumb grasps the index and middle fingers together. 2. Five-finger grasp (POG): The thumb grasps the other four fingers together. 3. Three-finger wrap-around grasp (TEG): The thumb, index finger, and middle finger form a wrap-around contact grasp; 4. Five-finger full palm grip (FPG): All fingers and palms participate in the grip.

[0015] Through the above semantic segmentation, the high-dimensional continuous grasping parameters of the five-finger dexterity hand are mapped to low-dimensional category labels with clear physical meaning, providing a stable foundation for subsequent supervised learning.

[0016] For a given aligned RGB image and the corresponding depth map The crawling semantic labeling system uses pixel-level crawling semantic representation as the prediction target of a lightweight multi-task generative crawling detection network. Defined as: ; in, To capture a quality image, this is used to describe the feasibility of using each pixel as the capture center; To capture the width image, provide a priori information for the corresponding capture width. This is a gesture mapping diagram used to describe the pixel-level distribution of gesture categories. These correspond to four types of grabbing gestures and background categories, respectively. In this way, the high-dimensional grasping detection of dexterous hands in multi-object scenes is reconstructed into three low-dimensional sub-tasks: grasping quality prediction, grasping width regression, and grasping gesture classification.

[0017] Step 2: Construct a soft mask multi-color grasping region representation based on the original grasping rectangle annotation, perform continuous quality modeling on the grasping region, and simultaneously generate grasping quality labels, grasping width labels, and grasping gesture labels; To accommodate the five-finger dexterity hand grasping representation, this invention designs a soft-mask multi-color grasping region representation method, S-MCRR, based on the two-finger gripper rectangular bounding box annotation, and generates training labels based on this method. For example... Figure 2 As shown, based on the original rotated rectangle, keeping the center and short side unchanged, the long side is shortened to half its original length to obtain a new rectangle. Then, multiple rectangles representing the same target are merged into a single region as the final grasping area, represented as follows: ; in, This represents a multi-color capture region using a soft mask. To capture the area, For width, The color coding indicates the grab gesture category, with TPP, POG, TEG, and FPG identified by pink, green, yellow, and purple, respectively.

[0018] The method for constructing quality labels is as follows: A continuous quality map is generated within the capture area using an adaptive Gaussian distribution, and for any pixel... Its tag definition is: ; in, To capture the center pixel of the region, the standard deviation , For adjustment coefficients, This represents the area of ​​the grabbing region; the design makes the center response peak 1 and then smoothly decays towards the edges.

[0019] The method for constructing the grasp width label is as follows: considering the redundant opening amount during dexterity hand execution, the initial grasp width is set to 1.5 times the shorter side of the smallest bounding rectangle of the grasp area, and it is normalized by mapping it to the physical executable range [20mm, 120mm] for any pixel. Its tag definition is: ; in, To capture the short side of the smallest bounding rectangle of the capture area; and during training, width supervision is only performed within the effective capture area to reduce the interference of the background area on the width regression learning.

[0020] The grasping gesture tags are constructed as follows: a five-channel grasping gesture map is constructed using One-Hot encoding. The first four channels correspond to four types of grasping gestures, and the fifth channel represents the background category for any pixel. Its tag definition is: ; in, To capture the gesture image in the first On each channel, pixel position The value at that location, This indicates the area to be crawled. The corresponding gesture category number, Indicates the channel index.

[0021] Figure 3 The above design provides a visual example of three types of grasping labels. Through this design, the high-dimensional grasping posture information of the dexterous hand is mapped into learnable discrete semantic labels, which together with continuous quality supervision and width regression constitute a unified multi-task supervision signal.

[0022] Step 3: Design a grasping-oriented channel-space-geometric attention mechanism and construct a lightweight multi-task generative grasping detection network PAG-GraspNet. The multimodal data obtained in Step 1 is used as input, and the grasping quality map, grasping width map and grasping gesture map are jointly output through shared feature extraction and multi-scale feature fusion. This invention designs a lightweight, multi-task generative grasping and detection network for five-fingered dexterous hands, with only 0.91M network parameters. The overall architecture is as follows: Figure 5 As shown, this network takes aligned RGB-D images as input and employs an encoder-decoder architecture. It jointly generates a grasping quality map through lightweight coordinate-aware feature extraction, hybrid orientation depth convolutional blocks, a grasping-oriented channel-spatial-geometric attention mechanism, and a multi-task output constraint module. 1. Capture the width image and the image of the grabbing gesture .

[0023] In the encoder section, this invention constructs a lightweight coordinate-aware feature extraction unit. This unit uses depthwise separable convolution (DepthSepConv) as the basic operator. First, it extracts local spatial texture, object edges, and the structure of the grasping contact area within each channel through depthwise convolution. Then, it fuses information between channels through pointwise convolution, thereby extracting local geometric features related to grasping detection with a relatively low parameter count. Simultaneously, to avoid weakening spatial position information during the parameter reduction process of lightweight convolution, this invention introduces a coordinate attention (CA) mechanism in this feature extraction unit, enabling the network to retain horizontal and vertical position information during channel feature modeling. Therefore, the network can more accurately perceive the grasping center position, the major and minor axes of the object, the grasping contact area, and the spatial layout relationships in multi-object scenes. The calculation process is as follows: ; ; ; in, For the input features, BN represents batch normalization. and These represent depthwise convolution and pointwise convolution, respectively. This represents the SiLU activation function. This indicates a coordinate attention mechanism. This represents the lightweight feature output after coordinate-aware enhancement. Through the above structure, the encoder can enhance its ability to express the position, contour boundary and contact geometry of the grasping area while maintaining a low number of parameters and computational cost.

[0024] In this embodiment, the input features are processed through a 3×3 two-dimensional convolutional layer and two lightweight coordinate-aware feature extraction units to obtain shallow semantic features. Then, the mid-level semantic features are obtained through two lightweight coordinate-aware feature extraction units. Then, deep semantic features are obtained through two lightweight coordinate-aware feature extraction units. ; To further enhance the network's ability to model the directional structure of the grasping region of a five-finger dexterity hand, this invention focuses on deep semantic features within the encoder. Next, a mixed-direction depthwise convolution structure, MixDepConv, was set up. This structure decomposes the conventional two-dimensional large-kernel convolution into horizontal and vertical depthwise convolutions, extracting the spatial responses of the grasping region in different directions respectively, and then performing channel fusion through pointwise convolution. Since five-finger dexterity hand grasping detection not only needs to determine the grasping center position, but also needs to estimate the width information related to the object scale and grasping opening, the major and minor axis directions, boundary orientation, and local structural changes of the grasping region have a significant impact on the detection results. MixDepConv can expand the effective receptive field without significantly increasing the number of parameters, enabling the network to perceive grasping structural features in both horizontal and vertical directions, improving the stability of grasping center localization and width prediction. Its calculation process is as follows: ; ; ; ; in, This represents the features extracted by depthwise convolution in the horizontal direction. This represents the features extracted by depthwise convolution in the vertical direction. This indicates the result of directional feature fusion. This represents the output features obtained after passing through a single layer of mixed-direction deep convolutional structures, and the output features are obtained after passing through multiple layers of mixed-direction deep convolutional structures. By jointly modeling the horizontal and vertical directions, MixDepConv can enhance the network's ability to recognize the boundaries of the grasping area and changes in the scale of objects, making it particularly suitable for complex grasping scenarios with multiple objects in proximity. In this embodiment, deep semantic features The output features are obtained through a five-layer mixed-direction depthwise convolutional structure. ; In multi-object scenarios, relying solely on local features is easily affected by background noise from neighboring objects and local texture noise. Meanwhile, spatial and directional information of the environment is also crucial for detection. To enhance the network's global semantic understanding of graspable targets, highlight graspable regions, capture edge and directional information, and maintain training stability and feature integrity, this invention proposes a grasping-oriented channel-spatial-geometric attention mechanism (GCSGA), such as... Figure 4 As shown, the extracted features are further enhanced by introducing the algorithm at the encoder output. GCSGA constructs a "global-local-direction" ternary geometric feature representation through three branches: channel attention, spatial attention, and orientation geometry enhancement, and uses residual fusion to enhance the output features. First, the channel context branch aggregates the semantic context information of the entire image through global average pooling and global max pooling. The two branches share the same MLP to obtain channel-level intermediate features, which are then normalized by the activation function to obtain the channel attention map. This branch enhances local salient features while preserving global semantics, suppresses irrelevant channels, and retains important channel information. The calculation process is as follows: ; ; ; ; in, and These represent global average pooling and global max pooling, respectively. and These are the global channel descriptions after global average pooling and global max pooling, respectively. Represents a 1×1 convolution. Constructing a lightweight MLP, and for and Channel-level intermediate features obtained by lightweight MLP This represents the Sigmoid activation function. This is the output channel attention map; Simultaneously, the spatial attention branch calculates a two-dimensional spatial feature map by aggregating the mean and maximum values ​​along the channel dimension. This map is then concatenated along the channel dimension and further processed by a 7×7 convolution to generate a spatial attention map, thereby suppressing background noise and highlighting the graspable region. The calculation process is as follows: ; ; in, and These represent the mean and maximum values ​​along the channel dimension, respectively. and These are the mean and maximum feature maps along the channel dimension, respectively. The output is a spatial attention map; The directional geometry enhancement branch uses 1×5 and 5×1 depthwise separable convolutions to extract geometric features in the horizontal and vertical directions, respectively. These features are then added together and followed by channel fusion and nonlinear activation via a 1×1 convolution to obtain directional geometry enhancement features. This supplements the directional geometric information that channel + spatial attention cannot directly express. The calculation process is as follows: ; in, For directional geometric enhancement features; Apply channel attention and spatial attention This yields the channel-space joint enhancement features, which are then added to the directional geometric enhancement features to obtain the complete attention features. Finally, these features are combined with... The deepest output features of the encoder are obtained through residual connection fusion. The calculation process is as follows: ; ; in, This indicates channel-by-channel multiplication. This approach enhances features through a combination of channel-space joint enhancement. By organically integrating dual-pooling channel attention, spatial attention, and axial depth convolution in the GCSGA module, it effectively enhances the capture of relevant features, suppresses background interference, and improves the accuracy of dense capture detection. The design of the residual structure also preserves the original features, ensuring training stability and providing a more stable shared feature foundation for subsequent multi-scale decoding and multi-task prediction.

[0025] In the decoder section, a top-down multi-scale feature fusion structure is adopted to progressively recover the spatial resolution of the feature maps, meeting the requirements of pixel-level dense prediction in grasping detection. Specifically, , , , Cross-scale fusion is performed between different levels, enabling the network to utilize global semantic information from deep features while preserving edges, contours, and local geometric details from shallow features. Each decoding unit consists of channel concatenation, CSPBlock feature fusion, and bilinear interpolation upsampling. Channel concatenation integrates feature information from different scales, CSPBlock achieves feature fusion and channel compression while reducing redundant computation, and bilinear interpolation upsampling progressively restores the feature map spatial dimensions. Through this structure, the network can generate shared decoding features that simultaneously capture local geometric information and global semantic information. This process can be represented as: ; ; ; in, This indicates a bilinear interpolation upsampling operation. This refers to the CSPBlock module. This represents feature concatenation along the channel dimension; The final output features of the decoder; The multi-task output constraint module includes three parallel task output heads: Quality Head, Width Head, and Gesture Head, which are used to generate the capture quality map. 1. Capture the width image and grab gesture diagram The QualityHead predicts the feasibility of each pixel location as a grasping center. The Width Head predicts the dexterity hand prior grasping width at the corresponding pixel location, and the Gesture Head predicts the five-finger dexterity hand grasping gesture category at the corresponding pixel location. Each output head includes two 3×3 convolutional layers. The first convolutional layer is followed by batch normalization and a SiLU activation function, and Dropout is introduced to alleviate overfitting during training. The second layer generates the final prediction result for the corresponding task, and its outputs are as follows: ; ; ; in, , , These represent the grab quality, width, and gesture output head, respectively. , , These represent the predicted grasping quality map, grasping width map, and grasping gesture map, respectively. The quality map is used to determine the candidate grasping center, the width map is used to extract the prior grasping width at the candidate center, and the gesture map is used to extract the five-finger dexterity hand grasping type at the candidate center. Together, they constitute the grasping parameters at the same spatial location.

[0026] Compared to methods that only employ general lightweight convolutional networks for image feature extraction, the PAG-GraspNet described in this invention constructs a task-related collaborative feature generation mechanism for five-finger dexterity hand grasping detection. Specifically, depthwise separable convolutions reduce the computational cost of local feature extraction, coordinate attention preserves the spatial location information of the grasping center and contact area, hybrid directional depthwise convolutions enhance the major and minor axes and boundary orientation features of the grasping area, a grasping-oriented channel-space-geometric attention mechanism enhances grasping-related features and suppresses background interference, a multi-scale decoding structure restores the spatial resolution required for pixel-level grasping prediction, and three task output heads output grasping quality, grasping width, and grasping gesture respectively based on the same shared features. These structures work together to ensure that the grasping center, grasping width, and grasping gesture output by the network are consistent in spatial location, thereby improving the accuracy and real-time performance of five-finger dexterity hand grasping detection in multi-object scenarios.

[0027] Step 4: Construct an effective grasping region based on the grasping quality label, and train the lightweight multi-task generative grasping detection network using a multi-task adaptive weighted loss function based on the effective region constraint. Supervise each task only within the effective grasping region to suppress background interference, and design an adaptive weighting strategy based on task uncertainty to dynamically balance the multi-task joint learning process. To improve the stability of multi-task joint training and enhance the ability to locate the center of the target object, this invention proposes a multi-task adaptive weighted loss function, ERAM Loss, based on effective region constraints. Its total loss... Loss of grasping quality Crawling width loss Grab gesture classification loss And position-aided loss for further improving the positioning accuracy of the gripping center. Composition; First, based on the captured quality label Construct an effective capture region mask : ; in, These are crawling quality tags with values ​​in the range [0,1]. This is a threshold value, which is set to 0.05 in this invention. It is an indicator function. Represents pixels If it is a valid capture area, otherwise it is considered a background area. To effectively capture the number of pixels in the region; By using masking constraints, the loss of each subtask is limited to the effective region for computation, thus shielding the network optimization from background noise. ; in, , and Both use Huber loss for regression tasks. Multi-class cross-entropy loss is used for gesture classification. To further enhance the network's ability to focus on the capture center, location-assisted loss is introduced. A two-dimensional Gaussian heatmap is constructed based on the center of the captured region. And during the training phase, the positional responses within the network are supervised: ; in, This is an auxiliary location response map predicted during the training phase. This auxiliary term only enhances the center point response during the training phase and is not used as an independent output during the inference phase. It effectively complements the continuous quality distribution supervision.

[0028] Given the differences in numerical scale and convergence speed among multiple tasks, fixed weights can easily lead to a single task dominating the overall optimization process. Therefore, this invention introduces an adaptive weighting strategy based on task uncertainty to dynamically balance the joint learning process. The total loss is expressed as: ; in, Indicates the first The learnable uncertainty parameters corresponding to the task loss are optimized synchronously with the network. The network can dynamically adjust the loss weights according to the learning difficulty of each task, thereby improving the convergence stability of the model.

[0029] Step 5: In the inference phase, based on the grasp quality map, a multi-target grasp joint parsing method is used to obtain the candidate grasp center position. At the corresponding positions in the grasp width map and grasp gesture map, the grasp width prediction value and grasp gesture category prediction value are read to obtain the final grasp detection result.

[0030] After predicting the capture quality map, capture width map, and capture gesture map, the network output is parsed for capture detection to generate capture results suitable for single-object or multi-object scenarios.

[0031] For single-grab target detection scenarios, the pixel position with the largest response value is searched in the grab quality map and this position is determined as the grab center; at the same time, the category probability of this grab center in the grab gesture map is read, and the category corresponding to the highest probability is taken as the grab gesture prediction result, thereby generating a single grab detection result. Figure 6 This is a schematic diagram illustrating the grasping and detection results in a single-object scene according to the present invention. The process can be represented as follows: ; ; ; in, This indicates the optimal center point for capturing. This indicates the corresponding initial grab width. Indicates the corresponding grab gesture category. This represents the channel index of the grab gesture image (corresponding to different gesture categories).

[0032] For multi-object scenarios, to accurately locate the grasping center of each object, this invention proposes a multi-object grasping detection method based on local peak search. This method uses network prediction results... and gesture mapping table As input, output a list of multiple object capture detections. The specific implementation steps of the algorithm are as follows: Grab gesture diagram After performing probability normalization, pixel-level probability distribution maps of each grabbing gesture category are obtained. ; For the quality of the captured image Perform local peak detection to extract a set of candidate grab center points that meet the quality threshold and minimum peak spacing. ; For the candidate crawl center point set Each candidate grab center point According to the capture width map In position The predicted width value at that location is used as the crawl width parameter. According to the probability distribution diagram of the grasping gesture Read in position The probability vector at a given point is used to determine the gesture type, with the gesture category corresponding to the highest probability being selected. And record this probability value as the gesture confidence level. According to the capture quality map In position The quality response value at the location is used as the confidence level of the crawl center. ; Combined with the confidence level of the crawl center Gesture confidence and the capture width parameter By jointly screening the candidate grab center points, multi-target grab detection results that meet the constraints are obtained.

[0033] Ultimately, by limiting the number of detection outputs, it is possible to achieve parallel output of grasping information from multiple objects in the scene or output of a single grasping information with the highest global peak value in the scene. Figure 7The results of parallel grasping detection in multi-object scenes are demonstrated. As shown in the figure, the present invention can form significant response regions with multiple physical intervals in the quality map, thereby effectively outputting parallel grasping configurations corresponding to multiple objects. Furthermore, a real-time test environment was built based on an Intel RealSense D435i depth camera to simulate the "single-target" serial grasping mode in real robot operations, outputting only the single grasping configuration with the highest global peak value in each round of detection. Figure 8 The results of real-time serial grasping and detection of objects in multi-object scenarios are demonstrated by the present invention.

[0034] Furthermore, to verify the application capability of this invention in downstream tasks, a target region extraction application verification based on grasping cues was conducted. Specifically, the grasping detection result was used as a foreground cue, and combined with the scene RGB image input to the lightweight instance segmentation model MobileSAM, enabling it to adaptively generate target object masks corresponding to the current grasping intent in multi-object scenes, thereby achieving target object segmentation and providing a reliable visual perception foundation for subsequent tasks such as target point cloud extraction and dexterous hand grasping posture generation. Figure 9 The application effect of the present invention based on target area extraction using grabbing prompts is demonstrated.

[0035] To verify the effectiveness of the five-finger dexterity hand multi-object grasping and detection method described in this invention, multi-dimensional tests were conducted on single-object / multi-object scenes and on a real depth camera.

[0036] 1. Verification of the synergistic effect of key technology modules on grasping and detection performance To verify the contribution of each key module in this invention to the overall grasping and detection performance, an ablation experiment was conducted on a self-built five-finger dexterous hand grasping dataset. The experimental results are shown in Table 1 below.

[0037] Table 1 Impact of core modules on grasping and detection performance ; As shown in Table 1, the average accuracy of crawling detection is improved after introducing S-MCRR soft mask region representation on the basis of the baseline, verifying the gain of continuous semantic supervision on crawling center localization. On this basis, the accuracy is further improved after introducing MixDepConv, GCSGA and ERAM Loss respectively, which fully demonstrates the promoting effect of each technical module on network detection performance. When all modules work together, the best detection performance is achieved, with an accuracy of 96.7%, which proves the effectiveness of the improvement and also shows that there is a significant synergistic effect among the technical modules of this invention.

[0038] 2. Performance Verification of Real-Time Depth Camera Detection in Multi-Object Scenes To verify the feasibility of this invention in real-world engineering scenarios and the comprehensive advantages of PAG-GraspNet in multi-object scenes, a real-world experimental platform was built using an Intel RealSense D435i depth camera to test the grasping and detection performance in multi-object scenes. Test metrics included average localization accuracy, average gesture accuracy, average joint accuracy, and single-frame processing time, and performance was compared with other mainstream grasping and detection algorithms. Figure 10 As shown in (a) and (b), PAG-GraspNet achieves an average localization accuracy of 96.5%, an average gesture accuracy of 93.0%, and an average joint accuracy of 89.6%. The average processing time for a single round of real-time detection is approximately 45.9 ms. These results further demonstrate that the present invention not only meets the operational requirements of five-finger dexterity hands in terms of grasping detection accuracy but also possesses excellent real-time processing capabilities. It can effectively solve the problem of accurate and fast grasping prediction in complex multi-object scenarios for five-finger dexterity hands.

[0039] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning, characterized in that, Includes the following steps: Step 1: Acquire aligned RGB and depth images of the target scene to form multimodal input data; discretize the dexterous hand grasping posture into four typical grasping gestures: three-finger grasping, i.e., the thumb grasps the index and middle fingers; five-finger grasping, i.e., the thumb grasps the other four fingers; three-finger wrapping grasping, i.e., the thumb, index, and middle fingers form a wrapping contact grasping; and five-finger full palm grasping, i.e., all fingers and the palm participate in grasping. Step 2: Construct a soft mask multi-color grasping region representation based on the original grasping rectangle annotation, perform continuous quality modeling on the grasping region, and simultaneously generate grasping quality labels, grasping width labels, and grasping gesture labels; Step 3: Design a grasping-oriented channel-space-geometric attention mechanism and construct a lightweight multi-task generative grasping detection network PAG-GraspNet. The multimodal data obtained in Step 1 is used as input, and the grasping quality map, grasping width map and grasping gesture map are jointly output through shared feature extraction and multi-scale feature fusion. Step 4: Construct an effective grasping region based on the grasping quality label, and train the lightweight multi-task generative grasping detection network using a multi-task adaptive weighted loss function based on the effective region constraint. Supervise each task only within the effective grasping region to suppress background interference, and design an adaptive weighting strategy based on task uncertainty to dynamically balance the multi-task joint learning process. Step 5: In the inference phase, the candidate grab center position is obtained by using the multi-target grab joint parsing method based on the grab quality map. At the corresponding positions in the grab width map and grab gesture map, the grab width prediction value and grab gesture category prediction value are read to obtain the final grab detection result.

2. The five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning according to claim 1, characterized in that, The specific method for constructing the soft mask multi-color grasping region representation in step 2 is as follows: Based on the original rotated rectangle, keeping the center and short side of the rectangle unchanged, the long side is shortened to half of its original length to obtain a new rectangle. Then, multiple rectangles of the same target are merged into a whole region as the final grasping region, which is represented as follows: ; in, This represents a multi-color capture region using a soft mask. To capture the area, For width, This indicates the category of the grab gesture corresponding to the color code; The method for constructing the captured quality labels is as follows: A continuous quality map is generated within the captured area using an adaptive Gaussian distribution, and for any pixel... Its tag definition is: ; in, To capture the center pixel of the region, the standard deviation , For adjustment coefficients, Indicates the area to be captured; The method for constructing the grab width label is as follows: the initial grab width is set to 1.5 times the shorter side of the smallest bounding rectangle of the grab area, and it is mapped to the physically executable range [20mm, 120mm] and normalized for any pixel. Its tag definition is: ; in, To capture the short side of the smallest bounding rectangle of the region; The grasping gesture tags are constructed as follows: a five-channel grasping gesture map is constructed using One-Hot encoding. The first four channels correspond to four types of grasping gestures, and the fifth channel represents the background category for any pixel. Its tag definition is: ; in, To capture the gesture image in the first On each channel, pixel position The value at that location, This indicates the area to be crawled. The corresponding gesture category number, Indicates the channel index.

3. The five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning according to claim 1, characterized in that, The lightweight multi-task generative crawling detection network PAG-GraspNet in step 3 adopts an encoder-decoder architecture based on a hybrid deep convolutional structure and a channel-spatial-geometric attention mechanism. Finally, it outputs a crawling quality map through the joint output of multi-task output heads.

1. Capture the width image and the image of the grabbing gesture ; The encoder part constructs a lightweight coordinate-aware feature extraction unit using depthwise separable convolution and coordinate attention mechanisms. Through the design of multiple layers of lightweight coordinate-aware feature extraction units, shallow semantic features are obtained sequentially. Mid-level semantic features and deep semantic features In deep semantic features The subsequent design incorporates a mixed-direction depthwise convolution structure, MixDepConv, the specific process of which is as follows: ; ; ; ; in, This represents the features extracted by depthwise convolution in the horizontal direction. This represents the features extracted by depthwise convolution in the vertical direction. This indicates the result of directional feature fusion. This represents the output features obtained after passing through a single layer of mixed-direction depthwise convolutional structure, and the output features are obtained after passing through multiple layers of mixed-direction depthwise convolutional structures. ; The encoder output end is designed with a gripping-guided channel-space-geometric attention module (GCSGA). The calculation process of the channel attention branch is as follows: ; ; ; ; in, and These represent global average pooling and global max pooling, respectively. and These are the global channel descriptions after global average pooling and global max pooling, respectively. Represents a 1×1 convolution. and for and Channel-level intermediate features obtained by lightweight MLP This represents the Sigmoid activation function. This is the output channel attention map; Meanwhile, the calculation process for the spatial attention branch is as follows: ; ; in, and These represent the mean and maximum values ​​along the channel dimension, respectively. and These are the mean and maximum feature maps along the channel dimension, respectively. The output is a spatial attention map; The calculation process for directional geometric enhancement branch is as follows: ; in, For directional geometric enhancement features; Apply channel attention and spatial attention This yields the channel-space joint enhancement features, which are then added to the directional geometric enhancement features to obtain the complete attention features. Finally, these features are combined with... The deepest output features of the encoder are obtained through residual connection fusion. The calculation process is as follows: ; ; in, This indicates channel-by-channel multiplication. For channel-space joint enhancement features; The decoder section employs a top-down multi-scale feature fusion structure, which... , , , Cross-scale fusion is performed between levels. Each decoding unit consists of channel splicing, CSPBlock feature fusion, and bilinear interpolation upsampling. The calculation process is as follows: ; ; ; in, This indicates a bilinear interpolation upsampling operation. This refers to the CSPBlock module. This indicates feature concatenation along the channel dimension; These are the shared features in the final output of the decoder; The multi-task output constraint module contains three parallel task output heads: the Quality Head is used to predict the capture quality map. Width Head is used to predict the crawl width map. Gesture Head is used to predict grasping gestures. Each output header contains two 3×3 convolutional layers. The first convolutional layer is followed by batch normalization and the SiLU activation function, and Dropout is introduced. The second convolutional layer is used to generate the final prediction result. Finally, only the quality image needs to be captured. Find the position with the largest response as the grab center, and then in the grab width map and grab gesture diagram The system reads the grab width and gesture category corresponding to the position to obtain the grab parameters at the same spatial position.

4. The five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning according to claim 1, characterized in that, The total loss of the multi-task adaptive weighted loss function ERAMLoss based on effective region constraints in step 4 is... Loss of grasping quality Crawling width loss Grab gesture classification loss and location-assisted loss Composition; First, based on the captured quality label Construct an effective capture region mask : ; in, It is a threshold. It is an indicator function. Represents pixels If it is a valid capture area, otherwise it is considered a background area. To effectively capture the number of pixels in the region; The losses for each subtask are calculated within the effective region: ; in, , and Both use Huber loss. Employ multi-class cross-entropy loss; The location-assisted loss A two-dimensional Gaussian heatmap is constructed based on the center of the captured region. And during the training phase, the positional responses within the network are supervised: ; in, Auxiliary position response map predicted during the training phase; The total loss is calculated using an adaptive weighted strategy based on task uncertainty. ; in, Indicates the first The learnable uncertainty parameter corresponding to the loss of each task.

5. The five-finger dexterity hand grasping detection method based on soft mask region representation and multi-task learning according to claim 1, characterized in that, Step 5 uses a multi-target grasping joint analysis method to generate multi-target grasping detection results, specifically including: Grab gesture diagram After performing probability normalization, pixel-level probability distribution maps of each grabbing gesture category are obtained. ; For the quality of the captured image Perform local peak detection to extract a set of candidate grab center points that meet the quality threshold and minimum peak spacing. ; For the candidate crawl center point set Each candidate grab center point According to the capture width map In position The predicted width value at the location determines the grab width parameter. According to the probability distribution diagram of the grasping gesture In position The probability vector at a given point is used to determine the gesture type, with the gesture category corresponding to the highest probability being selected. According to the capture quality map In position The quality response value at the location determines the confidence level of the capture center. ; Combined with the confidence level of the crawl center Gesture confidence and the capture width parameter By jointly screening the candidate grab center points, multi-target grab detection results that meet the constraints are obtained.