A single-view three-dimensional human hand reconstruction method, device and readable storage medium

Through the convolutional neural network to extract the deep features of the human hand and the two-dimensional joint heat map, combined with upsampling and mesh shape optimization, the problem of inaccurate posture and shape features in single-view human hand reconstruction is solved, and more accurate three-dimensional human hand reconstruction is achieved.

CN115170762BActive Publication Date: 2025-09-12SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210518645.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-12
Publication Date
2025-09-12
Estimated Expiration
2042-05-12

AI Technical Summary

Technical Problem

The existing single-view hand reconstruction method is difficult to obtain accurate three-dimensional hand posture and shape features, which affects the hand reconstruction effect.

Method used

A convolutional neural network is used to extract deep hand features and two-dimensional joint heatmaps, which are gradually upsampled and fused with hand posture features to reconstruct a three-dimensional hand mesh model with a preset number of vertices. The mesh shape is optimized using a spatial mesh convolution method based on neighborhood sorting.

Benefits of technology

It improves the perception ability of human hand posture information and improves the posture accuracy and reconstruction precision of the three-dimensional human hand model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115170762B_ABST
    Figure CN115170762B_ABST
Patent Text Reader

Abstract

The present invention relates to a single-view three-dimensional human hand reconstruction method, device and readable storage medium, which includes the following steps: inputting a human hand RGB image into a convolutional neural network to obtain deep human hand features and a two-dimensional joint heat map; extracting human hand posture features based on the two-dimensional joint heat map; gradually upsampling the deep human hand features and fusing them with the human hand posture features until a three-dimensional human hand mesh model with a preset number of vertices is reconstructed. The present invention relates to a single-view three-dimensional human hand reconstruction method, device and readable storage medium. Since the human hand posture features are extracted from the two-dimensional joint heat map and fused with the deep human hand features to improve the model's perception of human hand posture information, the posture accuracy of the reconstructed hand model is improved. In addition, the reconstruction difficulty is reduced by gradually upsampling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a single-view three-dimensional human hand reconstruction method, device and readable storage medium. Background Art

[0002] Currently, three-dimensional human hand reconstruction is a hot topic in the fields of computer vision and human-computer interaction, and has wide applications in virtual reality, intelligent control, and terminal devices.

[0003] In related technologies, two-dimensional RGB images contain the approximate structural information and color texture information of the human hand. Due to the occlusion problem and lack of depth information, the single-view based hand reconstruction method is difficult to obtain accurate three-dimensional human hand posture and shape features, which in turn affects the hand reconstruction effect.

[0004] Therefore, it is necessary to design a new single-view three-dimensional human hand reconstruction method, device and readable storage medium to overcome the above problems. Summary of the Invention

[0005] Embodiments of the present invention provide a single-view three-dimensional human hand reconstruction method, device and readable storage medium to solve the problem in related technologies that single-view based human hand reconstruction methods are difficult to obtain accurate three-dimensional human hand posture and shape features, affecting the hand reconstruction effect.

[0006] In a first aspect, a single-view three-dimensional human hand reconstruction method is provided, which includes the following steps: inputting an RGB image of a human hand into a convolutional neural network to obtain deep human hand features and a two-dimensional joint heat map; extracting human hand posture features based on the two-dimensional joint heat map; gradually upsampling the deep human hand features and fusing them with the human hand posture features until a three-dimensional human hand mesh model with a preset number of vertices is reconstructed.

[0007] In some embodiments, the method of inputting the RGB image of the human hand into the convolutional neural network to obtain deep human hand features and a two-dimensional joint heat map includes: inputting the RGB image of the human hand into the convolutional neural network to extract deep human hand features; inputting the deep human hand features into the convolutional neural network to predict a two-dimensional joint heat map.

[0008] In some embodiments, the extracting of hand posture features based on the two-dimensional joint heat map includes: obtaining a hand distribution map based on the two-dimensional joint heat map; determining a hand distribution area based on the value range of all joint coordinates in the hand distribution map; and performing a pooling operation on the hand distribution area to complete the extraction of hand posture features.

[0009] In some embodiments, obtaining a hand distribution map based on the two-dimensional joint heat map includes: retaining the pixel with the highest confidence in each of the two-dimensional joint heat maps, and setting the remaining pixels to zero to obtain a single joint distribution map; and summing all the joint distribution maps to obtain a hand distribution map.

[0010] In some embodiments, the gradually upsampling the deep hand features and fusing them with the hand posture features until a three-dimensional hand mesh model with a preset number of vertices is reconstructed includes: upsampling the deep hand features and fusing them with the hand posture features to obtain hand mesh features; predicting a rough three-dimensional hand mesh coordinate vector based on the hand mesh features, and performing mesh shape optimization to obtain an optimized three-dimensional hand mesh coordinate vector; repeating the above steps until a three-dimensional hand mesh model with a preset number of vertices is reconstructed.

[0011] In some embodiments, predicting a coarse three-dimensional hand mesh coordinate vector based on the hand mesh features and performing mesh shape optimization to obtain an optimized three-dimensional hand mesh coordinate vector includes: constructing a reconstruction network using a neighborhood sorting-based spatial mesh convolution method; inputting the hand mesh features into the reconstruction network to predict a coarse three-dimensional hand mesh coordinate vector and a two-dimensional hand mesh vertex coordinate vector; optimizing the coarse three-dimensional hand mesh coordinate vector based on the two-dimensional hand mesh vertex coordinate vector and the hand mesh features to obtain an optimized three-dimensional hand mesh coordinate vector.

[0012] In some embodiments, the coarse three-dimensional hand mesh coordinate vector is optimized based on the two-dimensional hand mesh vertex coordinate vector and the hand mesh feature to obtain an optimized three-dimensional hand mesh coordinate vector, including: fusing the two-dimensional hand mesh vertex coordinate vector with the hand mesh feature to establish non-local connections between network vertices; obtaining a three-dimensional offset correction mesh coordinate vector based on feature prediction of the non-local connection; and adding the three-dimensional offset correction mesh coordinate vector to the coarse three-dimensional hand mesh coordinate vector to obtain an optimized three-dimensional hand mesh coordinate vector.

[0013] In some embodiments, after predicting a rough three-dimensional hand grid coordinate vector based on the hand grid features and performing grid shape optimization to obtain an optimized three-dimensional hand grid coordinate vector, the method further includes: fusing the optimized three-dimensional hand grid coordinate vector with the deep hand features.

[0014] In a second aspect, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one program code, and the program code is loaded and executed by the processor to implement the above-mentioned single-view three-dimensional human hand reconstruction method.

[0015] In a third aspect, a computer-readable storage medium is provided, characterized in that at least one program code is stored in the computer-readable storage medium, and the program code is loaded and executed by a processor to implement the above-mentioned single-view three-dimensional human hand reconstruction method.

[0016] The beneficial effects brought about by the technical solution provided by the present invention include:

[0017] Embodiments of the present invention provide a single-view 3D hand reconstruction method, device, and readable storage medium. By extracting hand posture features from a 2D joint heatmap and fusing them with deep-layer hand features, the method enhances the model's perception of hand posture information and improves the accuracy of the reconstructed hand model. Furthermore, through gradual upsampling, the reconstruction difficulty is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 A schematic flow chart of a single-view 3D human hand reconstruction method provided by an embodiment of the present invention;

[0020] Figure 2 for Figure 1 Schematic diagram of the S3 process. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0022] An embodiment of the present invention provides a single-view three-dimensional human hand reconstruction method, which can solve the problem in related technologies that single-view based human hand reconstruction methods are difficult to obtain accurate three-dimensional human hand posture and shape features, affecting the hand reconstruction effect.

[0023] See also Figure 1 As shown, a single-view 3D human hand reconstruction method provided by an embodiment of the present invention may include the following steps:

[0024] S1: Input the RGB image of the human hand into the convolutional neural network to obtain deep human hand features and two-dimensional joint heat map.

[0025] In step S1, a single RGB image of a hand can be used as input to the convolutional neural network. As the image is downsampled and encoded into features through the convolutional neural network, the feature map resolution decreases as the number of network layers increases, while the semantic information contained increases. Generally, shallow hand features contain some spatial image information (such as color, texture, edges, and corners), while deep hand features, after being encoded through more layers of the neural network, contain more abstract information about the image as a whole, namely semantic information. In this embodiment, a convolutional neural network is used to encode the RGB image of a hand, and the final layer is used to obtain the deep hand features.

[0026] S2: Extracting hand posture features based on the two-dimensional joint heat map.

[0027] In step S2, the two-dimensional joint heat map can be input into the joint encoding module, and the human hand posture features can be extracted according to the input two-dimensional joint heat map.

[0028] S3: Gradually up-sample the deep human hand features and fuse them with the human hand posture features until a three-dimensional human hand mesh model with a preset number of vertices is reconstructed.

[0029] In step S3, the hand mesh model to be predicted in the end generally contains many vertices. It is difficult to directly predict so many parameters. Therefore, this embodiment adopts the process of upsampling-reconstruction-upsampling-reconstruction to realize the reconstruction of multiple vertices of the hand. Specifically, the number of feature points is upsampled to increase the number of vertices. Among them, taking 778 vertices as an example, the reconstruction part can adopt a four-layer network (of course, a three-layer or five-layer network can also be used according to actual needs), and the input features of each layer of the network come from the previous layer of the network. The input of the first-layer reconstruction network is the deep hand feature, which is first converted into a 98-vertex mesh feature through upsampling, and then the subsequent reconstruction is performed; the input of the second-layer reconstruction network is the 98-vertex feature, which becomes a 195-vertex feature after upsampling, and so on.

[0030] Furthermore, in step S1, the step of inputting the RGB image of the hand into the convolutional neural network to obtain deep hand features and a two-dimensional joint heat map may include: inputting the RGB image of the hand into the convolutional neural network to extract deep hand features; and inputting the deep hand features into the convolutional neural network to predict a two-dimensional joint heat map. In this embodiment, the input is a single RGB image of the hand I∈R 224×224×3, we can use the residual network ResNet as the feature extractor and construct the Stacked Hourglass Networks (SHN) according to the stacked hourglass structure framework to extract the hand features and predict the joint heat map. The deep hand features extracted by the residual module in the model (N0 represents the number of hand mesh vertices, C represents the number of feature channels) is used for subsequent 3D mesh reconstruction, and the entire network finally predicts and outputs 21 2D hand joint heat maps H by inputting the RGB image of the hand and the deep hand features. j ∈R 224×224 (J=1,2,…,21 represents the joint number), the pixel value of the heat map indicates the confidence of the joint position, and this process can be expressed as: The formula here indicates that the deep hand features are extracted by the residual module, and the two-dimensional joint heat map is obtained by passing the RGB image of the hand through the complete SHN network.

[0031] In this embodiment, a neural network using the SHN architecture can be viewed as a stack of multiple neural network modules. With all the network modules stacked together, an RGB hand image is fed into the multi-layered neural network. The penultimate level of the stacked structure outputs deep hand features, which are then passed through the final neural network level to produce a two-dimensional hand joint heatmap. All network outputs are related to the RGB hand image input. This means that the entire network predicts a two-dimensional hand joint heatmap based on the RGB hand image input, or that the final neural network level in the stacked architecture predicts a two-dimensional hand joint heatmap based on the deep hand features input.

[0032] Among them, when performing hand feature extraction, the hand image feature extraction part can be replaced by a convolutional neural network backbone of any structure and any type. Commonly used network structures include ResNet, Densenet, VGG, EfficienNet, etc.

[0033] Furthermore, in step S2, the extraction of hand posture features based on the two-dimensional joint heat map may include: obtaining a hand distribution map based on the two-dimensional joint heat map; determining a hand distribution area based on the value range of all joint coordinates in the hand distribution map; and performing a pooling operation on the hand distribution area to complete the extraction of hand posture features. The hand distribution area may be used as a RoI-Box area (i.e., a region of interest, represented by coordinates), and the RoI-Align pooling operation may be used on the area to complete the extraction of hand posture features. The RoI-Align algorithm may pool the pixels within the region of interest and downsample them to a feature map of lower resolution.

[0034] In some embodiments, obtaining a hand distribution map based on the two-dimensional joint heat map may include: retaining the pixel with the highest confidence in each of the two-dimensional joint heat maps and setting the remaining pixels to zero to obtain a single joint distribution map. The zero setting means taking the pixel value as 0; summing up all the joint distribution maps to obtain the hand distribution map H hand ∈ R 224×224 .

[0035] in,

[0036]

[0037] In the above formula h jxy They represent the hand distribution map of the J-th joint and the pixel with coordinates (x, y) in the heat map respectively.

[0038] In this embodiment, the data volume of a single joint distribution map obtained by retaining the pixel with the highest confidence and setting the remaining pixels to zero is less than the data volume of 21 two-dimensional joint heat maps of the human hand (that is, all two-dimensional joint heat maps, taking a total of 21 as an example), and it can still indicate the distribution of human hand joints. If the two-dimensional joint heat maps of the human hand are directly summed to obtain a single human hand distribution map to reduce the data volume, the resulting human hand distribution map cannot indicate the distribution of human hand joints. A two-dimensional joint heat map of the human hand has one for each joint. The pixel value of each map represents the probability value of the pixel being the corresponding joint. The pixel with the largest value is most likely to be a joint, but the probability values ​​of many pixels around the maximum value are also relatively high. Therefore, a large area around the location of the corresponding joint in the two-dimensional joint heat map has a certain value (in fact, the pixel with the largest value is most likely to be the actual location of the joint). Generally speaking, the pixel with the largest value in the center is most likely to be the actual location of the joint. Because the data of 21 heat maps is too large, if all the two-dimensional joint heat maps are directly added together, large areas of pixel values ​​in different two-dimensional joint heat maps may overlap, and the location of the joint cannot be distinguished after superposition. In the hand distribution map obtained using the above method, only the largest point in each map has a value, while the others are all zero. Therefore, when adding together, the zero pixel cannot affect the distribution position of other joints. Only the pixel values ​​of the nodes are not zero. Therefore, a single map can indicate the position of 21 joints.

[0039] Furthermore, preferably, the two vertices (x min ,y min ) and (x max ,y max) to define the location of the RoI-Box area, and use the RoI-Align pooling operation on the area to finally obtain the human hand posture feature F pose ∈R 15×15 This process can be expressed as:

[0040] F pose =RoIAlign(RoIBox(H hand ,((x min ,y min ),(x max ,y max )))).

[0041] In this embodiment, the RoI-Box region is located based on the maximum and minimum values ​​of all joint coordinates in the hand distribution map. This minimizes the impact of irrelevant information within the RoI-Box region. The region of interest should be as small as possible while still encompassing the hand joints. Of course, in other embodiments, the RoI-Box region can be larger, encompassing more pixels. The corresponding values ​​would need to be smaller than the minimum joint coordinate value and larger than the maximum joint coordinate value.

[0042] See also Figure 2 As shown, in some embodiments, in step S3, gradually upsampling the deep human hand features and fusing them with the human hand posture features until a three-dimensional human hand mesh model with a preset number of vertices is reconstructed may include the following steps:

[0043] S301: Up-sampling the deep hand feature and fusing it with the hand posture feature to obtain a hand mesh feature.

[0044] S302: Predicting a rough three-dimensional hand grid coordinate vector according to the hand grid feature, and performing grid shape optimization to obtain an optimized three-dimensional hand grid coordinate vector.

[0045] S303: Repeat steps S301 to S302 until a three-dimensional hand mesh model with a preset number of vertices is reconstructed.

[0046] In this embodiment, after the rough three-dimensional human hand mesh coordinate vector is predicted, the mesh shape of the three-dimensional rough human hand mesh is optimized. After the optimization, the coordinates of the reconstructed three-dimensional human hand mesh model are more accurate.

[0047] Among them, upsampling the deep human hand features can gradually increase the number of mesh vertices. Specifically, this embodiment takes four upsampling as an example. In the process of four-level reconstruction, the number of mesh vertices can be increased to 98, 195, 389, and 778 respectively. Then, the human hand posture feature F is connected in a feature connection manner.pose Propagate to each vertex feature The hand grid features fused with posture information are obtained This process can be expressed as:

[0048]

[0049] Furthermore, in step S302, the step of predicting a rough three-dimensional hand grid coordinate vector based on the hand grid features and performing grid shape optimization to obtain an optimized three-dimensional hand grid coordinate vector may include the following steps: constructing a reconstruction network ISM using the neighborhood sorting-based spatial grid convolution method SpiralConv++; inputting the hand grid features into the reconstruction network to predict a rough three-dimensional hand grid coordinate vector. And the two-dimensional hand mesh vertex coordinate vector This process can be expressed as: Then, the rough three-dimensional hand mesh coordinate vector can be optimized based on the two-dimensional hand mesh vertex coordinate vector and the hand mesh features to obtain an optimized three-dimensional hand mesh coordinate vector. In this embodiment, a neighborhood-ordered spatial mesh convolution method is used to construct the reconstruction network ISM. Of course, in other embodiments, other methods can also be used to achieve the same effect, such as using a graph convolution layer to predict the hand mesh vertex coordinates, or using a multi-layer perceptron (MLP) to predict the hand mesh vertex coordinates. At the same time, using a neighborhood-ordered spatial mesh convolution layer alone can also achieve the function of inputting hand features and predicting vertex coordinates.

[0050] In some optional embodiments, the optimizing of the rough three-dimensional hand grid coordinate vector to obtain the optimized three-dimensional hand grid coordinate vector according to the two-dimensional hand grid vertex coordinate vector and the hand grid feature may include: firstly, the two-dimensional hand grid vertex coordinate vector F 2D With the human hand grid feature Fusion, establish non-local connections between network vertices; after fusion, the self-attention module can be used to obtain the non-local association between each grid vertex in the three-dimensional space and obtain features containing contextual information This process can be expressed as: Then, the 3D offset correction grid coordinate vector can be obtained based on the feature prediction of non-local connection; Among them, the spatial grid convolution operation SpiralConv++ based on neighborhood sorting can be used to construct a single-layer 3D grid vertex regressor, and then according to the feature Predict a 3D offset-corrected grid coordinate vector Finally, the three-dimensional offset correction grid coordinate vector and the rough three-dimensional human hand grid coordinate vector F coarse3D Add together to get the optimized three-dimensional hand grid coordinate vector This process can be expressed as: In this embodiment, a shift-corrected mesh is predicted by learning the non-local connections between vertices of the hand mesh to improve the shape accuracy of the reconstructed hand model.

[0051] Preferably, after predicting the coarse 3D hand grid coordinate vector based on the hand grid features and performing grid shape optimization to obtain the optimized 3D hand grid coordinate vector, the process may further include fusing the optimized 3D hand grid coordinate vector with the deep hand features. That is, in this embodiment, after optimization, the vector is fused again with the deep hand features to form new deep hand features for the next level of hand reconstruction.

[0052] The four-level reconstruction process is as follows:

[0053] The first level directly uses the deep hand features before the reconstruction network, then fuses the hand posture features with the deep hand features to obtain the hand mesh features. The coarse three-dimensional hand mesh coordinate vector is then optimized and fused with the deep hand features as the input of the next level of deep hand features; and the three-dimensional hand mesh coordinate vector of 98 vertices is predicted.

[0054] The second level uses the deep hand features fused in the first level. The deep hand features are upsampled to increase the number of vertices to 195. The hand posture features are then fused with the deep hand features to obtain the hand mesh features. The coarse three-dimensional hand mesh coordinate vector is then optimized and fused with the deep hand features as the input for the next level of deep hand features. The three-dimensional hand mesh coordinate vector of 195 vertices is then predicted.

[0055] The third level uses the deep hand features fused in the second level. The deep hand features are upsampled to increase the number of vertices to 389. The hand posture features are then fused with the deep hand features to obtain the hand mesh features. The coarse three-dimensional hand mesh coordinate vector is then optimized and fused with the deep hand features as the input for the next level of deep hand features. The three-dimensional hand mesh coordinate vector of 389 vertices is then predicted.

[0056] The fourth stage uses the deep hand features fused in the third stage. These features are upsampled to increase the number of vertices to 778. The hand pose features are then fused with the deep hand features to create a hand mesh. The coarse 3D hand mesh coordinate vectors are then optimized and the 778-vertex 3D hand mesh coordinate vectors are predicted (completely reconstructing the mesh). Throughout the reconstruction process, the deep hand features continuously change due to the integration of other features.

[0057] This method proposes an efficient hand joint encoding method to extract hand posture information, and fully integrates it with deep hand features to improve the accuracy of reconstructed posture. At the same time, the accuracy of the reconstructed hand model is further improved by establishing connections between hand mesh vertices; the single-view three-dimensional hand reconstruction method based on posture feature fusion and hand shape optimization can better obtain hand posture and shape information, and further obtain more accurate and robust three-dimensional hand reconstruction results.

[0058] An embodiment of the present invention further provides a computer device comprising a processor and a memory, wherein the memory stores at least one program code, and the program code is loaded and executed by the processor to implement the above-mentioned single-view three-dimensional human hand reconstruction method.

[0059] An embodiment of the present invention further provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores at least one program code, and the program code is loaded and executed by a processor to implement the above-mentioned single-view three-dimensional human hand reconstruction method.

[0060] Computer-readable storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes RAM, ROM, EPROM, EEPROM, flash memory or other solid-state storage technologies, CD-ROM, DVD or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices.

[0061] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0062] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0063] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.

[0064] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.

[0065] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0066] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0067] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A single-view 3D human hand reconstruction method, characterized in that: It includes the following steps: Input the RGB image of the hand into the convolutional neural network to obtain deep hand features and two-dimensional joint heat map; Extracting hand posture features according to the two-dimensional joint heat map; Gradually upsampling the deep human hand features and fusing them with the human hand posture features until a three-dimensional human hand mesh model with a preset number of vertices is reconstructed; The step of gradually upsampling the deep human hand features and fusing them with the human hand posture features until a three-dimensional human hand mesh model with a preset number of vertices is reconstructed comprises: Upsampling the deep hand features and fusing them with the hand posture features to obtain hand mesh features; Predicting a rough three-dimensional hand grid coordinate vector based on the hand grid features, and performing grid shape optimization to obtain an optimized three-dimensional hand grid coordinate vector; Repeat the above steps until a three-dimensional human hand mesh model with a preset number of vertices is reconstructed; The method of predicting a rough three-dimensional hand grid coordinate vector based on the hand grid features and performing grid shape optimization to obtain an optimized three-dimensional hand grid coordinate vector includes: The reconstruction network is constructed using the spatial grid convolution method based on neighborhood sorting; Inputting the hand mesh features into the reconstruction network to predict a rough three-dimensional hand mesh coordinate vector and a two-dimensional hand mesh vertex coordinate vector; Optimizing the rough three-dimensional hand grid coordinate vector according to the two-dimensional hand grid vertex coordinate vector and the hand grid feature to obtain an optimized three-dimensional hand grid coordinate vector; The step of optimizing the rough three-dimensional hand grid coordinate vector according to the two-dimensional hand grid vertex coordinate vector and the hand grid feature to obtain an optimized three-dimensional hand grid coordinate vector includes: fusing the vertex coordinate vectors of the two-dimensional hand mesh with the hand mesh features to establish non-local connections between network vertices; The three-dimensional offset correction grid coordinate vector is obtained based on the feature prediction of the non-local connection; The three-dimensional offset correction grid coordinate vector is added to the rough three-dimensional human hand grid coordinate vector to obtain an optimized three-dimensional human hand grid coordinate vector.

2. The single-view 3D human hand reconstruction method according to claim 1, wherein: The RGB image of the human hand is input into the convolutional neural network to obtain deep human hand features and two-dimensional joint heat maps, including: Input the RGB image of the hand into the convolutional neural network to extract deep hand features; The deep human hand features are input into a convolutional neural network to predict a two-dimensional joint heat map.

3. The single-view 3D human hand reconstruction method according to claim 1, wherein: The extracting of hand posture features according to the two-dimensional joint heat map includes: Obtaining a hand distribution map according to the two-dimensional joint heat map; Determining a hand distribution area according to the value ranges of all joint coordinates in the hand distribution map; A pooling operation is performed on the hand distribution area to complete the hand posture feature extraction.

4. The single-view 3D human hand reconstruction method according to claim 3, wherein: The obtaining of a hand distribution map according to the two-dimensional joint heat map includes: The pixel with the highest confidence in each of the two-dimensional joint heat maps is retained, and the remaining pixels are set to zero to obtain a single joint distribution map; The hand distribution map is obtained by summing up all the joint distribution maps.

5. The single-view 3D human hand reconstruction method according to claim 1, wherein: After predicting a rough three-dimensional hand grid coordinate vector based on the hand grid features and performing grid shape optimization to obtain an optimized three-dimensional hand grid coordinate vector, the method further includes: The optimized three-dimensional hand grid coordinate vector is fused with the deep hand feature.

6. A computer device, characterized in that: The computer device includes a processor and a memory, wherein at least one program code is stored in the memory, and the program code is loaded and executed by the processor to implement the single-view three-dimensional human hand reconstruction method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one program code, and the program code is loaded and executed by the processor to implement the single-view three-dimensional human hand reconstruction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Three-dimensional human body posture estimation method based on feature fusion and sample enhancement

    CN111428586A

  • Image-based step-by-step generation type human body reconstruction method and device

    CN114419277A