A method, device and storage medium for generating a hand image based on a single-frame input image

By extracting hand identity features and geometric models from single-frame input images, optimizing feature vectors and geometric models, and performing Gaussian sputtering rendering, the information loss and geometric deformation problems generated by hand images in the prior art are solved, and high-quality hand 360-degree view rendering is achieved.

CN119006702BActive Publication Date: 2025-05-06SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411006404.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-05-06
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-quality realistic hand images from single-frame input images, especially in the reconstruction of the interactive area inside the hand and between the hand, and the problems of complex geometric deformation.

Method used

By obtaining the input image and target hand posture, processing the image to obtain hand identity feature maps and geometric models, optimizing feature vectors and geometric models to enhance the reconstruction effect of the interactive area, and performing Gaussian sputtering rendering to generate high-quality hand rendering results.

Benefits of technology

High-quality hand 360-degree rendering is achieved, which can maintain gesture consistency, enhance the reconstruction effect of hand interactive areas, improve rendering quality, and overcome the limitations of existing hand reconstruction methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119006702B_ABST
    Figure CN119006702B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and storage medium for generating a hand image based on a single-frame input image. The hand image generation method includes the steps of processing the input image to obtain a first hand identity feature map, reconstructing the hand geometry of the target hand posture to obtain a first hand geometry model, processing the first hand identity feature map and the first hand geometry model to obtain a hand feature map, processing and optimizing the hand feature map to obtain a second hand feature vector, optimizing the first hand geometry model to obtain a second hand geometry model, rendering according to the second hand feature vector and the second hand geometry model to obtain a hand rendering result, etc. The present invention can better utilize the data-based hand prior, and enhance the reconstruction effect of the hand interaction area, improve the image rendering quality of the intra-hand and inter-hand interaction areas, and overcome the limitations of the existing hand reconstruction methods. The present invention is widely used in the field of computer vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device and storage medium for generating a hand image based on a single-frame input image. Background Art

[0002] Recent advances in 3D reconstruction and differentiable rendering techniques have greatly improved the creation of 3D hand models and related applications. However, creating an interactive virtual model of a hand from a single image remains challenging. The limited input views do not provide enough geometric and texture information for accurate reconstruction. In addition, the interactions within and between hands exacerbate information loss and introduce complex geometric deformations. To address these issues, early methods relied on parametric geometric models of the hand (e.g., MANO) for geometric modeling and utilized UV maps, vertex colors, or image space rendering for appearance modeling. Although these methods are efficient in rendering, they fail to achieve realistic rendering results due to the combination of coarse mesh resolution and simple hand appearance and geometry. Summary of the invention

[0003] In view of the technical problems existing in the hand image generation technology, the object of the present invention is to provide a hand image generation method, device and storage medium based on a single frame input image.

[0004] On the one hand, an embodiment of the present invention includes a method for generating a hand image according to a single frame input image, and the method for generating a hand image according to a single frame input image includes the following steps:

[0005] Get input image and target hand pose;

[0006] Processing the input image to obtain a first hand identity feature map;

[0007] Performing hand geometry reconstruction on the target hand posture to obtain a first hand geometry model;

[0008] Processing the first hand identity feature map and the first hand geometric model to obtain a hand feature map;

[0009] Processing the hand feature map to obtain a first hand feature vector;

[0010] detecting an interaction area in the first hand geometric model;

[0011] Optimizing a portion of the first hand feature vector corresponding to the interaction area to obtain a second hand feature vector;

[0012] Optimizing the first hand geometric model according to the second hand feature vector to obtain a second hand geometric model;

[0013] Rendering is performed according to the second hand feature vector and the second hand geometric model to obtain a hand rendering result; the hand rendering result is used as a hand image.

[0014] Furthermore, the step of processing the input image to obtain a first hand identity feature map includes:

[0015] Establish a two-dimensional convolutional mapping network;

[0016] Inputting the input image into the two-dimensional convolutional mapping network for processing;

[0017] An output result of the two-dimensional convolutional mapping network is obtained as the first hand identity feature map.

[0018] Furthermore, the processing of the first hand identity feature map and the first hand geometric model to obtain a hand feature map includes:

[0019] Establish identity feature encoding network, geometric feature encoding network and feature decoding network;

[0020] Using the identity feature encoding network to encode the first hand identity feature map to obtain hand appearance information;

[0021] Encoding the first hand geometric model using the geometric feature encoding network to obtain hand geometric information;

[0022] Decoding the hand appearance information using the feature decoding network to obtain a neural texture map;

[0023] Decoding the hand geometry information using the feature decoding network to obtain a geometry feature map;

[0024] The neural texture map and the geometric feature map are merged to obtain the hand feature map.

[0025] Furthermore, the processing of the hand feature map to obtain a first hand feature vector includes:

[0026] According to the target hand posture, obtaining hand geometric surface points;

[0027] Perform bilinear interpolation on the hand feature map and the hand geometric surface points to obtain the first hand feature vector.

[0028] Further, the detecting the interaction area in the first hand geometric model includes:

[0029] Acquire a non-interaction hand geometric model corresponding to the first hand geometric model;

[0030] Calculating a difference in adjacent points between the first hand geometric model and the non-interaction hand geometric model;

[0031] Set a difference threshold;

[0032] An area in the first hand geometric model where the difference between the adjacent points is greater than the difference threshold is determined as the interaction area.

[0033] Further, the optimizing the part of the first hand feature vector corresponding to the interaction area to obtain a second hand feature vector includes:

[0034] Build a self-attention network layer;

[0035] Inputting the portion of the first hand feature vector corresponding to the interaction area into the self-attention network layer for processing;

[0036] Obtain an output result of the self-attention network layer as the second hand feature vector.

[0037] Further, optimizing the first hand geometric model according to the second hand feature vector to obtain a second hand geometric model includes:

[0038] Establish a linear prediction layer;

[0039] Inputting the first hand geometric model and the second hand feature vector into the linear prediction layer for processing;

[0040] Obtaining an output result of the linear prediction layer;

[0041] Determining the validity of vertices in the first hand geometric model according to an output result of the linear prediction layer;

[0042] Setting effectiveness thresholds;

[0043] The vertices whose corresponding vertex validity in the first hand geometric model is less than the validity threshold are removed, and vertices are added to the positions where the corresponding vertex validity in the first hand geometric model is greater than the validity threshold to obtain the second hand geometric model.

[0044] Further, the performing rendering according to the second hand feature vector and the second hand geometric model to obtain a hand rendering result includes:

[0045] Decoding the second hand feature vector and the second hand geometric model to obtain rendering attribute information;

[0046] Gaussian sputtering rendering is performed according to the rendering attribute information to obtain the hand rendering result.

[0047] On the other hand, an embodiment of the present invention also includes a computer device, including a memory and a processor, the memory is used to store at least one program, and the processor is used to load at least one program to execute the method for generating a hand image based on a single frame input image in the embodiment.

[0048] On the other hand, an embodiment of the present invention also includes a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute the method for generating a hand image based on a single frame input image in the embodiment.

[0049] The beneficial effects of the present invention are: through the hand image generation method based on a single-frame input image in the embodiment, better use is made of the data-based hand prior, and the reconstruction effect of the hand interaction area is enhanced, which is easy to animate and quickly predict. The generated hand rendering result, that is, the hand image, is a high-quality 360-degree view of the hand and can drive the 360-degree view of the hand. For different input images, the gesture of the hand rendering result can be kept consistent with the target hand posture, which can enhance the reconstruction effect of the hand interaction area, improve the image rendering quality of the intra-hand and inter-hand interaction areas, and overcome the limitations of existing hand reconstruction methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A schematic diagram of the steps of a method for generating a hand image according to a single-frame input image in an embodiment;

[0051] Figure 2 Schematic diagram of the framework of the method for generating a hand image according to a single-frame input image in an embodiment;

[0052] Figure 3 Schematic diagram of the steps of a method for generating a hand image based on a single-frame input image including a training step in an embodiment. DETAILED DESCRIPTION

[0053] With the remarkable success of neural radiance fields, several studies have adopted models based on neural radiance fields for implicit modeling. These methods usually require scene optimization for each new identity using dense and annotated images with intrinsic and extrinsic parameters, which leads to expensive training costs. Generalizable neural radiance fields get rid of the training of each scene by leveraging image alignment features to achieve reconstruction from several or even a single frame of view. However, their reliance on image alignment features also limits their performance under large viewing angles or large pose changes. However, these methods are not suitable for application scenarios such as generating hand images from a single frame of input images, because they do not include any modules to detect and handle interactions, and cannot achieve real-time rendering.

[0054] Based on the above principle, in this embodiment, a method for generating a hand image based on a single frame input image is provided.

[0055] Reference Figure 1 , the hand image generation method according to a single frame input image comprises the following steps:

[0056] S1. Obtain input image and target hand posture;

[0057] S2. Processing the input image to obtain a first hand identity feature map;

[0058] S3. reconstructing the hand geometry of the target hand posture to obtain a first hand geometry model;

[0059] S4. Processing the first hand identity feature map and the first hand geometric model to obtain a hand feature map;

[0060] S5. Processing the hand feature map to obtain a first hand feature vector;

[0061] S6. Detecting an interaction region in the first hand geometric model;

[0062] S7. Optimizing the portion of the first hand feature vector corresponding to the interaction area to obtain a second hand feature vector;

[0063] S8. Optimizing the first hand geometric model according to the second hand feature vector to obtain a second hand geometric model;

[0064] S9. Rendering is performed according to the second hand feature vector and the second hand geometric model to obtain a hand rendering result; the hand rendering result is used as a hand image.

[0065] In this embodiment, steps S1-S9 can be executed using a computer.

[0066] In this embodiment, the process and principle of steps S1-S9 are as follows: Figure 2 shown.

[0067] In step S1, a single frame of input image can be obtained by photographing both hands of the subject. A target hand posture is set, and the target hand posture is used to represent the hand posture of the hand rendering result that is expected to be generated in the end, such as the hand movement, finger direction, etc. of the hand rendering result that is generated in the end.

[0068] In this embodiment, when executing step S2, that is, processing the input image to obtain the first hand identity feature map, the following steps may be specifically performed:

[0069] S201. Establish a two-dimensional convolutional mapping network;

[0070] S202. Input the input image to the two-dimensional convolutional mapping network for processing;

[0071] S203. Obtain an output result of the two-dimensional convolutional mapping network as a first hand identity feature map.

[0072] Reference Figure 2 , Steps S201-S203 extract features from the input image by using a two-dimensional convolutional mapping network, thereby obtaining a first hand identity feature map.

[0073] In this embodiment, when executing step S3, that is, reconstructing the hand geometry of the target hand posture to obtain the first hand geometry model, a drivable parameterized hand model can be used to reconstruct the hand geometry of the target hand posture.

[0074] In this embodiment, the drivable parameterized hand model used when executing step S3 may be MANO (Metric-Affine Hand Model). MANO is a real-time 3D hand model loader implemented based on PyTorch. Specifically, according to the target hand posture, the hand posture parameters are determined. and the average hand model , set the MANO model parameters , by running MANO, the first hand geometry model is generated. The process can be expressed as:

[0075]

[0076] in is the first hand geometry model. This provides the geometric information of the hand for subsequent steps.

[0077] In this embodiment, when executing step S4, that is, processing the first hand identity feature map and the first hand geometric model to obtain the hand feature map, the following steps may be specifically performed:

[0078] S401. Establish identity feature encoding network, geometric feature encoding network and feature decoding network;

[0079] S402. Encode the first hand identity feature map using an identity feature encoding network to obtain hand appearance information;

[0080] S403. Encode the first hand geometry model using a geometric feature encoding network to obtain hand geometry information;

[0081] S404. Decode the hand appearance information using a feature decoding network to obtain a neural texture map;

[0082] S405. Decode the hand geometry information using a feature decoding network to obtain a geometric feature map;

[0083] S406. Merge the neural texture map and the geometric feature map to obtain a hand feature map .

[0084] Reference Figure 2 In steps S402-S403, the first hand identity feature map and the first hand geometry model are encoded using the identity feature encoding network and the geometry feature encoding network, respectively, to obtain hand appearance information and hand geometry information. The first hand identity feature map provides hand appearance information for subsequent steps, and the hand appearance information is determined by the person being captured in the input image (a single shot image containing both hands).

[0085] Reference Figure 2 ,In steps S404-S405, the hand appearance information and the feature decoding network are respectively decoded to obtain the neural ,texture map and the geometric feature map.

[0086] Reference Figure 2 In step S406, the neural texture map obtained in step S404 and the geometric feature map obtained in step S405 are merged to obtain a hand feature map .

[0087] In this embodiment, by executing steps S401-S406, the input image that needs to be reconstructed is decomposed into an identity feature map, a geometric feature map, and a neural texture map. Such decomposition can provide reliable priors for posture, shape, and texture, and effectively adjust the reconstructed appearance of the hand according to the input image.

[0088] In this embodiment, when executing step S5, that is, processing the hand feature map to obtain the first hand feature vector, the following steps may be specifically performed:

[0089] S501. Obtain hand geometric surface points according to the target hand posture;

[0090] S502. Perform bilinear interpolation on the hand feature map and the hand geometric surface points to obtain a first hand feature vector.

[0091] Reference Figure 2 In step S501, the hand geometric surface point is determined according to the target hand posture obtained in step S1. .

[0092] Reference Figure 2 In step S502, the hand geometric surface points obtained in step S501 are and the hand feature map obtained in step S406 Perform bilinear interpolation to obtain the corresponding hand feature vector, which is the first hand feature vector, that is, .

[0093] In this embodiment, when executing step S6, that is, the step of detecting the interaction area in the first hand geometric model, the following steps may be specifically performed:

[0094] S601. Obtain a non-interaction hand geometric model corresponding to the first hand geometric model;

[0095] S602. Calculate the difference between the adjacent points of the first hand geometric model and the non-interaction hand geometric model;

[0096] S603. Set a difference threshold;

[0097] S604. Determine the area in the first hand geometric model where the difference between the corresponding adjacent points is greater than the difference threshold as the interaction area.

[0098] In step S601, a non-interaction hand geometric model corresponding to the first hand geometric model can be generated by setting the hand posture parameters of MANO, wherein the non-interaction hand geometric model has no intra-hand or inter-hand interaction.

[0099] In step S602, all regions in the first hand geometric model are traversed. For example, for a region in the first hand geometric model, a corresponding region can be determined in the non-interactive hand geometric model, and the difference in adjacent points between the two regions can be calculated. By traversing all groups of two corresponding regions, the difference in adjacent points between any pair of corresponding regions between the first hand geometric model and the non-interactive hand geometric model can be obtained.

[0100] In step S603, a difference threshold is set. The difference threshold can measure the size of the difference between adjacent points. For example, if the difference between adjacent points is greater than the difference threshold, then the difference between adjacent points can be determined to be large, and if the difference between adjacent points is less than the difference threshold, then the difference between adjacent points can be determined to be small.

[0101] Since if the difference in adjacent points between a region in the first hand geometric model and a corresponding region in the non-interactive hand geometric model is small, it indicates that the probability of interaction between the two regions is low; if the difference in adjacent points between a region in the first hand geometric model and a corresponding region in the non-interactive hand geometric model is large, it indicates that the probability of interaction between the two regions is high. Therefore, in step S604, the region in which the difference in corresponding adjacent points in the first hand geometric model is greater than the difference threshold is determined as the interaction region, and the region in which the difference in corresponding adjacent points in the first hand geometric model is less than the difference threshold is not determined as the interaction region.

[0102] In this embodiment, when executing step S7, that is, optimizing the portion of the first hand feature vector corresponding to the interaction area to obtain the second hand feature vector, the following steps may be specifically performed:

[0103] S701. Establish a self-attention network layer;

[0104] S702. Input the part of the first hand feature vector corresponding to the interaction area into the self-attention network layer for processing;

[0105] S703. Obtain the output result of the self-attention network layer as the second hand feature vector.

[0106] In this embodiment, an interactive perception attention module can be set to execute step S7, namely steps S701-S703.

[0107] In this embodiment, by executing steps S701-S703 and using the self-attention network layer to optimize the first hand feature vector, surface points involved in intra-hand and inter-hand interactions can be identified to enhance their features through the attention mechanism. This enhancement enables the network to better simulate geometric deformations and fine-grained textures caused by interactions (such as wrinkles). By executing steps S701-S703, the first hand feature vector can be optimized, that is, the second hand feature vector is equivalent to the optimized first hand feature vector.

[0108] In this embodiment, when executing step S8, that is, optimizing the first hand geometric model according to the second hand feature vector to obtain the second hand geometric model, the following steps may be specifically performed:

[0109] S801. Establish a linear prediction layer;

[0110] S802. Input the first hand geometry model and the second hand feature vector into a linear prediction layer for processing;

[0111] S803. Obtain the output result of the linear prediction layer;

[0112] S804. Determine the validity of vertices in the first hand geometry model according to the output result of the linear prediction layer;

[0113] S805. Set validity threshold;

[0114] S806. Remove vertices whose corresponding vertex validity in the first hand geometric model is less than the validity threshold, and add vertices to positions where the corresponding vertex validity in the first hand geometric model is greater than the validity threshold to obtain a second hand geometric model.

[0115] In this embodiment, a Gaussian refinement module may be provided to execute step S8, namely steps S801 to S806.

[0116] In steps S801-S804, by using a linear prediction layer (for example, a Multi Layer Perceptron, MLP, may be used as a linear prediction layer), the validity of each vertex in the first hand geometric model may be predicted according to the second hand feature vector.

[0117] In step S805, a validity threshold is set. The validity threshold can measure the validity of each vertex in the first hand geometric model. For example, if the validity of a vertex in the first hand geometric model is greater than the validity threshold, then the vertex can be determined to be valid. If the validity of a vertex in the first hand geometric model is less than the validity threshold, then the vertex can be determined to be invalid.

[0118] In step S806, the vertices whose corresponding vertex validity in the first hand geometric model is less than the validity threshold are invalid vertices. Invalid vertices can be considered as redundant vertices. Therefore, the invalid vertices in the first hand geometric model are removed, thereby reducing the number of vertices in the first hand geometric model and reducing unnecessary data processing.

[0119] In step S806, the positions whose corresponding vertex validity in the first hand geometric model is greater than the validity threshold are valid positions, which indicates that there should be vertices at these positions in the first hand geometric model. If there are no vertices, it means that the vertices at these positions in the first hand geometric model are too sparse, and additional vertices are added at the valid positions.

[0120] By executing step S806, invalid vertices are removed from the first hand geometric model and valid vertices are added to convert the first hand geometric model into a second hand geometric model. Compared with the first hand geometric model, the vertex distribution of the second hand geometric model is more suitable for subsequent rendering, that is, the second hand geometric model is equivalent to the optimized first hand geometric model.

[0121] In this embodiment, by executing steps S801-S806 and using the Gaussian refinement module to optimize the first hand geometry model, the rough first hand geometry model can be optimized into a more refined second hand geometry model by learning to eliminate redundant surface points and allocate additional points in complex texture and deformation areas. Specifically, the positions of the vertices in the first hand geometry model are optimized (an offset value is predicted for each vertex in the first hand geometry model), so that the optimized second hand geometry model is closer to the geometry of the target hand. By executing steps S801-S806, the image rendering quality of the intra-hand and inter-hand interaction areas can be improved, overcoming the limitations of existing hand reconstruction methods.

[0122] In this embodiment, when executing step S9, that is, the step of performing rendering according to the second hand feature vector and the second hand geometric model to obtain a hand rendering result, the following steps may be specifically performed:

[0123] S901. Decode the second hand feature vector and the second hand geometric model to obtain rendering attribute information;

[0124] S902. Perform Gaussian sputtering rendering according to the rendering attribute information to obtain a hand rendering result.

[0125] Reference Figure 2 In step S901, the second hand feature vector and the second hand geometric model are decoded to obtain rendering attribute information required for Gaussian sputtering. Specifically, the rendering attribute information includes information such as the center position corresponding to the rendering point, the covariance matrix, the transparency, and the spherical harmonic coefficients.

[0126] Reference Figure 2 In step S902, Gaussian sputtering rendering is performed according to the rendering attribute information obtained in step S901 to obtain a hand rendering result. The hand rendering result obtained in step S902 is executed as the hand image to be finally obtained. The hand image is an image generated based on the input image obtained in step S1, and has a hand posture described by the target hand posture obtained in step S1.

[0127] In this embodiment, by executing steps S901-S902, the vertices of the second hand geometric model (ie, the surface points of the second hand geometric model) are directly used for rendering, which further reduces the computing resources and running time required for rendering.

[0128] In this embodiment, the principle of steps S1-S9 is: when the number of input images is limited (for example, only a single frame of input image), the hand posture of the input image is ever-changing or even blocked, the existing hand reconstruction method for single image input often produces unrealistic hand rendering results; in order to meet these challenges, steps S1-S9 apply a hand reconstruction framework based on three-dimensional Gaussian sputtering rendering, which utilizes data-based hand priors and enhances the reconstruction effect of the hand interaction area, which is easy to animate and quickly predict. In order to better handle different hand postures, by decomposing the input image that needs to be reconstructed into an identity feature map, a geometric feature map, and a neural texture map, reliable priors can be provided for posture, shape, and texture, and the reconstructed appearance of the hand can be effectively adjusted according to the input image. In order to enhance the reconstruction effect of the hand interaction area, steps S1-S9 also use an interaction-aware attention module and an adaptive Gaussian refinement module, which improve the image rendering quality of the intra-hand and inter-hand interaction areas, overcoming the limitations of existing hand reconstruction methods.

[0129] By executing the hand image generation method according to a single-frame input image, i.e., steps S1-S9, better use is made of the data-based hand prior, and the reconstruction effect of the hand interaction area is enhanced, which is easy to animate and quickly predict. The generated hand rendering result, i.e., the hand image, is a high-quality 360-degree view of the hand and can drive the 360-degree view of the hand. For different input images, the gesture of the hand rendering result can be kept consistent with the target hand gesture; the hand gestures used in steps S1-S9 are Figure 2 The framework shown enhances the reconstruction of hand interaction regions. An interaction-aware attention module and an adaptive Gaussian refinement module improve the image rendering quality of intra-hand and inter-hand interaction regions, overcoming the limitations of existing hand reconstruction methods.

[0130] In this embodiment, refer to Figure 3 , based on executing steps S1-S9, the following steps may also be executed:

[0131] S10. Obtaining a target image corresponding to the input image;

[0132] S11. According to the target image and the hand rendering result, the loss function is used to calculate the gradient, obtain the gradient parameters and transmit them back to update the network parameters;

[0133] S12. When the network parameters converge, save the network parameters; otherwise, jump back to step S1 and restart the execution.

[0134] In step S10, a target image corresponding to the input image is obtained. The target image is an image obtained by photographing the hand of the person whose input image is captured after the person makes a corresponding gesture according to the target hand posture. The target image can also be called a real image.

[0135] In step S11, the rendering result of the hand image is compared with the target image (real image), and the loss function is used to perform gradient calculation to obtain the corresponding error loss.

[0136] In step S11, the loss functions used include L1 image reconstruction loss and VGG perceptual loss, and the overall loss is shown in formula (1):

[0137] (1)

[0138] in, Represents the hand image rendering result, represents the target image (real image), represents the L1 distance, represents the VGG perceptual loss function, and is the coefficient, It is the total loss calculated according to the loss function.

[0139] In step S11, according to the value of the loss function Perform gradient calculation, obtain gradient parameters and return them, so as to Figure 2 The network parameters of the two-dimensional convolutional mapping network, identity feature encoding network, geometric feature encoding network, decoding network, etc. are updated.

[0140] In step S12, it is determined whether the network parameters of the two-dimensional convolutional mapping network, the identity feature encoding network, the geometric feature encoding network, the decoding network, etc. have converged. Specifically, if the difference (absolute value) of the network parameters of each network before and after executing step S11 is less than the respective thresholds, then it can be determined that the network parameters have converged, otherwise it is determined that the network parameters have not converged.

[0141] In step S12, if it is determined that the network parameters have converged, then the network parameters of the two-dimensional convolutional mapping network, the identity feature encoding network, the geometric feature encoding network, the decoding network, etc. can be saved to implement the training of the two-dimensional convolutional mapping network, the identity feature encoding network, the geometric feature encoding network, the decoding network, etc.; if it is determined that the network parameters have not converged, then jump back to step S1, obtain a new input image, target hand posture and target image, and restart steps S1-S12.

[0142] By executing steps S10-S12, Figure 2The two-dimensional convolutional mapping network, identity feature encoding network, geometric feature encoding network, decoding network and other networks are trained to enable it to have the ability to render hand rendering results.

[0143] A computer program that executes the method for generating a hand image based on a single-frame input image in this embodiment can be written and written into a computer device or a storage medium. When the computer program is read out and executed, the method for generating a hand image based on a single-frame input image in this embodiment is executed, thereby achieving the same technical effect as the method for generating a hand image based on a single-frame input image in the embodiment.

[0144] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. In addition, the descriptions of up, down, left, right, etc. used in the present disclosure are only relative to the relative positional relationship of the components of the present disclosure in the accompanying drawings. The singular forms of "a", "" and "the" used in the present disclosure are also intended to include the plural forms, unless the context clearly indicates other meanings. In addition, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as those generally understood by those skilled in the art. The terms used in the specification of this embodiment are only for describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this embodiment includes any combination of one or more related listed items.

[0145] It should be understood that, although the term first, second, third etc. may be adopted to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish the same type of elements from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as" etc.) provided by the present embodiment is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention.

[0146] It should be appreciated that embodiments of the present invention may be implemented or enforced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method may be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may be implemented in assembly or machine language. In any case, the language may be a compiled or interpreted language. In addition, the program may be run on a programmed dedicated integrated circuit for this purpose.

[0147] In addition, the operations of the process described in this embodiment may be performed in any suitable order, unless otherwise indicated in this embodiment or otherwise clearly contradicted by the context. The process described in this embodiment (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as codes (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or a combination thereof. A computer program includes multiple instructions that may be executed by one or more processors.

[0148] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, a RAM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or part thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above steps in conjunction with a microprocessor or other data processor, the invention of this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.

[0149] The computer program can be applied to input data to perform the functions of the present embodiment, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.

[0150] The above are only preferred embodiments of the present invention. The present invention is not limited to the above embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation methods may have various modifications and changes.

Claims

1. A method for generating a hand image based on a single frame input image, characterized in that: The method for generating a hand image according to a single-frame input image comprises: Get input image and target hand pose; Processing the input image to obtain a first hand identity feature map; Performing hand geometry reconstruction on the target hand posture to obtain a first hand geometry model; Processing the first hand identity feature map and the first hand geometric model to obtain a hand feature map; Processing the hand feature map to obtain a first hand feature vector; detecting an interaction area in the first hand geometric model; Optimizing a portion of the first hand feature vector corresponding to the interaction area to obtain a second hand feature vector; Optimizing the first hand geometric model according to the second hand feature vector to obtain a second hand geometric model; Rendering is performed according to the second hand feature vector and the second hand geometric model to obtain a hand rendering result; the hand rendering result is used as a hand image; The step of processing the input image to obtain a first hand identity feature map includes: Establish a two-dimensional convolutional mapping network; Inputting the input image into the two-dimensional convolutional mapping network for processing; An output result of the two-dimensional convolutional mapping network is obtained as the first hand identity feature map.

2. The method for generating a hand image according to a single frame input image according to claim 1, characterized in that: The processing of the first hand identity feature map and the first hand geometric model to obtain a hand feature map includes: Establish identity feature encoding network, geometric feature encoding network and feature decoding network; Using the identity feature encoding network to encode the first hand identity feature map to obtain hand appearance information; Encoding the first hand geometric model using the geometric feature encoding network to obtain hand geometric information; Decoding the hand appearance information using the feature decoding network to obtain a neural texture map; Decoding the hand geometry information using the feature decoding network to obtain a geometry feature map; The neural texture map and the geometric feature map are merged to obtain the hand feature map.

3. The method for generating a hand image according to a single frame input image according to claim 1, characterized in that: The step of processing the hand feature map to obtain a first hand feature vector includes: According to the target hand posture, obtaining hand geometric surface points; Perform bilinear interpolation on the hand feature map and the hand geometric surface points to obtain the first hand feature vector.

4. The method for generating a hand image according to a single frame input image according to claim 1, characterized in that: The detecting the interaction area in the first hand geometric model includes: Acquire a non-interaction hand geometric model corresponding to the first hand geometric model; Calculating a difference in adjacent points between the first hand geometric model and the non-interaction hand geometric model; Set a difference threshold; An area in the first hand geometric model where the difference between the adjacent points is greater than the difference threshold is determined as the interaction area.

5. The method for generating a hand image according to a single frame input image according to claim 1, characterized in that: The optimizing the part of the first hand feature vector corresponding to the interaction area to obtain a second hand feature vector includes: Build a self-attention network layer; Inputting the portion of the first hand feature vector corresponding to the interaction area into the self-attention network layer for processing; Obtain an output result of the self-attention network layer as the second hand feature vector.

6. The method for generating a hand image according to a single frame input image according to claim 1, characterized in that: The step of optimizing the first hand geometric model according to the second hand feature vector to obtain a second hand geometric model includes: Establish a linear prediction layer; Inputting the first hand geometric model and the second hand feature vector into the linear prediction layer for processing; Obtaining an output result of the linear prediction layer; Determining the validity of vertices in the first hand geometric model according to an output result of the linear prediction layer; Setting effectiveness thresholds; The vertices whose corresponding vertex validity in the first hand geometric model is less than the validity threshold are removed, and vertices are added to the positions where the corresponding vertex validity in the first hand geometric model is greater than the validity threshold to obtain the second hand geometric model.

7. The method for generating a hand image according to a single frame input image according to any one of claims 1 to 6, characterized in that: The performing rendering according to the second hand feature vector and the second hand geometric model to obtain a hand rendering result includes: Decoding the second hand feature vector and the second hand geometric model to obtain rendering attribute information; Gaussian sputtering rendering is performed according to the rendering attribute information to obtain the hand rendering result.

8. A computer device, characterized in that: It includes a memory and a processor, the memory is used to store at least one program, and the processor is used to load at least one program to execute the hand image generation method based on a single-frame input image as described in any one of claims 1-7.

9. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor executable program is used to execute the hand image generation method based on a single frame input image as described in any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Real-time reconstruction method and device for interaction process between hand and object based on physics

    CN115239906A

  • Lightweight hand reconstruction and driving method

    CN117994480A