Gaussian Hand-Object Interaction Rendering Denoising Method
By initializing, reconstructing, and pre-training Gaussian representations of hand poses, and combining diffusion models and geometric perception methods, the problems of hand pose distortion and clipping in virtual reality and augmented reality are solved, achieving high-precision hand-object interaction rendering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2025-08-07
- Publication Date
- 2026-06-30
AI Technical Summary
In virtual reality and augmented reality applications, when capturing hand poses using 3D motion tracking technology, the motion capture results often contain noise that causes hand pose distortion and clipping. Furthermore, Gaussian representations are difficult to reasonably construct the relationship between the hand and the object, and cannot remove erroneous pose inputs.
By performing Gaussian initialization and reconstruction on rigid object images, combined with Gaussian initialization and pre-training processing of the hand on the surface of a 3D parametric model, a representation of the hand bone points to the Gaussian surface is constructed. Then, the diffusion model and geometric perception method are used to denoise and correct the hand movements to eliminate clipping phenomena.
It effectively corrects hand posture distortion and clipping phenomena, improves the accuracy of hand posture and hand-object interaction, and enhances the perception of geometric relationships between the hand and objects.
Smart Images

Figure CN120931520B_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the fields of computer graphics and virtual reality, and specifically to a Gaussian hand-object interaction rendering denoising method. Background Technology
[0002] Hand-object interaction is common in daily life and plays a crucial role in virtual reality (VR) and augmented reality (AR) applications. Currently, the common method for virtual hand capture is to use 3D motion tracking technology to capture the user's real hand posture to construct a virtual hand for interacting with virtual objects.
[0003] However, when using the above method for virtual hand capture, the following technical problems often arise:
[0004] Due to limitations in tracking hardware and computational precision, when capturing hand poses using 3D motion tracking technology, the motion capture results often contain noise, leading to hand pose distortion and clipping phenomena.
[0005] Gaussian representations differ from grid data, making it difficult to reasonably construct the relationship between the hand and the object.
[0006] High-precision hand poses are difficult to obtain in conventional AR or VR scenarios, which relies on high-end camera arrays. Hand-object interaction rendering methods based on 3D Gaussian cannot perceive the geometric relationship between the hand and the object's Gaussian plane, and therefore cannot remove incorrect pose inputs.
[0007] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0008] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0009] Some embodiments of this disclosure provide methods, apparatus, electronic devices, and computer-readable media for Gaussian hand-object interactive rendering denoising to solve one or more of the technical problems mentioned in the background section above.
[0010] In a first aspect, some embodiments of this disclosure provide a Gaussian hand-object interaction rendering denoising method, the method comprising: initializing a rigid object image with an object Gaussian to obtain an initialized object Gaussian; reconstructing the initialized object Gaussian based on a Gaussian training method to obtain a reconstructed object Gaussian; initializing a hand Gaussian on the registered 3D parametric model surface to obtain a standard hand Gaussian; pre-training the standard hand Gaussian to obtain a pre-trained hand Gaussian; jointly modeling the skeletal joints of the pre-trained hand Gaussian with the surface Gaussian scatter points of the reconstructed object Gaussian to construct a hand skeletal point-to-Gaussian surface representation; denoising the hand skeletal point-to-Gaussian surface representation based on a diffusion model to obtain a denoised hand skeletal point-to-Gaussian surface representation; and detecting the interaction clipping region corresponding to the denoised hand skeletal point-to-Gaussian surface representation based on a geometry-aware clipping removal method, and recalibrating the hand movement based on the interaction clipping region to obtain a clipped hand-object interaction image.
[0011] Secondly, some embodiments of this disclosure provide a Gaussian hand-object interaction rendering denoising apparatus, the apparatus comprising: an object Gaussian initialization unit configured to initialize the object Gaussian of a rigid object image to obtain an initialized object Gaussian; an object Gaussian reconstruction unit configured to reconstruct the initialized object Gaussian based on a Gaussian training method to obtain a reconstructed object Gaussian; a hand Gaussian initialization unit configured to initialize the hand Gaussian on the surface of a registered 3D parametric model to obtain a standard hand Gaussian; and a pre-training unit configured to pre-train the standard hand Gaussian to obtain a pre-trained hand Gaussian. The system includes: a pre-trained Gaussian hand modeling unit, configured to jointly model the skeletal joints of the pre-trained Gaussian hand model with the surface Gaussian scatter points of the reconstructed object Gaussian model to construct a hand skeletal point-to-Gaussian surface representation; a denoising unit, configured to denoise the hand skeletal point-to-Gaussian surface representation based on a diffusion model to obtain a denoised hand skeletal point-to-Gaussian surface representation; and a correction unit, configured to detect interactive clipping regions corresponding to the denoised hand skeletal point-to-Gaussian surface representation based on a geometry-aware clipping removal method, and to re-correct hand movements based on these interactive clipping regions to obtain a clipped hand-object interaction image.
[0012] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processes; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processes, the one or more processes implement the method described in any implementation of the first aspect above.
[0013] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is processed and executed, it implements the method described in any of the implementations of the first aspect above.
[0014] The various embodiments disclosed above have the following beneficial effects: Through the Gaussian hand-object interaction rendering denoising method of some embodiments of this disclosure, the representation of hand skeletal points to the Gaussian surface is reasonably constructed, correcting hand pose distortion and clipping phenomena. Specifically, the causes of hand pose distortion and clipping phenomena are: due to limitations in tracking hardware and computational accuracy, when capturing hand poses using 3D motion tracking technology, the motion capture results often contain noise, leading to hand pose distortion and clipping phenomena. Gaussian representation differs from mesh data, making it difficult to reasonably construct the relationship between the hand and the object. High-precision hand poses are difficult to obtain in conventional AR / VR scenarios, relying on high-end camera arrays. The 3D Gaussian-based hand-object interaction rendering method cannot perceive the geometric relationship between the hand and the object's Gaussian surface, thus failing to remove erroneous pose inputs. Based on this, the Gaussian hand-object interaction rendering denoising method of some embodiments of this disclosure first initializes the object Gaussian representation of the rigid object image; then, based on a Gaussian training method, reconstructs the initialized object Gaussian representation; finally, initializes the hand Gaussian representation on the registered 3D parametric model surface; and pre-trains the standard hand Gaussian representation. This constructs a structured 3D Gaussian representation of the hand and object. Next, a representation of the hand's skeletal points to the object's Gaussian surface is constructed. This structurally and quantitatively fuses the driving core of the hand movement (skeleton points) with the precise geometry (Gaussian surface) constraints of the target object at the spatial relationship level, reasonably constructing a representation of the hand's skeletal points to the Gaussian surface. Then, based on a trained diffusion model, denoising is performed on the aforementioned representation of the hand's skeletal points to the object's Gaussian surface. This eliminates obvious physical errors and initially optimizes the movement. Finally, based on a geometry-aware de-clipping method, the hand movement is re-corrected. This corrects the clipping phenomenon between the hand's skeletal points and the object's Gaussian surface. This implementation method reasonably constructs the representation of hand skeletal points to a Gaussian surface, correcting hand posture distortion and clipping phenomena. Attached Figure Description
[0015] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0016] Figure 1 This is a flowchart of some embodiments of the Gaussian hand-object interactive rendering denoising method according to the present disclosure;
[0017] Figure 2These are schematic diagrams of the structure of some embodiments of the Gaussian hand-object interactive rendering and denoising apparatus according to the present disclosure;
[0018] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0019] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0020] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0022] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0023] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0024] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0025] Figure 1 The flowchart 100 illustrates some embodiments of the Gaussian hand-object interactive rendering denoising method according to the present disclosure, which includes the following steps:
[0026] Step 101: Initialize the rigid object image with an object Gaussian to obtain the initialized object Gaussian.
[0027] In some embodiments, the execution entity (e.g., a server) of the Gaussian hand-object interactive rendering denoising method can initialize the object Gaussian on the rigid object image to obtain the initialized object Gaussian.
[0028] In practice, the aforementioned execution entity can perform 3D reconstruction based on a set of rigid object images to obtain an initialized Gaussian object. The algorithm for this 3D reconstruction can be a Structure from Motion (SFM) algorithm.
[0029] Step 102: Based on the Gaussian training method, reconstruct the Gaussian object after initialization to obtain the reconstructed Gaussian object.
[0030] In some embodiments, the aforementioned execution entity may reconstruct the Gaussian object based on the Gaussian training method to obtain the reconstructed Gaussian object.
[0031] In practice, the aforementioned execution entity can reconstruct the Gaussian object from the initialized object using the following steps based on the Gaussian training method, thereby obtaining the reconstructed Gaussian object:
[0032] The first step is to perform Gaussian rendering on the initialized objects based on the camera poses corresponding to the aforementioned set of rigid object images, thereby obtaining a rendered image set. In practice, firstly, the executing entity can adjust the Gaussian distribution in three-dimensional space to a direction that matches the camera poses corresponding to the aforementioned set of rigid object images, and project it onto a two-dimensional plane to obtain a two-dimensional plane projection set; then, the executing entity can generate the rendered image set using rasterization technology.
[0033] The second step involves determining the loss between the rendered image set and the corresponding real image set, and optimizing the Gaussian parameters based on the gradient descent algorithm and the aforementioned loss to obtain the reconstructed object Gaussian. In practice, the execution entity can determine the gradient of the loss with respect to the Gaussian parameters and adjust the Gaussian parameters along the direction of gradient reduction to obtain the reconstructed object Gaussian. The aforementioned loss can be an L2 loss function, which can be implemented using the following formula:
[0034]
[0035] in, This represents the L2 training loss function. N represents the total number of pixels in the image. j represents the pixel index. y j This represents the actual color value of the j-th pixel. This represents the rendered color value of the j-th pixel.
[0036] Step 103: Initialize the Gaussian hand shape on the registered 3D parametric model surface to obtain the standard Gaussian hand shape.
[0037] In some embodiments, the aforementioned execution entity can initialize the Gaussian hand on the registered three-dimensional parametric model (Metric-AffineHand Model, MANO) surface to obtain a standard Gaussian hand.
[0038] In practice, the aforementioned execution entity can project the Gaussian hand image onto the surface of the registered 3D parametric model after filtering or noise reduction to obtain a standard Gaussian hand image.
[0039] Step 104: Perform pre-training processing on the standard hand Gaussian to obtain the pre-trained hand Gaussian.
[0040] In some embodiments, the execution entity may pre-train the standard hand Gaussian to obtain a pre-trained hand Gaussian.
[0041] In practice, the aforementioned implementing entity can pre-train the standard hand Gaussian using the following steps to obtain a pre-trained hand Gaussian:
[0042] The first step is to label the hand pose using a hand pose estimation method based on the multi-view images, obtaining the labeled hand pose. In practice, firstly, the aforementioned execution entity can use a detection model to locate the hand region and output a hand region ROI (Region of Interest) image; secondly, the aforementioned execution entity can use a preset detection method to detect the predicted 2D positions of hand joints in the ROI image; finally, the aforementioned execution entity can use triangulation to reconstruct the 3D pose of the aforementioned hand joints. The aforementioned detection model can be YOLOv8-Hand or MediaPipeHands, and the aforementioned preset detection method can be HRNet or HigherHRNet. The triangulation formula is as follows:
[0043]
[0044] Where q represents the camera view index. Q represents the total number of camera views. k represents the joint index, ranging from 1 to the total number of skeletal joints n. b π() represents the camera projection function. K q Let R represent the intrinsic parameter matrix of the q-th camera viewpoint. q |t q [] represents the extrinsic parameters of the q-th camera viewpoint. P represents the coordinates of a homogeneous 3D point. This represents the two-dimensional pixel coordinates of joint k as observed from the q-th camera viewpoint. Represents the 3D pixel coordinates of joint k. || || represents taking the absolute value. argmin() represents the input value that minimizes the objective function. ∑ represents summation.
[0045] The second step involves transforming the standard hand Gaussian pose to pose space using a rigid kinematic chain, resulting in the pose-oriented hand Gaussian pose. In practice, this transformation can be performed using the following formula:
[0046]
[0047] Where θ represents the hand posture parameter. This represents the skeleton transformation function. This represents the bone transformation corresponding to pose θ. (B) k Let represent the rigid body transformation matrix of the k-th joint. Let i represent the index of the Gaussian distribution. c Represents the standard pose space. Represents the standard Gaussian model of the hand The positional attribute of the i-th Gaussian distribution in x. p Represents the target attitude space. This represents the position of the i-th Gaussian distribution in the target attitude space after deformation. This indicates the skin weight. Indicates position Skin weights for the k-th joint. This represents a weighted summation over all joints.
[0048] The third step involves rendering the Gaussian hand image of the aforementioned pose and training it by comparing it with the multi-view images to obtain a pre-trained Gaussian hand image. In practice, firstly, the executing agent can adjust the Gaussian hand image in 3D space to match the camera pose corresponding to the multi-view images and project it onto a 2D plane to obtain a 2D projection set of the hand image; then, the executing agent can generate a set of rendered hand images using rasterization techniques; finally, the executing agent can determine the gradient of the L2 loss with respect to the parameters of the Gaussian hand image and adjust the Gaussian hand image parameters along the direction of decreasing gradient to obtain the pre-trained Gaussian hand image.
[0049] Step 105: Jointly model the pre-trained hand Gaussian skeletal joints with the reconstructed object Gaussian surface Gaussian scatter points to construct a hand skeletal point to Gaussian surface representation.
[0050] In some embodiments, the aforementioned execution entity may jointly model the pre-trained Gaussian skeletal joints of the hand with the reconstructed Gaussian surface Gaussian scatter points through the following steps to construct a representation of hand skeletal points to a Gaussian surface:
[0051] The first step is to determine the dynamic curvature of the reconstructed Gaussian object based on a dynamic curvature sensing mechanism. In practice, the curvature weight κ of the reconstructed Gaussian object is... geoRepresented as:
[0052] k geo ={k(coord o )|p o ∈select()}.
[0053]
[0054] Where O represents the Gaussian point index of the object's Gaussian point. coord o This represents the coordinates of the Gaussian point O of the object. coord represents the coordinates of the object. o The neighborhood of κ(coord). o ) indicates coord o The curvature weight corresponding to N(coord). o ) indicates coord o The neighborhood set of n. coord indicates coord o The normal vector at coord. indicates coord o The normal vector at that location. coord indicating correction o The normal vector at that location. This represents the gradient of the symbolic distance field. Indicates the surface sharpness enhancement factor (e.g., ).
[0055] The second step, based on the aforementioned dynamic curvature weights, involves selecting an anchor point Gaussian subset from the reconstructed object Gaussian data through curvature-weighted sampling. In practice, this anchor point Gaussian subset... The anchor points are selected from the Gaussian distribution of the spatiotemporal trajectory near the hand skeleton using curvature-weighted sampling, and can be obtained using the following formula:
[0056]
[0057] Where, n a Indicates the number of anchor points. t c t represents the threshold. c =10mm. This represents the sequence of skeletal joints. The object is represented by Gaussian. This represents the union trajectory of the skeletal joint sequences. ||2 represents the Euclidean distance. `select()` is a conditional selection operator that outputs the subset that meets the given conditions. `.weighted` sample ( ) indicates the application of a weighted sampling function.
[0058]
[0059] Among them, weighted sample ( ) denotes the weighted sampling function. m represents the anchor index, ranging from 1 to n. a W() represents the weighted sampling probability corresponding to the weighted sampling function. τ represents the weighting intensity coefficient (e.g., it can be 0.3 or 1.0). a represents the index of the candidate Gaussian point. A represents the total number of candidate Gaussian points.
[0060] The third step is to construct a parametric joint camera representation from the pre-trained hand Gaussian image, thus obtaining the parametric joint camera representation:
[0061]
[0062] Where p0 represents the camera position (joint position), and p1 represents the target point (anchor Gaussian position). lookat(p0, p1) generates a rotation matrix that points the Z-axis from p0 to p1. -1 This indicates matrix inversion. `camera(p0, p1)` represents the camera extrinsic parameters represented by `p0` and `p1`. `F` represents the number of frames in the sequence. `f` represents the frame index, ranging from 1 to `F`. `m` represents the anchor index, ranging from 1 to `n`. a . This represents the 3D position of the k-th joint at frame f. m ·pos represents the position attribute of the m-th anchor point. v f,m,k This indicates a camera located at the k-th joint of frame f, pointing to the m-th anchor point.
[0063] The fourth step is to perform motion consistency detection on each parameterized joint camera representation in the above parameterized joint camera representation set based on the inter-frame camera reuse mechanism, so as to obtain the joint camera set.
[0064] v f,m,k =v f+1,m,k .
[0065]
[0066] in, This represents the three-dimensional position of the k-th joint at frame f. This represents the 3D position of the k-th joint at frame f-1. Θ represents an empirical threshold, which can be 5mm. This represents a set of joint cameras.
[0067] The fifth step involves jointly modeling the pre-trained Gaussian skeletal joints of the hand with the reconstructed Gaussian surface scatter points of the object to construct a representation of the hand's skeletal points to the Gaussian surface. In practice, the aforementioned execution entity can establish a correspondence or constraint between the hand's skeletal points and the object's surface Gaussian scatter points, and identify the contact relationship between the hand's skeletal points and the object's surface Gaussian points for joint modeling. The aforementioned representation of the hand's skeletal points to the Gaussian surface is defined as follows:
[0068]
[0069] in, This represents the sequence of skeletal joints. The object is represented by Gaussian. This represents the Gaussian subset of the anchor points. Represented by skeletal joint sequence Object Gauss and anchor point Gaussian subset The skeletal points are represented to the Gaussian surface. This represents a set of articulated cameras. d ( ) represents the Gaussian depth rendering function. unproject() represents the depth map back projection function. Indicates dependence on camera The screen is projected back to world coordinates.
[0070] Step 106: Based on the diffusion model, denoise the hand bone point to Gaussian surface representation to obtain the denoised hand bone point to Gaussian surface representation.
[0071] In some embodiments, the aforementioned execution entity can denoise the hand bone point-to-Gaussian surface representation based on a diffusion model to obtain a denoised hand bone point-to-Gaussian surface representation. The diffusion model can be a standard MotionDiffusion model.
[0072] In practice, the aforementioned execution entity can denoise the hand bone points onto the Gaussian surface representation using the following steps:
[0073] The first step is to perform multimodal feature enhancement based on the above hand skeleton point-to-Gaussian surface representation and depth map to obtain enhanced hand skeleton point-to-Gaussian surface representation, wherein the depth map is acquired by an RGB-D sensor.
[0074] In sub-step one, the aforementioned execution entity can use the Sobel operator to perform gradient calculations on the depth map DP(Horiz, Vert), generate surface normal maps, and map them to RGB space:
[0075]
[0076] Where DP represents the depth value corresponding to each pixel (Horiz, Vert) in the depth map. grad x This represents the horizontal gradient of the depth map. grad y Represents the vertical gradient of the depth map. * indicates a 2D convolution operation. grad 3D represents the 3D gradient vector. f represents the normalized normal vector. F represents the surface normal mapping to RGB space. x ,f y ,f z The three components of the surface normal vector are represented.
[0077] In sub-step two, the aforementioned execution entity can project the coordinates of the hand bone points to the Gaussian surface representation of the bone points onto the two-dimensional image space of the normal mapping, and perform bilinear interpolation sampling on each projection point to obtain the sampled normal mapping.
[0078]
[0079] ω rs =(1-|Horiz k -r|)(1-|Vert k -s|).
[0080]
[0081] Among them, Horiz k Vert k This represents the two-dimensional coordinates of the projection of the k-th hand bone point onto the Gaussian surface representing the bone point. Π() represents the perspective projection matrix, determined by the camera intrinsic parameters. k y k , z k ω represents the coordinates of the k-th hand bone point relative to the Gaussian surface representation of the bone point. rs This represents the interpolation weights. r and s represent the indices of the two-dimensional coordinates of the projection of the k-th hand bone point onto the Gaussian surface representation of the bone point, where r ∈ Horiz k ,s∈Vert k f k This represents the normal mapping of the sampled data.
[0082] Sub-step three involves concatenating the coordinates of the hand bone points represented by the Gaussian surface with the sampled normal mapping in the channel dimension to obtain the enhanced hand bone point-to-Gaussian surface representation. Here, the coordinates of the hand bone points represented by the Gaussian surface representation are expressed as the three-dimensional coordinates x of k bone points. k y k , z k
[0083]
[0084] in, This represents the enhanced hand skeletal points mapped to a Gaussian surface. k y k , z k This represents the coordinates of the k-th hand bone point to the bone point represented by the Gaussian surface. This represents the three-dimensional normal information from the k-th hand bone point to the bone point k represented by the Gaussian surface.
[0085] The second step involves using a diffusion model to diffuse the noise-free hand-object interaction representation into a whitened noise space, resulting in a noisy hand-object interaction representation. In practice, the aforementioned executing entity can use the following formula to diffuse the noise-free hand-object interaction representation into a whitened noise space, thus obtaining a noisy hand-object interaction representation:
[0086]
[0087] Where e represents the time step, e = 1, 2, ..., E. e This represents the hand-object interaction at time step e. When e = 1, x e-1 This represents a noise-free representation of hand-object interaction. β e β represents the noise scheduling parameter of e at time step e. e ∈(0, 1). ∈ e This represents standard Gaussian noise (independent noise with a mean of 0 and a covariance equal to the identity matrix).
[0088] The third step involves training the diffusion model based on the noiseless hand-object interaction representation and the noisy hand-object interaction representation described above, resulting in the trained diffusion model. In practice, firstly, the executing agent can use the noiseless hand-object interaction representation x0 and the hand-object interaction representation x at time step e. e The input is fed into the diffusion model, which outputs predicted noise. Secondly, the aforementioned execution entity can determine the mean squared error loss between the predicted noise and the actual noise. Finally, the aforementioned execution entity can optimize the model parameters through backpropagation to obtain the trained diffusion model.
[0089] The fourth step involves performing a blended fill process on the enhanced hand skeleton point-to-Gaussian surface representation to obtain an edge-enhanced hand skeleton point-to-Gaussian surface representation. In practice, for regions with continuous textures (such as skin and cloth) in the hand skeleton point-to-Gaussian surface representation, the execution subject can use a mirror fill method to complete the edges; for isolated object edges (such as fingers and tree branches), the execution subject can use a radial basis function (RBF) method to complete the edges. Finally, the execution subject obtains the edge-enhanced hand skeleton point-to-Gaussian surface representation.
[0090] Fifth step: Input the above edge-enhanced hand bone point to Gaussian surface representation into the above trained diffusion model to obtain the denoised hand bone point to Gaussian surface representation.
[0091] The first to fifth steps and related content described above, as an inventive point of this disclosure, solve the technical problem of "traditional skeletal point representation lacking the ability to physically analyze the object surface and exhibiting distortion at the model edges." Factors leading to hand pose distortion and clipping phenomena often include: traditional skeletal point representation lacking the ability to physically analyze the object surface and distortion at the model edges. Solving these factors enhances the diffusion model's ability to perceive the curvature of the object surface while reducing distortion of the open boundary contour of the hand. To achieve this effect, firstly, based on the aforementioned hand skeletal point-to-Gaussian surface representation and depth map, multimodal feature enhancement is performed to obtain an enhanced hand skeletal point-to-Gaussian surface representation. This enhances the subsequent diffusion model's ability to perceive the curvature of the object surface. Secondly, the diffusion model diffuses the noise-free hand-object interaction representation into a whitened noise space, resulting in a noisy hand-object interaction representation; based on the aforementioned noise-free and noisy hand-object interaction representations, the diffusion model is trained to obtain a trained diffusion model. This trains a diffusion model capable of eliminating noise. Third, the enhanced hand skeletal point-to-Gaussian surface representation is subjected to blending and filling processing to obtain an edge-enhanced hand skeletal point-to-Gaussian surface representation. This edge-enhanced representation is then input into the trained diffusion model to obtain a denoised hand skeletal point-to-Gaussian surface representation. This reduces the distortion of the open boundary contour of the hand. Ultimately, this enhances the diffusion model's ability to perceive the curvature of object surfaces while reducing distortion of the open boundary contour of the hand.
[0092] Step 107: Based on the geometry-aware de-clipping method, detect the interaction clipping region corresponding to the Gaussian surface representation of the denoised hand bone points, and recalibrate the hand movements based on the interaction clipping region to obtain the de-clipping hand-object interaction image.
[0093] In some embodiments, the aforementioned execution entity can detect the interaction clipping region corresponding to the Gaussian surface representation of the denoised hand skeleton points based on a geometry-aware clipping removal method, and recalibrate the hand movements based on the interaction clipping region to obtain a clipped hand-object interaction image.
[0094] In practice, the aforementioned implementers can recalibrate their hand movements through the following steps:
[0095] The first step is to extract the Euler angles of the root, mid-shaft, and distal phalanges corresponding to the five fingers, as the pre-correction Euler angle set. Based on this set, the clipping depth of each bone segment to the reconstructed Gaussian surface of the object is determined, thus obtaining the interactive clipping region. In practice, the execution entity can determine the clipping depth of each bone segment to the reconstructed Gaussian surface of the object using the following clipping depth calculation method to obtain the interactive clipping region:
[0096] Step one: The execution entity can perform spatial sampling on the reconstructed Gaussian surface of the object: extract the surface point set from the reconstructed Gaussian surface of the object.
[0097] S={S d}
[0098] Where d represents the surface point index. S d This represents the d-th surface point. Each point S... d Associative center location μ d covariance σ d and transparency α d .
[0099] Step two, the aforementioned executing entity can determine the clipping depth L from the point cloud P′ of the hand Gaussian to the object surface. u :
[0100]
[0101] Where x represents a point in three-dimensional space, and sign(x) represents the spatial position of point x. represents the Mahalanobis distance. u represents the point cloud index of the aforementioned hand Gaussian. p′ u Let p′u be the point representing the u-th hand Gaussian segment. ∈ P′. φ(x) represents the spatial distance function.
[0102] The second step involves flattening the root bone Euler angles, mid-skeletal bone Euler angles, and terminal bone Euler angles corresponding to each through-the-finger in the aforementioned interactive through-the-finger area to generate a skeletal Euler angle sequence, resulting in a skeletal Euler angle sequence set. In practice, flattening the Euler angles involves setting the pitch angle in the Euler angles to 180°. The aforementioned skeletal Euler angle sequence includes the root bone Euler angles, mid-skeletal bone Euler angles, and terminal bone Euler angles arranged in sequence.
[0103] Third, for each bone Euler angle sequence in the above set of bone Euler angle sequences, perform the following processing steps:
[0104] Step 1: Based on the first bone Euler angle in the bone Euler angle sequence, perform the following correction steps:
[0105] The first correction step is to increase the curvature of the first bone's Euler angle. In practice, the aforementioned execution entity can increase the z-axis angle ψ of the first bone's Euler angle by a preset angle. root The aforementioned preset angle is a pre-defined variation of the Euler angles of the skeleton. For example, the preset angle could be 0.5°. As another example, the preset angle could be 1°.
[0106] The second correction step involves determining the clipping depth of the first bone's Euler angle. In practice, the aforementioned execution entity can apply the clipping depth calculation method described above to determine the clipping depth.
[0107] The third correction step is to synchronously update the values in the pre-correction set of Euler angles in response to the contact between the bone corresponding to the Euler angle of the first bone, which represents the depth of the molding, and the reconstructed object Gaussian.
[0108] In the fourth correction step, in response to the empty bone Euler angle sequence after deleting the first bone Euler angle, the updated pre-correction bone Euler angle sets are used as the post-correction bone Euler angle set.
[0109] Step 2: In response to the fact that the bone Euler angle sequence after deleting the first bone Euler angle is not empty, the bone Euler angle sequence after deleting the first bone Euler angle is used as the bone Euler angle sequence, and the above correction steps are performed again.
[0110] The fourth step is to adjust the bone pose based on the corrected bone Euler angle set to obtain the Gaussian of the hand with the de-clipping pose.
[0111] The fifth step is to render the Gaussian image of the hand with the de-clipping pose and the reconstructed object Gaussian image together to obtain the de-clipping hand-object interaction image.
[0112] The first to fifth steps and related content described above, as an inventive point of this disclosure, solve the technical problem that "it is difficult to obtain high-precision hand poses in conventional AR / VR scenes, which relies on high-end camera arrays. The hand-object interaction rendering method based on 3D Gaussian cannot perceive the geometric relationship between the hand and the object's Gaussian surface, thus failing to remove erroneous pose inputs." Factors leading to hand pose distortion and clipping phenomena are often as follows: it is difficult to obtain high-precision hand poses in conventional AR / VR scenes, which relies on high-end camera arrays. The hand-object interaction rendering method based on 3D Gaussian cannot perceive the geometric relationship between the hand and the object's Gaussian surface, thus failing to remove erroneous pose inputs. If these factors are resolved, the effect of correcting hand pose distortion and clipping phenomena can be achieved. To achieve this effect, firstly, the Euler angles of the root, middle, and distal bones corresponding to the five fingers are extracted as a pre-correction Euler angle set, and based on this pre-correction Euler angle set, the clipping depth of each bone segment on the reconstructed object's Gaussian surface is determined. Thus, the clipping depth and clipping area between the hand's Gaussian surface and the object's Gaussian surface can be obtained. Second, flatten the root bone Euler angles, mid-segment bone Euler angles, and terminal bone Euler angles corresponding to each clipping finger in the aforementioned interactive clipping area to generate a bone Euler angle sequence. This ensures that the clipping finger does not touch the surface of the Gaussian object, and yields a set of bone Euler angle sequences for the clipping finger. Third, for each bone Euler angle sequence in the aforementioned set, perform the following processing steps: Step 1, based on the first bone Euler angle in the bone Euler angle sequence, perform the aforementioned correction step; Step 2, in response to the bone Euler angle sequence after deleting the first bone Euler angle not being empty, use the bone Euler angle sequence after deleting the first bone Euler angle as the bone Euler angle sequence, and perform the aforementioned correction step again. This yields a corrected set of bone Euler angles. Fourth, adjust the bone pose based on the corrected set of bone Euler angles. This yields a Gaussian hand with a clipped pose. Fifth, render the aforementioned Gaussian hand with a clipped pose and the reconstructed object Gaussian together. This yields a clipped hand-object interaction image. Ultimately, the distortion of hand posture and clipping issues were corrected.
[0113] The various embodiments disclosed above have the following beneficial effects: Through the Gaussian hand-object interaction rendering denoising method of some embodiments of this disclosure, the representation of hand skeletal points to the Gaussian surface is reasonably constructed, correcting hand pose distortion and clipping phenomena. Specifically, the causes of hand pose distortion and clipping phenomena are: due to limitations in tracking hardware and computational accuracy, when capturing hand poses using 3D motion tracking technology, the motion capture results often contain noise, leading to hand pose distortion and clipping phenomena. Gaussian representation differs from mesh data, making it difficult to reasonably construct the relationship between the hand and the object. High-precision hand poses are difficult to obtain in conventional AR / VR scenarios, relying on high-end camera arrays. The 3D Gaussian-based hand-object interaction rendering method cannot perceive the geometric relationship between the hand and the object's Gaussian surface, thus failing to remove erroneous pose inputs. Based on this, the Gaussian hand-object interaction rendering denoising method of some embodiments of this disclosure first initializes the object Gaussian representation of the rigid object image; then, based on a Gaussian training method, reconstructs the initialized object Gaussian representation; finally, initializes the hand Gaussian representation on the registered 3D parametric model surface; and pre-trains the standard hand Gaussian representation. This constructs a structured 3D Gaussian representation of the hand and object. Next, a representation of the hand's skeletal points to the object's Gaussian surface is constructed. This structurally and quantitatively fuses the driving core of the hand movement (skeleton points) with the precise geometry (Gaussian surface) constraints of the target object at the spatial relationship level, reasonably constructing a representation of the hand's skeletal points to the Gaussian surface. Then, based on a trained diffusion model, denoising is performed on the aforementioned representation of the hand's skeletal points to the object's Gaussian surface. This eliminates obvious physical errors and initially optimizes the movement. Finally, based on a geometry-aware de-clipping method, the hand movement is re-corrected. This corrects the clipping phenomenon between the hand's skeletal points and the object's Gaussian surface. This implementation method reasonably constructs the representation of hand skeletal points to a Gaussian surface, correcting hand posture distortion and clipping phenomena.
[0114] Continue to refer to Figure 2 As a response Figure 1 The implementation of the method shown in this disclosure provides some embodiments of a Gaussian hand-object interactive rendering denoising apparatus, which are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0115] like Figure 2As shown, the Gaussian hand-object interaction rendering denoising device 200 in some embodiments includes: an object Gaussian initialization unit 201, an object Gaussian reconstruction unit 202, a hand Gaussian initialization unit 203, a pre-training unit 204, a modeling unit 205, a denoising unit 206, and a correction unit 207. The object Gaussian initialization unit 201 is configured to initialize the object Gaussian of a rigid object image to obtain an initialized object Gaussian; the object Gaussian reconstruction unit 202 is configured to reconstruct the initialized object Gaussian based on a Gaussian training method to obtain a reconstructed object Gaussian; the hand Gaussian initialization unit 203 is configured to initialize the hand Gaussian on the registered 3D parametric model surface to obtain a standard hand Gaussian; the pre-training unit 204 is configured to pre-train the standard hand Gaussian to obtain a pre-trained hand Gaussian; the modeling unit 205 is configured to... The system is configured to jointly model the pre-trained Gaussian skeletal joints of the hand and the Gaussian scatter points of the reconstructed object Gaussian surface to construct a hand skeletal point-to-Gaussian surface representation; the denoising unit 206 is configured to denoise the hand skeletal point-to-Gaussian surface representation based on a diffusion model to obtain a denoised hand skeletal point-to-Gaussian surface representation; the correction unit 207 is configured to detect the interaction clipping region corresponding to the denoised hand skeletal point-to-Gaussian surface representation based on a geometry-aware de-clipping method, and re-correct the hand movement based on the interaction clipping region to obtain a de-clipping hand-object interaction image.
[0116] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.
[0117] Further reference Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0118] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0119] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0120] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0121] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0122] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs. When the aforementioned one or more programs are executed by the electronic device, the electronic device causes the following: it initializes the rigid object image with an object Gaussian model to obtain an initialized object Gaussian model; it reconstructs the initialized object Gaussian model based on a Gaussian training method to obtain a reconstructed object Gaussian model; it initializes the hand Gaussian model on the registered 3D parametric model surface to obtain a standard hand Gaussian model; it pre-trains the standard hand Gaussian model to obtain a pre-trained hand Gaussian model; it jointly models the skeletal joints of the pre-trained hand Gaussian model with the surface Gaussian scatter points of the reconstructed object Gaussian model to construct a hand skeletal point-to-Gaussian surface representation; it denoises the hand skeletal point-to-Gaussian surface representation based on a diffusion model to obtain a denoised hand skeletal point-to-Gaussian surface representation; and it detects the interactive clipping region corresponding to the denoised hand skeletal point-to-Gaussian surface representation based on a geometry-aware clipping removal method, and recalibrates the hand movements based on the interactive clipping region to obtain a clipped hand-object interaction image.
[0124] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0125] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0126] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an object Gaussian initialization unit, an object Gaussian reconstruction unit, a hand Gaussian initialization unit, a pre-training unit, a modeling unit, a denoising unit, and a correction unit. The names of these units do not necessarily limit the specific unit; for example, the object Gaussian initialization unit may also be described as "a unit that initializes a rigid object image with an object Gaussian representation to obtain an initialized object Gaussian representation."
[0127] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0128] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A Gaussian hand-object interaction rendering denoising method, comprising: Initialize the Gaussian shape of the rigid object image to obtain the initialized Gaussian shape; Based on the Gaussian training method, the initialized object Gaussian is reconstructed to obtain the reconstructed object Gaussian; On the registered 3D parametric model surface, the Gaussian hand is initialized to obtain the standard Gaussian hand. The standard hand Gaussian is pre-trained to obtain a pre-trained hand Gaussian. The pre-trained Gaussian skeletal joints of the hand are jointly modeled with the Gaussian scatter points of the reconstructed object Gaussian surface to construct a hand skeletal point-to-Gaussian surface representation. This joint modeling includes: Based on the dynamic curvature sensing mechanism, the dynamic curvature of the Gaussian of the reconstructed object is determined; Based on dynamic curvature weight, the anchor point Gaussian subset is selected by curvature-weighted sampling of the reconstructed object Gaussian. A parameterized joint camera is constructed from the pre-trained hand Gaussian to obtain a parameterized joint camera representation; Based on the inter-frame camera reuse mechanism, motion consistency detection is performed on each parameterized joint camera representation in the parameterized joint camera representation set to obtain the joint camera set. By jointly modeling the skeletal joints of the pre-trained hand Gaussian and the Gaussian scatter points of the reconstructed object Gaussian surface, a representation of the hand skeletal points to the Gaussian surface is constructed. Based on a diffusion model, the hand bone point-to-Gaussian surface representation is denoised to obtain a denoised hand bone point-to-Gaussian surface representation. The denoising process based on the diffusion model to obtain the denoised hand bone point-to-Gaussian surface representation includes: Based on the hand skeleton point-to-Gaussian surface representation and depth map, multimodal feature enhancement is performed to obtain an enhanced hand skeleton point-to-Gaussian surface representation. The noise-free hand-object interaction representation is diffused into the whitened noise space using a diffusion model to obtain a noisy hand-object interaction representation. Based on the noiseless hand-object interaction representation and the noisy hand-object interaction representation, the diffusion model is trained to obtain the trained diffusion model. The enhanced hand bone point-to-Gaussian surface representation is subjected to a hybrid filling process to obtain an edge-enhanced hand bone point-to-Gaussian surface representation. Input the edge-enhanced hand bone point to Gaussian surface representation into the trained diffusion model to obtain the denoised hand bone point to Gaussian surface representation. The geometry-aware clipping removal method detects the interaction clipping regions corresponding to the Gaussian surface representation of the denoised hand skeleton points, and recalibrates the hand movements based on these interaction clipping regions to obtain a clipped hand-object interaction image. The geometry-aware clipping removal method, which detects the interaction clipping regions corresponding to the Gaussian surface representation of the denoised hand skeleton points and recalibrates the hand movements based on these interaction clipping regions to obtain a clipped hand-object interaction image, includes: The Euler angles of the root bone, middle bone and end bone corresponding to the five fingers are taken out as the pre-correction Euler angle set, and the pre-correction Euler angle set is obtained. Based on the pre-correction Euler angle set, the clipping depth of each bone segment to the reconstructed Gaussian surface of the object is determined, and the interactive clipping area is obtained. Unfold the root bone Euler angle, middle bone Euler angle, and terminal bone Euler angle corresponding to each through-finger in the interactive through-finger area to generate a bone Euler angle sequence, thus obtaining a bone Euler angle sequence set; For each bone Euler angle sequence in the set of bone Euler angle sequences, the following processing steps are performed: Based on the first bone Euler angle in the bone Euler angle sequence, perform the following correction steps: Increase the curvature of the Euler angle of the first bone; Determine the clipping depth of the Euler angle of the first bone; In response to the contact between the bone corresponding to the Euler angle of the first bone representing the penetration depth and the Gaussian of the reconstructed object, the values in the Euler angle set of the bones before correction are updated synchronously. In response to the empty sequence of bone Euler angles after deleting the first bone Euler angle, the updated set of each pre-correction bone Euler angle group is used as the set of corrected bone Euler angle groups. In response to the fact that the bone Euler angle sequence after deleting the first bone Euler angle is not empty, the bone Euler angle sequence after deleting the first bone Euler angle is used as the bone Euler angle sequence, and the correction step is performed again. Based on the corrected skeleton Euler angle set, the skeleton pose is adjusted to obtain the Gaussian hand pose without clipping. The Gaussian image of the hand in the de-ghosting pose and the Gaussian image of the reconstructed object are rendered together to obtain a de-ghosting hand-object interaction image.
2. The method according to claim 1, wherein, The initialization of the rigid object image with Gaussian to obtain the initialized object Gaussian includes: Three-dimensional reconstruction is performed based on a set of rigid object images to obtain an initialized Gaussian object.
3. The method according to claim 2, wherein, The Gaussian training method reconstructs the Gaussian object from the initialized object to obtain the reconstructed Gaussian object, including: Based on the camera pose corresponding to the rigid object image set, the initialized object is Gaussian rendered to obtain a rendered image set. The loss between the rendered image set and the corresponding real image set is determined, and the Gaussian parameters are optimized based on the gradient descent algorithm and the loss to obtain the reconstructed object Gaussian.
4. The method according to claim 1, wherein, The initialization of the hand Gaussian on the registered 3D parametric model surface to obtain a standard hand Gaussian includes: After filtering or noise reduction, the Gaussian hand image is projected onto the surface of the registered 3D parametric model to obtain a standard Gaussian hand image.
5. The method according to claim 1, wherein, The pre-training process of the standard hand Gaussian to obtain a pre-trained hand Gaussian includes: Based on multi-view images, hand pose estimation methods are used for annotation to obtain the annotated hand poses; Based on the labeled hand posture, the standard hand Gaussian is deformed to the posture space using a rigid kinematic chain to obtain the posture hand Gaussian. The hand Gaussian image of the posture is rendered and trained by comparing it with the multi-view image to obtain a pre-trained hand Gaussian image.
6. A Gaussian hand-object interactive rendering and denoising device, comprising: The object Gaussian initialization unit is configured to initialize the rigid object image with an object Gaussian initialization to obtain the initialized object Gaussian. The object Gaussian reconstruction unit is configured to reconstruct the initialized object Gaussian based on the Gaussian training method to obtain the reconstructed object Gaussian. The hand Gaussian initialization unit is configured to initialize the hand Gaussian on the surface of the registered 3D parametric model to obtain a standard hand Gaussian. The pre-training unit is configured to pre-train the standard hand Gaussian to obtain a pre-trained hand Gaussian. The modeling unit is configured to jointly model the pre-trained Gaussian skeletal joints of the hand and the Gaussian scatter points of the reconstructed object Gaussian surface to construct a hand skeletal point-to-Gaussian surface representation, wherein the joint modeling of the pre-trained Gaussian skeletal joints of the hand and the Gaussian scatter points of the reconstructed object Gaussian surface to construct a hand skeletal point-to-Gaussian surface representation includes: Based on the dynamic curvature sensing mechanism, the dynamic curvature of the Gaussian of the reconstructed object is determined; Based on dynamic curvature weight, the anchor point Gaussian subset is selected by curvature-weighted sampling of the reconstructed object Gaussian. A parameterized joint camera is constructed from the pre-trained hand Gaussian to obtain a parameterized joint camera representation; Based on the inter-frame camera reuse mechanism, motion consistency detection is performed on each parameterized joint camera representation in the parameterized joint camera representation set to obtain the joint camera set. By jointly modeling the skeletal joints of the pre-trained hand Gaussian and the Gaussian scatter points of the reconstructed object Gaussian surface, a representation of the hand skeletal points to the Gaussian surface is constructed. The denoising unit is configured to denoise the hand bone point-to-Gaussian surface representation based on a diffusion model, obtaining a denoised hand bone point-to-Gaussian surface representation. The denoising of the hand bone point-to-Gaussian surface representation based on a diffusion model to obtain the denoised hand bone point-to-Gaussian surface representation includes: Based on the hand skeleton point-to-Gaussian surface representation and depth map, multimodal feature enhancement is performed to obtain an enhanced hand skeleton point-to-Gaussian surface representation. The noise-free hand-object interaction representation is diffused into the whitened noise space using a diffusion model to obtain a noisy hand-object interaction representation. Based on the noiseless hand-object interaction representation and the noisy hand-object interaction representation, the diffusion model is trained to obtain the trained diffusion model. The enhanced hand bone point-to-Gaussian surface representation is subjected to a hybrid filling process to obtain an edge-enhanced hand bone point-to-Gaussian surface representation. Input the edge-enhanced hand bone point to Gaussian surface representation into the trained diffusion model to obtain the denoised hand bone point to Gaussian surface representation. The correction unit is configured to use a geometry-aware de-clipping method to detect the interaction clipping regions corresponding to the Gaussian surface representations of the denoised hand skeleton points, and to re-correct the hand movements based on the interaction clipping regions to obtain a de-clipping hand-object interaction image. The geometry-aware de-clipping method, which detects the interaction clipping regions corresponding to the Gaussian surface representations of the denoised hand skeleton points and re-corrects the hand movements based on the interaction clipping regions to obtain the de-clipping hand-object interaction image, includes: The Euler angles of the root bone, middle bone and end bone corresponding to the five fingers are taken out as the pre-correction Euler angle set, and the pre-correction Euler angle set is obtained. Based on the pre-correction Euler angle set, the clipping depth of each bone segment to the reconstructed Gaussian surface of the object is determined, and the interactive clipping area is obtained. Unfold the root bone Euler angle, middle bone Euler angle, and terminal bone Euler angle corresponding to each through-finger in the interactive through-finger area to generate a bone Euler angle sequence, thus obtaining a bone Euler angle sequence set; For each bone Euler angle sequence in the set of bone Euler angle sequences, the following processing steps are performed: Based on the first bone Euler angle in the bone Euler angle sequence, perform the following correction steps: Increase the curvature of the Euler angle of the first bone; Determine the clipping depth of the Euler angle of the first bone; In response to the contact between the bone corresponding to the Euler angle of the first bone representing the penetration depth and the Gaussian of the reconstructed object, the values in the Euler angle set of the bones before correction are updated synchronously. In response to the empty sequence of bone Euler angles after deleting the first bone Euler angle, the updated set of each pre-correction bone Euler angle group is used as the set of corrected bone Euler angle groups. In response to the fact that the bone Euler angle sequence after deleting the first bone Euler angle is not empty, the bone Euler angle sequence after deleting the first bone Euler angle is used as the bone Euler angle sequence, and the correction step is performed again. Based on the corrected skeleton Euler angle set, the skeleton pose is adjusted to obtain the Gaussian hand pose without clipping. The Gaussian image of the hand in the de-ghosting pose and the Gaussian image of the reconstructed object are rendered together to obtain a de-ghosting hand-object interaction image.
7. An electronic device, comprising: One or more processes; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processes, the one or more processes perform the method as described in any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the program is processed and executed, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Deformable grabbing interaction method under virtual reality environment
CN108664126A
Real-time reconstruction method and device for hand and object interaction process
CN110007754A