Method, device, electronic device and storage medium for three-dimensional reconstruction of articulated objects

By obtaining the state images of multiple objects from a single perspective of articulated objects, extracting Gaussian representation and joint parameters, and combining data fusion technology for Gaussian splash reconstruction, the problems of high data demand and high computational cost in articulated objects are solved, and efficient and accurate three-dimensional reconstruction is achieved.

CN119832158BActive Publication Date: 2025-08-12BEIJING YUANLUO TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411917534.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-08-12
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

Existing three-dimensional reconstruction methods have high data requirements for articulated objects, high computational costs and difficult to effectively capture the dynamic and geometric characteristics of their joint parts, especially in real-time or near-real-time applications.

Method used

By obtaining multiple object state images from a single perspective of the target articulated object, the original three-dimensional Gaussian representation and joint parameter data are extracted, combined with data fusion technology of multi-object state, the three-dimensional reconstruction is used using Gaussian sputtering to reduce the data volume and calculation cost.

Benefits of technology

It significantly reduces the data volume requirement and calculation cost, improves the accuracy and efficiency of three-dimensional reconstruction, and is suitable for real-time or near-real-time application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832158B_ABST
    Figure CN119832158B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, electronic device, and storage medium for three-dimensional reconstruction of an articulated object, relating to the technical field of three-dimensional reconstruction. The present invention can obtain target images of at least two object states of a target articulated object to be reconstructed from a single perspective; extract the original three-dimensional Gaussian representation and target joint parameter data of the target articulated object from the target image; and perform three-dimensional reconstruction of the target articulated object in a specified object state based on the original three-dimensional Gaussian representation and target joint parameter data to obtain a target three-dimensional reconstructed image. By applying the joint parameter data to the three-dimensional reconstruction of the target articulated object and combining it with data fusion technology in multiple object states, the appearance, geometry, and dynamic information of the articulated object can be effectively captured from a small number of images from a single perspective, and efficient three-dimensional reconstruction can be performed based on Gaussian splashing, reducing the demand for data volume and computing cost, improving data utilization efficiency, and enhancing the accuracy and efficiency of three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction, and in particular to a method, device, electronic equipment and storage medium for three-dimensional reconstruction of an articulated object. Background Art

[0002] 3D reconstruction of articulated objects has always been a challenge in the fields of 3D modeling and computer vision. Articulated objects, such as furniture and robots, are difficult to accurately capture using traditional 3D reconstruction methods due to their complex geometry and dynamic characteristics. Existing methods typically require a large number of multi-view images, which is often impractical in practical applications and computationally expensive.

[0003] Furthermore, existing dynamic scene modeling methods, such as Neural Radiance Fields (NeRF) and its derivatives, typically rely on multilayer perceptrons (MLPs) to implicitly represent 3D objects, which are computationally expensive, limiting their applications. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, device, electronic device and storage medium for three-dimensional reconstruction of an articulated object, so as to reduce data requirements and computing costs and improve the accuracy and efficiency of three-dimensional reconstruction.

[0005] In a first aspect, the present invention provides a method for three-dimensional reconstruction of an articulated object, comprising:

[0006] Acquire target images of at least two object states of a target articulated object to be reconstructed under a single viewing angle; wherein the articulation states of at least one articulated structure of the target articulated object in different object states are different;

[0007] Extracting an original three-dimensional Gaussian representation and target joint parameter data of a target articulated object from a target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the target joint parameter data includes joint types and joint parameter information of an articulated structure of the target articulated object;

[0008] According to the original three-dimensional Gaussian representation and the target joint parameter data, the target articulated object is three-dimensionally reconstructed in the specified object state to obtain a target three-dimensional reconstructed image.

[0009] Furthermore, the original three-dimensional Gaussian representation of the target articulated object and target joint parameter data are extracted from the target image, including:

[0010] Perform feature extraction on the target image to obtain target image features;

[0011] The target image features are converted into a three-dimensional Gaussian representation and joint parameters are extracted to obtain the original three-dimensional Gaussian representation and target joint parameter data; wherein each original Gaussian sphere in the original three-dimensional Gaussian representation includes a three-dimensional covariance matrix, position information, joint probability information, transparency information and color information, the three-dimensional covariance matrix is used to represent the three-dimensional scale information and direction information of the corresponding Gaussian sphere in an object state, and the joint probability information includes a set of probability values for indicating that the corresponding Gaussian sphere belongs to each articulated structure and non-articulated structure; the target joint parameter data includes a translation vector of a translation joint, and / or, a rotation axis, a rotation center and a rotation angle of a rotation joint.

[0012] Furthermore, the target image includes an image sequence consisting of two-dimensional images in different object states; and feature extraction is performed on the target image to obtain target image features, including:

[0013] The first set of two-dimensional convolutional networks is used to extract and compress the image features of each two-dimensional image to obtain the initial image features corresponding to each two-dimensional image;

[0014] A set of two-dimensional convolutional long short-term memory neural networks is used to extract the temporal dynamic information of the initial image features corresponding to each two-dimensional image to obtain the target image features.

[0015] Furthermore, the target image features are converted into a three-dimensional Gaussian representation and joint parameters are extracted to obtain the original three-dimensional Gaussian representation and target joint parameter data, including:

[0016] The target image features are converted into a three-dimensional Gaussian representation through a second set of two-dimensional convolutional networks to obtain the original three-dimensional Gaussian representation;

[0017] The joint type and joint parameters in the target image features are extracted through a multi-layer perceptron to obtain the target joint parameter data.

[0018] Furthermore, a three-dimensional reconstruction of a specified object state is performed on the target articulated object based on the original three-dimensional Gaussian representation and the target joint parameter data to obtain a target three-dimensional reconstructed image, including:

[0019] According to the target joint parameter data, the original 3D Gaussian representation is adjusted to the specified object state to obtain the target 3D Gaussian representation;

[0020] The target 3D Gaussian representation is rendered at a specified viewing angle to obtain a target 3D reconstructed image.

[0021] Furthermore, the original Gaussian sphere includes a three-dimensional covariance matrix, position information, joint probability information, transparency information, and color information. According to the target joint parameter data, the original three-dimensional Gaussian representation is adjusted to the specified object state to obtain the target three-dimensional Gaussian representation, including:

[0022] According to the target joint parameter data, a set of SE3 transformation matrices is determined to transform the object state corresponding to the original three-dimensional Gaussian representation to the specified object state, and the SE3 transformation matrix corresponds to the articulated structure one by one;

[0023] According to the joint probability information of each original Gaussian sphere, a set of SE3 transformation matrices are transformed on the three-dimensional covariance matrix, position information, transparency information and color information of each original Gaussian sphere to obtain a set of target Gaussian spheres corresponding to each original Gaussian sphere;

[0024] Multiple groups of target Gaussian spheres corresponding to all original Gaussian spheres are determined as target three-dimensional Gaussian representations.

[0025] Furthermore, the target 3D Gaussian representation is rendered at a specified viewing angle to obtain a target 3D reconstructed image, including:

[0026] Project the target 3D Gaussian representation onto the 2D plane corresponding to the specified viewing angle;

[0027] The target Gaussian sphere projected onto the two-dimensional plane is rendered using the following transparency blending strategy: for each pixel to be rendered, the transparency and order of the overlapping Gaussian spheres in the projection direction of the pixel are calculated, and then the overlapping Gaussian spheres are traversed in order, and the final color of the pixel is obtained by combining the predicted color with the weighted transparency.

[0028] In a second aspect, the present invention further provides a device for three-dimensional reconstruction of an articulated object, comprising:

[0029] An image acquisition module is configured to acquire target images of at least two object states of a target articulated object to be reconstructed under a single viewing angle; wherein the articulation states of at least one articulated structure of the target articulated object in different object states are different;

[0030] a data extraction module, configured to extract an original three-dimensional Gaussian representation and target joint parameter data of a target articulated object from a target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the target joint parameter data includes joint types and joint parameter information of the articulated structure of the target articulated object;

[0031] The three-dimensional reconstruction module is used to perform three-dimensional reconstruction of the target articulated object in a specified object state according to the original three-dimensional Gaussian representation and the target joint parameter data to obtain a target three-dimensional reconstructed image.

[0032] In a third aspect, the present invention further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the method for three-dimensional reconstruction of an articulated object of the first aspect is implemented.

[0033] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for three-dimensional reconstruction of an articulated object according to the first aspect is executed.

[0034] The method, device, electronic device, and storage medium for three-dimensional reconstruction of an articulated object provided by the embodiments of the present invention can obtain target images of at least two object states of a target articulated object to be reconstructed from a single perspective; wherein the articulation state of at least one articulated structure of the target articulated object in different object states is different; extract the original three-dimensional Gaussian representation and target joint parameter data of the target articulated object from the target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the target joint parameter data includes the joint type and joint parameter information of the articulated structure of the target articulated object; based on the original three-dimensional Gaussian representation and the target joint parameter data, the target articulated object is three-dimensionally reconstructed in a specified object state to obtain a target three-dimensional reconstructed image. In this way, by applying the joint parameter data to the three-dimensional reconstruction of the target articulated object, combined with data fusion technology in multiple object states, it is possible to effectively capture the appearance, geometry, and dynamic information of the articulated object from a small number of images from a single perspective, and perform efficient three-dimensional reconstruction based on Gaussian splashing, significantly reducing the demand for data volume and computing cost, while improving data utilization efficiency and enhancing the accuracy and efficiency of three-dimensional reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0036] Figure 1 A schematic flow chart of a method for 3D reconstruction of an articulated object provided by an embodiment of the present invention;

[0037] Figure 2 A data flow diagram of a method for 3D reconstruction of an articulated object provided by an embodiment of the present invention;

[0038] Figure 3 A schematic diagram of joint parameters of a translation joint provided by an embodiment of the present invention;

[0039] Figure 4 A schematic diagram of joint parameters of a rotary joint provided by an embodiment of the present invention;

[0040] Figure 5A schematic structural diagram of a device for 3D reconstruction of an articulated object provided by an embodiment of the present invention;

[0041] Figure 6 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0043] For the three-dimensional reconstruction of articulated objects, traditional three-dimensional reconstruction methods usually require a large number of multi-view images. Among dynamic scene modeling methods, NeRF technology currently faces multiple challenges: it relies entirely on multi-layer perceptrons, resulting in huge computing resource requirements, and the model training and rendering processes are time-consuming. In recent years, Gaussian Splatting (3DGS) technology, as a method of three-dimensional reconstruction, has achieved impressive results by rasterizing a set of Gaussian ellipsoids to approximate the appearance of a three-dimensional scene. However, due to the geometric and dynamic complexity of articulated objects, there are certain difficulties in applying Gaussian Splatting technology to articulated objects.

[0044] Therefore, the main disadvantages of the prior art are:

[0045] 1. High data requirements: A large number of multi-view images are required, which is difficult to achieve in practical applications.

[0046] 2. High computational cost: The representation method that relies on neural radiation fields is computationally intensive and is not suitable for real-time or near-real-time application scenarios.

[0047] 3. Lack of joint representation: For the modeling of articulated objects, existing methods often cannot effectively capture the dynamic and geometric characteristics of their joint parts.

[0048] Furthermore, regarding the aforementioned NeRF technology, to improve training and rendering efficiency, many studies have turned to explicit feature representation methods, such as voxels, grids, point clouds, and hash tables, in an effort to address this issue. These methods can significantly reduce the time required to optimize the radiation field, but they also incur higher video memory consumption and, due to lower resolution, reduce the quality of the final rendered image. Furthermore, video memory limitations also limit the maximum resolution of renderable images.

[0049] At the same time, NeRF can only be used to render static objects. When faced with three-dimensional articulated objects, it is usually necessary to use multiple MLPs to represent each part of the object separately, and then use predefined axis parameters to model the articulated structure. Because each part of the object is stored independently in multiple MLPs, there is often a lack of effective supervision during training to store each part of the object in the corresponding MLP. During rendering, due to the need to independently sample multiple MLPs, the computational resource overhead increases linearly with the number of object parts. Therefore, current mainstream methods find it difficult to strike a balance between final rendering speed and rendering quality.

[0050] At the same time, NeRF-based methods are difficult to generalize effectively. When faced with new objects, they often require a large number of images from different perspectives for retraining. In general application scenarios, it is difficult to obtain these images.

[0051] Based on this, embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for 3D reconstruction of articulated objects. These methods effectively capture the appearance, geometry, and dynamic information of articulated objects from a small number of images from a single perspective, allowing for efficient 3D reconstruction. By explicitly modeling joints during the 3D reconstruction process, simultaneously capturing geometric, dynamic, and appearance information, these embodiments improve the accuracy and data efficiency of 3D reconstruction.

[0052] To facilitate understanding of this embodiment, a method for three-dimensional reconstruction of an articulated object disclosed in an embodiment of the present invention is first introduced in detail.

[0053] An embodiment of the present invention provides a method for 3D reconstruction of an articulated object, which can be performed by an electronic device with data processing capabilities. Figure 1 The flowchart of a method for 3D reconstruction of an articulated object is shown, and the method mainly includes the following steps S110 to S130:

[0054] Step S110 , obtaining target images of at least two object states of the target articulated object to be reconstructed under a single viewing angle; wherein the articulation states of at least one articulated structure of the target articulated object in different object states are different.

[0055] An articulated object refers to two or more rigid bodies connected together by one or more articulated structures such as joints in three-dimensional space. These rigid bodies can rotate or translate relative to each other within a limited range. The articulated object to be reconstructed is referred to here as a target articulated object. Different articulated states of an articulated structure refer to: for example, the opening or closing of a cabinet door, the pulling out and pushing back of a drawer, different cabinet door opening angles or different drawer pulling-out distances are regarded as different articulated states. If the articulated state of one or more articulated structures of the target articulated object changes, then the target articulated object is in different object states before and after the change.

[0056] In specific implementation, the target articulated object can be photographed from a single perspective by a camera installed in a fixed position. The object state of the target articulated object can be changed during the shooting process, and a two-dimensional image can be taken under different object states, thereby obtaining an image sequence composed of two-dimensional images under different object states. This image sequence is the target image.

[0057] Step S120, extracting the original three-dimensional Gaussian representation and target joint parameter data of the target articulated object from the target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the target joint parameter data includes the joint type and joint parameter information of the articulated structure of the target articulated object.

[0058] In some possible embodiments, the above step S120 may include the following steps S121 and S122:

[0059] Step S121 , extracting features from the target image to obtain target image features.

[0060] The above-mentioned target image may include an image sequence composed of two-dimensional images in different object states; based on this, the image features of each two-dimensional image can be first extracted and compressed through a first group of two-dimensional convolutional networks to obtain the initial image features corresponding to each two-dimensional image; and then a group of two-dimensional convolutional long short-term memory neural networks can be used to extract the temporal dynamic information of the initial image features corresponding to each two-dimensional image to obtain the target image features.

[0061] Step S122 , performing three-dimensional Gaussian representation conversion and joint parameter extraction on the target image features to obtain the original three-dimensional Gaussian representation and target joint parameter data.

[0062] Among them, each original Gaussian sphere in the original three-dimensional Gaussian representation includes a three-dimensional covariance matrix, position information, joint probability information, transparency information and color information. The three-dimensional covariance matrix is used to represent the three-dimensional scale information and direction information of the corresponding Gaussian sphere in an object state. The joint probability information includes a set of probability values indicating that the corresponding Gaussian sphere belongs to each articulated structure and non-articulated structure. The color information can be represented by a spherical harmonic function for representing anisotropic colors; the target joint parameter data includes the translation vector of the translation joint, and / or, the rotation axis, rotation center and rotation angle of the rotation joint.

[0063] A set of probability values for an original Gaussian sphere may include a probability value belonging to each joint and a probability value belonging to the body (non-joint). Assuming that the target articulated object has two joints, a set of probability values may include a probability of 0.8 belonging to the body, a probability of 0.1 belonging to joint 1, and a probability of 0.1 belonging to joint 2. The target joint parameter data may include a set of joint parameter data, wherein each joint parameter data corresponds to an articulated structure, and each joint parameter data may include joint type and joint parameter information of the corresponding articulated structure, wherein the number of articulated structures may be given in the form of a hyperparameter.

[0064] In specific implementation, the target image features can be converted into a three-dimensional Gaussian representation through a second set of two-dimensional convolutional networks to obtain the original three-dimensional Gaussian representation; the joint type and joint parameters in the target image features can be extracted through a multi-layer perceptron to obtain the target joint parameter data.

[0065] Step S130 , performing three-dimensional reconstruction of the target articulated object in a specified object state according to the original three-dimensional Gaussian representation and the target joint parameter data, to obtain a target three-dimensional reconstructed image.

[0066] The above-mentioned specified object state may be one of the object states corresponding to the target image, or may not be the object state corresponding to the target image.

[0067] In some possible embodiments, the above step S130 may include the following steps S131 and S132:

[0068] Step S131 : According to the target joint parameter data, the original three-dimensional Gaussian representation is adjusted to the specified object state to obtain the target three-dimensional Gaussian representation.

[0069] The original Gaussian sphere may include a three-dimensional covariance matrix, position information, joint probability information, transparency information, and color information. Based on this, a set of SE3 transformation matrices can be determined based on the target joint parameter data to transform the object state corresponding to the original three-dimensional Gaussian representation to the specified object state. The SE3 transformation matrices correspond one-to-one to the articulated structure. Based on the joint probability information of each original Gaussian sphere, the three-dimensional covariance matrix, position information, transparency information, and color information of each original Gaussian sphere are transformed by a set of SE3 transformation matrices to obtain a set of target Gaussian spheres corresponding to each original Gaussian sphere. The multiple sets of target Gaussian spheres corresponding to all original Gaussian spheres are determined as the target three-dimensional Gaussian representation. One joint parameter data corresponds to one SE3 transformation matrix. The SE3 transformation matrix can describe the transformation of an object from one object state to another. Specifically, the SE3 transformation matrix can describe the transformation of the corresponding articulated structure from one articulated state to another.

[0070] Step S132 , rendering the target 3D Gaussian representation at a specified viewing angle to obtain a target 3D reconstructed image.

[0071] The above-mentioned specified perspective may not be the perspective corresponding to the target image. Optionally, the target three-dimensional Gaussian representation can be first projected onto a two-dimensional plane corresponding to the specified perspective; the target Gaussian sphere projected onto the two-dimensional plane is rendered using the following transparency blending strategy: for each pixel to be rendered, the transparency and order of the overlapping Gaussian spheres in the projection direction of the pixel are calculated, and then the overlapping Gaussian spheres are traversed in order, and the final color of the pixel is obtained by combining the predicted color with the weighted transparency.

[0072] The method for three-dimensional reconstruction of an articulated object provided by an embodiment of the present invention can obtain target images of at least two object states of a target articulated object to be reconstructed from a single perspective; wherein the articulation state of at least one articulated structure of the target articulated object in different object states is different; extract the original three-dimensional Gaussian representation and target joint parameter data of the target articulated object from the target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the target joint parameter data includes the joint type and joint parameter information of the articulated structure of the target articulated object; based on the original three-dimensional Gaussian representation and the target joint parameter data, the target articulated object is three-dimensionally reconstructed in a specified object state to obtain a target three-dimensional reconstructed image. In this way, by applying the joint parameter data to the three-dimensional reconstruction of the target articulated object, combined with data fusion technology in multiple object states, it is possible to effectively capture the appearance, geometry, and dynamic information of the articulated object from a small number of images from a single perspective, and perform efficient three-dimensional reconstruction based on Gaussian splashing, significantly reducing the demand for data volume and computing cost, while improving data utilization efficiency and improving the accuracy and efficiency of three-dimensional reconstruction.

[0073] For easier understanding, see Figure 2 An embodiment of the present invention discloses a system for implementing a three-dimensional reconstruction method of an articulated object based on Gaussian splashing, the system including an image encoder 200, a timing information processor 210, a Gaussian decoder 220, a joint predictor 230, a Gaussian editor 240, and a Gaussian renderer 250.

[0074] The image encoder 200 is used to process an input image sequence, extract its features and output them to the temporal information processor 210. Different image frames share the same image encoder weights.

[0075] The TIP 210 captures and processes temporal dynamic information in image sequences and outputs it to the Gaussian decoder 220 and joint predictor 230. This is achieved by integrating a convolutional long short-term memory network (LSTM), which can process spatiotemporal data to understand the temporal motion and changes of objects in image sequences. The output of the TIP 210 provides critical timing information for subsequent joint prediction and Gaussian editing.

[0076] The Gaussian decoder 220 is responsible for converting the feature information processed by the temporal information processor 210 into a three-dimensional Gaussian representation in 3D space. This process involves converting the image information into a 3D representation, which is achieved by predicting the parameters of each Gaussian sphere. These parameters can include the position, scale, orientation, and color of the Gaussian sphere.

[0077] The joint predictor 230 is a core component of the system, responsible for predicting joint types and parameters. This can include the translation vector for translational joints and the axis, center, and angle of rotation for rotational joints. The predictions from the joint predictor 230 directly influence how the 3D Gaussian representation is edited in the Gaussian editor 240.

[0078] Gaussian Editor 240 receives the joint parameters from Joint Predictor 230 and the 3D Gaussian representation from Gaussian Decoder 220 and adjusts the position and orientation of the Gaussian representation based on these parameters. This process involves applying a corresponding geometric transformation to each Gaussian sphere to simulate the motion of the joint. Gaussian Editor 240 also uses the joint parameters from Joint Predictor 230 to ensure that these transformed Gaussian spheres accurately reflect the actual motion of the object.

[0079] The Gaussian renderer 250 is the last component of the entire system, which is responsible for rendering the edited Gaussian spheres into the final image. This process involves calculating the contribution of each Gaussian sphere to the final image and compositing them based on their color, opacity and other properties.

[0080] The design of the entire system allows the joint prediction and 3D reconstruction processes to mutually enhance each other, improving the accuracy and efficiency of reconstruction. In this way, the dynamic behavior of articulated objects can be accurately reconstructed and simulated within a limited field of view. The innovation of this method lies in its ability to simultaneously process the appearance, geometry, and dynamic information of an object, which is difficult to achieve with traditional 3D reconstruction methods. In addition, by introducing the concept of joints, the present invention can also provide richer and more accurate 3D models for fields such as robotics, virtual reality, and augmented reality.

[0081] For ease of understanding, the following is a detailed introduction to the above-mentioned method for 3D reconstruction of articulated objects. The method includes the following steps:

[0082] Step 0: First, obtain images of multiple states of the target to be reconstructed under one or more viewing angles, extract the two-dimensional image features of the target to be reconstructed in the image of each state, and compress the features of the image of each state into a feature map (i.e., initial image features).

[0083] The different states of the target (i.e., the target articulated object) refer to different states of the articulated object, such as whether a door is open or closed, or whether a drawer is pulled out or pushed back in. Different door opening angles or different drawer pull-out distances are considered different articulated states.

[0084] Two-dimensional image features can be meaningful information extracted from an image, which can describe certain attributes or relationships of objects in the image. Two-dimensional image features can include one or more of edges, corners, spots, and texture patterns in the image; feature maps can be used to convert two-dimensional image features in an image into a two-dimensional numerical matrix.

[0085] It can be understood that the embodiment of the present invention achieves dimensionality reduction by extracting the two-dimensional image features of the target to be reconstructed in the image of each state and compressing these features into a feature map, which significantly improves the computing efficiency and storage efficiency, and can ensure that the extracted feature vector has high recognition accuracy and robustness.

[0086] The feature extraction in step 0 above can be achieved through a set of two-dimensional convolutional networks.

[0087] Step 1: Generate a set of new feature maps by extracting features from the feature maps of images in different states.

[0088] It is understandable that by comprehensively analyzing the feature maps of an image in different states, it is possible to effectively supplement location information that is difficult to observe in a single state. For example, when a cabinet door is open, its image can provide relevant information about the cabinet interior for the closed door image, which is not available in a single state.

[0089] Furthermore, by comparing and analyzing feature maps in different states, embodiments of the present invention can more accurately obtain articulation information about articulated objects. For example, by comparing images of a cabinet door in its open and closed states, information about the door panel, cabinet body, and hinge can be revealed. This comparative analysis method facilitates a deeper understanding of the structure and motion characteristics of articulated objects.

[0090] The above step 1 can be implemented by a set of two-dimensional convolutional long short-term memory neural networks.

[0091] Step 2: By processing the new feature map obtained in step 1, a set of Gaussian spheres representing the target to be reconstructed is obtained.

[0092] It can be understood that the new feature map obtained in step 1 already contains information in different states, so the Gaussian sphere obtained in step 2 also contains information about the object in different states, such as position, scale, direction, and color in different states.

[0093] The Gaussian sphere predicted in this step contains a set of probabilities, which are used to indicate the probability that the current Gaussian sphere belongs to a specific part of the articulated object, including the probability of belonging to each joint and the probability of belonging to the body (non-joint).

[0094] Each Gaussian sphere is represented by a three-dimensional covariance matrix Σ, which is expressed as the combination of a scaling matrix S and a rotation matrix Q, where S is represented by a three-dimensional vector s and Q is represented by a rotation quaternion q: Σ = QSS T Q T .

[0095] The above step 2 can be implemented by a set of two-dimensional convolutional networks.

[0096] Step 3: By processing the new feature map obtained in step 1, multiple sets of joint parameters for representing the articulated structure of the target to be reconstructed are obtained.

[0097] like Figure 3 As shown, for a translation joint, such as a drawer, the drawer's moving direction axis can be predicted p and moving distance δ p Among them, axis p It can be obtained by the components u in the three coordinate axis directions of the coordinate system where the drawer is located p 、v p 、w p Composition, δ p It may include distance values at different times t0, t1, t2, etc.

[0098] like Figure 4 As shown, for a rotation joint, such as a cabinet door, the normal vector axis of the rotation plane of the cabinet door can be predicted r , rotation center O r and the rotation angle δ r Among them, axis r The components u in the three coordinate axis directions of the coordinate system where the cabinet door is located can be obtained r 、v r 、w r Composition, δ r It can include angle values at different times t0, t1, t2, etc. r The components x in the three coordinate axis directions of the three-dimensional coordinate system of the target to be reconstructed can be r 、y r 、z r constitute.

[0099] The above step 3 can be implemented by a multilayer perceptron, which can be a feedforward artificial neural network composed of multiple layers of neurons, including an input layer, a hidden layer, and an output layer.

[0100] Step 4: Convert the joint parameters obtained in step 3 into an SE3 transformation matrix, and then use the obtained SE3 transformation matrix to process the set of Gaussian spheres obtained in step 2.

[0101] Among them, the SE3 transformation matrix is a representation of the special Euclidean group, which is used to describe the rigid body transformation in three-dimensional space, including rotation and translation. It can be expressed as a 4×4 homogeneous matrix in the following form:

[0102]

[0103] Here, R is a 3×3 rotation matrix representing the SO(3) group, t is a 3×1 translation vector, and 0 is a 1×3 zero vector. The SO(3) group is a special orthogonal group in three-dimensional space consisting of all possible rotations that preserve the length of the vector and maintain the right-hand rule (i.e., they do not introduce mirroring).

[0104] Existing translation direction vector The translation distance δ, then the translation transformation matrix R can be written as:

[0105]

[0106] Normal vector of the existing rotation plane The rotation center O = (x0, y0, z0), the rotation angle δ, then the rotation transformation matrix R can be written as:

[0107] R=T -1 R r T;

[0108] in,

[0109] R r =I+(sin(δ))K+(1-cos(δ))K 2 ;

[0110]

[0111] When using the SE3 transformation matrix to process the Gaussian sphere obtained in step 2, the Q constituting Σ (i.e., the three-dimensional covariance matrix), the xyz used to represent the position of the Gaussian sphere, and the spherical harmonics used to represent the anisotropic color can be SE3 transformed through a set of probability values (i.e., joint probability information).

[0112] Through SE3 transformation, multiple groups of target Gaussian spheres under specified states can be obtained. The number of groups of target Gaussian spheres is the same as the number of probability values (for example, 3 probability values correspond to 3 groups of target Gaussian spheres). Among them, an original Gaussian sphere can be transformed into multiple target Gaussian spheres through a set of probability values. Different target Gaussian spheres may differ in transparency, position and spherical harmonic functions.

[0113] Step 5: Render each target Gaussian sphere obtained in step 4 into an image at a specified viewing angle.

[0114] The target Gaussian sphere projected onto the two-dimensional plane corresponding to a specified viewing angle can be rendered using a transparency α blending strategy: for each pixel to be rendered, the transparency α corresponding to the overlapping Gaussian spheres in the projection direction of the pixel and their order are calculated. The overlapping Gaussian spheres are then traversed in order, and the final color of the pixel is obtained by combining the predicted color c with the weighted transparency. In the embodiments of the present invention, multiple sets of target Gaussian spheres obtained through the SE3 transformation are considered during rendering.

[0115] The network models used in the above methods (e.g., the first set of 2D convolutional networks, the first set of 2D convolutional long short-term memory neural networks, the second set of 2D convolutional networks, and the multilayer perceptron) are trained. During training, the labels of the sample images include information such as object state, viewpoint, and joint representation. It should be noted that the viewpoint specified during training can be different from the viewpoint specified in subsequent applications.

[0116] Compared with the prior art, the embodiments of the present invention have achieved significant improvements and advantages in many aspects. First, the embodiments of the present invention improve the accuracy of joint prediction by integrating joint parameters into the 3DGS (3D Geographic Scene) rendering process in a differentiable way, so that the network model can more accurately capture the dynamic and geometric characteristics of articulated objects. Secondly, combined with multimodal data fusion technology (i.e., data fusion technology under multi-object states), the embodiments of the present invention effectively capture the appearance, geometry and dynamic information of articulated objects from a small number of images from a single perspective, significantly reducing the demand for data volume while improving data utilization efficiency. In terms of computational cost, the embodiments of the present invention achieve a fast three-dimensional reconstruction process through optimized architectural design, reducing the computational cost of the model, making it more suitable for real-time or near real-time application scenarios. In summary, the embodiments of the present invention not only improve the accuracy and real-time performance of three-dimensional reconstruction, but also enhance the generalization ability of the network model and the quality of the data set, providing an efficient and accurate solution for the field of three-dimensional reconstruction of articulated objects, which has important practical application value and broad market prospects.

[0117] Corresponding to the above-mentioned method for 3D reconstruction of an articulated object, an embodiment of the present invention further provides a device for 3D reconstruction of an articulated object. Figure 5The schematic diagram of the structure of a device for 3D reconstruction of an articulated object is shown, the device comprising:

[0118] An image acquisition module 501 is configured to acquire target images of at least two object states of a target articulated object to be reconstructed from a single viewing angle; wherein the articulation states of at least one articulated structure of the target articulated object in different object states are different;

[0119] A data extraction module 502 is configured to extract an original three-dimensional Gaussian representation and target joint parameter data of a target articulated object from a target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the target joint parameter data includes joint types and joint parameter information of the articulated structure of the target articulated object;

[0120] The 3D reconstruction module 503 is used to perform 3D reconstruction of the target articulated object in a specified object state according to the original 3D Gaussian representation and the target joint parameter data to obtain a target 3D reconstructed image.

[0121] The apparatus for three-dimensional reconstruction of an articulated object provided by an embodiment of the present invention can obtain target images of at least two object states of a target articulated object to be reconstructed from a single perspective; wherein the articulation state of at least one articulated structure of the target articulated object in different object states is different; extract the original three-dimensional Gaussian representation and target joint parameter data of the target articulated object from the target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the target joint parameter data includes the joint type and joint parameter information of the articulated structure of the target articulated object; based on the original three-dimensional Gaussian representation and the target joint parameter data, the target articulated object is three-dimensionally reconstructed in a specified object state to obtain a target three-dimensional reconstructed image. In this way, by applying the joint parameter data to the three-dimensional reconstruction of the target articulated object, combined with data fusion technology in multiple object states, it is possible to effectively capture the appearance, geometry, and dynamic information of the articulated object from a small number of images from a single perspective, and perform efficient three-dimensional reconstruction based on Gaussian splashing, significantly reducing the demand for data volume and computing cost, while improving data utilization efficiency and improving the accuracy and efficiency of three-dimensional reconstruction.

[0122] Furthermore, the above-mentioned data extraction module 502 is specifically used to: perform feature extraction on the target image to obtain target image features; perform three-dimensional Gaussian representation conversion and joint parameter extraction on the target image features to obtain original three-dimensional Gaussian representation and target joint parameter data; wherein, each original Gaussian sphere in the original three-dimensional Gaussian representation includes a three-dimensional covariance matrix, position information, joint probability information, transparency information and color information, the three-dimensional covariance matrix is used to represent the three-dimensional scale information and direction information of the corresponding Gaussian sphere in an object state, and the joint probability information includes a set of probability values for indicating that the corresponding Gaussian sphere belongs to each articulated structure and non-articulated structure; the target joint parameter data includes the translation vector of the translation joint, and / or, the rotation axis, rotation center and rotation angle of the rotation joint.

[0123] Furthermore, the target image includes an image sequence consisting of two-dimensional images in different object states; the data extraction module 502 is also used to: extract and compress image features of each two-dimensional image through a first group of two-dimensional convolutional networks to obtain initial image features corresponding to each two-dimensional image; extract temporal dynamic information of the initial image features corresponding to each two-dimensional image through a group of two-dimensional convolutional long short-term memory neural networks to obtain target image features.

[0124] Furthermore, the above-mentioned data extraction module 502 is also used to: convert the target image features into a three-dimensional Gaussian representation through a second group of two-dimensional convolutional networks to obtain the original three-dimensional Gaussian representation; extract the joint type and joint parameters in the target image features through a multi-layer perceptron to obtain target joint parameter data.

[0125] Furthermore, the above-mentioned 3D reconstruction module 503 is specifically used to: adjust the original 3D Gaussian representation to the specified object state according to the target joint parameter data to obtain the target 3D Gaussian representation; render the target 3D Gaussian representation at a specified perspective to obtain the target 3D reconstructed image.

[0126] Furthermore, the above-mentioned original Gaussian sphere includes a three-dimensional covariance matrix, position information, joint probability information, transparency information and color information; the three-dimensional reconstruction module 503 is also used to: determine a set of SE3 transformation matrices for transforming the object state corresponding to the original three-dimensional Gaussian representation to the specified object state based on the target joint parameter data, and the SE3 transformation matrix corresponds one-to-one to the articulated structure; according to the joint probability information of each original Gaussian sphere, the three-dimensional covariance matrix, position information, transparency information and color information of each original Gaussian sphere are transformed by a set of SE3 transformation matrices to obtain a set of target Gaussian spheres corresponding to each original Gaussian sphere; and determine multiple groups of target Gaussian spheres corresponding to all original Gaussian spheres as the target three-dimensional Gaussian representation.

[0127] Furthermore, the above-mentioned three-dimensional reconstruction module 503 is also used to: project the target three-dimensional Gaussian representation onto a two-dimensional plane corresponding to the specified viewing angle; the target Gaussian sphere projected onto the two-dimensional plane is rendered through the following transparency blending strategy: for each pixel point to be rendered, calculate the transparency corresponding to the overlapping Gaussian spheres in the projection direction of the pixel point and their order, and then traverse the overlapping Gaussian spheres in order, and obtain the final color of the pixel point by weighting the transparency and combining the predicted color.

[0128] The device provided in this embodiment has the same implementation principle and technical effects as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned method embodiment.

[0129] like Figure 6 As shown, an embodiment of the present invention provides an electronic device 600, including: a processor 601, a memory 602 and a bus. The memory 602 stores a computer program that can be run on the processor 601. When the electronic device 600 is running, the processor 601 and the memory 602 communicate through the bus, and the processor 601 executes the computer program to implement the above-mentioned three-dimensional reconstruction method of articulated objects.

[0130] Specifically, the memory 602 and processor 601 can be general-purpose memories and processors, which are not specifically limited here.

[0131] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program that, when executed by a processor, executes the method for 3D reconstruction of an articulated object described in the preceding method embodiments. The computer-readable storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), RAM, a magnetic disk, or an optical disk.

[0132] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0133] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not limiting, and thus other examples of the exemplary embodiments may have different values.

[0134] The flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0135] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0136] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0137] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for 3D reconstruction of an articulated object, characterized in that: include: Acquire target images of at least two object states of a target articulated object to be reconstructed under a single viewing angle; wherein the articulation states of at least one articulated structure of the target articulated object in different object states are different; By extracting time-dynamic information, an original three-dimensional Gaussian representation and target joint parameter data of the target articulated object are extracted from the target image; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, and the original Gaussian spheres include a three-dimensional covariance matrix, position information, joint probability information, transparency information, and color information. The three-dimensional covariance matrix is used to represent the three-dimensional scale information and direction information of the corresponding Gaussian sphere in an object state. The joint probability information includes a set of probability values for indicating that the corresponding Gaussian sphere belongs to each articulated structure and non-articulated structure. The target joint parameter data includes joint type and joint parameter information of the articulated structure of the target articulated object; Performing three-dimensional reconstruction of a specified object state on the target articulated object according to the original three-dimensional Gaussian representation and the target joint parameter data to obtain a target three-dimensional reconstructed image; The method of performing three-dimensional reconstruction of a specified object state on the target articulated object according to the original three-dimensional Gaussian representation and the target joint parameter data to obtain a target three-dimensional reconstructed image includes: Determine, based on the target joint parameter data, a set of SE3 transformation matrices for transforming the object state corresponding to the original three-dimensional Gaussian representation to the specified object state, wherein the SE3 transformation matrices correspond one-to-one to the articulated structure; According to the joint probability information of each original Gaussian sphere, the three-dimensional covariance matrix, position information, transparency information and color information of each original Gaussian sphere are transformed by a set of SE3 transformation matrices to obtain a set of target Gaussian spheres corresponding to each original Gaussian sphere; Determine the target three-dimensional Gaussian representations of the plurality of target Gaussian spheres corresponding to all the original Gaussian spheres; The target three-dimensional Gaussian representation is rendered at a specified viewing angle to obtain the target three-dimensional reconstructed image.

2. The method according to claim 1, characterized in that The step of extracting the original three-dimensional Gaussian representation and target joint parameter data of the target articulated object from the target image comprises: Performing feature extraction on the target image to obtain target image features; The target image features are subjected to three-dimensional Gaussian representation conversion and joint parameter extraction to obtain the original three-dimensional Gaussian representation and the target joint parameter data; wherein the target joint parameter data includes the translation vector of the translation joint, and / or the rotation axis, rotation center and rotation angle of the rotation joint.

3. The method according to claim 2, characterized in that The target image includes an image sequence composed of two-dimensional images in different object states; and extracting features from the target image to obtain target image features includes: Extracting and compressing image features of each of the two-dimensional images through a first set of two-dimensional convolutional networks to obtain initial image features corresponding to each of the two-dimensional images; The target image features are obtained by extracting temporal dynamic information from the initial image features corresponding to each of the two-dimensional images through a group of two-dimensional convolutional long short-term memory neural networks.

4. The method according to claim 2, characterized in that The step of performing three-dimensional Gaussian representation conversion and joint parameter extraction on the target image features to obtain the original three-dimensional Gaussian representation and the target joint parameter data includes: Converting the target image features into a three-dimensional Gaussian representation through a second set of two-dimensional convolutional networks to obtain the original three-dimensional Gaussian representation; The joint type and joint parameters in the target image features are extracted through a multi-layer perceptron to obtain the target joint parameter data.

5. The method according to claim 1, wherein The rendering of the target three-dimensional Gaussian representation at a specified viewing angle to obtain the target three-dimensional reconstructed image includes: Projecting the target three-dimensional Gaussian representation onto a two-dimensional plane corresponding to the specified viewing angle; The target Gaussian sphere projected onto the two-dimensional plane is rendered using the following transparency blending strategy: for each pixel to be rendered, the transparency and order of the overlapping Gaussian spheres in the projection direction of the pixel are calculated, and then the overlapping Gaussian spheres are traversed in order, and the final color of the pixel is obtained by weighting the transparency and combining the predicted color.

6. A device for 3D reconstruction of an articulated object, characterized in that: include: an image acquisition module, configured to acquire target images of at least two object states of a target articulated object to be reconstructed under a single viewing angle; wherein the articulation states of at least one articulated structure of the target articulated object in different object states are different; a data extraction module, configured to extract, from the target image, an original three-dimensional Gaussian representation and target joint parameter data of the target articulated object by extracting temporal dynamic information; wherein the original three-dimensional Gaussian representation includes a set of original Gaussian spheres corresponding to the target articulated object, the original Gaussian spheres including a three-dimensional covariance matrix, position information, joint probability information, transparency information, and color information; the three-dimensional covariance matrix is used to represent the three-dimensional scale information and orientation information of the corresponding Gaussian sphere in an object state; the joint probability information includes a set of probability values indicating that the corresponding Gaussian sphere belongs to each articulated structure and non-articulated structure; and the target joint parameter data includes joint type and joint parameter information of the articulated structure of the target articulated object; a three-dimensional reconstruction module, configured to perform three-dimensional reconstruction of a specified object state of the target articulated object based on the original three-dimensional Gaussian representation and the target joint parameter data, to obtain a target three-dimensional reconstructed image; The three-dimensional reconstruction module is specifically used to: determine a set of SE3 transformation matrices for transforming the object state corresponding to the original three-dimensional Gaussian representation to the specified object state based on the target joint parameter data, and the SE3 transformation matrix corresponds one-to-one to the articulated structure; according to the joint probability information of each original Gaussian ball, perform a set of SE3 transformation matrix conversions on the three-dimensional covariance matrix, position information, transparency information and color information of each original Gaussian ball to obtain a set of target Gaussian balls corresponding to each original Gaussian ball; determine multiple groups of target Gaussian balls corresponding to all the original Gaussian balls as the target three-dimensional Gaussian representation; render the target three-dimensional Gaussian representation at a specified perspective to obtain the target three-dimensional reconstructed image.

7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor executes the computer program, the method for three-dimensional reconstruction of an articulated object according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for three-dimensional reconstruction of an articulated object according to any one of claims 1 to 5 is executed.

Citation Information

Patent Citations

  • Method for establishing three-dimensional model of object and electrical equipment

    CN118447161A