Hinge object digital twin model construction method and device, terminal and medium

By utilizing multi-view images of hinged objects and a 3D Gaussian splash model, combined with a self-supervised deformation network, the problem of insufficient reconstruction accuracy of hinged objects in sparse states is solved, achieving high-precision hinge logic estimation and geometric reconstruction.

CN119251407BActive Publication Date: 2025-11-25SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411439048.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-15
Publication Date
2025-11-25
Estimated Expiration
2044-10-15

AI Technical Summary

Technical Problem

Existing digital twin construction technologies for hinged objects cannot accurately determine the motion state and hinge logic under sparse state input, and the reconstruction accuracy is insufficient.

Method used

By acquiring multi-view images of a hinged object in two states, a digital twin model of the hinged object is constructed using a 3D Gaussian splash model and a self-supervised deformation network. This includes deformation network prediction, self-supervised training, and joint optimization, to identify movable parts and hinge logic.

Benefits of technology

It achieves high-accuracy estimation of the motion state and hinge logic of hinged objects in sparse states, improving the geometric and visual accuracy of the reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251407B_ABST
    Figure CN119251407B_ABST
Patent Text Reader

Abstract

The application provides a hinge object digital twin model construction method and device, a terminal and a medium. The method comprises the following steps: constructing a three-dimensional Gaussian splash model in a starting state by using a first multi-view picture set of the hinge object in the starting state, identifying an initial movable Gaussian model based on a three-dimensional Gaussian splash model after deformation predicted by a deformation network, realizing complete self-supervised optimization of the deformation network based on a second multi-view picture set of the hinge object in an ending state and the model after deformation, and jointly optimizing the three-dimensional Gaussian splash model, a target movable Gaussian model obtained by re-segmenting the initial movable Gaussian model, and a target hinge motion parameter obtained by applying estimated initial hinge motion parameters to the initial movable Gaussian model to optimize the motion parameters, so as to construct a digital twin of the hinge object. In a completely self-supervised manner, the digital twin of the hinge object can be accurately reconstructed by using the sparse states of the hinge object, and the hinge logic can be accurately estimated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital twin model construction, in particular to a hinge object digital twin model construction method, device, terminal and medium. BACKGROUND

[0002] At present, constructing a digital twin model of a hinge object, accurately reconstructing and estimating the hinge logic, often plays an important role in the fields of robots, simulation, animation production, etc.

[0003] However, most of the traditional hinge object modeling methods are dependent on prior knowledge, which requires a large amount of labeled three-dimensional data as pre-training materials, which is very laborious and material-consuming, and when the hinge object to be reconstructed is not in the pre-training data set, it may not be able to estimate the correct hinge logic, and the existing pure self-supervised algorithm uses neural radiation field (NeRF) for reconstruction, and the unified optimization method leads to whether the result is correct or not strongly depends on the initialization of the parameters, resulting in a high error rate, and for the traditional reconstruction dynamic scene three-dimensional Gaussian algorithm, it cannot estimate the hinge logic, resulting in that in the case of sparse state input, the motion state of the object cannot be accurately interpolated.

[0004] For example, in the field of articulated object reconstruction, there is a way to reconstruct surfaces from a series of point clouds, but this way has the limitation of representing articulated objects with a unified surface, and the way of reconstructing articulated objects at the component level also has the problem of high dependence on the density of frames, where when the frames are sparse, the Gaussian motion process between the frames cannot be accurately predicted; or using three-dimensional Gaussian splash (3D-GS) technology to simulate organic articulated objects such as humans, body parts and animals, but the motion allowed by these objects is usually structurally complex, with a relatively large number of joint numbers connecting the topology and degrees of freedom, resulting in increased complexity of construction, and most previous work on the construction of digital twins of articulated objects has focused on the regeneration or construction of three-dimensional geometric representations of articulated objects, which involves estimating articulated parameters or functions to animate articulated objects to different articulated states, usually based on mesh-based articulated object representations, but mesh-based articulated object representations have relatively low geometric fidelity and visual realism, on the one hand, current reconstruction in a supervised learning manner usually requires consuming a large number of three-dimensional data sets with articulated information annotations, thereby also limiting the generalization ability to unseen object categories, and in self-supervised work with RGB observations as input, a neural radiation field (NeRF) is used as an intermediate representation and outputs a mesh model that compromises on geometric and appearance accuracy, but it exhibits instability and suboptimality problems during optimization; on the other hand, existing deformable three-dimensional Gaussian splash reconstruction work requires state-dense input or relies on human interaction for deformation and does not have the ability to automatically understand deformation logic.

[0005] In summary, the related technology of existing articulated object digital twin construction cannot accurately determine the motion state of the articulated object in the case of sparse state input, and cannot accurately estimate the correct articulated logic.

[0006] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. SUMMARY

[0007] The technical problem to be solved by the present application is to provide an articulated object digital twin model construction method, device, terminal and medium, which can accurately estimate the articulated logic of the articulated object and improve the geometric precision, visual precision and articulated structure estimation precision of the articulated object reconstruction, in view of the above defects of the prior art.

[0008] The technical solution adopted by the present application to solve the technical problem is as follows:

[0009] An articulated object digital twin model construction method, wherein the method comprises:

[0010] Obtain a first multi-view image set of the movable part in the hinge object when it is in the initial state before moving and a second multi-view image set of the movable part when it is in the final state after moving. Use the first multi-view image set to construct a three-dimensional Gaussian splash model of the hinge object when it is in the initial state.

[0011] The rigid transformation of each Gaussian point in the 3D Gaussian splash model relative to the final state is predicted using a pre-constructed deformable network to obtain the corresponding deformed 3D Gaussian splash model. Based on the deformed 3D Gaussian splash model, the initial movable Gaussian model corresponding to the movable part in the hinge object is identified. The deformable network is a neural network obtained by self-supervised training and optimization based on the second multi-view image set and the deformed 3D Gaussian splash model.

[0012] The pre-estimated initial hinge motion parameters are applied to the initial movable Gaussian model to optimize the initial hinge motion parameters to obtain the corresponding target hinge motion parameters, and the initial movable Gaussian model is re-segmented to obtain the target movable Gaussian model.

[0013] A joint optimization operation is performed based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable parts and the corresponding hinge logic.

[0014] In one implementation, the step of using a pre-constructed deformation network to predict the rigid transformation of each Gaussian point in the 3D Gaussian splash model relative to the final state to obtain the corresponding deformed 3D Gaussian splash model includes:

[0015] A preset gradient truncation operation is performed on the Gaussian point positions corresponding to each Gaussian point in the three-dimensional Gaussian splash model to obtain the processed Gaussian point positions.

[0016] The processed Gaussian point positions are input into a pre-constructed deformation network to obtain the deformation parameters of each Gaussian point relative to the final state, which are predicted and output by the deformation network.

[0017] Based on the deformation parameters, the corresponding deformed three-dimensional Gaussian splash model is determined.

[0018] The process of identifying the initial movable Gaussian model corresponding to the movable component in the hinge object based on the deformed three-dimensional Gaussian splash model includes:

[0019] The deformation parameters corresponding to the deformed three-dimensional Gaussian splash model are normalized to obtain the processed deformation parameters.

[0020] Determine whether the processed deformation parameter is less than a first preset threshold;

[0021] When the processed deformation parameter is not less than the first preset threshold, the processed deformation parameter is determined as the target deformation parameter, and the Gaussian model composed of the Gaussian points corresponding to the target deformation parameter is determined as the initial movable Gaussian model corresponding to the movable part in the hinge object.

[0022] In one implementation, the self-supervised training optimization based on the second multi-view image set and the deformed 3D Gaussian splash model further includes:

[0023] The deformed 3D Gaussian splash model is input into a pre-constructed Gaussian differentiable rendering framework to obtain the corresponding first rendered image;

[0024] Calculate the L1 loss and D_SSIM loss between the first rendered image and the first target image in the corresponding viewpoint of the second multi-view image set;

[0025] Based on the L1 loss and the D_SSIM loss, the corresponding first appearance loss is determined, and the corresponding ARAP loss is also determined.

[0026] Based on the first appearance loss and the ARAP loss, a corresponding deformation loss is determined so that the deformation network can be optimized through self-supervised training according to the deformation loss.

[0027] In one implementation, applying the pre-estimated initial hinge motion parameters to the initial movable Gaussian model to optimize the initial hinge motion parameters to obtain the corresponding target hinge motion parameters includes:

[0028] The pre-estimated initial hinge motion parameters are applied to the initial movable Gaussian model to control the initial movable Gaussian model to undergo motion deformation, resulting in a new deformed three-dimensional Gaussian splash model.

[0029] The new deformed 3D Gaussian splash model is input into the pre-constructed Gaussian differentiable rendering framework to obtain the corresponding second rendered image;

[0030] Calculate the loss between the second rendered image and the second target image under the corresponding view in the second multi-view image set to obtain the corresponding second appearance loss;

[0031] The coordinates of each Gaussian point corresponding to the new deformed three-dimensional Gaussian splash model are determined to obtain the first set of Gaussian point coordinates. Based on the coordinates of each Gaussian point corresponding to the three-dimensional Gaussian splash model and the deformation parameters of each Gaussian point relative to the final state predicted by the deformation network, the second set of Gaussian point coordinates is determined.

[0032] Calculate the distance between the coordinates of the first set of Gaussian points and the coordinates of the second set of Gaussian points to obtain the corresponding geometric loss;

[0033] Based on the appearance loss and the geometric loss, the initial hinge motion parameters are optimized to obtain the corresponding target hinge motion parameters.

[0034] In one implementation, the joint optimization operation based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, having the movable component and corresponding hinge logic, includes:

[0035] The target hinge motion parameters are applied to the target movable Gaussian model to determine the state of the three-dimensional Gaussian splash model after the target movable Gaussian model moves, thereby obtaining a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable component and the corresponding hinge logic.

[0036] In one implementation, the method for constructing a digital twin model of the hinged object further includes:

[0037] The first target 3D Gaussian digital twin model of the hinge object in the initial state is input into the pre-constructed Gaussian differentiable rendering framework to obtain the corresponding third rendered image;

[0038] Calculate the third appearance loss between the third rendered image and the third target image in the corresponding viewpoint of the first multi-view image set;

[0039] The second target three-dimensional Gaussian digital twin model of the hinge object in the final state is input into the pre-constructed Gaussian differentiable rendering framework to obtain the corresponding fourth rendered image;

[0040] Calculate the fourth appearance loss between the fourth rendered image and the fourth target image in the corresponding viewpoint of the second multi-view image set;

[0041] Dynamically assign corresponding weights to the third appearance loss and the fourth appearance loss to obtain the first target model construction loss and the second target model construction loss;

[0042] Based on the loss constructed from the first objective model and the loss constructed from the second objective model, the construction of the target three-dimensional Gaussian digital twin model corresponding to the hinge object is iteratively optimized.

[0043] In one implementation, when the hinge object is a hinge object with multiple movable parts, it further includes:

[0044] Obtain a third multi-view image set of another movable part in the hinge object when it is in the final state after the movement, and determine the constructed target three-dimensional Gaussian digital twin model corresponding to the hinge object as a new three-dimensional Gaussian splash model of the other movable part in the hinge object when it is in the initial state before the movement.

[0045] Based on the new 3D Gaussian splash model and the third multi-view image set, the steps of predicting the rigid transformation of each Gaussian point in the 3D Gaussian splash model relative to the final state using a pre-built deformation network to obtain the corresponding deformed 3D Gaussian splash model and its subsequent steps are re-executed to construct a new target 3D Gaussian digital twin model corresponding to the hinge object, having the movable part and the other movable part and the corresponding hinge logic.

[0046] The present invention also discloses a digital twin model construction device for a hinged object, wherein the device comprises:

[0047] The image acquisition module is used to acquire a first multi-view image set of the movable part in the hinge object when it is in the initial state before moving and a second multi-view image set of the movable part when it is in the final state after moving.

[0048] The first model building module is used to construct a three-dimensional Gaussian splash model of the hinge object when it is in the initial state using the first multi-view image set.

[0049] The model deformation module is used to predict the rigid transformation of each Gaussian point in the three-dimensional Gaussian splash model relative to the final state using a pre-constructed deformation network to obtain the corresponding deformed three-dimensional Gaussian splash model.

[0050] The moving model recognition module is used to identify the initial movable Gaussian model corresponding to the movable part in the hinge object based on the deformed three-dimensional Gaussian splash model; wherein, the deformed network is a neural network obtained by self-supervised training and optimization based on the second multi-view image set and the deformed three-dimensional Gaussian splash model;

[0051] The motion parameter optimization module is used to apply the pre-estimated initial hinge motion parameters to the initial movable Gaussian model to optimize the initial hinge motion parameters and obtain the corresponding target hinge motion parameters.

[0052] The movable model segmentation module is used to re-segment the initial movable Gaussian model to obtain the target movable Gaussian model;

[0053] A digital twin construction module is used to perform joint optimization operations based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable parts and corresponding hinge logic.

[0054] The present invention also discloses a terminal, comprising: a memory, a processor, and a digital twin model building program for a hinge object stored in the memory and executable on the processor, wherein the digital twin model building program for the hinge object, when executed by the processor, implements the steps of the digital twin model building method for the hinge object as described above.

[0055] The present invention also discloses a computer-readable storage medium storing a computer program that can be executed to implement the steps of the method for constructing a digital twin model of a hinged object as described above.

[0056] This invention provides a method, apparatus, terminal, and medium for constructing a digital twin model of a hinged object. The method for constructing the digital twin model of the hinged object includes: acquiring a first multi-view image set of a movable part in the hinged object in an initial state before movement and a second multi-view image set of the movable part in an final state after movement; constructing a three-dimensional Gaussian splash model of the hinged object in the initial state using the first multi-view image set; predicting the rigid transformation of each Gaussian point in the three-dimensional Gaussian splash model relative to the final state using a pre-constructed deformation network to obtain a corresponding deformed three-dimensional Gaussian splash model; and identifying the movable part in the hinged object based on the deformed three-dimensional Gaussian splash model. An initial movable Gaussian model corresponding to the moving part; wherein, the deformable network is a neural network obtained by self-supervised training and optimization based on the second multi-view image set and the deformed three-dimensional Gaussian splash model; the pre-estimated initial hinge motion parameters are applied to the initial movable Gaussian model to optimize the initial hinge motion parameters to obtain the corresponding target hinge motion parameters, and the initial movable Gaussian model is re-segmented to obtain the target movable Gaussian model; a joint optimization operation is performed based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, having the movable part and the corresponding hinge logic. Therefore, this invention, based on 3D Gaussian splashing technology, only requires multi-view images of the hinge object in two states to achieve self-supervised construction of the digital twin of the hinge object. That is, based on multi-view images of the hinge object in a sparse state, it uses a fully self-supervised approach and hierarchical optimization to reconstruct the hinge object stably and with high quality, obtaining a digital twin model of the hinge object. At the same time, it can also estimate the hinge logic of the movable parts in the hinge object with high accuracy, thereby improving the geometric accuracy, visual accuracy, and hinge structure estimation accuracy of the reconstructed hinge object. Attached Figure Description

[0057] Figure 1 This is a flowchart of a preferred embodiment of the method for constructing a digital twin model of a hinged object in this invention;

[0058] Figure 2 This is a schematic diagram of a deformation field estimation process disclosed in this invention;

[0059] Figure 3 This is a schematic diagram of a self-supervised hinge object reconstruction disclosed in this invention;

[0060] Figure 4 This is a schematic diagram of a multi-moving component reconstruction method disclosed in this invention;

[0061] Figure 5This is a comparative effect diagram disclosed in this invention;

[0062] Figure 6 This is a functional principle block diagram of a preferred embodiment of the digital twin model construction device for hinged objects in this invention;

[0063] Figure 7 This is a functional principle block diagram of a preferred embodiment of the terminal in this invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0065] Please see Figure 1 , Figure 1 This is a flowchart of the method for constructing a digital twin model of a hinged object in this invention. For example... Figure 1 As shown, the method for constructing a digital twin model of a hinged object according to an embodiment of the present invention includes:

[0066] Step S11: Obtain a first multi-view image set of the movable part in the hinge object when it is in the initial state before moving and a second multi-view image set of the movable part when it is in the final state after moving. Use the first multi-view image set to construct a three-dimensional Gaussian splash model of the hinge object when it is in the initial state.

[0067] In this embodiment, multi-view images of the hinge object in two states I0 and I1 are obtained to obtain corresponding image sets. Specifically, the first multi-view image set is obtained when a movable part of the hinge object is in the initial state I0 before moving, and the second multi-view image set is obtained when a movable part of the hinge object is in the final state I1 after moving. Then, the static three-dimensional Gaussian splash model in the initial state is reconstructed using the first multi-view image set. The three-dimensional Gaussian splash model is used as the underlying representation. By utilizing its differentiable rendering and explicit representation, the training efficiency and rendering quality under the new perspective can be improved.

[0068] It should be noted that achieving self-supervised reconstruction of a hinge object by using only multi-angle images of the two states before and after the hinge structure change is consistent with the behavior pattern of hinge object reconstruction in daily life. Hinge objects mainly include two types: translational hinge objects, such as drawers and paper cutters, and rotational hinge objects, such as refrigerators, ovens, and laptops. The former's hinge principle is that the moving part translates along a certain direction 'a' in a three-dimensional space, while the latter's is that the moving part rotates along an axis in space, where the axis is determined by a point 'p' in space.

[0069] Step S12: Using a pre-constructed deformable network, predict the rigid transformation of each Gaussian point in the three-dimensional Gaussian splash model relative to the final state to obtain the corresponding deformed three-dimensional Gaussian splash model, and identify the initial movable Gaussian model corresponding to the movable part in the hinge object based on the deformed three-dimensional Gaussian splash model; wherein, the deformable network is a neural network obtained by self-supervised training and optimization based on the second multi-view image set and the deformed three-dimensional Gaussian splash model.

[0070] In this embodiment, a pre-constructed deformation network is used to predict the rigid transformation of each Gaussian point in the 3D Gaussian splash model relative to the final state to obtain the corresponding deformed 3D Gaussian splash model. Then, based on the deformed 3D Gaussian splash model, the initial movable Gaussian model corresponding to the movable part in the hinge object is identified. Specifically, a preset gradient truncation operation is performed on the Gaussian point positions corresponding to each Gaussian point in the 3D Gaussian splash model to obtain the processed Gaussian point positions. The processed Gaussian point positions are input to the pre-constructed deformation network to obtain the deformation parameters of each Gaussian point relative to the final state predicted by the deformation network. Based on the deformation parameters, the corresponding deformed 3D Gaussian splash model is determined. The deformation parameters corresponding to the deformed 3D Gaussian splash model are normalized to obtain the processed deformation parameters. It is determined whether the processed deformation parameters are less than a first preset threshold. If the processed deformation parameters are not less than the first preset threshold, the processed deformation parameters are determined as the target deformation parameters, and the Gaussian model composed of the Gaussian points corresponding to the target deformation parameters is determined as the initial movable Gaussian model corresponding to the movable part in the hinge object. Understandably, a pre-constructed deformation network is obtained by using a second multi-view image set to pre-train the deformation network. The deformation network is then used to predict the deformation parameters of each Gaussian point in the 3D Gaussian splash model relative to the final state, thereby determining the deformed 3D Gaussian splash model. Based on these deformation parameters, the movable parts of the hinged object can also be identified, that is, the movable Gaussian parts in the 3D Gaussian splash model can be identified, thus obtaining the initial movable Gaussian model. In other words, the deformation network is used to predict the deformation field of the 3D Gaussian splash model, thereby roughly identifying the initial movable region on the 3D Gaussian splash model. The deformation parameters can include translational motion parameters and rotational motion parameters. For the deformed 3D Gaussian splash model, the motion type can also be classified according to the degree of deformation to determine whether the movable part in the deformed 3D Gaussian splash model is a translational hinge structure or a rotational hinge structure.

[0071] It should be noted that the initially identified movable region (initial movable Gaussian model) does not need to be very precise, because the pre-estimated motion parameters can be further refined in the subsequent joint optimization. The initial movable Gaussian model can provide a good approximation as an initialization for the subsequent joint optimization.

[0072] For example, see Figure 2 As shown, the deformable network uses the position x of each Gaussian point in the 3D Gaussian splash model. i As input, and output the predicted Gaussian point corresponding to... and Right now:

[0073]

[0074] in, Gaussian g i Translation, Gaussian g i The rotation.

[0075] It should be noted that, in order to predict the deformation field without changing the original position of each Gaussian point, we can do so at position x. i Add a gradient cutoff operation sg(·) to the network, and the network architecture of the deformed network can be a multi-layer perceptron (MLP), such as a four-layer perceptron, or other neural network structures, or heuristic methods such as simulated annealing.

[0076] In identifying the movable Gaussian component in the 3D Gaussian splash model, the displacement ‖δx‖2 and rotation ‖δr‖2 are first normalized to between 0 and 1, and the Gaussian components with ‖δx‖2≥θ or ‖δr‖2≥θ are identified as the movable Gaussian model G. m .

[0077] It should be noted that choosing a low first preset threshold θ often includes a static Gaussian portion, which may lead to suboptimal results when optimizing motion parameters later. Therefore, by setting a higher first preset threshold θ, a Gaussian portion with higher confidence can be obtained, resulting in an initial movable Gaussian model that can be used for motion parameter prediction of the entire portion. For example, the first preset threshold θ can be set to 0.3 as a threshold to identify whether the Gaussian portion composed of Gaussian points is a movable Gaussian portion.

[0078] Furthermore, the process of self-supervised training optimization of the deformable network under the supervision of the final state I1 can specifically include: inputting the deformed 3D Gaussian splash model into a pre-constructed Gaussian differentiable rendering framework to obtain the corresponding first rendered image; calculating the L1 loss and D_SSIM loss between the first rendered image and the first target image in the corresponding viewpoint of the second multi-view image set; determining the corresponding first appearance loss based on the L1 loss and the D_SSIM loss, and determining the corresponding ARAP loss; determining the corresponding deformation loss based on the first appearance loss and the ARAP loss so that the deformable network can perform self-supervised training optimization according to the deformation loss.

[0079] For example, the deformed 3D Gaussian splash model is input into a Gaussian differentiable rendering framework to obtain the rendered image. Then, the L1 loss and D_SSIM loss between the rendered image and the image at the corresponding viewpoint in the final state I1 are calculated. Finally, the L1 loss and D_SSIM loss are processed through a linear combination to obtain the corresponding first appearance loss (gradient error), i.e.:

[0080] L app =(1-λ)L1+λL D_SSIM ;

[0081] Among them, L app L1 represents the first appearance loss, and L2 represents the L1 loss. D_SSIM Let λ represent the D_SSIM loss, λ represent the weights corresponding to the D_SSIM loss, and 1-λ represent the weights corresponding to the L1 loss.

[0082] It should be noted that the expected deformation should be prior knowledge of local stiffness, so an additional ARAP (As-Rigid-As-Possible) loss L is added. arap ARAP loss is a regularized loss function used to learn deformable shape generators to encourage local rigidity in the motion of Gaussian points, where L is the ARAP loss. arap for:

[0083]

[0084] Where S represents the number of Gaussians, knn i,k This represents the k nearest Gaussians around Gaussian i, where k can be set to 20. Let ω represent the ARAP loss function between Gaussian i and Gaussian j. i,j Let x represent the weighting factor of the Gaussian pair. j Let x represent the three-dimensional coordinates of Gaussian j. i Let R represent the three-dimensional coordinates of Gaussian i. iThis represents the rotation matrix of Gaussian coordinate i relative to world coordinates. This represents the three-dimensional coordinates after Gaussian i-deformation to the final state. This represents the three-dimensional coordinates after the Gaussian J-deformation reaches its final state. Let λ represent the rotation matrix relative to world coordinates after the Gaussian i-transformation reaches its final state. ω This represents a hyperparameter, which can be set to 20.

[0085] Therefore, the final deformation loss is determined based on the first appearance loss and the ARAP loss, and then this deformation loss is passed back to the deformation network to achieve self-supervised training. For example, the deformation loss can be obtained by processing the first appearance loss and the ARAP loss through linear combination, i.e.:

[0086] L deform =L app +λ arap L arap ;

[0087] Wherein, the weight λ corresponding to the ARAP loss arap It can be set to 1.

[0088] Step S13: Apply the pre-estimated initial hinge motion parameters to the initial movable Gaussian model to optimize the initial hinge motion parameters to obtain the corresponding target hinge motion parameters, and re-segment the initial movable Gaussian model to obtain the target movable Gaussian model.

[0089] In this embodiment, the movable Gaussian portion in the three-dimensional Gaussian splash model is identified, and the initial movable Gaussian model G is obtained. m Then, a global rigid transformation can be applied to the initial movable Gaussian model G. mTo optimize the estimated motion parameters, the movable Gaussian part deforms along with the motion parameters. By adhering to the basic principle of maintaining visual and geometric consistency between the two states, more accurate motion parameters can be obtained. Specifically, the pre-estimated initial hinge motion parameters are applied to the initial movable Gaussian model to control its motion deformation, resulting in a new deformed 3D Gaussian splash model. This new deformed 3D Gaussian splash model is then input into a pre-constructed Gaussian differentiable rendering framework to obtain a corresponding second rendered image. The loss between the second rendered image and the second target image in the corresponding viewpoint of the second multi-view image set is calculated to obtain a corresponding second appearance loss. The coordinates of each Gaussian point corresponding to the new deformed 3D Gaussian splash model are determined to obtain a first set of Gaussian point coordinates. A second set of Gaussian point coordinates is determined based on the coordinates of each Gaussian point corresponding to the 3D Gaussian splash model and the deformation parameters of each Gaussian point predicted by the deformation network relative to the final state. The distance between the first set of Gaussian point coordinates and the second set of Gaussian point coordinates is calculated to obtain a corresponding geometric loss. The initial hinge motion parameters are optimized based on the appearance loss and the geometric loss to obtain the corresponding target hinge motion parameters.

[0090] It should be noted that when fitting translational motion with rotational motion parameters, the axes of the rotational motion parameters are usually very far from the object to ensure that the movable Gaussian part moves in a near-translational manner during the motion. Therefore, in the 3D Gaussian splash model, the movable Gaussian part is a rotational change, and the pivot point ‖p‖ is greater than the spatial radius R. The motion is considered translational, and the optimization parameters and motion function will be reinitialized to adapt to the translational motion, where R can be set to 4.

[0091] In this embodiment, the initial movable Gaussian model is re-segmented to obtain the target movable Gaussian model. It can be understood that re-segmenting the Gaussian parts of the moving components yields a more rigorous segmentation result for the movable components on the Gaussian model, which is then used for joint optimization. For example, a smaller second preset threshold, such as τ = 0.1, can be used to re-segment the movable Gaussian portion of the 3D Gaussian splash model to obtain a more rigorous segmentation of the movable Gaussian portion, resulting in a more accurate segmentation result, i.e., the target movable Gaussian model.

[0092] Step S14: Perform a joint optimization operation based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable parts and corresponding hinge logic.

[0093] In this embodiment, a more accurate 3D Gaussian splash model, a more precise movable Gaussian part, and target hinge motion parameters are obtained. These can be optimized together to find the optimal values. Specifically, the target hinge motion parameters are applied to the target movable Gaussian model to determine the state of the 3D Gaussian splash model after the target movable Gaussian model has moved. This yields a target 3D Gaussian digital twin model corresponding to the hinge object, containing the movable part and corresponding hinge logic. It is understood that the more accurate target hinge motion parameters are applied to the target movable Gaussian model to obtain its corresponding state after movement, while for the immovable part of the 3D Gaussian splash model, the initial and final states are the same.

[0094] It should be noted that after constructing the target 3D Gaussian digital twin of the hinged object in its final state, the target 3D Gaussian digital twin can be equipped with a high-quality colored appearance, thereby enabling it to play a greater role in downstream applications, such as transferring embodied intelligence from simulators to the real world during training.

[0095] It should also be noted that during the Gaussian densification and deletion process, the classification inaccuracy within the two masks can be regarded as Gaussian noise. This noise will be automatically updated during the Gaussian densification and deletion process, thereby achieving accurate segmentation results for the movable parts. Furthermore, since using only the initial state to supervise the optimization of Gaussian distribution may lead to overfitting of the Gaussian distribution to the initial state, the weights of the loss function on the initial and final states can be dynamically set to prevent the final Gaussian distribution from overfitting to the initial state. Specifically, the first target 3D Gaussian digital twin model of the hinge object in the initial state is input into a pre-built Gaussian differentiable rendering framework to obtain the corresponding third rendered image; the third appearance loss between the third rendered image and the third target image in the corresponding viewpoint of the first multi-view image set is calculated; the second target 3D Gaussian digital twin model of the hinge object in the final state is input into the pre-built Gaussian differentiable rendering framework to obtain the corresponding fourth rendered image; the fourth appearance loss between the fourth rendered image and the fourth target image in the corresponding viewpoint of the second multi-view image set is calculated; the corresponding weights are dynamically set for the third appearance loss and the fourth appearance loss to obtain the first target model construction loss and the second target model construction loss; the construction of the target 3D Gaussian digital twin model corresponding to the hinge object is iteratively optimized based on the first target model construction loss and the second target model construction loss.

[0096] For example, define the third appearance loss The weight is ω 0 Define the fourth appearance loss For ω1 , where ω 0 and ω 1 Both can be set to 0.5, i.e., ω 0 =ω 1 =0.5.

[0097] Furthermore, when the hinge object is a hinge object with multiple movable parts, it may further include: acquiring a third multi-view image set of another movable part in the hinge object when it is in the final state after movement, and determining the constructed target three-dimensional Gaussian digital twin model corresponding to the hinge object as a new three-dimensional Gaussian splash model of the other movable part in the hinge object when it is in the initial state before movement; based on the new three-dimensional Gaussian splash model and the third multi-view image set, re-executing the step of using a pre-constructed deformation network to predict the rigid transformation of each Gaussian point in the three-dimensional Gaussian splash model relative to the final state to obtain the corresponding deformed three-dimensional Gaussian splash model and its subsequent steps, so as to construct a new target three-dimensional Gaussian digital twin model corresponding to the hinge object with the movable part, the other movable part and the corresponding hinge logic.

[0098] Understandably, reconstructing a hinged object with multiple movable parts involves breaking down the entire reconstruction process into multiple sub-processes of modeling individual objects, and then solving each sub-process sequentially. In other words, it utilizes a sub-task approach to progressively reconstruct the hinged object with multiple movable parts. (See [link to relevant documentation]). Figure 3 As shown, throughout the reconstruction process, the same set of Gaussian distributions is maintained. The starting state of subsequent subprocesses is aligned with a certain state of the previous subprocess, ensuring that only one new movable part is introduced in a new subprocess. Before entering the next subprocess, the movable parts in the existing Gaussian distribution are configured to the starting state of the subsequent subprocess using the motion parameters derived from the previous subprocess. In each subprocess, a new identifier can be used to mark the movable mask, while the non-movable Gaussian distribution is optimized to update the new movable mask. Since the movable mask can be densified and pruned simultaneously with the Gaussian distribution, and the Gaussian distribution marked by the movable mask in the previous subprocess can also be optimized in the next new subprocess, the step of pre-training the initial Gaussian distribution can be skipped for new subprocesses, and the deformable network can be directly entered.

[0099] As can be seen, in this embodiment of the invention, the digital twin of the hinge object can be constructed by using multi-view images of the hinge object in two states based on the 3D Gaussian splashing technology. That is, based on the multi-view images of the hinge object in the sparse state, the hinge object is reconstructed stably and with high quality using a fully self-supervised method and a hierarchical optimization method, resulting in a digital twin model of the hinge object. At the same time, it can also estimate the hinge logic of the movable parts in the hinge object with high accuracy, thereby improving the geometric accuracy, visual accuracy and hinge structure estimation accuracy of the hinge object reconstruction.

[0100] It should be noted that simultaneously optimizing motion parameters, the movable Gaussian model (movable region), and the Gaussian representation (appearance) can lead to a tendency to seek local optima or computational failures. In other words, for movable parts of an object, rigid body transformations based on motion parameters can significantly influence the learning of the movable part's geometry and visual appearance, and vice versa. The optimization of movable parts and motion parameters can be significantly affected by rendering quality. Therefore, optimizing the overall geometry, appearance, and motion parameters of an object becomes extremely difficult. To address this, the overall objective can be processed in steps, with each step aiming to obtain initial parameter estimates close to the true values, thereby increasing the likelihood of reaching the optimal solution.

[0101] By constructing a digital twin of a hinged object using 3D Gaussian splashing technology, the method avoids limitations imposed by prior object types through self-supervised training. This involves acquiring multi-view RGB views of the hinged object in two states and estimating the moving parts and hinge motion parameters in a self-supervised manner without relying on any prior knowledge. Simultaneously, a 3D Gaussian splashed digital twin of the object is constructed to achieve the goal of synthesizing new perspectives. For example, see... Figure 4 As shown, two states I0 and I1 of the hinge object are determined. First, a static Gaussian model is trained using only the multi-view map in the first state I0. Then, deformation field estimation is implemented, i.e., supervised by the multi-view map in the second state I1, and the deformation field of the 3D Gaussian splash model is estimated using a deformation network. That is, based on the multi-view map in the second state I1, the deformation field of the Gaussian model is fitted using a deformation network without directly optimizing the motion parameters. Then, the estimation of the hinge structure is entered. That is, once the deformation field of the 3D Gaussian splash model is obtained, the movable parts on the 3D Gaussian splash model can be roughly identified, and the hinge motion parameters are estimated according to the deformation. Then, the rigid motion is projected onto the movable Gaussian model to optimize the hinge motion parameters to obtain the target hinge motion parameters. The movable parts on the roughly identified 3D Gaussian splash model are re-segmented to obtain a more accurate segmentation result. Finally, the Gaussian model representation, segmentation results and target hinge motion parameters are jointly optimized to obtain the final reconstruction result of the hinge object and the accurate hinge logic.

[0102] Experiments were conducted on the publicly available PARIS dataset using an open-source algorithm. Additionally, multiple objects were selected from the PartNet dataset to create a dataset for further experiments. (See [link to related documentation]). Figure 5 As shown, the results indicate that the technical solution of this application has significantly improved the geometric accuracy, visual accuracy, and hinge structure estimation accuracy of the reconstruction. It has also achieved a realistic effect by constructing real-life objects as experimental verification, and the training efficiency is also higher than that of traditional methods.

[0103] Furthermore, the application scenarios of the technical solution of this application may include, but are not limited to, the reconstruction of the hinge structure of a single human joint, the reconstruction of the hinge structure of a robotic arm, etc.

[0104] In one embodiment, such as Figure 6 As shown, based on the above-described method for constructing a digital twin model of a hinged object, the present invention also provides a corresponding apparatus for constructing a digital twin model of a hinged object, comprising:

[0105] Image acquisition module 11 is used to acquire a first multi-view image set when the movable part in the hinge object is in the initial state before movement and a second multi-view image set when the movable part is in the final state after movement.

[0106] The first model construction module 12 is used to construct a three-dimensional Gaussian splash model of the hinge object when it is in the initial state using the first multi-view image set.

[0107] Model deformation module 13 is used to predict the rigid transformation of each Gaussian point in the three-dimensional Gaussian splash model relative to the final state using a pre-constructed deformation network to obtain the corresponding deformed three-dimensional Gaussian splash model.

[0108] The moving model recognition module 14 is used to identify the initial movable Gaussian model corresponding to the movable part in the hinge object based on the deformed three-dimensional Gaussian splash model; wherein, the deformed network is a neural network obtained by self-supervised training and optimization based on the second multi-view image set and the deformed three-dimensional Gaussian splash model;

[0109] The motion parameter optimization module 15 is used to apply the pre-estimated initial hinge motion parameters to the initial movable Gaussian model to optimize the initial hinge motion parameters and obtain the corresponding target hinge motion parameters.

[0110] The movable model segmentation module 16 is used to re-segment the initial movable Gaussian model to obtain the target movable Gaussian model.

[0111] The digital twin construction module 17 is used to perform joint optimization operations based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable parts and corresponding hinge logic.

[0112] Figure 7 A schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal may include:

[0113] The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.

[0114] When the processor 502 executes the program, it implements the method for constructing a digital twin model of a hinge object provided in the above embodiments.

[0115] Furthermore, the terminal also includes:

[0116] Communication interface 503 is used for communication between memory 501 and processor 502.

[0117] The memory 501 is used to store computer programs that can run on the processor 502.

[0118] The memory 501 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0119] If the memory 501, processor 502, and communication interface 503 are implemented independently, they can be interconnected via a bus to communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in the diagram, but this does not imply that there is only one bus or one type of bus.

[0120] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.

[0121] Processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0122] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for constructing a digital twin model of a hinged object.

[0123] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0124] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0125] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0126] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can read and execute instructions from or in conjunction with such an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). In addition, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically by optically scanning paper or other media, followed by editing, interpreting or otherwise processing as necessary, and then stored in computer memory.

[0127] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0128] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.

[0129] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0130] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

[0131] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for constructing a digital twin model of a hinged object, characterized in that, The method includes: Obtain a first multi-view image set of the movable part in the hinge object when it is in the initial state before moving and a second multi-view image set of the movable part when it is in the final state after moving. Use the first multi-view image set to construct a three-dimensional Gaussian splash model of the hinge object when it is in the initial state. The rigid transformation of each Gaussian point in the 3D Gaussian splash model relative to the final state is predicted using a pre-constructed deformable network to obtain the corresponding deformed 3D Gaussian splash model. Based on the deformed 3D Gaussian splash model, the initial movable Gaussian model corresponding to the movable part in the hinge object is identified. The deformable network is a neural network obtained by self-supervised training and optimization based on the second multi-view image set and the deformed 3D Gaussian splash model. The pre-estimated initial hinge motion parameters are applied to the initial movable Gaussian model to optimize the initial hinge motion parameters to obtain the corresponding target hinge motion parameters, and the initial movable Gaussian model is re-segmented to obtain the target movable Gaussian model. A joint optimization operation is performed based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable parts and the corresponding hinge logic.

2. The method for constructing a digital twin model of a hinged object according to claim 1, characterized in that, The step of using a pre-constructed deformation network to predict the rigid transformation of each Gaussian point in the 3D Gaussian splash model relative to the final state to obtain the corresponding deformed 3D Gaussian splash model includes: A preset gradient truncation operation is performed on the Gaussian point positions corresponding to each Gaussian point in the three-dimensional Gaussian splash model to obtain the processed Gaussian point positions. The processed Gaussian point positions are input into a pre-constructed deformation network to obtain the deformation parameters of each Gaussian point relative to the final state, which are predicted and output by the deformation network. Based on the deformation parameters, the corresponding deformed three-dimensional Gaussian splash model is determined. The process of identifying the initial movable Gaussian model corresponding to the movable component in the hinge object based on the deformed three-dimensional Gaussian splash model includes: The deformation parameters corresponding to the deformed three-dimensional Gaussian splash model are normalized to obtain the processed deformation parameters. Determine whether the processed deformation parameter is less than a first preset threshold; When the processed deformation parameter is not less than the first preset threshold, the processed deformation parameter is determined as the target deformation parameter, and the Gaussian model composed of the Gaussian points corresponding to the target deformation parameter is determined as the initial movable Gaussian model corresponding to the movable part in the hinge object.

3. The method for constructing a digital twin model of a hinged object according to claim 2, characterized in that, The self-supervised training optimization based on the second multi-view image set and the deformed 3D Gaussian splash model also includes: The deformed 3D Gaussian splash model is input into a pre-constructed Gaussian differentiable rendering framework to obtain the corresponding first rendered image; Calculate the L1 loss and D_SSIM loss between the first rendered image and the first target image in the corresponding viewpoint of the second multi-view image set; Based on the L1 loss and the D_SSIM loss, the corresponding first appearance loss is determined, and the corresponding ARAP loss is also determined. Based on the first appearance loss and the ARAP loss, a corresponding deformation loss is determined so that the deformation network can be optimized through self-supervised training according to the deformation loss.

4. The method for constructing a digital twin model of a hinged object according to claim 3, characterized in that, The step of applying the pre-estimated initial hinge motion parameters to the initial movable Gaussian model to optimize the initial hinge motion parameters and obtain the corresponding target hinge motion parameters includes: The pre-estimated initial hinge motion parameters are applied to the initial movable Gaussian model to control the initial movable Gaussian model to undergo motion deformation, resulting in a new deformed three-dimensional Gaussian splash model. The new deformed 3D Gaussian splash model is input into the pre-constructed Gaussian differentiable rendering framework to obtain the corresponding second rendered image; Calculate the loss between the second rendered image and the second target image under the corresponding view in the second multi-view image set to obtain the corresponding second appearance loss; The coordinates of each Gaussian point corresponding to the new deformed three-dimensional Gaussian splash model are determined to obtain the first set of Gaussian point coordinates. Based on the coordinates of each Gaussian point corresponding to the three-dimensional Gaussian splash model and the deformation parameters of each Gaussian point relative to the final state predicted by the deformation network, the second set of Gaussian point coordinates is determined. Calculate the distance between the coordinates of the first set of Gaussian points and the coordinates of the second set of Gaussian points to obtain the corresponding geometric loss; Based on the appearance loss and the geometric loss, the initial hinge motion parameters are optimized to obtain the corresponding target hinge motion parameters.

5. The method for constructing a digital twin model of a hinged object according to claim 1, characterized in that, The joint optimization operation based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, including the movable component and corresponding hinge logic, comprises: The target hinge motion parameters are applied to the target movable Gaussian model to determine the state of the three-dimensional Gaussian splash model after the target movable Gaussian model moves, thereby obtaining a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable component and the corresponding hinge logic.

6. The method for constructing a digital twin model of a hinged object according to claim 1, characterized in that, Also includes: The first target 3D Gaussian digital twin model of the hinge object in the initial state is input into the pre-constructed Gaussian differentiable rendering framework to obtain the corresponding third rendered image; Calculate the third appearance loss between the third rendered image and the third target image in the corresponding viewpoint of the first multi-view image set; The second target three-dimensional Gaussian digital twin model of the hinge object in the final state is input into the pre-constructed Gaussian differentiable rendering framework to obtain the corresponding fourth rendered image; Calculate the fourth appearance loss between the fourth rendered image and the fourth target image in the corresponding viewpoint of the second multi-view image set; Dynamically assign corresponding weights to the third appearance loss and the fourth appearance loss to obtain the first target model construction loss and the second target model construction loss; Based on the loss constructed from the first objective model and the loss constructed from the second objective model, the construction of the target three-dimensional Gaussian digital twin model corresponding to the hinge object is iteratively optimized.

7. The method for constructing a digital twin model of a hinged object according to any one of claims 1 to 6, characterized in that, When the hinge object is a hinge object with multiple movable parts, it further includes: Obtain a third multi-view image set of another movable part in the hinge object when it is in the final state after the movement, and determine the constructed target three-dimensional Gaussian digital twin model corresponding to the hinge object as a new three-dimensional Gaussian splash model of the other movable part in the hinge object when it is in the initial state before the movement. Based on the new 3D Gaussian splash model and the third multi-view image set, the steps of predicting the rigid transformation of each Gaussian point in the 3D Gaussian splash model relative to the final state using a pre-built deformation network to obtain the corresponding deformed 3D Gaussian splash model and its subsequent steps are re-executed to construct a new target 3D Gaussian digital twin model corresponding to the hinge object, having the movable part and the other movable part and the corresponding hinge logic.

8. A device for constructing a digital twin model of a hinged object, characterized in that, The device includes: The image acquisition module is used to acquire a first multi-view image set of the movable part in the hinge object when it is in the initial state before moving and a second multi-view image set of the movable part when it is in the final state after moving. The first model building module is used to construct a three-dimensional Gaussian splash model of the hinge object when it is in the initial state using the first multi-view image set. The model deformation module is used to predict the rigid transformation of each Gaussian point in the three-dimensional Gaussian splash model relative to the final state using a pre-constructed deformation network to obtain the corresponding deformed three-dimensional Gaussian splash model. The moving model recognition module is used to identify the initial movable Gaussian model corresponding to the movable part in the hinge object based on the deformed three-dimensional Gaussian splash model; wherein, the deformed network is a neural network obtained by self-supervised training and optimization based on the second multi-view image set and the deformed three-dimensional Gaussian splash model; The motion parameter optimization module is used to apply the pre-estimated initial hinge motion parameters to the initial movable Gaussian model to optimize the initial hinge motion parameters and obtain the corresponding target hinge motion parameters. The movable model segmentation module is used to re-segment the initial movable Gaussian model to obtain the target movable Gaussian model; A digital twin construction module is used to perform joint optimization operations based on the three-dimensional Gaussian splash model, the target movable Gaussian model, and the target hinge motion parameters to construct a target three-dimensional Gaussian digital twin model corresponding to the hinge object, which has the movable parts and corresponding hinge logic.

9. A terminal, characterized in that, include: The device includes a memory, a processor, and a digital twin model building program for a hinged object stored in the memory and executable on the processor. When executed by the processor, the digital twin model building program for the hinged object implements the steps of the digital twin model building method for a hinged object as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed to implement the steps of the method for constructing a digital twin model of a hinged object as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for constructing relightable three-dimensional digital human based on multi-view video

    CN116934948A

  • Hinged object modeling method and system, electronic equipment and storage medium

    CN118037957A