Learning Articulated Shape Reconstruction from Images
Through the computing system, the input image and grid model is processed, the camera model and object deformation data are used to render the rendered image, and the model parameters are adjusted through the loss function, the problem of inaccurate 3D modeling in the existing technology under low data conditions is solved, and accurate 3D shape reconstruction under finite data is realized.
Patent Information
- Application Number
- CN202080102368.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-12-21
AI Technical Summary
Existing 3D modeling methods are difficult to accurately reconstruct 3D shapes under low data conditions, especially when image observation is poor, it is prone to produce inaccurate 3D structural hallucinations.
A computer-implemented method is adopted to obtain the input image and the current grid model through a computing system, and to process the input image using the camera model to obtain camera parameters and object deformation data. The rendered image of the object can be differentially rendered based on these data, and the camera model and grid model are modified through the gradient of the loss function.
This method can accurately reconstruct 3D shapes in limited image data, avoiding dependence on prior shape templates and category information, and extending the scope of objects that can generate accurate 3D models.
Smart Images

Figure CN115769259B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to three-dimensional (“3D”) reconstruction. More specifically, the present disclosure relates to systems and methods for reconstructing a model of an object (e.g., a non-rigid object) from imagery (e.g., RGB input images). Background Art
[0002] Modeling a 3D entity is the process of developing a mathematical representation of a three-dimensional object (e.g., the surface of an object). Modeling the dynamics of a 3D entity can involve using data describing the object to construct a 3D mesh shape of the object that can be deformed into various poses.
[0003] Some standard 3D modeling methods rely on 3D supervision, such as synthetic rendering and depth scanning. However, due to current sensor designs, depth data is often difficult to obtain and even more difficult to scale up. Other standard 3D modeling methods rely on inferring 3D shapes from point trajectories of multiple static images. These standard models are able to achieve high accuracy on benchmarks with rich training labels; however, they do not generalize well within low-data regimes. Additionally, this method often produces hallucinations of inaccurate 3D structures when image observations are poor.
[0004] Although progress has been made in the field by leveraging multi-view data recordings without relying on strong shape priors, such results are limited to static scenes. Summary of the Invention
[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0006] One example aspect of the present disclosure relates to a computer-implemented method for determining the 3D object shape from imagery. The method includes a computing system obtaining one or more computing devices, an input image depicting the object, and a current mesh model of the object. The method includes the computing system processing the input image using a camera model to obtain camera parameters of the input image and object deformation data. The camera parameters describe the camera pose of the input image. The object deformation data describes one or more deformations of the current mesh model relative to the shape of the object shown in the input image. The method includes the computing system differentiably rendering a rendered image of the object based on the camera parameters, the object deformation data, and the current mesh model. The method includes the computing system evaluating a loss function that compares one or more characteristics of the input image of the object with one or more characteristics of the rendered image of the object. The method includes the computing system modifying one or more values of one or both of the camera model and the current mesh model based on the gradient of the loss function.
[0007] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0008] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and together with the description serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] A detailed discussion of embodiments directed to those of ordinary skill in the art is set forth in the specification with reference to the accompanying drawings, in which:
[0010] Figure 1A A block diagram of an example computing system performing Learning Articulated Shape Reconstruction (LASR) in accordance with an example embodiment of the present disclosure is depicted.
[0011] Figure 1B A block diagram of an example computing device performing articulated shape reconstruction in accordance with an example embodiment of the present disclosure is depicted.
[0012] Figure 1C A block diagram of an example computing device performing articulated shape reconstruction in accordance with an example embodiment of the present disclosure is depicted.
[0013] Figure 2 A block diagram of an example 3D reconstruction technique in accordance with an example embodiment of the present disclosure is depicted.
[0014] Figure 3 A flowchart outlining a method for learning articulated shape reconstruction is depicted.
[0015] Figure 4 An example of a coarse-to-fine reconstruction is shown.
[0016] Figure 5 An example of a visual comparison of mesh reconstructions of humans and animals using various mesh reconstruction methods is shown.
[0017] Figure 6 Another example of a visual comparison of mesh reconstructions of humans and animals using various mesh reconstruction methods is shown.
[0018] Figure 7 An example of key point transfer using various key point transfer methods is shown.
[0019] Figure 8 An example of shape and articulation reconstruction results at different timestamps using various shape and articulation reconstruction methods is shown.
[0020] Figure 9Shows a visual comparison of the reconstruction of a nearly rigid video sequence between the COLMAP and LASR methods.
[0021] Figure 10 Shows an example of the results of an ablation study of camera and rigid shape optimization using various methods and articulated shape optimization using various methods.
[0022] Figure 11 Depicts a flowchart of an example method for performing LASR according to an example embodiment of the present disclosure.
[0023] Reference numerals repeated across multiple figures are intended to identify the same features in various embodiments. Detailed Description
[0024] Overview
[0025] In general, the present disclosure relates to a computing system and method that can be used to reconstruct the 3D shape of an object from images of the object, such as, for example, a monocular video of the object. Specifically, the present disclosure provides a general pipeline for learning articulated shape reconstruction (which may be referred to as LASR) from one or more images. The pipeline can reconstruct models of rigid or non-rigid 3D shapes. Specifically, the example pipeline described herein can automatically decompose non-rigidly deformed shapes into rigid motions near a rigid skeleton. The pipeline incorporates a synthetic analysis strategy and forward renders contours, optical flow, and color images that can be compared with video observations to adjust the internal parameters of the model. By inverting the rendering pipeline and incorporating image analysis techniques such as optical flow, the pipeline can recover the mesh of the 3D model from one or more images input by the user.
[0026] More specifically, the example 3D modeling pipeline can perform a synthetic analysis task in which a machine-learned mesh model of an object can be jointly learned with a machine-learned camera model by minimizing a loss function that evaluates the difference between one or more input images of the object and one or more rendered images of the object. Additionally, a shape model library can be constructed from a single set of one or more images of the object. The pipeline can solve the inverse graphics problem of recovering the 3D object shape (e.g., spatio-temporal deformation) and camera trajectory (e.g., intrinsic) in order to fit video or image frame observations such as contours, raw pixels, and optical flow. As another example, a shape model library can be constructed by performing the pipeline on multiple images depicting multiple objects.
[0027] An example method of a model-free method for 3D shape learning from one or more images may include obtaining an input image depicting an object and a current mesh model of the object. Specifically, ground truth may be included in a set of one or more images. For example, the ground truth may be one or more monocular sequences, such as a video captured by a monocular camera. As another example, the monocular sequence may have a segmentation of the foreground object.
[0028] The input image may be processed with a machine-learned camera model. The machine-learned camera model may predict information about the ground truth data. Specifically, the information may include camera parameters and / or object deformation data. The camera parameters may describe the camera pose of the input image (e.g., relative to a reference position and / or orientation). The object deformation data may describe one or more deformations of the current mesh model. For example, the deformation of the current mesh model may be a relative change between the shape of the current mesh model and the object shown in the image.
[0029] A rendered image of the object may be differentiably rendered (e.g., using differentiable rendering techniques). The rendered image may be based on the camera parameters, the object deformation data, and the current mesh model. The rendered image may depict the current mesh model deformed according to the object deformation data and the camera pose described by the camera parameters.
[0030] A loss function may be evaluated that compares one or more characteristics of the input image of the object with one or more characteristics of the rendered image of the object. One or more values of one or both of the machine-learned camera model and the current mesh model may be modified based on the loss function. For example, modifying one or both of the machine-learned camera model and the current mesh model may be at least partially based on a gradient signal, where the gradient signal describes the gradient of the loss function with respect to the parameters of the model.
[0031] In some embodiments, evaluating the loss function can include evaluating the difference between one or more input images of an object and one or more rendered images of the object. Specifically, the camera pose at a particular frame can be included in the loss function evaluation. Even more specifically, the rotation of a particular bone about its parent joint can be included in the loss function evaluation. Even more specifically, the 3D coordinates of the vertices of the rest shape can be included in the loss function evaluation. For example, the motion regularization terms used in evaluating the loss function can include a temporal smoothness term, a minimum motion term, and an as-rigid-as-possible term. As yet another example, the shape regularization terms used in evaluating the loss function can include a Laplacian smoothness term and a normalization term to remove the ambiguity of multiple solutions up to a rigid transformation. One or more rendered images can include images rendered based on a machine-learned mesh model combined with camera parameters generated by a machine-learned camera model. The data generated by the machine-learned camera model can be derived from one or more input images. The pipeline can further instruct the system to receive an additional set of camera parameters. The pipeline can again further instruct the system to render additional rendered images of the object at least partially based on the machine-learned mesh model and the additional set of camera parameters.
[0032] In some embodiments, evaluating the loss function can include determining a first flow (e.g., using one or more optical flow techniques, etc.) and a second flow (e.g., based on known changes across image rendering). The first flow can be used for the input images, while the second flow can be used for the rendered images. The loss function can be evaluated at least in part based on a comparison of the first flow and the second flow.
[0033] In some embodiments, evaluating the loss function can include determining a first contour (e.g., using one or more segmentation techniques, etc.) and a second contour (e.g., based on the known position of the object within the rendered image). The first contour can be used for the input images, while the second contour can be used for the rendered images. The loss function can be evaluated at least in part based on a comparison of the first contour and the second contour.
[0034] In some embodiments, evaluating the loss function can include determining first texture data (e.g., using raw pixel data and / or various feature extraction techniques) and second texture data (e.g., using known texture data from the rendered images). The first texture data can be used for the input images, while the second texture data can be used for the rendered images. The loss function can be evaluated at least in part based on a comparison of the first texture data and the second texture data.
[0035] As an example, evaluating a loss function can include generating a gradient signal. By comparing one or more characteristics of an input image of an object with one or more characteristics of a rendered image of the object, a gradient signal can be generated for the loss function. As an example, by comparing a first stream of the input image with a second stream of the rendered image, a gradient signal can be generated for the loss function. As another example, by comparing a first contour of the input image with a second contour of the rendered image, a gradient signal can be generated for the loss function. As yet another example, by comparing first texture data associated with the input image with second texture data associated with the rendered image, a gradient signal can be generated for the loss function.
[0036] In some embodiments, obtaining an input image depicting the object can include selecting a canonical image. Specifically, the canonical image can be an image frame from a video. The canonical image can be selected automatically or manually. As an example for selecting a canonical input image, one or more candidate frames can be selected. The loss of each candidate frame among the candidate frames can be evaluated. The candidate frame having the lowest final loss can be selected as the canonical frame.
[0037] In some embodiments, a mesh model can include various shapes for constructing the mesh model. For example, the mesh model can be a polygonal mesh. The polygonal mesh can include a set of vertices, a plurality of joints, a plurality of blend skinning weights of the plurality of joints relative to the plurality of vertices, and / or edges and faces defining the shape of a polyhedral object. Specifically, the faces of the polygonal mesh can be composed of concave polygons, polygon with holes, simple convex polygons, and other more specific structures (e.g., triangles, quadrilaterals, etc.). As another example, the mesh model can be initialized to a subdivided icosahedron projected onto a sphere. In some embodiments, a linear blend skinning algorithm can be used to deform the mesh model. In some embodiments, the plurality of joints and the plurality of blend skinning weights are learnable.
[0038] In some embodiments, camera parameters can describe the object-to-camera transformation of an input image. As an example, different views of a 3D object can be created by applying a rigid 3D transformation matrix to a matrix of object-centered coordinates. By applying the object-to-camera transformation, the matrix of object-centered coordinates can be transformed to camera-centered coordinates. Even more specifically, the object for which the transformation is computed can have a known geometric model. Calibration of the camera can start by capturing an image of a real-world object and locating a set of fiducial points in the image. Any suitable technique can be used to find the position (i.e., pose) of the fiducial points in the image.
[0039] In some embodiments, a machine-learned camera model can include a convolutional neural network. For example, the convolutional neural network can estimate camera pose. Specifically, the convolutional neural network can represent the camera pose using its position vector and orientation quaternion. The convolutional neural network can be trained to determine the camera pose by being trained to minimize the loss between the ground truth data and the estimated pose. As another example, the convolutional neural network can predict camera extrinsic factors (e.g., the position of the camera in the world, what direction the camera is pointing, etc.). Specifically, the camera extrinsic factors can be at least partially based on camera calibration.
[0040] In some embodiments, camera parameters can describe intrinsic camera parameters (e.g., focal length, image center, aspect ratio, etc.). The intrinsic camera parameters can be described for the input image. In particular, the intrinsic camera parameters can be at least partially based on camera calibration.
[0041] In some embodiments, the dynamics of the skeleton can be shared. For example, if one skeleton reaches a determined threshold of similarity to another skeleton with more data or a better 3D model (e.g., in a library), the system can apply information from one 3D model to the other 3D model to improve the second 3D model (e.g., if there is not enough data to create the second 3D model with the same accuracy).
[0042] In some embodiments, keypoint constraints can be incorporated. Additionally, shape template prior knowledge can speed up the inference and improve the accuracy.
[0043] Accordingly, the present disclosure provides a template-free method for 3D shape learning from one or more images (e.g., a single video). Example embodiments employ a synthetic analysis strategy and forward-render silhouettes, optical flow, and / or color images, which are compared with the video observations to adjust the model's camera, shape, and / or motion parameters. The proposed technique is capable of accurately reconstructing rigid and non-rigid 3D shapes (e.g., categories of humans, animals, and in the wild) without relying on category or 3D shape prior knowledge.
[0044] The systems and methods of the present disclosure provide a number of technical effects and benefits. As an example technical effect, the proposed techniques are capable of performing articulated shape reconstruction from limited image data (e.g., monocular video) without relying on prior templates or category information. Specifically, example embodiments leverage two-frame optical flow to overcome the inherent incompleteness and motion estimation problems of non-rigid structures. By enabling model reconstruction from limited data and without relying on object- or category-specific prior knowledge, the techniques described herein are able to expand the range of objects for which accurate 3D models can be generated. Specifically, many existing non-rigid shape reconstruction methods rely on prior shape templates, such as SMPL for humans, SMAL for quadrupeds, and other category-specific 3D scans. In contrast, the proposed systems and methods can jointly recover the camera, shape, and articulation from monocular video of an object without using a shape template or category information. By reducing the reliance on prior knowledge, the proposed systems and methods can be applied to a wider range of non-rigid shapes and better fit the data.
[0045] As another example technical effect, some example embodiments automatically recover non-rigid shapes under the constraint of a rigid skeleton under linear blend skinning. Example embodiments may combine coarse-to-fine remeshing with soft symmetry constraints to recover a high-quality mesh.
[0046] Example experiments further described herein and conducted on example embodiments of the proposed techniques demonstrate state-of-the-art reconstruction performance on the BADJA animal video dataset, strong performance against model-based methods for humans, and higher accuracy for two animated animals compared to A-CSM and SMALify that use shape templates.
[0047] Example devices and systems
[0048] Figure 1A A block diagram of an example computing system 100 that performs articulated shape reconstruction in accordance with an example embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0049] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0050] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, magnetic disks, etc. and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0051] In some embodiments, the user computing device 102 can store or include one or more 3D reconstruction models 120. For example, the 3D reconstruction model 120 can be or can additionally include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. Neural networks can include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Refer to Figure 2 Discuss example 3D reconstruction model 120.
[0052] In some embodiments, one or more machine learning models 120 can be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some embodiments, the user computing device 102 can implement multiple parallel instances of a single 3D reconstruction model 120 (e.g., perform parallel 3D reconstruction across multiple instances of input images).
[0053] More specifically, the 3D reconstruction model can jointly recover the camera, shape, and articulation from a sequence of images of an object without using a shape template or category information.
[0054] Additionally or alternatively, one or more machine learning models 140 can be included in or stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning model 140 can be implemented by the server computing system 140 as part of a web service (e.g., a streaming service). Thus, one or more models 120 can be stored and implemented at the user computing device 102, and / or one or more models 140 can be stored and implemented at the server computing system 130.
[0055] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include microphones, traditional keyboards, or other devices that a user may use to provide user input.
[0056] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operably connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0057] In some embodiments, the server computing system 130 includes or is implemented by one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0058] As described above, the server computing system 130 may store or otherwise include one or more 3D reconstruction models 140. For example, the model 140 may be or may include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Refer to Figure 2 for a discussion of example models 140.
[0059] The user computing device 102 and / or the server computing system 130 may train the model 120 and / or 140 via interaction with a training computing system 150 communicatively coupled through a network 180. The training computing system 150 may be separate from the server computing system 130 or may be a part of the server computing system 130.
[0060] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash devices, disks, etc. and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some embodiments, the training computing system 150 includes or is implemented by one or more server computing devices.
[0061] The training computing system 150 can include a model trainer 160 that uses various training or learning techniques, such as, for example, backpropagation of error, to train the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130. For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over multiple training iterations.
[0062] In some embodiments, performing backpropagation of error can include performing truncated backpropagation over time. The model trainer 160 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.
[0063] Specifically, the model trainer 160 can train the 3D reconstruction model 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, a set of one or more images. In some embodiments, one or more of the images can be directed to an object of interest. In some embodiments, one or more of the images can be strung together to form a video. In some embodiments, the video can be a monocular video.
[0064] In some embodiments, if the user has provided consent, training examples can be provided by the user computing device 102. Thus, in such embodiments, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In certain cases, this process can be referred to as personalizing the model.
[0065] The model trainer 160 includes computer logic for providing the required functionality. The model trainer 160 can be implemented with hardware, firmware, and / or software that controls a general-purpose processor. For example, in some embodiments, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, the model trainer 160 includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium.
[0066] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 180 can be carried via any type of wired and / or wireless connection using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL).
[0067] Figure 1A An example computing system that can be used to implement the present disclosure is shown. Other computing systems can also be used. For example, in some embodiments, the user computing device 102 can include the model trainer 160 and the training data set 162. In such embodiments, the model 120 can be trained and used locally at the user computing device 102. In some such embodiments, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0068] Figure 1B A block diagram of an example computing device 10 executing in accordance with an example embodiment of the present disclosure is depicted. The computing device 10 can be a user computing device or a server computing device.
[0069] The computing device 10 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and a model of machine learning. For example, each application can include a model of machine learning. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.
[0070] As Figure 1B shown, each application can communicate with multiple other components of the computing device, such components being, for example, one or more sensors, a context manager, a device status component, and / or additional components. In some embodiments, each application can communicate with each device component using an API (e.g., a common API). In some embodiments, the API used by each application is specific to that application.
[0071] Figure 1C Depicts a block diagram of an example computing device 50 executing in accordance with an example embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.
[0072] The computing device 50 includes a plurality of applications (e.g., Application 1 through N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some embodiments, each application may communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a common API across all applications).
[0073] The central intelligence layer includes a number of machine learning models. For example, as Figure 1C shown, a corresponding machine learning model may be provided for each application, and the corresponding machine learning model is managed by the central intelligence layer. In other embodiments, two or more applications may share a single machine learning model. For example, in some embodiments, the central intelligence layer may provide a single model for all applications. In some embodiments, the central intelligence layer is included in or implemented by the operating system of the computing device 50.
[0074] The central intelligence layer may communicate with a central device data layer. The central device data layer may be a centralized data repository of the computing device 50. As Figure 1C shown, the central device data layer may communicate with a number of other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some embodiments, the central device data layer may communicate with each device component using an API (e.g., a private API).
[0075] Example Model Setup
[0076] Figure 2 Depicts a block diagram of an example 3D reconstruction pipeline 200 in accordance with an example embodiment of the present disclosure. In some embodiments, the 3D reconstruction pipeline 200 may be executed to receive a collection of one or more images 204 depicting an object of interest (e.g., a monocular video directed at the object of interest), and as a result of receiving the images 204, provide a reconstructed 3D model 206 of the object of interest. In some embodiments, the 3D reconstruction pipeline 200 may include performing inverse graphics optimization 202, which includes solving an inverse graphics optimization problem that may jointly recover the static shape, skinning weights, articulation, and / or camera parameters of the object through video-based optimization.
[0077] Example Methods
[0078] As Figure 3 shown, given an input of one or more images (such as a monocular video {I t}), some example embodiments of the present disclosure utilize certain methods of solving non-rigid 3D shape and motion estimation problems as synthesis analysis tasks. The methods described below can solve for "low-level" shape and motion at a certain scale by giving appropriate video measurements, although the problem has an underconstrained nature. Figure 3 An example embodiment showing the basic steps of a computing system 600 is presented, where one or more images of an object of interest (e.g., an object for which a user wishes to create a 3D model) 602, more specifically a monocular video {I t}, can first be input into the computing system. The object of interest can be indicated by a segmentation mask {S t} 622.
[0079] The computing system can solve the inverse graphics problem to jointly recover the static shape S 604 of the object, the skinning weights W 606, the time-varying articulation, and the object-camera transformation D t 608 and / or the camera parameters, also referred to as the camera intrinsics K t 610, by an optimization method (e.g., video-based optimization). The method can be iteratively repeated, and in each iteration, multiple consecutive image frames can be sampled. For example, C = 8 pairs of consecutive frames can be randomly sampled. It will be understood that other numbers of consecutive frames can alternatively be used. In some embodiments, the sampled frames can be non-consecutive. Some frames in the video can be skipped, e.g., every other frame or every two frames can be taken.
[0080] The randomly sampled frames can be fed into a convolutional neural network. The convolutional neural network can predict the time-varying camera and motion parameters. The static shape S 604, also referred to as the average shape, can undergo a linear blend skinning process 614. The linear blend skinning process 614 can be carried out according to further details discussed below. Given certain parameters (e.g., the predicted articulation parameters D t 608, the skinning weights W 606, etc.), the linear blend skinning process 614 can output the articulated static shape 612.
[0081] Next, the computing system can use a differentiable renderer 616 to forward-render texture, optical flow, and contour images. According to further details discussed below, the forward rendering can be carried out using the differentiable renderer. The forward rendering can output a rendering 618, and the rendering 618 can be input 620 into a loss function 628. The ground truth pixels, the ground truth optical flow 624, the ground truth segmentation {S t} 622 are also input 626 into the loss function 628.
[0082] The loss function 628 can be evaluated to generate one or more gradients 630. The one or more gradients 630 can be used to update the camera K t 610, the shape S 604, and the articulation parameters D t 608. The one or more gradients 630 can be used to update K t 610, the shape S 604, and the articulation parameters D t 608 to minimize the difference between the rendered output Y = f(X) and the ground truth video measurements Y at test time. To handle the fundamental ambiguities in object shape S 604, deformation, and camera motion, the following disclosure can leverage a "low-level" but expressive parameterization of deformation, the rich constraints provided by optical flow and raw pixels, and appropriate regularization of object shape deformation and camera motion. *
[0083] Example forward synthesis model
[0084] Continuing with the example steps of the above computing system, in some embodiments, the computing system can use a differentiable renderer 616 to forward render texture, optical flow, and silhouette images. Given a frame index t and model parameters X, measurements for the corresponding frame pair {t, t+1} can be synthesized, including color image rendering object silhouette rendering and forward-backward optical flow rendering
[0085] In some embodiments, the object shape can be represented as a mesh with a fixed topology having N colored vertices and M faces. The mesh can be a triangular mesh. The time-varying articulation D t can be modeled by where △V t can be a per-vertex motion field applied to the stationary vertices and G 0,t =(R0|T0) t can be the object-camera transformation matrix (index 0 can be used to distinguish from the 1-indexed bone transformations in the deformation modeling utilized by the computing system). Finally, a perspective projection K t can be applied before rasterization, where it can be assumed that the principal point (p x , p y ) is constant and the focal length f t varies over time to handle scaling.
[0086] In some embodiments, a differentiable renderer can be used to render object silhouettes and color images. Given the per-vertex appearance C and a constant ambient light, the color image can be rendered. The color image can be rendered by obtaining the surface position V corresponding to each pixel in frame tt , calculate their positions V in the next frame t+1 , and then obtain the difference of their projections to complete the synthesis of the forward flow For example:
[0087]
[0088] where P (i) can represent the i-th row of the projection matrix P
[0089] Example Deformation Modeling
[0090] As described above, in some embodiments, a computing system can construct a deformation model of an object of interest. The deformation model can utilize multiple computational processes. Computational processes utilized for deformation modeling can include linear blend skinning (continuing with the example steps of the computing system above) and parametric skinning. The number of unknowns and constraints for solving the inverse problem can be analyzed. Given T frames of a video
[0091]
[0092] which can grow linearly with the number of vertices. Thus, an expressive but low-level representation of shape and motion can be generated
[0093] Continuing with the example steps of the computing system above, in some embodiments, the computing system can utilize linear blend skinning. Some embodiments of modeling deformation can use the modeled deformation as the per-vertex motion △V t . In other embodiments, the linear blend skinning model can constrain vertex motion by blending B rigid "bone" transformations {G1,..., G B}, which can reduce the number of parameters and make optimization easier. In addition to bone transformations, the LBS model can also define a skinning weight matrix that attaches the vertices of the rest shape vertices to the set of bones. Each vertex can be transformed by linearly combining weighted bone transformations in the object coordinate system and then transforming the vertex to the camera coordinate system, e.g.:
[0094] V i,t = G 0,t (∑ j W j,i G j,t )V i
[0095] where i can be the vertex index and j can be the bone index. In some embodiments, the skinning weights and time-varying bone transformations can be jointly learned
[0096] In some embodiments, the computing system may utilize parametric skinning. The skinning weights may be modeled as a Gaussian mixture, e.g.:
[0097]
[0098] where may be the position of the j-th bone, Q j may be the corresponding precision matrix that determines the orientation and radius of the Gaussian, and C may be a normalization factor that ensures that the sum of the probabilities of assigning vertices to different bones is one. In particular, W → {Q, J} may be optimized. Notably, in some embodiments, the mixture of Gaussian models reduces the number of parameters of the skinning weights from NB to 9B. In further embodiments, the mixture of Gaussian models may also ensure smoothness. The number of shape and motion parameters can now be expressed as:
[0099]
[0100] which may grow linearly with respect to the number of frames and bones.
[0101] Self-supervised learning from video examples
[0102] In some embodiments, rich supervision signals from dense optical flow and raw pixels may be utilized. Additionally, in some embodiments, shape and motion regularizers may be utilized to further constrain the problem.
[0103] In some embodiments, inverse graphics losses may be utilized. For example, the supervision of the synthesis analysis pipeline may include contour loss, texture loss, and optical flow loss. The contour loss compares the rendered texture with the measured contour, e.g., using L2 loss. The texture loss compares the rendered texture with the measured texture, e.g., using L1 loss and / or perceptual distance. The optical flow loss compares the rendered optical flow with the measured optical flow, e.g., using L2 loss. For example, given a pair of rendered outputs and measurements (S t , I t , u t ), the inverse graphics loss can be calculated as,
[0104]
[0105] where {β1,..., β4} may be weights selected empirically, σ t may be a normalized confidence map for flow measurements, and pdist(·, ·) may be the perceptual distance. In some embodiments, applying L1 loss to optical flow may be better than L2 loss. For example, since the L1 flow loss is more tolerant to outliers (e.g., non-rigid motion).
[0106] In some embodiments, shape and motion regularization can be utilized. For example, general shape and temporal regularization can be used to further constrain the problem. Laplacian smoothness operations can be used to enhance surface smoothness, such as:
[0107]
[0108] Motion regularization can include one or more of a minimum motion term, an ARAP (as rigid as possible) deformation term, and a temporal smoothness term. The minimum motion term can promote an articulated shape to remain close to the rest shape, and can be based on the difference between the mesh vertices of the object and the rest vertices of the object, such as:
[0109]
[0110] This can effectively solve the shape deformation ambiguity, that is, modifying the shape can be expressed as applying a bone transformation to the original shape. The ARAP term can be used to promote natural deformation, which can be based on the difference between the distances between vertices in consecutive frames, such as:
[0111]
[0112] In some embodiments, first-order temporal smoothing can be applied to the camera rotation (j = 0) and the bone rotations (j = 1,…,B), such as:
[0113]
[0114] where geodesic distance can be used to compare the rotations.
[0115] In some embodiments, soft symmetry constraints can be utilized. For example, reflection symmetry structures exhibited in common object categories can be utilized. For example, for both the rest shape and the skinning weights, soft symmetry constraints can be set in the object frame along the y-z plane (i.e., (n0, d0) = (1, 0, 0, 0, 0)).
[0116] In some cases, the rest shape and the reflected rest shape can be similar.
[0117]
[0118] where can be a Householder reflection matrix, and the Chamfer distance can be calculated as a two-way pixel-to-face distance. Similarly, for the rest bone J, it can be calculated by the following formula
[0119]
[0120] Finally, the normalization terms can be applied,
[0121]
[0122] where t * can be a canonical frame, and n * can be a symmetry plane in the frame. For example, the canonical camera pose can be offset to align with the symmetry plane. The symmetry plane can be initialized with an approximation and optimized. The total loss can be a weighted sum of all losses, with weights chosen empirically and kept constant across all experiments.
[0123] Example implementation details
[0124] In some embodiments, implementation details can be obtained using cameras and poses. In some embodiments, the time-varying parameters {D t , K t} can be directly optimized. In some embodiments, given an input image I t , the time-varying parameters {D t , K t} can be parameterized as predictions from a convolutional network,
[0125] ψ w (I t ) = (K, G0, G1, G2,..., G B ) t ,
[0126] In cases where one parameter can be predicted for the focal length, multiple parameters (e.g., four) can be predicted for each bone rotation parameterized by a quaternion, and multiple parameters (e.g., three) can be predicted for each translation. These amounts can be added to a total of 1 + 7(B + 1) amounts per frame. The predicted camera and pose predictions can be used to synthesize videos that are compared with the original measurements Y * that generate gradients to update the weights w. The network can learn a joint basis for the camera and pose that is easier to optimize than the original parameters.
[0127] In some embodiments, implementation details can be obtained using contour and flow measurements. It can be assumed that a reliable segmentation of the foreground object is provided. Instance segmentation and tracking methods can be used to manually annotate or estimate the segmentation. Reasonable optical flow estimation can be utilized, which can be provided by state-of-the-art flow estimators trained on a mixture of datasets. It is worth noting that learning articulated shape reconstruction can recover from some bad flow initializations and obtain better long-term correspondences.
[0128] In some embodiments, implementation details can be obtained using coarse-to-fine reconstruction, as Figure 4As shown. A coarse-to-fine strategy can be utilized to reconstruct a high-quality mesh. For S0, 702, a rigid object can be assumed, and the rest shape and camera {S, G 0,t , K t} can be optimized for L epochs. For S1 - S3, 704, 706, and 708, all parameters {S, D t , K t} can be jointly optimized, and remeshing can be performed after every L epochs, which can be repeated multiple times. (For example, it can be repeated three times, the first remeshing 704, the second remeshing 706, and the third remeshing 708). After each remeshing, both the number of vertices and the number of bones increase, as Figure 4 shown.
[0129] In some embodiments, implementation details can be obtained using initialization. The rest shape can be initialized as a subdivided icosahedron projected onto a sphere at S0 702. The rest bones can be initialized by running K-means on the coordinates of the vertices at S1 - S3 704, 706, and 708. The first frame of the video can be selected as the canonical frame, and the canonical symmetry plane n * can be given manually (by providing one of the y - z plane or the x - y plane), or selected from eight hypotheses whose azimuth and elevation angles are evenly spaced on a hemisphere, which is achieved by running S0 702 in parallel for each hypothesis and selecting the hypothesis with the lowest final loss.
[0130] Example 2D keypoint transfer on animal videos
[0131] For example, a user can use the computing system on an animal video dataset, which can provide multiple real animal videos with 2D keypoint and mask annotations (e.g., nine real animal videos). This data can be obtained from a video segmentation dataset or online stock footage. It can include many videos of many animals such as dogs (e.g., three videos of dogs), horses (e.g., two videos of horses), and camels, cows, bears, and impalas (e.g., one each of camels, cows, bears, and impalas).
[0132] To approximate the accuracy of 3D shape and articulated recovery, the percentage of correct keypoint transfer (PCK - T) can be used. Given a reference and target image pair with 2D keypoint annotations, the reference keypoints can be transferred to the target image, and if the transferred keypoints are within a certain threshold distance from the target keypoints If it is within, it is marked as "correct", where |S| can be the area of the ground truth contour. Given the articulated shape and camera pose estimation, the transferred points can be transferred by re-projecting from the reference frame to the target frame. If the back-projected key point is outside the reconstructed mesh, its nearest neighboring point that intersects the mesh can be re-projected. The accuracy can be averaged over all T(T - 1) pairs of frames.
[0133] The classification of alternative animal reconstruction methods that can be used as a baseline for comparison purposes is shown in Table 1 below. (1) refers to model-based shape optimization. (2) refers to model-based regression. (3) refers to class-specific reconstruction. (4) refers to template-free methods. S refers to single view. V refers to video or multi-view data. I refers to image. J2 refers to 2D joints. J3 refers to 3D joints. M refers to 2D mask. V3 refers to 3D mesh. C refers to camera matrix. O refers to optical flow. Quad refers to quadruped animals. Only the representative classes listed are referred to. * refers to unavailable implementations. SMALST is a model-based regressor trained for zebras. It takes an image as input and predicts the shape, pose, and texture for the SMAL model. UMR is a class-specific shape estimator trained for several classes, including birds, horses, and other classes with large sets of annotated images. Since models for other animal classes are not available, the performance of the horse model is reported. A-CSM learns class-specific canonical surface mappings and articulations from an image collection. At test time, it takes an image as input and predicts the articulation parameters for an assembled template mesh. It provides 3D templates for 27 animal classes and an articulation model for horses, which is used throughout the experiment. SMALify is a model-based optimization method that matches one of five classes of the SMAL model (including cats, dogs, horses, cows, and hippos) to a video or a single image. All video frames are provided with ground truth key points and mask annotations. Finally, a detection-based method, OJA, is included, which trains a hourglass network to detect animal key points (indicated by the detector) and post-processes the joint cost map with the proposed optimal assignment algorithm.
[0134]
[0135] Table 1: Related work on non-rigid shape reconstruction.
[0136] In Figure 5 example qualitative results of 3D shape reconstruction are shown. Figure 5Shows example 3D shape reconstruction results using the reference image 802 based on camel and human data from LASR and competitors. At 804, the shape reconstruction results from LASR are shown at different times of a sequence of images (e.g., a video). At 806, further shape reconstruction results from LASR at 0° rotation and 60° rotation are shown. At 808, further shape reconstruction results from UMR - Horse at 0° rotation and 60° rotation are shown. At 810, further shape reconstruction results from A - CSM (camel template) at 0° rotation and 60° rotation are shown. At 812, further shape reconstruction results from SMALify Horse at 0° rotation and 60° rotation are shown. At 814, further shape reconstruction results from LASR at 0° rotation and 60° rotation are shown, specifically showing a humanoid form. At 816, further shape reconstruction results from PIFuHD at 0° rotation and 60° rotation are shown. At 818, further shape reconstruction results from SMPLify - X at 0° rotation and 60° rotation are shown. At 820, further shape reconstruction results from VIBE at 0° rotation and 60° rotation are shown. LASR can jointly recover the camera, shape, and articulation from one or more images of an object (e.g., a monocular video) without using shape templates or category information. By relying on less prior knowledge, LASR can be applied to a wider range of non - rigid shapes and better fit the data. The results from LASR 806 recover two humps that are lost in the results from other methods 808, 810, and 812. Additionally, the cloth ribbon 822 of the dancer can be reconstructed from the results of LASR 814 and PIFuHD 816, but SMPLify - X 818 and VIBE 820 are confused with the right arm of the dancer.
[0137] At Figure 6Another example showing the qualitative results of 3D shape reconstruction, where we compare with UMR, A-CSM, and SMALify on bear and dog data 902 (e.g., bear and dog videos). The reconstruction of the first frame of the video is shown from two perspectives. Compared with UMR, which also does not use a shape template, LASR reconstructs a finer-grained geometry. Compared with A-CSM and SMALify, which use shape templates, LASR recovers instance-specific details, such as the furry tail of the dog and a more natural pose. Example shape reconstruction results from LASR are shown at 0° rotation 904 and 60° rotation 914. Example shape reconstruction results from UMR horse are shown at 0° rotation 906 and 60° rotation 914. Example shape reconstruction results from A-CSM (wolf template) are shown at 0° rotation 908 and 60° rotation 916. Example shape reconstruction results from SMALify dog are shown at 0° rotation 910 and 60° rotation 918.
[0138] The quantitative results of keypoint transfer are shown in Table 2 below. Assuming that all 3D reconstruction baselines are category-specific and may not provide accurate models for some categories (such as camels), the best model or template was selected for each animal video. Compared with the 3D reconstruction baselines, LASR is better for all categories, even on the categories for which the baselines were trained (e.g., on vault horse, LASR: 49.3 vs UMR: 32.4). Replacing the ground-truth segmentation mask with the object segmenter PointRend, the performance of LASR drops, but it is still better than all reconstruction baselines. Compared with detection-based methods, we have higher accuracy on the vault horse video and are close to the baseline on other videos. Compared with the initial optical flow, LASR also shows a large improvement (81.9% vs 47.9% for camels). (2) refers to model-based regression. (3) refers to category-specific reconstruction. (4) refers to free-form reconstruction. Refers to methods that do not reconstruct 3D shapes. * Refers to methods not specified for this category. If reconstructing 3D shapes, the best results are shown underlined and in bold.
[0139]
[0140] Table 2: 2D Keypoint Transfer Accuracy
[0141] Compared with the initial optical flow, LASR shows a large improvement, especially between long-range frames as Figure 7 shown. Figure 7An example key-point transfer between frame 2 and frame 70 of a sample camel video is shown. The distance between the transferred key-points and the target annotations is represented by the radius of the circles. The correct transfers are marked with solid circles 1014, and the incorrect transfers are marked with dashed lines 1016. A reference image with the LASR flow overlaid on top 1002 is shown. An example image with key-point transfers between frame 2 and frame 70 is shown using LASR 1004. An example image with key-point transfers between frame 2 and frame 70 is shown using the VCN flow 1006. An example image with key-point transfers between frame 2 and frame 70 is shown using A-CSM (camel template) 1008. An example image with key-point transfers between frame 2 and frame 70 is shown using SMALST - zebra 1010. An example image with key-point transfers between frame 2 and frame 70 is shown using UMR - horse 1012.
[0142] Example mesh reconstruction on articulated objects
[0143] For an example mesh reconstruction on articulated objects, to evaluate the accuracy of the mesh reconstruction, a video dataset of five articulated objects with ground-truth meshes and articulations was used, including a dancer video, a German shepherd video, a horse video, an eagle video, and a stone giant video. Rigid objects, Keenan's points, were also included to evaluate the performance of rigid object reconstruction and ablation in the S0 stage.
[0144] Most previous mesh reconstruction works assume given camera parameters. However, in some cases where LASR can be modeled, both the camera and the geometry are unknown, which leads to ambiguities in the evaluation, including scale ambiguity (present for all monocular reconstructions) and depth ambiguity (present for weak perspective cameras used in, e.g., UMR, A-CSM, VIBE, etc.). To factor out the unknown camera matrix, two meshes are aligned by a 3D similarity transformation solved by iterative closest point. Then, the bidirectional chamfer distance is adopted as the evaluation metric. 10k points are randomly sampled uniformly from the surfaces of the predicted mesh and the ground-truth mesh, and the average distance between the nearest neighbors of each point in the corresponding point clouds is calculated.
[0145] In addition to A-CSM, SMALify, and UMR for animal reconstruction, SMPLify-X, VIBE, and PiFUHD are also compared with LASR for human reconstruction. SMPLify-X is a model-based optimization method for human expression capture. The female SMPL model of the dancer sequence is used, and the keypoints estimated from OpenPose are provided as input. VIBE
[19] is a state-of-the-art model-based video regressor for human pose and shape inference. PIFuHD is a state-of-the-art free-form 3D shape estimator designed for clothed humans. It takes a single image as input and predicts an implicit shape representation, which is converted to a mesh by the marching cube algorithm. To compare with SMALify for dogs and horses, 18 keypoints are manually annotated per frame and initialized with the corresponding shape templates.
[0146] In Figure 5 and Figure 8 a visual comparison of humans and animals is shown. Figure 8 Shape and articulated reconstruction results at different timestamps on our synthetic dog and horse sequences are shown. Example shape and articulated reconstructions of the dog at t = 0, t = 5, and t = 10 are shown using GT 1102. Example shape and articulated reconstructions of the dog at t = 0, t = 5, and t = 10 are shown using LASR 1104. Example shape and articulated reconstructions of the dog at t = 0, t = 5, and t = 10 are shown using A-CSM (wolf template) 1106. Example shape and articulated reconstructions of the dog at t = 0, t = 5, and t = 10 are shown using SMALify-dog 1108. Example shape and articulated reconstructions of the horse at t = 0, t = 5, and t = 10 are shown using GT 1110. Example shape and articulated reconstructions of the horse at t = 0, t = 5, and t = 10 are shown using LASR 1112. Example shape and articulated reconstructions of the horse at t = 0, t = 5, and t = 10 are shown using UMR 1114. Example shape and articulated reconstructions of the horse at t = 0, t = 5, and t = 10 are shown using A-CSM-horse 1116. Example shape and articulated reconstructions of the horse at t = 0, t = 5, and t = 10 are shown using SMALify-horse 1118. The reference figure is shown at the upper left corner of each reconstruction 1122. The template meshes used are shown in the lower right 1120. Compared with the template-based method (UMR 1114), LASR 1112 successfully reconstructs the four legs of the horse. Compared with the template-based methods (A-CSM 1106 and SMALify 1108), LASR 1104 successfully reconstructs instance-specific details (the ears and tail of the dog) and recovers a more natural articulation.
[0147] The quantitative results are shown in Table 3 below. On the dog videos, LASR is better than all baselines (0.28 vs A-CSM: 0.38). LASR may be better because A-CSM and UMR are not specifically trained for dogs (although A-CSM uses a wolf template), and SMALify cannot reconstruct natural 3D shapes from limited keypoint and contour annotations. For the horse videos, LASR is slightly better than A-CSM which uses a horse shape template and outperforms the remaining baselines. For the dancer sequence, LASR is less accurate than the baseline methods (0.35 vs VIBE: 0.22), which is expected because all baselines either use well-designed human models or have been trained using 3D human mesh data, while LASR has no access to 3D human data. For the stone giant video, LASR is the only method that reconstructs a meaningful shape. Although the stone giant has a similar shape to humans, OpenPose cannot correctly detect the joints, resulting in the failure of SMALify-X, VIBE, and PiFUHD. The best results are shown in bold. "-" refers to methods that are not applicable to a specific sequence.
[0148]
[0149] Table 3: Mesh reconstruction error according to chamfer distance on the animated object dataset.
[0150] To examine the performance on arbitrary real-world objects, five videos were used, including dance spin, scooter, soap box, car turn, wild duck flight, and cat videos. The videos were segmented. The comparison with the template-free SfM-MVS pipeline COLMAP is as Figure 9 shown. The representative input frame is shown on the left 1202. Figure 9 The results of COLMAP using the scooter video are shown 1204. Figure 9 The results of LASR using the scooter video are shown 1206. Figure 9 The results of COLMAP using the soap box video are shown 1208. Figure 9 The results of LASR using the soap box video are shown 1210. Figure 9 The results of COLMAP using the car turn video are shown 1212. Figure 9 The results of LASR using the car turn video are shown 1214. These comparisons were made between sequences that are close to rigid. COLMAP only reconstructs the visible rigid parts, while LASR reconstructs both rigid objects and near-rigid people.
[0151] The effects of different design choices on the rigid cow and animated dog sequences were studied. At T = 15 frames, ambient light was used and the camera was rotated horizontally around the object (a full circle for the cow and Render the video. In addition to the color image, the contours and optical flow are also rendered as supervision. The results are as Figure 10 shown. Figure 10 Shows the rendering 1302 of the contours and optical flow as supervision for the cow using GT at t = 0 and t = 5. Figure 10 Shows the rendering 1304 of the contours and optical flow as supervision for the cow using the reference at t = 0 and t = 5. Figure 10 Shows the rendering 1304 of the contours and optical flow as supervision for the cow using GT at t = 0 and t = 5. Figure 10 Shows the rendering 1306 of the contours and optical flow without flow for the cow at t = 0 and t = 5 with optical flow as the supervision signal. Figure 10 Shows the rendering 1308 of the contours and optical flow without L_can for the cow at t = 0 and t = 5 with a normalized symmetric plane. Figure 10 Shows the rendering 1310 of the contours and optical flow without CNN for the cow at t = 0 and t = 5 with a convolutional neural network (CNN) having an implicit representation as the camera parameters. Figure 10 Shows the rendering 1312 of the contours and optical flow as supervision for the dog using GT at t = 8 and α = 0° and α = 60°. Figure 10 Shows the rendering 1314 of the contours and optical flow as supervision for the dog using the reference at t = 8 and α = 0° and α = 60°. Figure 10 Shows the rendering 1316 of the contours and optical flow without LBS for the dog at t = 8 and α = 0° and α = 60° with linear blend skinning. Figure 10 Shows the rendering 1318 of the contours and optical flow without C2F for the dog at t = 8 and α = 0° and α = 60° with coarse-to-fine remeshing. Figure 10Shows the rendering 1312 of the supervised contours and optical flow for a dog without Gaussian mixture model (GMM) at t = 8 and α = 0° and α = 60° with a parametric skinning model. 1302, 1304, 1306, 1308, and 1310 show ablation studies for camera and rigid shape optimization. Removing the optical flow loss introduces large errors in camera pose estimation and thus the overall geometry cannot be recovered. Removing the normalization loss results in worse camera pose estimation and thus the symmetry shape constraint is not properly enforced. Finally, if the camera pose is directly optimized without using a convolutional network, it converges much slower and does not produce the desired shape within the same number of iterations. 1312, 1314, 1316, 1318, 1320 show ablation studies for articulated shape optimization. Specifically, the articulated shapes reconstructed at the middle frame (t = 8) from two viewpoints are shown. Without the LBS model, although the reconstruction seems reasonable from the visible views, the reconstruction cannot recover the complete geometry due to redundant deformation parameters and lack of constraints. Without coarse-to-fine remeshing, the fine-grained details cannot be recovered. Replacing the GMM skinning weights (9xB parameters) with an NxB matrix results in additional limbs and tails during reconstruction.
[0152] The quantitative results are reported in Table 4 below. In terms of camera parameter optimization and rigid shape reconstruction (S0), (1) refers to the optical flow as a supervision signal, (2) refers to the normalization of the symmetry plane, and (3) refers to the CNN as an implicit representation of the camera parameters. For articulated shape reconstruction (S1 - S3), (1) refers to linear blend skinning, (2) refers to coarse-to-fine remeshing, and (3) refers to the parametric skinning model.
[0153]
[0154] Table 4: Ablation study with mesh reconstruction error.
[0155] Example method
[0156] Figure 11 Depicts a flowchart of an example method performed in accordance with an example embodiment of the present disclosure. Although, for purposes of illustration and discussion, Figure 11 depicts steps performed in a specific order, the methods of the present disclosure are not limited to the specifically illustrated order or arrangement. Without departing from the scope of the present disclosure, the various steps of method 1400 may be omitted, rearranged, combined, and / or adapted in various ways.
[0157] At 1402, a computing system may obtain an input image depicting an object and a current mesh model of the object. The input image may be one or more images. Additionally, the input image may be multiple images attached together in the form of a video. The video may be a monocular video. The object depicted in the input image may be an object of interest. Additionally, the object may be any entity of interest, such as an animal, a human, or an inanimate object.
[0158] At 1404, the computing system may process the input image using a machine-learned camera model to obtain camera parameters of the input image and object deformation data. The camera parameters may describe the camera pose of the input image. The object deformation data describes one or more deformations of the current mesh model relative to the shape of the object shown in the image. The input image may be further processed to obtain a rest shape, skin weights, and articulations.
[0159] At 1406, the computing system may differentiably render a rendered image of the object based at least in part on the camera parameters. The computing system may also differentiably render a rendered image of the object based at least in part on the object deformation data. The computing system may also differentiably render a rendered image of the object based at least in part on the current mesh model. Differentiably rendering a rendered image of the object may include articulating a rest shape under linear blend skinning. Given predicted articulation parameters and skin weights, the computing system may articulate the rest shape under linear blend skinning.
[0160] At 1408, the computing system may evaluate a loss function that compares one or more characteristics of the input image of the object with one or more characteristics of the rendered image of the object. One or more characteristics of the rendered image of the object that may be compared with one or more characteristics of the input image of the object may include pixels, optical flow, and segmentation.
[0161] At 1410, the computing system may modify one or more values of one or both of the machine-learned camera model and the current mesh model based on the gradient of the loss function. One or more values of one or both of the machine-learned camera model and the current mesh model may include camera, shape, or articulation parameters.
[0162] Additional disclosure
[0163] The techniques discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from these systems. The inherent flexibility of computer-based systems allows for a variety of possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0164] Although the subject matter has been described in detail with respect to its various specific example embodiments, each example is provided by way of explanation and not limitation of the disclosure. Those skilled in the art, having obtained an understanding of the foregoing, can readily generate alterations, variations, and equivalents of such embodiments. Thus, as will be apparent to those of ordinary skill in the art, the disclosure does not exclude such modifications, variations, and / or additions to the subject matter. For example, features shown or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, the disclosure is intended to cover these alterations, variations, and equivalents.
Claims
1. A computer-implemented method for determining the shape of a 3D object from an image, the method comprising: Obtain an input video depicting an object and a current mesh model of the object by a computing system including one or more computing devices, where the input video includes a plurality of input images; Process the input video by the computing system using a camera model to obtain camera parameters of the input video and object deformation data, where the camera parameters describe the camera pose of the input video, and where the object deformation data describes one or more deformations of the current mesh model relative to the shape of the object shown in the input video; Differentiably render a rendered video of the object by the computing system based on the camera parameters, the object deformation data, and the current mesh model, the rendered video including a plurality of rendered images; Evaluate a loss function by the computing system, the loss function comparing one or more characteristics of the input video of the object with one or more characteristics of the rendered video of the object, where evaluating the loss function includes: Determine a first flow of the input video; Determine a second flow of the rendered video; and Evaluate the loss function at least in part based on a comparison of the first flow and the second flow; and Modify one or more values of one or both of the camera model and the current mesh model by the computing system based on the gradient of the loss function.
2. The computer-implemented method according to claim 1, wherein, Evaluating the loss function includes: Determine a first contour of the input image; Determine a second contour of the rendered image; Evaluate the loss function at least in part based on a comparison of the first contour and the second contour.
3. The computer-implemented method according to claim 1, wherein, Evaluating the loss function includes: evaluating the loss function at least in part based on a comparison of first texture data associated with the input video and second texture data associated with the rendered video.
4. The computer-implemented method according to claim 1, wherein, The mesh model includes a triangular mesh model.
5. The computer-implemented method according to claim 1, wherein, The camera parameters describe the object-to-camera transformation of the input video.
6. The computer-implemented method according to claim 1, wherein, The camera model includes a convolutional neural network.
7. The computer-implemented method according to claim 1, wherein, The camera parameters further describe the focal length of the input video.
8. The computer-implemented method according to claim 1, wherein, The mesh model is initialized to project onto a subdivided icosahedron of a sphere.
9. The computer-implemented method according to claim 1, wherein, Modifying one or both of the camera model and the current mesh model by the computing system based on the gradient of the loss function includes modifying both the camera model and the current mesh model.
10. The computer-implemented method according to claim 1, wherein, The current mesh model includes a plurality of vertices, a plurality of joints, and a plurality of blend skinning weights of the plurality of joints relative to the plurality of vertices, and where the plurality of joints and the plurality of blend skinning weights are learnable.
11. The computer-implemented method according to claim 1, wherein, Differentiably rendering the rendered video of the object based on the camera parameters, the object deformation data, and the current mesh model includes: rendering the current mesh model deformed according to the object deformation data and deformed from the camera pose according to the camera parameters.
12. The computer-implemented method according to claim 1, further comprising: Perform the method according to claim 1 on a plurality of videos depicting a plurality of objects to build a shape model library from the plurality of videos.
13. The computer-implemented method according to any one of claims 1-12, wherein, Obtaining the input video depicting the object includes: manually selecting a canonical image frame from the input video.
14. The computer-implemented method according to any one of claims 1 to 12, wherein, Obtaining the input video depicting the object includes: Select a plurality of candidate frames of the input video; Evaluate the loss of each of the candidate frames; and Select the candidate frame with the lowest final loss as the canonical frame.
15. A computer-implemented method according to any one of claims 1 to 12, wherein, Wherein, The first flow is a first optical flow, and the second flow is a second optical flow, and wherein the loss function is evaluated at least in part based on a comparison of the first optical flow and the second optical flow.
16. A computer system, comprising: One or more processors; And One or more non-transitory computer-readable media that collectively store computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 15.
17. One or more non-transitory computer-readable media storing instructions that, when executed by a computing system including one or more computing devices, cause the one or more computing devices to perform the method according to any one of claims 1 to 15.