Saving Geometry Details in a Series of Tracking Meshes

The electronic device employs a neural network model to generate displacement maps, addressing the challenge of creating high-quality, high-frame-rate tracking meshes by transferring fine geometric details from 3D scans, thereby enhancing efficiency and reducing manual processing time.

JP7690117B2Active Publication Date: 2025-06-09SONY GROUP CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024512110
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-25
Filing Date
2022-08-16
Publication Date
2025-06-09
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

Conventional methods struggle to generate a series of tracking meshes with both high frame rates and high quality, while preserving geometric details, which is time-consuming and cumbersome.

Method used

An electronic device and method that utilize a neural network model to generate displacement maps for tracking meshes, allowing the transfer of fine geometric details from high-quality 3D scans to low-quality tracking meshes.

Benefits of technology

This approach enables the creation of high-quality, high-frame-rate tracking meshes without the need for manual extraction and processing of geometric details, thus improving efficiency and reducing time consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007690117000001
    Figure 0007690117000001
  • Figure 0007690117000002
    Figure 0007690117000002
  • Figure 0007690117000003
    Figure 0007690117000003
Patent Text Reader

Abstract

An electronic device and method for preserving geometry details in a tracking mesh includes: acquiring a three-dimensional (3D) scan of an object of interest and a set of tracking meshes; the set of tracking meshes includes a set of tracking meshes that correspond in time to the set of 3D scans; generating a set of displacement maps based on differences between the set of tracking meshes and the set of 3D scans; calculating a plurality of vectors, each of which includes a surface tension value associated with a mesh vertex of a corresponding tracking mesh; training a neural network for a displacement map generation task based on the set of displacement maps and the corresponding set of vectors; applying the trained neural network model to the plurality of vectors to generate a displacement map; and updating each tracking mesh based on the corresponding displacement map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] 〔Cross - Reference to Related Applications / Incorporation by Reference〕 This application claims the benefit of priority of U.S. Patent Application No. 17 / 411,432, filed on August 25, 2021, with the United States Patent and Trademark Office. Each of the above applications is hereby incorporated by reference in its entirety into this specification.

[0002] Various embodiments of the present disclosure relate to 3D computer graphics, animation, and artificial neural networks. Specifically, various embodiments of the present disclosure relate to an electronic device and a method for storing geometry details in a sequence of tracked meshes.

Background Art

[0003] Advances in the field of computer graphics have led to the development of various three - dimensional (3D) modeling and mesh tracking techniques for generating a sequence of tracked 3D meshes. In computer graphics, mesh tracking is widely used in the movie and video game industries to generate 3D computer graphics (CG) characters. Usually, multiple video cameras can be used to generate the tracked meshes. Mesh tracking is a difficult task because it is difficult to generate a tracked mesh while preserving geometric details. Many movie studios and video game studios use only video cameras to generate the tracked meshes and use digital cameras alone to capture fine geometric details. Although CG artists can manually clarify the geometric details and then add them to the 3D CG character model, this can sometimes be time - consuming.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Those skilled in the art will appreciate the limitations and disadvantages of conventional and customary techniques by comparing the described system with some aspects of the present disclosure shown with reference to the drawings in the remainder of this application. **Means for Solving the Problem**

[0005] Provide an electronic device and method for storing geometric details in a series of tracking meshes, substantially illustrated and / or described in relation to at least one figure and further shown more fully in the claims.

[0006] These and other features and advantages of the present disclosure can be understood by considering the following detailed description of the present disclosure with reference to the accompanying drawings, in which like elements are designated by like reference numerals throughout. **Brief Description of the Drawings**

[0007]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0008] In an electronic device and method for storing the geometry details of a series of disclosed tracking meshes, an implementation described below can be found. Exemplary aspects of the present disclosure can provide a hybrid dataset, i.e., an electronic device configured to acquire a 3D scan set (such as high-quality, low-frame-rate raw scans) and a series of tracking meshes (such as high-frame-rate tracking meshes of objects of interest). By way of example and not limitation, a series of tracking meshes can be acquired using a video camera and an existing mesh tracking tool. Similarly, a 3D scan set can be acquired using a digital still camera or a high-end video camera and an existing photogrammetry tool.

[0009] First, calculate the difference between the tracking mesh and the corresponding 3D scan, and the difference can be baked into the texture using the UV coordinates of each vertex of the tracking mesh. By this operation, a displacement map (a 3-channel image) that can store the XYZ displacement of the vertices in the RGB image channels can be obtained. The displacement map can be generated for each 3D scan in the 3D scan set. Then, calculate the surface tension value of each vertex of the tracking mesh. Usually, when the face deforms from a neutral expression to another expression, some areas of the face may appear stretched or compressed. After calculating the surface tension values for all vertices of the tracking mesh, such values can be stacked into the column vector of the tracking mesh. By this operation, a vector (also called the surface tension vector) for each tracking mesh can be obtained. In some tracking meshes, both the surface tension vector and the displacement map can be utilized, while in other tracking meshes of the same sequence, only the vector (the surface tension vector) can be utilized. Therefore, the purpose is to generate the displacement map for the tracking meshes where the displacement map may be missing. To generate the displacement map, a vector-to-image function in the form of a neural network model can be trained for the displacement map generation task. The neural network model can receive the surface tension vector as input and output a displacement map containing all available pairs of surface tension vectors and displacement maps.

[0010] After training the neural network, the surface tension vectors of all tracking meshes can be supplied to the trained neural network to generate the displacement map for each tracking mesh. Finally, the generated displacement map can be applied to each tracking mesh (such as a low-quality, high-frame-rate tracking mesh). By this operation, a series of high-quality, high-frame-rate tracking meshes can be obtained.

[0011] In conventional methods, it may be difficult to generate a series of tracking meshes that can have both a frame rate higher than a threshold (with respect to polycount or fine geometric details) and high quality. Conventionally, geometric details such as skin geometry or microgeometry can be manually extracted using software tools and then processed for application to the tracking mesh. This process can be time-consuming and cumbersome. In contrast, in the present disclosure, it is possible to train a neural network model for a displacement map generation task. After training the neural network model based on a displacement map set and a plurality of vectors (i.e., surface tension vectors), the trained neural network model can generate a displacement map for each tracking mesh. By making the displacement map of each tracking mesh available, the fine geometric details of the raw 3D scan can be transferred onto the tracking mesh. Thus, the displacement map generated by the trained neural network model can be applied to the obtained series of tracking meshes to update the series of tracking meshes with the fine geometric details included in the 3D scan set. This can eliminate the need to manually extract and process geometric details using known software tools.

[0012] FIG. 1 is a block diagram showing an exemplary network environment for storing geometric details in a tracking mesh according to an embodiment of the present disclosure. FIG. 1 shows a network environment 100. The network environment 100 can include an electronic device 102, a capture system 104, a first imaging device 106, a second imaging device 108, and a server 110. The network environment 100 can further include a communication network 112 and a neural network model 114. The network environment 100 can further include an object of interest such as a person 116. The first imaging device 106 can capture a first series of image frames 118 related to the person 116, and the second imaging device 108 can capture a second series of image frames 120 related to the person 116.

[0013] The electronic device 102 can include suitable logic, circuitry, and interfaces configured to obtain a three-dimensional (3D) scan set of an object of interest and a series of tracking meshes of the object of interest. The 3D scan set (also referred to as a four-dimensional (4D) scan) can be assumed to be a high-quality raw scan, and the series of tracking meshes (also referred to as a 4D tracking mesh) can be assumed to be a low-quality tracking mesh. The electronic device 102 can transfer fine geometric details from the high-quality raw scan (i.e., the 3D scan set) to the low-quality tracking mesh (i.e., the series of tracking meshes). The high-quality raw scan can be obtained from the first imaging device 106, and the low-quality tracking mesh can be obtained from an image captured through the second imaging device. The quality of the raw scan and the tracking mesh can depend on factors such as polycount. Examples of the electronic device 102 can include, but are not limited to, a computer device, a smartphone, a cellular phone, a mobile phone, a gaming device, a mainframe machine, a server, a computer workstation, and / or a consumer electronics (CE) device.

[0014] The capture system 104 can include suitable logic, circuitry, and interfaces configured to control one or more imaging devices, such as the first imaging device 106 and the second imaging device 108, to capture one or more image sequences of an object of interest from one or more viewpoints. All imaging devices, including the first imaging device 106 and the second imaging device 108, can be time synchronized. In other words, each of such imaging devices can be triggered to capture images almost simultaneously or exactly simultaneously. Accordingly, there can be a temporal correspondence between frames that can be captured at the same instant by two or more imaging devices.

[0015] In one embodiment, the imaging system 104 can include a dome-shaped lighting rig having sufficient space to include the person 116 or at least the face portion of the person 116. The first imaging device 106 and the second imaging device 108 can be attached at specific positions on the dome-shaped lighting rig.

[0016] The first imaging device 106 can include suitable logic, circuitry, and interfaces configured to capture a first series of image frames 118 related to an object of interest (such as the person 116) in a dynamic state. A dynamic state can mean that the object of interest or at least a portion thereof is in motion. In one embodiment, the first imaging device 106 can be a digital still camera capable of capturing a still image at a resolution higher than that of most high-frame-rate video cameras. In another embodiment, the first imaging device 106 can be a video camera capable of capturing a set of still image frames of the object of interest at a higher resolution and a lower frame rate compared to a high-frame-rate and low-resolution video camera.

[0017] The second imaging device 108 can include suitable logic, circuitry, and interfaces configured to capture a second series of image frames 120 related to an object of interest (such as the person 116) in a dynamic state. According to one embodiment, the first imaging device 106 can capture the first series of image frames 118 at a frame rate lower than the frame rate at which the second imaging device 108 can capture the second series of image frames 120. The image resolution of each frame of the first series of image frames 118 can be higher than the image resolution of each frame of the second series of image frames 120. Examples of the second imaging device 108 can include, but are not limited to, a video camera, an image sensor, a wide-angle camera, an action camera, a digital camera, a camcorder, a camera-equipped mobile phone, a time-of-flight camera (ToF camera), and / or other imaging devices.

[0018] Server 110 can include suitable logic, circuitry, and interfaces and / or code configured to generate a 3D scan set 122 and a series of tracking meshes 124 based on a first series of image frames 118 and a second series of image frames 120, respectively. In an exemplary implementation, server 110 can host a 3D modeling application, a graphics engine, and a 3D animation application. Such applications can include functions such as 3D mesh tracking and photogrammetry. Server 110 can receive the first series of image frames 118 to generate the 3D scan set 122 and receive the second series of image frames 120 to generate the series of tracking meshes 124. Server 110 can execute operations via, for example, a web application, a cloud application, an HTTP request, a repository operation, and a file transfer. Examples of implementations of server 110 include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, or a cloud computing server.

[0019] In at least one embodiment, server 110 can be implemented as a plurality of distributed cloud-based resources by using a plurality of techniques well known to those skilled in the art. Those skilled in the art will understand that the scope of the present disclosure can be implemented not limited to server 110 and electronic device 102 as two independent entities. In some embodiments, without departing from the scope of the present disclosure, the functions of server 110 can be wholly or at least partially incorporated into electronic device 102.

[0020] The communication network 112 can include a communication medium that enables the electronic device 102, the capture system 104, and the server 110 to communicate with each other. The communication network 114 can be either a wired connection or a wireless connection. Examples of the communication network 112 include, but are not limited to, the Internet, a cloud network, a cellular or wireless mobile network (such as Long-Term Evolution or 5G New Radio), a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). The various devices within the network environment 100 can be configured to connect to the communication network 112 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols include, but are not limited to, Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE802.11, Light Fidelity (Li-Fi), 802.16, IEEE802.11s, IEEE802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and at least one of the Bluetooth (BT) communication protocols.

[0021] The neural network model 114 can be a computational network or system of artificial neurons or nodes that can be arranged in multiple layers. The multiple layers of the neural network model 114 can include an input layer, one or more hidden layers, and an output layer. Each layer of the multiple layers can include one or more nodes (or artificial neurons represented, for example, by circles). The outputs of all the nodes in the input layer can be coupled to at least one node of the (single or multiple) hidden layer(s). Similarly, the inputs of each hidden layer can be coupled to the outputs of at least one node in other layers of the neural network model 114. The output of each hidden layer can be coupled to the input of at least one node in other layers of the neural network model 114. The (single or multiple) nodes of the final layer can receive inputs from at least one hidden layer and output results. The number of layers and the number of nodes in each layer can be determined from the hyperparameters of the neural network model 114. Such hyperparameters can be set before, during, or after training of the neural network based on a training dataset.

[0022] Each node of the neural network model 114 can correspond to a mathematical function (e.g., sigmoid function or rectified linear unit) having a set of parameters that can be adjusted during training of the network. The set of parameters can include, for example, weight parameters and regularization parameters. Each node can calculate an output using the mathematical function based on one or more inputs from nodes of other (single or multiple) layers of the neural network model 114 (e.g., the previous (single or multiple) layer). All or some of the nodes of the neural network model 114 can correspond to the same or different mathematical functions.

[0023] In the training of the neural network model 114, one or more parameters of each node of the neural network model 114 can be updated based on whether the output of the final layer (such as a displacement map) for a given input from a training data set (such as a vector related to the surface tension value) matches the correct result based on the loss function of the neural network model 114. The above process can be repeated for the same or different inputs until the minimum value of the loss function is achieved and the training error is minimized. In the art, multiple training methods are known, such as the gradient descent method, the stochastic gradient descent method, the batch gradient descent method, the gradient boosting method, and the meta-heuristic method.

[0024] The neural network model 114 can include electronic data that can be implemented as a software component of an application executable on, for example, the electronic device 102. The neural network model 114 can rely on a library, an external script, or other logic / instructions for execution by a processing device such as the electronic device 102. The neural network model 114 can include code and routines that enable a computer device such as the electronic device 102 to perform one or more operations. For example, such operations can be related to the generation of a plurality of displacement maps from a given surface tension vector associated with a mesh vertex. In this example, the neural network model 114 can be referred to as a vector-to-image function. In addition to or instead of this, the neural network model 114 can also be implemented using hardware including a processor, a microprocessor (for example, performing or controlling one or more operations), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network model 114 can also be implemented using a combination of both hardware and software.

[0025] Examples of the neural network model 114 include, but are not limited to, deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN), CNN-recurrent neural network (CNN-RNN), R-CNN, Fast R-CNN, Faster R-CNN, artificial neural network (ANN), (You Only Look Once) YOLO network, long short-term memory (LSTM) network-based RNN, CNN+ANN, LSTM+ANN, gated recurrent unit (GRU)-based RNN, fully connected neural network, Connectionist Temporal Classification (CTC)-based RNN, deep Bayesian neural network, adversarial generative network (GAN), and / or combinations of these networks. In some embodiments, the learning engine can include numerical algorithms using data flow graphs. In some embodiments, the neural network model 114 can be based on a hybrid architecture of multiple deep neural networks (DNNs).

[0026] Each 3D scan of the 3D scan set 122 can be a high-resolution raw 3D mesh that can include a plurality of polygons such as triangles. Each 3D scan can correspond to a particular instant and can capture the 3D shape / geometry of the object of interest at that particular instant. For example, if a person 116 changes their facial expression from a neutral expression to another expression within a 3-second duration, the 3D scan set 122 can capture the 3D shape / geometric details related to the facial expressions at different instants within the duration. The 3D scan set 122 can be obtained based on a first series of image frames 118. Such frames can be captured by a first imaging device 106 (such as a digital still camera).

[0027] A series of tracking meshes 124 can be a 4D mesh that captures dynamic changes to a 3D mesh over a plurality of discrete instants. The series of tracking meshes 124 can be obtained from video (such as a second series of image frames 118) by using suitable mesh tracking software and a parametric 3D model of the object of interest.

[0028] During operation, the electronic device 102 can synchronize a first imaging device 106 (such as a digital still camera) with a second imaging device 108 (such as a video camera). The electronic device 102 can control the first imaging device 106 to capture a first series of image frames 118 of an object of interest, such as a person 116, after synchronization. Similarly, the electronic device 102 can control the second imaging device 108 to capture a second series of image frames 120 of the object of interest. While both the first imaging device 106 and the second imaging device 108 are being controlled, it can be assumed that the object of interest is in a dynamic state. The dynamic state can correspond to a state in which at least a part of the object of interest or the entire object is moving. For example, if the object of interest is an animate object, the dynamic state can correspond to, but is not limited to, articular or non-articular movement of body parts including joints of bones, clothing, eyes, changes in the face (related to changes in facial expressions), hair, or other body parts. The first series of image frames 118 and the second series of image frames 120 can be temporally synchronized by synchronization.

[0029] According to an embodiment, the first imaging device 106 can capture the first series of image frames 118 at a frame rate lower than the frame rate at which the second imaging device 108 captures the second series of image frames 120. The image resolution of the image frames in the first series of image frames 118 can be higher than the image resolution of the image frames in the second series of image frames 120. Details of the capture of the first series of image frames 118 and the second series of image frames 120 are further shown, for example, in FIG. 3A.

[0030] The electronic device 102 can acquire a 3D scan set 122 of an object of interest such as the person 116. The 3D scan set 122 can be acquired from the server 110 or generated locally. In some embodiments, the electronic device 102 can be configured to execute a first set of operations including a photogrammetry operation for acquiring the 3D scan set 122. The execution of such operations can be based on a first series of image frames 118. Details of the acquisition of the 3D scan set 122 are further shown, for example, in FIG. 3A.

[0031] The electronic device 102 can further acquire a series of tracking meshes 124 of an object of interest (such as the head / face of the person 116). The acquired series of tracking meshes 124 can include a set of tracking meshes that can be temporally corresponding to the acquired 3D scan set 122. The temporal correspondence can exist due to the synchronization of two imaging devices (i.e., the first imaging device 106 and the second imaging device 108). Since the first series of image frames 118 can be captured at a relatively low frame rate (but high resolution) compared to the second series of image frames 120, such a correspondence can exist only for the set of tracking meshes.

[0032] The electronic device 102 can acquire a series of tracking meshes 124 from the server 110 or generate a series of tracking meshes 124 locally. According to some embodiments, the electronic device 102 can be configured to execute a second set of operations including a mesh tracking operation for acquiring the series of tracking meshes 124. This execution can be based on the captured second series of image frames 120 and a parametric 3D model of the object of interest. Details of the acquisition of the series of tracking meshes 124 are further shown, for example, in FIG. 3A.

[0033] The electronic device 102 can generate a displacement map set based on the difference between the tracking mesh set and the acquired 3D scan set 122. For example, the electronic device 102 can determine the difference between the 3D coordinates of the mesh vertices of each tracking mesh in the tracking mesh set and the corresponding 3D coordinates of the mesh vertices of each 3D scan in the 3D scan set 122. Details of the generation of the displacement map set are further shown, for example, in FIG. 4.

[0034] The electronic device 102 can calculate a plurality of vectors each including a surface tension value associated with the mesh vertices of the corresponding tracking mesh among a series of tracking meshes 124. All the surface tension values can be flattened into a 1-D array to obtain a 1-D vector of the surface tension values. According to an embodiment, the surface tension value can be determined based on a comparison between a reference value of the mesh vertices of one or more first tracking meshes among the series of tracking meshes 124 and a reference value of the corresponding mesh vertices of one or more second tracking meshes among the series of tracking meshes 124. One or more first tracking meshes can be associated with a neutral facial expression, and one or more second tracking meshes can be associated with a facial expression different from the neutral facial expression. As a non-limiting example, the tracking mesh of the face or head can represent a smiling expression. The tracking mesh can include regions (e.g., cheeks and lips) that appear stretched or compressed compared to the tracking mesh of a neutral face. The reference value of the mesh vertices of the neutral facial expression can be set to zero (representing a zero surface tension value). When the mesh vertices are stretched or compressed for any facial expression different from the neutral facial expression, the surface tension value of such mesh vertices can be set to a floating-point value between -1 and +1. Details of the calculation of the plurality of vectors are further shown, for example, in FIG. 5.

[0035] The neural network model 114 can be trained with respect to the displacement map generation task based on the calculated displacement map set and the corresponding vector set among the calculated plurality of vectors. According to an embodiment, the electronic device 102 can construct a training data set to include input value-output value pairs. Each input value-output value pair can include a set of calculated vectors and a displacement map of the generated displacement map set. The vectors can be provided as inputs for the training of the neural network model 114, and the displacement maps can be used as the ground truth for the output of the neural network model 114. Details of the training of the neural network model 114 are further shown, for example, in FIG. 6.

[0036] Once trained, the electronic device 102 can apply the trained neural network model 114 to the calculated plurality of vectors to generate a plurality of displacement maps. Each vector can be provided as an input simultaneously to the trained neural network model 114 to generate a corresponding displacement map as the output of the trained neural network model 114 for the input vector. The trained neural network model 114 can help predict missing displacement maps that are not initially available due to the difference between the number of 3D scans (high resolution) and the number of tracking meshes (although low-poly but numerous). Therefore, after the application of the trained neural network model 114, a one-to-one correspondence can exist between the plurality of displacement maps and the series of tracking meshes 124. Details of the generation of the plurality of displacement maps are further shown, for example, in FIG. 7.

[0037] The electronic device 102 can update each tracking mesh of the acquired series of tracking meshes 124 based on the corresponding displacement map among the plurality of generated displacement maps. By this update, a series of high-quality, high-frame-rate tracking meshes can be obtained. Since the displacement map of each tracking mesh is available, the fine geometric details of the (raw) 3D scan can be transferred onto the tracking mesh. Therefore, the displacement map generated by the trained neural network model 114 can be applied to the acquired series of tracking meshes 124 to update the series of tracking meshes 124 with the fine geometric details included in the 3D scan set. Thereby, the need to manually extract and process geometric details by using known software tools can be eliminated. Details of the update of the acquired series of tracking meshes 124 are further shown, for example, in FIG. 8.

[0038] FIG. 2 is a block diagram showing an exemplary electronic device for storing geometry in a tracking mesh according to an embodiment of the present disclosure. The description of FIG. 2 is made in relation to the elements of FIG. 1. FIG. 2 shows a block diagram 200 of the electronic device 102. The electronic device 102 can include a circuit 202, a memory 204, an input / output (I / O) device 206, and a network interface 208. The circuit 202 can be communicatively coupled to the memory 204, the I / O device 206, and the network interface 208. In some embodiments, the memory 204 can include the neural network model 114.

[0039] Circuit 202 can include suitable logic, circuitry, and interfaces configured to execute program instructions related to different operations performed by electronic device 102. Circuit 202 can include one or more special processing units that can be implemented as an independent processor. In certain embodiments, one or more special processing units can be implemented as an integrated processor or group of processors that collectively execute the functions of the one or more special processing units. Circuit 202 can be implemented based on a plurality of processor technologies well known in the art. Examples of implementations of circuit 202 can be, but are not limited to, an X86-based processor, a graphics processing unit (GPU), a reduced instruction set computing (RISC) processor, an application specific integrated circuit (ASIC) processor, a complex instruction set computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other control circuitry.

[0040] Memory 204 can include suitable logic, circuitry, and interfaces configured to store program instructions executed by circuit 202. Memory 204 can be configured to store neural network model 114. Memory 204 can also be configured to store 3D scan set 122, a series of tracking meshes 124, a plurality of displacement maps, a plurality of vectors, and an updated series of tracking meshes. Examples of implementations of memory 204 can include, but are not limited to, random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), hard disk drive (HDD), solid state drive (SSD), CPU cache, and / or secure digital (SD) card.

[0041] The I / O device 206 can include suitable logic, circuitry, and interfaces configured to receive input from a user and provide output based on the received input. The I / O device 206, which can include various input and output devices, can be configured to communicate with the circuitry 202. For example, the electronic device 102 can receive user input via the I / O device 206 to obtain a 3D scan set 122 and a series of tracking meshes 124 from the server 110. The I / O device 206, such as a display, can render the updated series of tracking meshes. Examples of the I / O device 206 can include, but are not limited to, a touch screen, a display device, a keyboard, a mouse, a joystick, a microphone, and a speaker.

[0042] The network interface 208 can include suitable logic, circuitry, and interfaces configured to facilitate communication between the circuit 202, the capture system 104, and the server 110 via the communication network 112. The network interface 208 can be implemented to support wired or wireless communication between the electronic device 102 and the communication network 112 using various known techniques. The network interface 208 can include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuit. The network interface 208 can be configured to communicate wirelessly with networks such as the Internet, an intranet, or wireless networks such as a cellular telephone network, a wireless local area network (LAN), and a metropolitan area network (MAN). The wireless communication can use one or more of a plurality of communication standards, protocols, and technologies such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), 5G NR, Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (WiFi) (such as IEEE802.11a, IEEE802.11b, IEEE802.11g, or IEEE802.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), protocols for email, instant messaging, and Short Message Service (SMS).

[0043] The functions or operations executed by the electronic device 102 as described with reference to FIG. 1 can be executed by the circuit 202. The operations executed by the circuit 202 will be described in detail, for example, with reference to FIGS. 4A, 4B, 5, 6, 7, and 8.

[0044] FIG. 3A is a diagram showing an exemplary three-dimensional (3D) scan and tracking mesh set of a face or head according to an embodiment of the present disclosure. The description of FIG. 3A is made in relation to the elements of FIGS. 1 and 2. FIG. 3A shows a diagram 300A showing an exemplary series of tracking meshes 302 and a 3D scan set 304. Here, the operation of acquiring both the series of tracking meshes 302 and the 3D scan set 304 will be described.

[0045] The circuit 202 can be configured to synchronize a first imaging device 106 (such as a digital still camera) with a second imaging device 108 (such as a video camera). The first imaging device 106 and the second imaging device 108 can be temporally synchronized such that both the first imaging device 106 and the second imaging device 108 are triggered to capture frames at a common instant. After synchronization, the circuit 202 can be configured to control the first imaging device 106 to capture a first series of image frames 118 of an object of interest (in a dynamic state). Further, the circuit 202 can control the second imaging device 108 to capture a second series of image frames 120 of the object of interest (in a dynamic state). As shown in the figure, for example, the object of interest can be the face portion of a person 116. The dynamic state of the object of interest can correspond to changes in the facial expression of the person 116 over a certain period.

[0046] According to an embodiment, the first imaging device 106 can capture the first series of image frames 118 at a frame rate lower than the frame rate at which the second imaging device 108 captures the second series of image frames 120. For example, the frame rate at which the first imaging device 106 captures the first series of image frames 118 can be 10 frames per second. The frame rate at which the second imaging device 108 captures the second series of image frames 120 can be 30 frames per second. In such a case, the number of images of the second series of image frames 120 that can be captured simultaneously when the first series of image frames 118 are captured is reduced. On the other hand, for the remaining images of the second series of image frames 120, there may be no corresponding images within the first series of image frames 118.

[0047] According to an embodiment, the image resolution of each of the first series of image frames 118 can be higher than the image resolution of each of the image frames of the second series of image frames 120. Each of the first series of image frames 118 captured by the first imaging device 106 (such as a digital still camera) can include detailed geometric details related to the skin geometry or microgeometry of the face portion, the facial hair of the face portion of the person 116, the pores of the face portion of the person 116, and other facial features of the person 116.

[0048] According to an embodiment, circuit 202 can be configured to execute a first set of operations including a photogrammetry operation to obtain 3D scan set 304. The execution can be based on the first series of captured image frames 118. The 3D scan set 304 can include, for example, 3D scans 304A, 3D scans 304B, 3D scans 304C, and 3D scans 304D. The first set of operations can include corresponding searching for related images from the first series of image frames 118 to obtain the 3D scan set 304. Based on the corresponding search, circuit 202 can detect overlapping areas within the related images from the first series of image frames 118. Such areas can be removed from the first series of image frames 118.

[0049] The photogrammetry operation can include feature extraction from the first series of image frames 118. For example, the positions of the encoded markers within the first series of image frames 118 can be utilized for feature extraction. The photogrammetry operation further includes triangulation. Triangulation can provide 3D coordinates of the mesh vertices of each 3D scan (such as 3D scan 304A) of the 3D scan set 304. The first set of operations can further include post-processing the 3D scans obtained from triangulation to obtain the 3D scan set 304. For example, post-processing of the 3D scans can include removing floating artifacts, background noise, holes, and irregularities from the 3D scans.

[0050] According to an embodiment, circuit 202 can be further configured to perform a second series of operations including a mesh tracking operation to obtain a series of tracking meshes 302. The execution can be based on the captured second series of image frames 120 and a 3D parametric model of an object of interest (e.g., a face or a head). The series of tracking meshes 302 can include tracking mesh 302A, tracking mesh 302B, tracking mesh 302C, tracking mesh 302D, tracking mesh 302E, tracking mesh 302F, tracking mesh 302G, tracking mesh 302H, tracking mesh 302I, and tracking mesh 302J.

[0051] The second series of operations can include an operation to recover geometry and motion information (across the frames) from the second series of image frames 120. The recovered geometry and motion information can be used to deform a parametric 3D model of an object of interest (e.g., a face or a head) such that a mesh tracking tool generates a series of tracking meshes 302 (also referred to as 4D tracking meshes). The second series of operations can also include post-processing of the series of tracking meshes 302 to remove floating artifacts, background noise, holes, and irregularities from the series of tracking meshes 302.

[0052] In an exemplary embodiment, tracking mesh 302A can temporally correspond to 3D scan 304A, tracking mesh 302D can temporally correspond to 3D scan 304B, tracking mesh 302G can temporally correspond to 3D scan 304C, and tracking mesh 302J can temporally correspond to 3D scan 304D. Thus, the set of tracking meshes that temporally correspond to the acquired 3D scan set 304 includes tracking mesh 302A, tracking mesh 302D, tracking mesh 302G, and tracking mesh 302J. According to an embodiment, the polycount of each tracking mesh of the acquired set of tracking meshes 302 can be smaller than the polycount of each 3D scan of the acquired 3D scan set 304. In other words, the acquired 3D scan set 304 can be a high-quality scan that can capture the complex or detailed geometric details of the face portion of person 116.

[0053] FIG. 3B is a diagram showing an exemplary full-body three-dimensional (3D) scan and a set of tracking meshes of a person wearing an exemplary garment, according to an embodiment of the present disclosure. The description of FIG. 3B is made in relation to the elements of FIGS. 1, 2, and 3A. FIG. 3B shows a diagram 300B showing a series of tracking meshes 306 and a 3D scan set 308 of a full body of a person wearing a garment.

[0054] Circuit 202 can capture a first series of image frames 118 related to person 116 and can obtain a 3D scan set 308 based on the execution of a first series of operations including photogrammetry operations. The 3D scan set 308 can include 3D scans 308A, 308B, 308C, and 308D. Similarly, circuit 202 can further capture a second series of image frames 120 of person 116 and can obtain a series of tracking meshes 306 based on the execution of a second series of operations including, but not limited to, mesh tracking operations. The series of tracking meshes 306 can include tracking meshes 306A, 306B, 306C, 306D, 306E, 306F, 306G, 306H, 306I, and 306J. As shown, tracking mesh 306A can temporally correspond to 3D scan 308A, tracking mesh 306D can temporally correspond to 3D scan 308B, tracking mesh 306G can temporally correspond to 3D scan 308C, and tracking mesh 306J can temporally correspond to 3D scan 308D.

[0055] FIG. 4A is a diagram showing a displacement map generation operation according to an embodiment of the present disclosure. The description of FIG. 4A is made in relation to the elements of FIGS. 1, 2, 3A, and 3B. FIG. 4A shows a diagram 400A showing a displacement map set generation operation.

[0056] At 406, differential calculation can be performed. Circuit 202 can calculate the difference between the tracking mesh set 402 and the acquired 3D scan set 304. The difference can be calculated in the form of pairs between the tracking mesh and the 3D scan. The tracking mesh set 402 can include a tracking mesh 302A that can temporally correspond to the 3D scan 304A. Similarly, the tracking mesh 306D can temporally correspond to the 3D scan 304B, the tracking mesh 302G can temporally correspond to the 3D scan 304C, and the tracking mesh 302J can temporally correspond to the 3D scan 304D. Details of the differential calculation are further shown, for example, in FIG. 4B.

[0057] At 408, a texture baking operation can be performed. Circuit 202 can bake the calculated texture difference between the tracking meshes of the tracking mesh set 402 and the corresponding 3D scans of the 3D scan 304. The difference can be baked by using the UV coordinates of each mesh vertex of the tracking mesh. All the tracking meshes have the same UV coordinates. Circuit 202 can generate a displacement map set 404 based on the baking operation. Specifically, by the operations at 406 and 408, for each pair of the tracking mesh and the corresponding 3D scan, a displacement map (which is a 3-channel image) that stores the XYZ displacement of the vertices in the RGB image channel can be obtained. As shown in the figure, for example, the displacement map set 404 includes a displacement map 404A corresponding to the tracking mesh 302A, a displacement map 404B corresponding to the tracking mesh 302D, a displacement map 404C corresponding to the tracking mesh 302G, and a displacement map 404D corresponding to the tracking mesh 302J. Since the UV coordinates are the same in all frames, the displacement of a specific area, such as the eye area of the face, can always be stored in the same area of the displacement map.

[0058] FIG. 400A shows discrete operations such as 406 and 408. In some embodiments, without compromising the essence of the disclosed embodiments, such discrete operations can be further divided into additional operations, combined into fewer operations, or deleted according to specific embodiments.

[0059] FIG. 4B is a diagram showing an operation of calculating a difference between a tracking mesh and a 3D scan according to an embodiment of the present disclosure. The description of FIG. 4B is made in relation to the elements of FIGS. 1, 2, 3A, 3B, and 4A. FIG. 4B shows FIG. 400B including a tracking mesh 302A and a 3D scan 304A. The tracking mesh 302A includes a plurality of mesh vertices. The 3D coordinates of the first mesh vertex 410 of the tracking mesh 302A can be x1, y1, and z1. The UV coordinates of the first mesh vertex 410 of the tracking mesh 302A can be u1 and v1. A surface normal vector 412 can be drawn from the first mesh vertex 410 to the corresponding 3D scan 304A. Since the polygon count of the 3D scan 304A is more than the polygon count of the tracking mesh 302A, it may be necessary to establish a one-to-one correspondence to determine the mesh vertices of the 3D scan 304A corresponding to the mesh vertices of the tracking mesh 302A. For this purpose, the surface normal vector 412 can be drawn from the first mesh vertex 410. The surface normal vector 412 can intersect the first mesh vertex 414 of the 3D scan 304A. The 3D coordinates of the first mesh vertex 414 of the 3D scan 304A can be x2, y2, and z2.

[0060] The circuit 202 can calculate the difference between the 3D coordinates (x2, y2, z2) of the first mesh vertex 414 and the 3D coordinates (x1, y1, z1) of the first mesh vertex 410. The graphical representation 416 shows the representation of the difference in the UV space. The X-axis of the graphical representation 416 can correspond to the "u" coordinate of the UV coordinates, and the Y-axis of the graphical representation 416 can correspond to the "v" coordinate of the UV coordinates. The circuit 202 can determine the resulting 3D coordinates 418 (represented by dx, dy, dz) at the position "u1, v1" within the graph 416. The resulting 3D coordinates 418 (represented by dx, dy, dz) can be derived by calculating the difference between the 3D coordinates (x1, y1, z1) of the first mesh vertex 410 and the 3D coordinates (x2, y2, z2) of the first mesh vertex 414. Similarly, the resulting 3D coordinates of other mesh vertices of the tracking mesh 302A can also be determined. A displacement map 404A can be generated based on the resulting 3D coordinates of the tracking mesh 302A.

[0061] FIG. 5 is a diagram showing the calculation of a plurality of vectors according to an embodiment of the present disclosure. The description of FIG. 5 is made in relation to the elements of FIGS. 1, 2, 3A, 3B, 4A, and 4B. FIG. 5 shows a diagram 500 including a first tracking mesh 502 and a second tracking mesh 504 among a series of tracking meshes 302.

[0062] Circuit 202 can be configured to calculate a plurality of vectors each including surface tension values associated with the mesh vertices of corresponding tracking meshes of a series of tracking meshes 302. Usually, when the mesh deforms from a neutral state (such as a neutral facial expression) to another state (such as another facial expression), some areas or regions of the mesh may appear to be stretched or compressed. For example, as shown in the figure, the stretched area is indicated by white and the compressed area is indicated by black. As an example, specific regions of the facial tracking mesh may be stretched or compressed by facial expressions such as an angry expression, a laughing expression, and a frowning expression. The surface tension values associated with the mesh vertices of such a mesh can be calculated with respect to a reference value of the mesh in a neutral state such as a neutral facial expression of the face or head mesh. When a series of tracking meshes are associated with different body parts (other than the head or face), the surface tension values can be determined from the surface deformation of the mesh observed in such a mesh with respect to the neutral state of the mesh. In models other than the face, a neutral state can be defined with respect to states such as body posture, cloth wrinkles, hair position, skin or muscle deformation, and joint position or orientation.

[0063] According to one embodiment, circuit 202 can determine one or more first tracking meshes among a series of tracking meshes 302 as being related to a neutral facial expression. For example, the first tracking mesh 502 among the series of tracking meshes 302 is shown as being related to a neutral facial expression. Circuit 202 can be configured to further determine the surface tension value associated with each mesh vertex of the determined one or more first tracking meshes to be zero (0.0). The surface tension values of each mesh vertex such as mesh vertex 502A, mesh vertex 502B, mesh vertex 502C, mesh vertex 502D, and mesh vertex 502E of the first tracking mesh 502 can be related to a neutral facial expression. After the surface tension values are determined for all mesh vertices, such values can be stacked into a column stack vector. By this operation, a reference vector 506 (also called a surface tension vector) is obtained. Circuit 202 can generate the reference vector 506 based on the surface tension value (0.0).

[0064] In one embodiment, circuit 202 can further determine one or more second tracking meshes among a series of tracking meshes 302 as being related to a facial expression different from the neutral facial expression. For example, the second tracking mesh 504 among the series of tracking meshes 302 is shown as being related to a facial expression different from the neutral facial expression. Circuit 202 can compare each mesh vertex of the determined one or more second tracking meshes with a reference value of the mesh vertices of the neutral facial expression. As an example, the mesh vertices such as mesh vertex 504A, mesh vertex 504B, mesh vertex 504C, mesh vertex 504D, and mesh vertex 504E of the second tracking mesh 504 can be compared with the corresponding reference values of the mesh vertices (such as mesh vertex 502A, mesh vertex 502B, mesh vertex 502C, mesh vertex 502D, and mesh vertex 502E) of the first tracking mesh 502.

[0065] Based on this comparison, circuit 202 can determine the surface tension values associated with each mesh vertex of the determined one or more second tracking meshes. For example, the surface tension values associated with the mesh vertices of the second tracking mesh 504 can be determined as floating point values that can each fall within the range of -1.0 to 1.0. For example, the surface tension value of a mesh vertex corresponding to a stretched area of the face portion can be a positive number between 0 and 1. The surface tension value of a mesh vertex corresponding to a compressed area of the face portion can be a negative number between 0 and -1.

[0066] In an exemplary scenario, circuit 202 can compare the reference value of mesh vertex 502A of the first tracking mesh 502 with the corresponding mesh vertex 504A of the second tracking mesh 504. The surface tension value of mesh vertex 504A can be determined as "0" based on the determination that mesh vertex 504A is part of the neutral area (i.e., the forehead area that is neither stretched nor compressed) of the second tracking mesh 504. The surface tension value of mesh vertex 504B can be determined as "-0.8" based on the determination that mesh vertex 504B belongs to the compressed area of the second tracking mesh 504. The surface tension value of mesh vertex 504C can be determined as "0.7" based on the determination that mesh vertex 504C belongs to the stretched area of the second tracking mesh 504. Whether a mesh vertex should be determined as one of stretched, compressed, or neutral, and the degree of stretching or compression can also be determined using a threshold (or reference value). The above process can be repeatedly iterated for all the remaining mesh vertices of the second tracking mesh 504 to determine the surface tension values of all the mesh vertices of the second tracking mesh 504. Circuit 202 can calculate a vector 508 associated with the second tracking mesh 504. The calculated vector 508 can include the surface tension values in a column vector. The operations described in FIG. 5 can be repeated for each tracking mesh of a series of tracking meshes 302 to calculate a plurality of vectors.

[0067] FIG. 6 is a diagram showing an exemplary operation of training a neural network model with respect to a displacement map generation task according to an embodiment of the present disclosure. The description of FIG. 6 is made in relation to the elements of FIGS. 1, 2, 3A, 3B, 4A, 4B, and 5. FIG. 6 shows a diagram 600 including a vector set 602, a displacement map 604, and a loss function 606.

[0068] From FIGS. 3A - 3B, 4, and 5, it can be observed that some of the tracking meshes have both vectors (i.e., surface tension vectors) and displacement maps, while some of the tracking meshes have only vectors. The purpose here is to generate a displacement map for the tracking meshes that do not have a displacement map. Such meshes are meshes for which there is no corresponding 3D scan. The reason such meshes occur is that the 3D scan set 304 can be obtained by using a first imaging device 106 (such as a digital still camera) that captures at a frame rate relatively lower than that of a video camera. In contrast, a series of tracking meshes 302 can be obtained using a second imaging device 108 (such as a video camera) that captures at a frame rate relatively higher than that of the digital still camera. As a result, the series of tracking meshes 302 have a higher frame rate but a lower polygon count compared to the 3D scan set 304. To generate a displacement map, a vector - image function in the form of a neural network model 114 can be trained as described herein.

[0069] Circuit 202 can be configured to train neural network model 114 with respect to the displacement map generation task based on the calculated displacement map set 404 and the corresponding vector set 602 of the calculated plurality of vectors. Prior to training, circuit 202 can construct a training data set that includes input - output value pairs. Each pair of input - output value pairs can include, as an input to neural network model 114, the vectors of the calculated vector set 602 corresponding to the tracking mesh set 402. The calculated vector set 602 can be provided as an input to neural network model 114 during training. Each pair of input - output value pairs can further include the displacement maps of the generated displacement map set 404 as the ground truth for the output of neural network model 114. In each pass (i.e., the forward pass and the backward pass), neural network model 114 can be trained to output a displacement map (such as displacement map 604) based on the vectors of vector set 602 as the input. Neural network model 114 can be trained with respect to the input - output value pairs over a number of epochs until the loss between the ground truth and the output of neural network model 114 falls below a threshold.

[0070] In an exemplary scenario, the vector set 602 can include a vector 602A corresponding to the tracking mesh 302A, a vector 602B corresponding to the tracking mesh 302D, a vector 602C corresponding to the tracking mesh 302G, and a vector 602D corresponding to the tracking mesh 302J. Each vector of the vector set 602 can be input one by one into the neural network model 114 to receive an output as the displacement map 604. For example, when the vector 602A can be input into the neural network model 114, the displacement map 404A can be regarded as ground truth. The circuit 202 can determine the loss function 606 based on the comparison between the displacement map 404A and the received displacement map 604. The circuit 202 can update the weights of the neural network model 114 when the loss function 606 may exceed a threshold. The circuit 202 can input the vector 602B into the neural network model 114 after the update. The circuit 202 can determine the loss function 606 based on the comparison between the displacement map 404B and the received displacement map 604. Similarly, the circuit 202 can train the neural network model 114 with respect to the input value-output value pairs over a number of epochs until the determined loss function 606 can fall below a threshold.

[0071] FIG. 7 is a diagram showing the generation of a plurality of displacement maps according to an embodiment of the present disclosure. The description of FIG. 7 is made in relation to the elements of FIGS. 1, 2, 3A, 3B, 4A, 4B, 5, and 6. FIG. 7 shows a diagram 700 including a plurality of vectors 702 and a plurality of displacement maps 704. The circuit 202 can be configured to apply the trained neural network model 114 to the calculated plurality of vectors 702 to generate a plurality of displacement maps 704. The plurality of vectors 702 can include vectors corresponding to each tracking mesh of a series of tracking meshes 302. The trained neural network model 114 can be applied to each vector such as vector 602A corresponding to tracking mesh 302A, vector 602B corresponding to tracking mesh 302D, and vector 602J corresponding to tracking mesh 302J among the series of tracking meshes 302. As a result of the application, the trained neural network model 114 can generate a plurality of displacement maps 704 such as displacement map 704A corresponding to tracking mesh 302A, displacement map 704D corresponding to tracking mesh 302D, and displacement map 704J corresponding to tracking mesh 302J.

[0072] FIG. 8 is a diagram showing an operation of updating a series of tracking meshes based on the displacement maps generated in FIG. 7 according to an embodiment of the present disclosure. The description of FIG. 8 is made in relation to the elements of FIGS. 1, 2, 3A, 3B, 4A, 4B, 5, 6, and 7. FIG. 8 shows a diagram 800 including an updated series of tracking meshes 802 that can be obtained based on the execution of one or more operations as described herein.

[0073] The circuit 202 can be configured to update each of the acquired series of tracking meshes 302 based on the corresponding displacement map of the plurality of generated displacement maps 704. For example, the plurality of displacement maps 704 can include a displacement map 704A corresponding to the tracking mesh 302A, a displacement map 704B corresponding to the tracking mesh 302B, and a displacement map 704J corresponding to the tracking mesh 302J. By this update, a series of high-quality, high-frame-rate tracking meshes can be obtained. Since the displacement map of each tracking mesh is available, the fine geometric details of the (raw) 3D scan can be transferred onto the tracking mesh. Therefore, the displacement map generated by the trained neural network model 114 can be applied to the acquired series of tracking meshes 302 to update the series of tracking meshes 302 with the fine geometric details included in the 3D scan set 304. This can eliminate the need to manually extract and process the geometric details using known software tools.

[0074] In one embodiment, before each tracking mesh is updated, circuit 202 can be configured to resample each of the plurality of displacement maps 704 that have been generated until the resolution of the resampled displacement map corresponds to the polygon count of the 3D scan of the 3D scan set 304 for which the resampled displacement map was obtained. For example, the number of pixels in the resampled displacement map can match the number of vertices or points of the 3D scan. In such a case, the update of the series of tracking meshes 302 can include a first application of the resampling operation to each tracking mesh of the series of tracking meshes 302 until the polygon count of each tracking mesh of the series of tracking meshes 302 matches the polygon count of the corresponding 3D scan of the 3D scan set 304. Also, the update of the series of tracking meshes 302 can include a second application of applying each resampled displacement map of the resampled displacement map set 404 to the corresponding resampled mesh of the series of resampled tracking meshes 302 to obtain a series of updated tracking meshes 802. For example, applying the resampled displacement map 704A to the resampled mesh 302A to obtain the updated mesh 802A, applying the resampled displacement map 704B to the resampled mesh 302B to obtain the updated mesh 802B, and applying the resampled displacement map 704J to the resampled mesh 302J to obtain the updated mesh 802J can be done.

[0075] According to one embodiment, the series of updated tracking meshes 802 can correspond to a set of blend - shapes for animation. For example, the set of blend - shapes can be used to animate the face portion of the person 116 (such as bringing changes to the facial expression).

[0076] FIG. 9 is a flowchart showing an exemplary method for storing geometry in a tracking mesh according to an embodiment of the present disclosure. The description of FIG. 9 is made in relation to the elements of FIGS. 1, 2, 3A, 3B, 4A, 4B, 5, 6, 7, and 8. FIG. 9 shows a flowchart 900. The method shown in flowchart 900 can be executed by any computer system such as the electronic device 102 or the circuit 202. The method can start from 902 and proceed to 904.

[0077] At 904, a 3D scan set 304 of the object of interest can be obtained. According to an embodiment, the circuit 202 can be configured to obtain a 3D scan set 304 of an object of interest such as the face portion of the person 116. Details of the acquisition of the 3D scan set 304 are further shown, for example, in FIG. 3A.

[0078] At 906, a series of tracking meshes 302 of the object of interest can be obtained. According to an embodiment, the circuit 202 can be configured to obtain a series of tracking meshes 302 of an object of interest such as the face portion of the person 116. The obtained series of tracking meshes 302 can include a tracking mesh set 402 that can be temporally corresponding to the obtained 3D scan set 304. Details of the acquisition of the series of tracking meshes 302 are further shown, for example, in FIG. 3A.

[0079] At 908, a displacement map set 404 can be generated. According to an embodiment, the circuit 202 can be configured to generate a displacement map set 404 based on the difference between the tracking mesh set 402 and the obtained 3D scan set 304. Details of the generation of the displacement map set 404 are further shown, for example, in FIGS. 4A and 4B.

[0080] At 910, a plurality of vectors 702 can be calculated. According to one embodiment, circuit 202 can be configured to calculate a plurality of vectors 702 each including surface tension values associated with the mesh vertices of the corresponding tracking meshes of a series of tracking meshes 302. Details of the calculation of the plurality of vectors 702 are further shown, for example, in FIG. 5.

[0081] At 912, the neural network model 114 can be trained. According to one embodiment, circuit 202 can be configured to train the neural network model 114 regarding the displacement map generation task based on the calculated displacement map set 404 and the corresponding vector set 602 of the calculated plurality of vectors 702. Details of the training of the neural network model 114 are further shown, for example, in FIG. 6.

[0082] At 914, the trained neural network model 114 can be applied to the calculated plurality of vectors 702 to generate a plurality of displacement maps 704. According to one embodiment, circuit 202 can be configured to apply the trained neural network model 114 to the calculated plurality of vectors 702 to generate a plurality of displacement maps 704. Details of the generation of the plurality of displacement maps 704 are further shown, for example, in FIG. 7.

[0083] At 916, each tracking mesh of the obtained series of tracking meshes 302 can be updated based on the corresponding displacement map of the generated plurality of displacement maps 704. According to one embodiment, circuit 202 can be configured to update each tracking mesh of the obtained series of tracking meshes 302 based on the corresponding displacement map of the generated plurality of displacement maps 704. Details of the update of each tracking mesh of the series of tracking meshes 302 are further shown, for example, in FIG. 8. The control can proceed to termination.

[0084] Although flowchart 900 shows discrete operations such as 902, 904, 906, 908, 910, 912, 914, and 916, the present disclosure is not so limited. Thus, in some embodiments, such discrete operations can be further divided into additional operations, combined into fewer operations, or deleted, depending on the particular embodiment, without detracting from the essence of the disclosed embodiments.

[0085] Various embodiments of the present disclosure can provide a non-transitory computer-readable medium and / or a storage medium storing instructions executable by a machine and / or a computer to operate an electronic device (such as the electronic device 102). These instructions can cause the machine and / or the computer to perform operations including obtaining a three-dimensional (3D) scan set (such as the 3D scan set 122) of an object of interest (such as the person 116). The operations can further include obtaining a series of tracking meshes (such as the series of tracking meshes 124) of the object of interest. The obtained series of tracking meshes 124 can include a set of tracking meshes (such as the set of tracking meshes 402) that temporally corresponds to the obtained 3D scan set 122. The operations can further include generating a set of displacement maps (such as the set of displacement maps 404) based on the difference between the set of tracking meshes 402 and the obtained 3D scan set 122. The operations can further include calculating a plurality of vectors (such as the plurality of vectors 702) each including a surface tension value associated with a mesh vertex of a corresponding tracking mesh of the series of tracking meshes 124. The operations can further include training a neural network model (such as the neural network model 114) with respect to a displacement map generation task based on the calculated set of displacement maps 404 and a corresponding set of vectors of the calculated plurality of vectors 702. The operations can further include applying the trained neural network model 114 to the calculated plurality of vectors 702 to generate a plurality of displacement maps (such as the plurality of displacement maps 704). The operations can further include updating each tracking mesh of the obtained series of tracking meshes 124 based on a corresponding displacement map of the generated plurality of displacement maps 704.

[0086] Exemplary aspects of the present disclosure can provide an electronic device (such as the electronic device 102 of FIG. 1) including a circuit (such as the circuit 202). The circuit 202 can be configured to acquire a three-dimensional (3D) scan set (such as the 3D scan set 122) of an object of interest (such as the person 116). The circuit 202 can be further configured to acquire a series of tracking meshes (such as the series of tracking meshes 124) of the object of interest. The acquired series of tracking meshes 124 can include a set of tracking meshes (such as the tracking mesh set 402) that temporally corresponds to the acquired 3D scan set 122. The circuit 202 can be further configured to generate a set of displacement maps (such as the displacement map set 404) based on the difference between the set of tracking meshes 402 and the acquired 3D scan set 122. The circuit 202 can be further configured to calculate a plurality of vectors (such as the plurality of vectors 702) each including a surface tension value associated with a mesh vertex of a corresponding tracking mesh of the series of tracking meshes 124. The circuit 202 can be further configured to train a neural network model (such as the neural network model 114) with respect to a displacement map generation task based on the calculated displacement map set 404 and a corresponding vector set of the calculated plurality of vectors 702. The circuit 202 can be further configured to apply the trained neural network model 114 to the calculated plurality of vectors 702 to generate a plurality of displacement maps (such as the plurality of displacement maps 704). The circuit 202 can be further configured to update each tracking mesh of the acquired series of tracking meshes 124 based on a corresponding displacement map of the generated plurality of displacement maps 704.

[0087] According to one embodiment, the circuit 202 can be further configured to synchronize the second imaging device 108 with the first imaging device 106. After synchronization, the circuit 202 can control the first imaging device 106 to capture a first series of image frames 118 of the object of interest in the dynamic state. The circuit 202 can further control the second imaging device 108 to capture a second series of image frames 120 of the object of interest in the dynamic state.

[0088] According to one embodiment, the first imaging device 106 can be a digital still camera, and the second imaging device 108 can be a video camera. The first imaging device 106 can capture the first series of image frames 118 at a frame rate lower than the frame rate at which the second imaging device 108 captures the second series of image frames 120.

[0089] According to one embodiment, the image resolution of each of the first series of image frames 118 can be higher than the image resolution of each of the image frames of the second series of image frames 120.

[0090] According to one embodiment, the circuit 202 can be further configured to execute a first set of operations including a photogrammetry operation for obtaining a 3D scan set 122. The execution can be based on the first series of image frames 118 that have been captured.

[0091] According to one embodiment, the circuit 202 can be further configured to execute a second set of operations including a mesh tracking operation for obtaining a series of tracking meshes 124. The execution can be based on the second series of image frames 120 that have been captured and a parametric 3D model of the object of interest.

[0092] According to one embodiment, the polycount in each of the obtained series of tracking meshes 124 can be smaller than the polycount in each of the 3D scans of the obtained 3D scan set 122.

[0093] According to one embodiment, circuit 202 can be further configured to construct a training data set to include input value-output value pairs each including a vector of the calculated vector set 602 corresponding to the tracking mesh set 402 as an input to the neural network model 114. Each of the input value-output value pairs can further include a displacement map of the generated displacement map set 404 as a ground truth for the output of the neural network model 114. The neural network model 114 can be trained over a number of epochs with respect to the input value-output value pairs until the loss between the ground truth and the output of the neural network model 114 is less than a threshold value.

[0094] According to one embodiment, the object of interest is a facial portion of a person.

[0095] According to one embodiment, circuit 202 can be further configured to determine one or more first tracking meshes of the series of tracking meshes 124 as being related to a neutral facial expression. Circuit 202 can determine the surface tension value associated with each mesh vertex of the determined one or more first tracking meshes to be zero. Circuit 202 can determine one or more second tracking meshes of the series of tracking meshes 124 as being related to a facial expression different from the neutral facial expression. Circuit 202 can further compare each mesh vertex of the determined one or more second tracking meshes to a reference value of the mesh vertices of the neutral facial expression. Circuit 202 can determine the surface tension value associated with each mesh vertex of the determined one or more second tracking meshes based on the comparison.

[0096] According to an embodiment, the circuit 202 can be further configured to resample each of the generated displacement maps 704 until the resolution of each resampled displacement map of the generated plurality of displacement maps 704 can correspond to the polygon count of the 3D scan of the 3D scan set 122 for which the resolution was obtained.

[0097] According to an embodiment, the update can include a first application that applies a resampling operation to each tracking mesh of the series of tracking meshes 124 until the polygon count of each resampled mesh of the resampled series of tracking meshes 124 matches the polygon count of the corresponding 3D scan of the 3D scan set 122. The update can further include a second application that applies each resampled displacement map of the resampled displacement map set 404 to the corresponding resampled mesh of the resampled series of tracking meshes 124 to obtain an updated series of tracking meshes 802.

[0098] According to an embodiment, the updated series of tracking meshes 802 can correspond to a blend shape set for animation.

[0099] The present disclosure can be implemented in hardware or in a combination of hardware and software. The present disclosure can be implemented in a centralized manner within at least one computer system or in a distributed manner in which different elements can be distributed across a plurality of interconnected computer systems. A computer system or other device adapted to execute the methods described herein can be suitable. The combination of hardware and software can be a general-purpose computer system including a computer program that can control the computer system to execute the methods described herein when loaded and executed. The present disclosure can be implemented in hardware including a part of an integrated circuit that also executes other functions.

[0100] This disclosure includes all features enabling the implementation of the methods described herein, and can also be incorporated into a computer program product capable of executing these methods when loaded into a computer system. A computer program in this context means any representation, in any language, code, or notation, of a set of instructions intended to cause a system with information processing capabilities to execute a particular function either directly or after performing either or both of a) conversion to another language, code, or notation, and b) reproduction in a different form of content.

[0101] Although the present disclosure has been described with reference to some embodiments, those skilled in the art will understand that various changes can be made without departing from the scope of the present disclosure, and equivalents can be substituted. Also, many modifications can be made to adapt a particular situation or content to the teachings of the present disclosure without departing from the scope of the present disclosure. Accordingly, the present disclosure is not intended to be limited to the specific embodiments disclosed, but is intended to include all embodiments falling within the scope of the appended claims.

Description of Reference Numerals

[0102] 902 Start 904 Obtain a three-dimensional (3D) scan set of the object of interest 906 Obtain a series of tracking meshes of the object of interest, including a first set of tracking meshes temporally corresponding to the obtained 3D scan set 908 Generate a first set of displacement maps based on the difference between the first set of tracking meshes and the obtained 3D scan set 910 Calculate a plurality of vectors each including a surface tension value related to the mesh vertices of the corresponding tracking meshes of the series of tracking meshes 912 Train a neural network model with respect to the displacement map generation task based on the calculated first set of displacement maps and the corresponding vector sets of the calculated plurality of vectors 914 Apply the trained neural network model to the calculated plurality of vectors to generate a plurality of displacement maps Update each tracking mesh of the obtained series of tracking meshes based on the corresponding displacement map among the plurality of generated displacement maps

Claims

Claim 1 An electronic device, obtains a three-dimensional (3D) scan set of an object of interest, obtains a series of tracking meshes of the object of interest, including a tracking mesh set temporally corresponding to the obtained 3D scan set, generates a displacement map set based on a difference between the tracking mesh set and the obtained 3D scan set, calculates a plurality of vectors each including a surface tension value related to a mesh vertex of a corresponding tracking mesh of the series of tracking meshes, trains a neural network model regarding a displacement map generation task based on the calculated displacement map set and a corresponding vector set among the calculated plurality of vectors, applies the trained neural network model to the calculated plurality of vectors to generate a plurality of displacement maps, updates each tracking mesh of the obtained series of tracking meshes based on a corresponding displacement map among the generated plurality of displacement maps, comprising a circuit configured as such, characterized in that it is an electronic device. Claim 2 The circuit further comprises: synchronizing a second imaging device with a first imaging device; after the synchronization, controlling the first imaging device to capture a first series of image frames of the object of interest in a dynamic state; controlling the second imaging device to capture a second series of image frames of the object of interest in the dynamic state, The electronic device according to claim 1, further configured as such. Claim 3 The first imaging device is a digital still camera, and the second imaging device is a video camera, wherein the first imaging device captures the first series of image frames at a frame rate lower than a frame rate at which the second imaging device captures the second series of image frames, The electronic device according to claim 2. Claim 4 The image resolution of each of the first series of image frames is higher than the image resolution of each image frame of the second series of image frames, The electronic device according to claim 2. Claim 5 The circuit is further configured to execute a first set of operations including a photogrammetry operation for obtaining the 3D scan set, wherein the execution is based on the captured first series of image frames, The electronic device according to claim 2. Claim 6 The circuit is further configured to execute a second set of operations including a mesh tracking operation for obtaining the series of tracking meshes, wherein the execution is based on the captured second series of image frames and the parametric 3D model of the object of interest, The electronic device according to claim 2.

7. The polycount in each of the obtained series of tracking meshes is smaller than the polycount in each of the obtained 3D scans of the 3D scan set, The electronic device according to claim 1.

8. The circuit is further configured to construct a training data set to include input value-output value pairs, each of the input value-output value pairs being, a vector of the calculated vector set corresponding to the tracking mesh set as an input to a neural network model, a displacement map of the generated displacement map set as a ground truth for an output of the neural network model, and is composed of, The neural network model is trained over a number of epochs with respect to the input value-output value pairs until the loss between the ground truth and the output of the neural network model falls below a threshold value. The electronic device according to claim 1.

9. The object of interest is a facial part of a person, The electronic device according to claim 1.

10. The circuit, determines one or more first tracking meshes among the series of tracking meshes as being related to a neutral facial expression, determines the surface tension value associated with each mesh vertex of the determined one or more first tracking meshes to be zero, determines one or more second tracking meshes among the series of tracking meshes as being related to a facial expression different from the neutral facial expression, compares each mesh vertex of the determined one or more second tracking meshes with a reference value of the mesh vertices of the neutral facial expression, and determines the surface tension value associated with each of the mesh vertices of the determined one or more second tracking meshes based on the comparison. The electronic device according to claim 9, further configured as such.

11. The circuit is further configured to resample each of the plurality of generated displacement maps until the resolution of each resampled displacement map of the plurality of generated displacement maps corresponds to the polygon count of the 3D scan of the acquired 3D scan set. The electronic device according to claim 1.

12. The update is a first application of applying a resampling operation to each of the series of tracking meshes until the polygon count of each resampled mesh of the series of resampled tracking meshes matches the polygon count of the corresponding 3D scan of the 3D scan set; a second application of applying each resampled displacement map of the resampled displacement map set to the corresponding resampled mesh of the series of resampled tracking meshes to obtain an updated series of tracking meshes; The electronic device according to claim 11, comprising:

13. The updated series of tracking meshes corresponds to a blend shape set for animation. The electronic device according to claim 1.

14. acquiring a three-dimensional (3D) scan set of an object of interest; acquiring a series of tracking meshes of the object of interest, including a set of tracking meshes temporally corresponding to the acquired 3D scan set; generating a set of displacement maps based on the difference between the set of tracking meshes and the acquired 3D scan set; calculating a plurality of vectors each including a surface tension value related to a mesh vertex of a corresponding tracking mesh of the series of tracking meshes; training a neural network model with respect to a displacement map generation task based on the calculated set of displacement maps and a corresponding set of vectors among the calculated plurality of vectors; applying the trained neural network model to the calculated plurality of vectors to generate a plurality of displacement maps; updating each of the acquired series of tracking meshes based on a corresponding displacement map among the generated plurality of displacement maps; A method characterized by including:

15. synchronizing a second imaging device with a first imaging device; controlling the first imaging device to capture a first series of image frames of the object of interest in a dynamic state after the synchronization; controlling the second imaging device to capture a second series of image frames of the object of interest in the dynamic state; The method according to claim 14, further comprising. **Claim 16** further comprising performing a first set of operations including a photogrammetry operation to obtain the 3D scan set, the performing being based on the first series of image frames captured; The method according to claim 15. **Claim 17** further comprising performing a second series of operations including a mesh tracking operation to obtain the series of tracking meshes, the performing being based on the second series of image frames captured and the parametric 3D model of the object of interest; The method according to claim 15. **Claim 18** further comprising constructing a training data set to include input value-output value pairs, each of the input value-output value pairs comprising: a vector of the calculated vector set corresponding to the tracking mesh set as an input to a neural network model; a displacement map of the generated displacement map set as a ground truth for an output of the neural network model; and being composed of; the neural network model is trained over a number of epochs with respect to the input value-output value pairs until a loss between the ground truth and the output of the neural network model falls below a threshold; The method according to claim 14. **Claim 19** determining one or more first tracking meshes of the series of tracking meshes as being related to a neutral facial expression; determining a surface tension value associated with each mesh vertex of the determined one or more first tracking meshes to be zero; determining one or more second tracking meshes of the series of tracking meshes as being related to a facial expression different from the neutral facial expression; comparing each mesh vertex of the determined one or more second tracking meshes with a reference value of the mesh vertices of the neutral facial expression; determining a surface tension value associated with each mesh vertex of the determined one or more second tracking meshes based on the comparison; The method according to claim 14, further comprising [

20. ] A non-transitory computer-readable medium storing computer-executable instructions, wherein when the computer-executable instructions are executed by an electronic device, the electronic device is caused to obtain a three-dimensional (3D) scan set of an object of interest; obtain a series of tracking meshes of the object of interest, including a tracking mesh set temporally corresponding to the obtained 3D scan set; generate a displacement map set based on a difference between the tracking mesh set and the obtained 3D scan set; calculate a plurality of vectors each including a surface tension value associated with a mesh vertex of a corresponding tracking mesh of the series of tracking meshes; train a neural network model with respect to a displacement map generation task based on the calculated displacement map set and a corresponding vector set of the calculated plurality of vectors; apply the trained neural network model to the calculated plurality of vectors to generate a plurality of displacement maps; update each tracking mesh of the obtained series of tracking meshes based on a corresponding displacement map of the generated plurality of displacement maps; A non-transitory computer-readable medium, characterized in that it causes an operation including

Citation Information

Patent Citations

  • Systems and methods for generating a skull surface for computer animation

    US11055892B1

  • Facial Performance Synthesis Using Deformation Driven Polynomial Displacement Maps

    US20090195545A1

  • Photo-video based spatial-temporal volumetric capture system

    WO2020132631A1