Computer-implemented method for generating a 3D body model, system for generating a 3D body model, and non-transitory machine-readable storage medium
By separating bone length variability from body features and adopting bone scaling and joint angle parameters, the complexity of the existing 3D body model generation system and the difficulty in adapting to new postures is solved, achieving more compact and intuitive model generation and control.
Patent Information
- Application Number
- CN202080078930.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-15
- Filing Date
- 2020-11-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-11-13
AI Technical Summary
The existing 3D body model generation system has complexity in generation and application, cannot be effectively implemented in mobile devices and daily applications, and it is difficult to control and adapt to new postures in real time.
By separating bone length variability from body features, a 3D model is generated, using bone scaling and joint angle as main parameters, enabling a more compact model and more intuitive control.
It realizes a more efficient 3D model generation process, reduces the input parameter set, improves the compactness and control intuitiveness of the model, and is suitable for the application of character animation and other 3D bodies.
Smart Images

Figure CN114981844B_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims the benefit of priority to U.S. Provisional Application No. 62 / 936,272, filed on November 15, 2019, the entire content of which is incorporated herein by reference. Technical Field
[0003] This disclosure relates to the generation of 3D body models and their use for image-driven character animation. Background Art
[0004] Modern user devices provide messaging applications that allow users to exchange messages with each other. Such messaging applications have recently begun to incorporate graphics into such communications. The graphics can include avatars or cartoons that mimic user actions. Brief Description of the Drawings
[0005] In the drawings, which are not necessarily drawn to scale, like reference numerals may describe similar components in different views. To facilitate identification of the discussion of any particular element or act, one or more of the most significant digits in the reference numeral refers to the figure number in which the element was first introduced. Some non-limiting examples are shown in the figures of the drawings, in which:
[0006] Figure 1 is a graphical representation of a networked environment in which the present disclosure may be deployed, in accordance with some examples.
[0007] Figure 2 is a graphical representation of a messaging system having functionality on both a client side and a server side, in accordance with some examples.
[0008] Figure 3 is a graphical representation of a data structure maintained in a database, in accordance with some examples.
[0009] Figure 4 is a graphical representation of a message, in accordance with some examples.
[0010] Figure 5 is a graphical representation of operations performed by a 3D body model generation system, in accordance with some examples.
[0011] Figure 6 shows a 3D body model generation system in accordance with some examples and more particularly shows the training of the 3D body model generation system.
[0012] Figure 7 shows the effect of bone length variation on a template, in accordance with some examples.
[0013] Figure 8 shows an example mesh synthesized by posing a template along one effective degree of freedom, in accordance with some examples.
[0014] Figure 9 is a graphical representation of a graphical user interface according to some examples.
[0015] Figure 10 is a flowchart showing an example operation of a messaging application server according to an example embodiment.
[0016] Figure 11 is a graphical representation of a machine in the form of a computer system according to some examples, in which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein.
[0017] Figure 12 is a block diagram showing a software architecture in which examples can be implemented. Detailed Description
[0018] The following description includes systems, methods, techniques, instruction sequences, and computer program products embodying illustrative embodiments of the present disclosure. In the following description, for purposes of explanation, numerous specific details are set forth to provide an understanding of the various embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these specific details. In general, well-known instruction instances, protocols, structures, and techniques are not shown in detail.
[0019] The volumetric representation of the human body forms a bridge between computer graphics and computer vision, facilitating a wide range of applications in motion capture, monocular 3D reconstruction, human synthesis, character animation, and augmented reality. Articulated human deformation can be captured through rigging modeling, in which bones animate the shape of a template (mesh). Current models (e.g., the Skinned Multi-Person Linear model (SMPL)) first synthesize a template mesh in a canonical pose through a linear-based extension. The bone joints are then estimated post hoc by regressing from the synthesized mesh to the joints. While such systems generally work well, their complexity makes them infeasible for implementation on mobile devices and in everyday applications. Additionally, since the models of such systems encode shape and pose information, such systems are not easily (especially not in real time) controllable and adaptable to new poses.
[0020] The present disclosure describes methods for generating 3D body models (specifically, bone-level skinned models that set bone scaling prior to template synthesis). Bone scaling can be manually specified by a user prior to template synthesis or can be determined in an automated process (e.g., using one or more machine learning techniques). Specifically, the disclosed embodiments provide an improved system that generates 3D models by separating bone length variability from body characteristics obtained depending on other factors such as exercise or eating habits. The disclosed embodiments first separately model bone length-driven mesh variability and then combine it with identity-specific updates to represent the complete distribution of the body. This separated representation results in a more compact model, enabling highly accurate reconstruction with a low parameter count. The disclosed embodiments model mesh synthesis as an ordered specification of identity-specific bone lengths, pose-specific joint angles, and identity-specific surface variations (bundled together by linear blend skinning). In one application, decoupling bone length from identity-specific variations can be used to reposition a costume onto a person, e.g., by scaling the length of the assembled costume to match the person's length while preserving the bone length-independent parts of the costume shape.
[0021] In this way, embodiments of the application provide a more efficient process for generating such 3D models. By first modeling the skeleton, the set of input parameters required can be reduced in size. The set of input parameters can be significantly reduced to include only bone scaling coefficients and joint angle coefficients. Additionally, the use of these parameters can be more intuitive for the user. By making the bone scaling coefficients explicitly controllable, more intuitive control over the output 3D body shape can be provided. Although the disclosed techniques are described in the context of 3D human bodies, the same or similar techniques apply to any other type of 3D body (e.g., an animal body or parts of a human or animal body).
[0022] According to a first aspect, the present disclosure describes a computer-implemented method for generating a 3D body model, the method comprising: receiving a plurality of bone scaling coefficients, each bone scaling coefficient corresponding to a respective bone of a bone model; receiving a plurality of joint angle coefficients that jointly define a pose of the bone model; generating the bone model based on the received bone scaling coefficients and the received joint angle coefficients; generating a base surface based on the plurality of bone scaling coefficients; generating an identity surface by deforming the base surface; and generating the 3D body model by mapping the identity surface onto the posed bone model.
[0023] The bone model can be based on a tree structure diagram. Generating the bone model can include: recursively generating a root bone element and a plurality of leaf bone elements.
[0024] The bone model can define a rest position for each bone element. The rest position of each bone element can be represented by a translation vector and a template rotation matrix for the respective element.
[0025] The skeletal model can further define the rotation factor and the scaling factor for each bone element. The rotation factor and the scaling factor are recursively applied to each bone element in sequence.
[0026] Each of the multiple joint angle coefficients can be restricted to a kinematically valid angular range. Each joint between two bone elements can be restricted to certain actual degrees of freedom. For example, the elbow can be restricted to a single degree of freedom. Each rotation can be further restricted to an actual range. For example, the joint angle coefficients can be restricted by mapping the corresponding unconstrained variables to the kinematically valid angular range.
[0027] According to an embodiment, there can be a total of 47 joint angle coefficients corresponding to the 47 degrees of freedom in the skeletal model. However, other skeletal models with more or fewer bone elements can be used. Thus, other skeletal models can have more or fewer joint angle coefficients.
[0028] Restricting the model in this way (restricting the number of joint angle coefficients) can make the processing more efficient. The constraints can be applied within the model itself. Thus, the method avoids the need to post - test the kinematic accuracy of the generated model, for example, by adversarial training methods. Compared with the post - checking process of kinematic feasibility, the method can similarly improve pose recognition, where the application of kinematic constraints improves the recognition accuracy.
[0029] Generating the base surface can include: generating an average surface and then applying a correction to the average surface based on the deformation related to the bone lengths. In this way, modeling the skeleton first simplifies the parameters required to generate a 3D body model. The changes in body shape can be automatically considered in the deformation related to the bone lengths. In this way, the characteristics of body shape and identity body changes can be decoupled. Embodiments of the method can thus reduce many parameters to the level of identity features, which can be mapped as modifications to the general body shape surface.
[0030] In some embodiments, a linear blend skinning process can be used to map the identity surface onto the posed skeletal model. The linear blend skinning process can include mapping surface points on the identity surface based on the bone elements of the skeletal model by: drawing the surface points relative to the rest position of the bone elements and transporting the drawn points based on the posed positions of the corresponding bone elements. The surface points on the identity surface can be mapped based on each bone element of the skeletal model in this way, and a weighting factor is calculated for each corresponding mapping.
[0031] In some implementations, 3D body model generation can be used in the process of character animation. An input 2D image can be provided to a device, where the input 2D image includes a representation of at least one body. Multiple bone scaling coefficients and multiple joint angle coefficients can be determined based on image processing of the input 2D image. For example, in some implementations, a deep convolutional neural network can be used to detect whether at least one body exists in the image, and estimate multiple bone scaling coefficients and multiple joint angle coefficients for the detected at least one body.
[0032] An output 2D image can thus be generated based on the input 2D image and the generated 3D body model, where the representation of the body in the input image is replaced or supplemented by the corresponding 2D representation of the 3D body model. For example, a character can be rigged with bones and joints of a set of bones. The bone rigging of the character can be adjusted based on the bone information generated from the model of the person depicted in the 2D image. In an example, the sizes of some components of the bone rigging can be reduced to match the bones of the corresponding bones of the person's model, while the sizes of other components remain the same or are increased. In this way, the character animation can match or correspond to the person depicted in the image. For example, if the arms of the person depicted in the image are longer than the bone rigging of the character animation, the arms of the bone rigging can be enlarged so that the character animation has arms that appear longer. As another example, if the person depicted in the image is overweight or has a large body weight, the size of the character animation can be increased similarly.
[0033] The 3D body model can be generated based on the multiple bone scaling coefficients and multiple joint angle coefficients determined from the input 2D image, and projected onto a 2D projection based on the input 2D image. The resulting 2D projection can be superimposed on the representation of the body in the input image.
[0034] The currently disclosed method of character animation can avoid unrealistic shape distortion when mapped to an image. By decoupling the body shape and identity body shape changes as described above, the character can be scaled according to the detected bone sizes, and the shape can be deformed accordingly for the actual body shape without affecting the recognition changes of the character.
[0035] Networked computing environment
[0036] Figure 1FIG. 0 is a block diagram showing an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. Messaging system 100 includes multiple instances of client devices 102, each client device 102 hosting several applications including a messaging client 104 and other external applications 109 (e.g., third-party applications). Each messaging client 104 is communicatively coupled via a network 112 (e.g., the Internet) to other instances of messaging clients 104 (e.g., hosted on corresponding other client devices 102), a messaging server system 108, and an external application server 110. Messaging client 104 can also communicate with locally hosted third-party applications 109 using an application programming interface (API).
[0037] Messaging client 104 is capable of communicating with and exchanging data with other messaging clients 104 and a messaging server system 108 via network 112. The data exchanged between messaging clients 104 and between messaging client 104 and messaging server system 108 includes functionality (e.g., commands to activate functionality) as well as payload data (e.g., text, audio, video, or other multimedia data).
[0038] Messaging server system 108 is capable of providing server-side functionality to a particular messaging client 104 via network 112. While certain functionality of messaging system 100 is described herein as being performed by messaging client 104 or by messaging server system 108, the location of certain functionality within messaging client 104 or messaging server system 108 can be a design choice. For example, it is technically preferred that certain technologies and functionality can be initially deployed within messaging server system 108 but later migrated to messaging client 104 on client devices 102 having sufficient processing power.
[0039] Messaging server system 108 supports various services and operations provided to messaging client 104. Such operations include sending data to messaging client 104, receiving data from messaging client 104, and processing data generated by messaging client 104. As an example, the data can include message content, client device information, geographical location information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. Activation and control of data exchange within messaging system 100 are performed through functions available via a user interface (UI) of messaging client 104.
[0040] Now specifically turning to the message server system 108, an Application Programming Interface (API) server 116 is coupled to the application server 114 and provides a programming interface to the application server 114. The application server 114 is communicatively coupled to a database server 120, which facilitates access to a database 126 that stores data associated with messages processed by the application server 114. Similarly, a web server 128 is coupled to the application server 114 and provides a web-based interface to the application server 114. To this end, the web server 128 processes incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.
[0041] The Application Programming Interface (API) server 116 receives and sends message data (e.g., commands and message payloads) between the client device 102 and the application server 114. Specifically, the Application Programming Interface (API) server 116 provides a set of interfaces (e.g., routines and protocols) that the message client 104 can call or query to activate the functions of the application server 114. The Application Programming Interface (API) server 116 exposes various functions supported by the application server 114, including: account registration; login functionality; sending messages from a particular message client 104 to another message client 104 via the application server 114; sending media files (e.g., images or videos) from the message client 104 to the message server 118 for possible access by another message client 104; setting a collection of media data (e.g., a story); retrieving a list of friends of the user of the client device 102; retrieving such collections; retrieving messages and content; adding and deleting entities (e.g., friends) in an entity graph (e.g., a social graph); locating friends in the social graph; and opening application events (e.g., related to the message client 104).
[0042] The application server 114 hosts several server applications and subsystems, including for example a message server 118, an image processing server 122, and a social network server 124. The message server 118 implements several message processing techniques and functions, particularly those related to the aggregation and other processing of the content (e.g., text and multimedia content) included in messages received from multiple instances of the message client 104. As will be described in further detail, text and media content from multiple sources can be aggregated into content collections (e.g., called stories or libraries). These collections are then made available to the message client 104. Given the hardware requirements for such processing, other processor and memory-intensive processing of the data can also be performed by the message server 118 on the server side.
[0043] The application server 114 also includes an image processing server 122, which is dedicated to performing various image processing operations, typically with respect to images or videos within the payloads of messages sent from or received at the message server 118. In combination with Figure 5 The detailed functions of the image processing server 122 are shown and described. The image processing server 122 is used to implement the 3D body model generation operation of the 3D body model generation system 230 ( Figure 2 ).
[0044] In one embodiment, the image processing server 122 detects a person in a 2D image. The image processing server 122 generates a set of landmarks representing the bones / joints of the person detected in the 2D image. The image processing server 122 adjusts the parameters of the model generated by the 3D body model generation system 230 based on the set of landmarks to generate a skeleton and template corresponding to the person detected in the 2D image. A character animation or avatar can be selected (for example, the user can manually select the desired avatar, or the avatar can be automatically retrieved and selected). Based on the skeleton and template generated by the 3D body model generation system 230, the skeleton rigging used to generate the character animation or avatar is adjusted to represent the physical attributes of the person depicted in the 2D image. For example, the arms of the skeleton rigging can be enlarged to match the arm bones of the person depicted in the image. Specifically, if a person's arms are longer than those of an average person represented by the avatar rigging, the arms of the avatar rigging can be extended. Different users with shorter arms can have the avatar rigging generated to have shorter arms. In this way, the avatar or character generated based on the skeleton rigging can more closely reflect the physical attributes of the person depicted in the 2D image. Then, the avatar can replace or supplement the person in the 2D image and / or video.
[0045] The social network server 124 supports various social networking functions and services and makes these functions and services available to the message server 118. To this end, the social network server 124 maintains and accesses the entity graph 308 within the database 126 (as Figure 3 shown). Examples of the functions and services supported by the social network server 124 include identifying other users of the messaging system 100 with whom a particular user has a relationship or "follows", and identifying the interests and other entities of a particular user.
[0046] Returning to the messaging client 104, the features and functionality of external resources (e.g., third-party applications 109 or applets) are made available to the user via the interface of the messaging client 104. The messaging client 104 receives a user selection of an option to initiate or access the features of an external resource (e.g., a third-party resource) (e.g., an external application 109). The external resource can be a third-party application (external app 109) (e.g., a "native app") installed on the client device 102, or a scaled-down version (e.g., an "applet") of a third-party application hosted on or located remotely from the client device 102 (e.g., on a third-party server 110). The scaled-down version of the third-party application includes a subset of the features and functionality of the third-party application (e.g., the full, native version of a third-party stand-alone application) and is implemented using a markup language document. In one example, the scaled-down version of the third-party application (e.g., an "applet") is a web-based markup language version of the third-party application and is embedded within the messaging client 104. In addition to using a markup language document (e.g., a.*ml file), an applet can incorporate a scripting language (e.g., a.*js file or a.json file) and a style sheet (e.g., a.*ss file).
[0047] In response to receiving a user selection of an option to initiate or access the features of an external resource (external application 109), the messaging client 104 determines whether the selected external resource is a web-based external resource or a locally installed external application. In some cases, an external application 109 locally installed on the client device 102 can be independent of the messaging client 104 and can be launched separately from the messaging client 104, for example, by selecting an icon corresponding to the external application 109 on the client's home screen. The scaled-down version of such an external application can be launched or accessed via the messaging client 104, and in some examples, no part (or only a limited part) of the scaled-down external application can be accessed outside of the messaging client 104. The scaled-down external application can be launched by receiving from the external application server 110 a markup language document associated with the scaled-down external application and processing such a document by the messaging client 104.
[0048] In response to determining that the external resource is a locally installed external application 109, the messaging client 104 instructs the client device 102 to launch the external application 109 by executing locally stored code corresponding to the external application 109. In response to determining that the external resource is a web-based resource, the messaging client 104 communicates with the external application server 110 to obtain a markup language document corresponding to the selected resource. The messaging client 104 then processes the obtained markup language document to render the web-based external resource within the user interface of the messaging client 104.
[0049] The messaging client 104 can notify a user of the client device 102 or other users associated with such a user (e.g., "friends") of activities occurring in one or more external resources. For example, the messaging client 104 can provide notifications to participants in a conversation (e.g., a chat session) in the messaging client 104 regarding current or recent use of an external resource by one or more members of a group of users. One or more users can be invited to join an external resource for an activity or to initiate (in a group of friends) an external resource that was recently used but is currently inactive. The external resource can provide to participants in the conversation (each using their respective messaging client 104) the ability of one or more members of a group of users to share items, conditions, statuses, or locations within the external resource into the chat session. The shared item can be an interactive chat card that the members of the chat can use to interact, e.g., to initiate the corresponding external resource, view specific information within the external resource, or take the members of the chat to a specific location or status within the external resource. Within a given external resource, a response message can be sent to the user on the messaging client 104. The external resource can selectively include different media items in the response based on the current context of the external resource.
[0050] The messaging client 104 can present a list of available external resources (e.g., third-party or external applications 109 or applets) to initiate or access a given external resource. The list can be presented in a context-sensitive menu. For example, icons representing different external applications 109 (or applets) can vary based on how the menu is initiated (e.g., from a conversation interface or from a non-conversation interface).
[0051] System Architecture
[0052] Figure 2 is a block diagram showing further details regarding the messaging system 100 according to some examples. Specifically, the messaging system 100 is shown as including a messaging client 104 and an application server 114. The messaging system 100 includes multiple subsystems that are supported by the messaging client 104 on the client side and by the application server 114 on the server side. These subsystems include, for example, an ephemeral timer system 202, a collection management system 204, an enhancement system 208, a map system 210, a game system 212, and an external resource system 220.
[0053] The short-lived timer system 202 is responsible for enforcing temporary or time-limited access to content by the message client 104 and the message server 118. The short-lived timer system 202 includes a number of timers that selectively enable access (e.g., for presentation and display) to messages and associated content via the message client 104 based on the duration and display parameters associated with a message or collection of messages (e.g., a story). Additional details regarding the operation of the short-lived timer system 202 are provided below.
[0054] The collection management system 204 is responsible for managing groups or collections of media (e.g., collections of text, image, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event libraries" or "event stories". Such collections can be made available for a specified period of time (e.g., the duration of the event to which the content pertains). For example, content related to a concert can be made available as a "story" for the duration of that concert. The collection management system 204 can also be responsible for publishing an icon that provides notice of the existence of a particular collection to the user interface of the message client 104.
[0055] In addition, the collection management system 204 also includes a curation interface 206 that enables a collection manager to manage and curate a particular collection of content. For example, the curation interface 206 enables an event organizer to curate a collection of content related to a particular event (e.g., delete inappropriate content or redundant messages). Additionally, the collection management system 204 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some embodiments, a user may be compensated for including user-generated content in a collection. In such cases, the collection management system 204 operates to automatically pay such users for the use of their content.
[0056] The enhancement system 208 provides various functions that enable a user to enhance (e.g., annotate or otherwise modify or edit) media content associated with a message. For example, the enhancement system 208 provides functions related to generating and publishing a media overlay for a message to be processed by the messaging system 100. The enhancement system 208 operably provides a media overlay or enhancement (e.g., an image filter) to the message client 104 based on the geographical location of the client device 102. In another example, the enhancement system 208 operably supplies a media overlay to the message client 104 based on other information such as the social network information of a user of the client device 102. The media overlay can include audio and visual content as well as visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. The audio and visual content or visual effects can be applied to a media content item (e.g., a photo) at the client device 102. For example, the media overlay can include text, graphic elements, or images that can be superimposed on top of a photo taken by the client device 102. In another example, the media overlay includes a location identifier overlay (e.g., Venice Beach), a live event name, or a merchant name overlay (e.g., Beach Café). In another example, the enhancement system 208 uses the geographical location of the client device 102 to identify a media overlay that includes the name of a merchant at the geographical location of the client device 102. The media overlay can include other markers associated with the merchant. The media overlay can be stored in the database 126 and accessed via the database server 120.
[0057] In some examples, the enhancement system 208 provides a user-based publishing platform that enables a user to select a geographical location on a map and upload content associated with the selected geographical location. The user can also specify the circumstances under which a particular media overlay should be provided to other users. The enhancement system 208 generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geographical location.
[0058] In other examples, the augmentation system 208 provides a merchant-based publishing platform that enables a merchant to select a specific media overlay associated with a geographical location via a bidding process. For example, the augmentation system 208 associates the media overlay of the highest bidding merchant with the corresponding geographical location for a predefined amount of time. The augmentation system 208 communicates with the image processing server 122 to automatically select and activate an augmented reality experience associated with an image captured by the client device 102. Once an augmented reality experience is selected when the user scans an image using a camera device in the user's environment, one or more images, videos, or augmented reality graphic elements are retrieved and presented as an overlay on top of the scanned image. In some cases, the camera device is switched to a front view (e.g., the front camera device of the client device 102 is activated in response to the activation of a specific augmented reality experience), and the image from the front camera device of the client device 102 instead of the rear camera device of the client device 102 starts to be displayed on the client device 102. One or more images, videos, or augmented reality graphic elements are retrieved and presented as an overlay on top of the image captured and displayed by the front camera device of the client device 102.
[0059] The map system 210 provides various geolocation functions and supports the presentation of map-based media content and messages by the messaging client 104. For example, the map system 210 is capable of displaying (e.g., stored in the profile data 316) user icons or avatars on a map to indicate the current or past locations of the user's "friends", as well as media content (e.g., a collection of messages including photos and videos) generated by these friends within the context of the map. For example, on the map interface of the messaging client 104, a message posted by a user from a specific geographical location to the messaging system 100 can be displayed to the "friends" of the specific user within the context of that specific location on the map. A user can also share their location and status information with other users of the messaging system 100 (e.g., using an appropriate status avatar) via the messaging client 104, and this location and status information is displayed to the selected users within the context of the map interface of the messaging client 104.
[0060] The gaming system 212 provides various gaming functions within the context of the messaging client 104. The messaging client 104 provides a gaming interface that provides a list of available games (e.g., web-based games or web-based applications) that can be launched by the user within the context of the messaging client 104 and played with other users of the messaging system 100. The messaging system 100 also enables a specific user to invite other users to participate in playing a specific game by sending an invitation to such other users from the messaging client 104. The messaging client 104 also supports messaging (e.g., chat) in both voice and text within the gaming context, provides a leaderboard for the game, and also supports providing in-game rewards (e.g., coins and items).
[0061] The external resource system 220 provides an interface for the messaging client 104 to communicate with the external application server 110 to initiate or access external resources. Each external resource (app) server 110 hosts a small-scale version of an application, such as an application based on a markup language (e.g., HTML5), or an external application (e.g., a game, utility, payment, or ride-sharing application outside of the messaging client 104). The messaging client 104 can initiate a web-based resource (e.g., an application) by accessing an HTML5 file from an external resource (app) server 110 associated with the web-based resource. In some examples, the applications hosted by the external resource server 110 utilize a software development kit (SDK) provided by the messaging server 118 for JavaScript programming. The SDK includes application programming interfaces (APIs) that have functions that can be called or activated by web-based applications. In some examples, the messaging server 118 includes a JavaScript library that provides access to certain user data of the messaging client 104 to a given third-party resource. HTML5 is used as an example technology for programming games, but applications and resources programmed based on other technologies can be used.
[0062] To integrate the functions of the SDK into a web-based resource, the SDK is downloaded by the external resource (app) server 110 from the messaging server 118 or otherwise received by the external resource (app) server 110. Once downloaded or received, the SDK will be included as part of the application code of the web-based external resource. Then, the code of the web-based resource can call or activate certain functions of the SDK to integrate the features of the messaging client 104 into the web-based resource.
[0063] The SDK stored on the messaging server 118 effectively provides a bridge between external resources (e.g., third-party or external applications 109 or applets) and the messaging client 104. This provides users with a seamless experience of communicating with other users on the messaging client 104 while also retaining the look and feel of the messaging client 104. To bridge the communication between external resources and the messaging client 104, in some examples, the SDK facilitates the communication between the external resource server 110 and the messaging client 104. In some examples, the WebViewJavaScriptBridge running on the client device 102 establishes two one-way communication channels between the external resource and the messaging client 104. Messages are sent asynchronously between the external resource and the messaging client 104 via these communication channels. Each SDK function call is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message using that callback identifier.
[0064] By using the SDK, not all information from the messaging client 104 is shared with the external resource server 110. The SDK restricts which information is shared based on the needs of the external resources. In some examples, each external resource server 110 provides an HTML5 file corresponding to a web-based external resource to the messaging server 118. The messaging server 118 can add a visual representation (such as a box design or other graphics) of the web-based external resource in the messaging client 104. Once the user selects the visual representation through the GUI of the messaging client 104 or indicates that the messaging client 104 accesses the features of the web-based external resource, the messaging client 104 obtains the HTML5 file and instantiates the resources required to access the features of the web-based external resource.
[0065] The messaging client 104 presents a graphical user interface for the external resource (e.g., a login page or a splash screen). During, before, or after presenting the login page or splash screen, the messaging client 104 determines whether the launched external resource has been previously authorized to access the user data of the messaging client 104. In response to determining that the launched external resource has been previously authorized to access the user data of the messaging client 104, the messaging client 104 presents another graphical user interface of the external resource, including the functions and features of the external resource. In response to determining that the launched external resource has not been previously authorized to access the user data of the messaging client 104, after a threshold time period (e.g., 3 seconds) of displaying the login page or splash screen of the external resource, the messaging client 104 slides up a menu (e.g., animates the menu to emerge from the bottom of the screen to the middle or other part of the screen) to authorize the external resource to access the user data. The menu identifies the type of user data that the external resource will be authorized to use. In response to receiving a user selection of the accept option, the messaging client 104 adds the external resource to the list of authorized external resources and allows the external resource to access the user data from the messaging client 104. In some examples, the external resource is authorized by the messaging client 104 to access the user data according to the OAuth 2 framework.
[0066] The messaging client 104 controls the type of user data shared with external resources based on the type of authorized external resources. For example, an external resource that includes a complete external application (e.g., a third-party or external application 109) is provided access to a first type of user data (e.g., a two-dimensional avatar of a user with or without different avatar characteristics). As another example, an external resource that includes a scaled-down version of an external application (e.g., a web-based version of a third-party application) is provided access to a second type of user data (e.g., payment information, a two-dimensional avatar of the user, a three-dimensional avatar of the user, and avatars with various avatar characteristics). Avatar characteristics include different ways of customizing the appearance and feel of an avatar (e.g., different poses, facial features, clothing, etc.).
[0067] The 3D body model generation system 230 generates a bone-level skinned model of a human skeleton. Specifically, the 3D body model generation system receives bone scaling, joint angles, and shape coefficients as inputs and returns an array of 3D vertex positions. The 3D body model generation system 230 operates along two streams, the results of which are combined in a final stage. The first stream (shown at the top of Figure 5 ) determines the internal skeleton by setting the bone scaling through the bone scaling coefficient c b to obtain the bound (rest) pose. This is then transformed into a new pose by specifying the joint angles θ, resulting in the final skeleton T(c b , θ). The second stream (shown at the bottom of Figure 5 ) models the person-specific template synthesis process: starting from a mesh corresponding to an average body shape , the effect of the bone scaling is absorbed by adding the shape correction term Vb. This is further enhanced by the identity-specific shape update Vs. The person template is represented by Equation 1 below:
[0068]
[0069] The results of the two streams are bundled together using linear blend skinning (LBS) to obtain the pose template
[0070] Data Architecture
[0071] Figure 3 is a schematic diagram showing a data structure 300 that can be stored in the database 126 of the messaging server system 108 according to certain examples. Although the contents of the database 126 are shown as including several tables, it should be understood that the data can be stored in other types of data structures (e.g., as an object-oriented database).
[0072] The database 126 includes message data stored in the message table 302. For any particular message, the message data includes at least message sender data, message recipient (or receiver) data, and a payload. Details regarding additional information that can be included in a message and that is included in the message data stored in the message table 302 are described below with reference to Figure 4 additional details regarding information that can be included in a message and that is included in the message data stored in the message table 302.
[0073] The entity table 306 stores entity data and is (e.g., referentially) linked to the entity graph 308 and the profile data 316. Entities for which records are maintained within the entity table 306 can include individuals, corporate entities, organizations, objects, locations, events, etc. Any entity for which the message server system 108 stores data regarding it can be an identified entity, regardless of the entity type. Each entity is provided with a unique identifier, as well as an entity type identifier (not shown).
[0074] The entity graph 308 stores information regarding relationships and associations between entities. By way of example only, such relationships can be social, professional (e.g., working in a common company or organization), interest-based, or activity-based.
[0075] The profile data 316 stores various types of profile data regarding a particular entity. Based on privacy settings specified by the particular entity, the profile data 316 can be selectively used and presented to other users of the messaging system 100. In the case where the entity is an individual, the profile data 316 includes, for example, a username, a telephone number, an address, settings (e.g., notification and privacy settings), and an avatar representation (or a collection of such avatar representations) selected by the user. A particular user can then selectively include one or more of these avatar representations in the content of messages transmitted via the messaging system 100 and on a map interface displayed by the message client 104 to other users. The collection of avatar representations can include “status avatars” that present a graphical representation of a status or activity that the user can select to communicate at a particular time.
[0076] In the case where the entity is a group, in addition to the group name, members, and various settings (e.g., notifications) of the related group, the profile data 316 of the group can similarly include one or more avatar representations associated with the group.
[0077] The database 126 also stores enhancement data, such as overlays or filters, in the enhancement table 310. The enhancement data is associated with videos (video data is stored in the video table 304) and images (image data is stored in the image table 312) and is applied to the videos and images.
[0078] In one example, the filter is an overlay that is displayed as an overlay on an image or video during presentation to the receiving user. The filter can be of various types, including a filter selected by the user from a set of filters presented to the sending user by the messaging client 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geo-filters), which can be presented to the sending user based on geolocation. For example, specific geolocation filters for nearby or special locations can be presented within the user interface by the messaging client 104 based on geolocation information determined by the global positioning system (GPS) unit of the client device 102.
[0079] Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client 104 based on other input or information collected by the client device 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed at which the sending user is traveling, the battery life of the client device 102, or the current time.
[0080] Other enhanced data that can be stored in the image table 312 includes augmented reality content items (e.g., corresponding to app lenses or augmented reality experiences). Augmented reality content items can be real-time special effects and sounds that can be added to an image or video. Each augmented reality experience can be associated with one or more marker images. In some embodiments, when a marker image is determined to match a query image received from the client device 102, the corresponding augmented reality experience (e.g., enhanced data) of the marker image is retrieved from the image table 312 and provided to the client device 102. In combination Figure 9 Various types of augmented reality experiences are shown and discussed.
[0081] As described above, enhanced data includes augmented reality content items, overlays, image transformations, AR images, and similar items that refer to modifications that can be applied to image data (e.g., video or images). This includes real-time modifications that modify an image when it is captured using the device sensors of client device 102 (e.g., one or more camera devices), and then the modified image is used for display on the screen of client device 102. This also includes modifications to stored content (e.g., video clips in a library that can be modified). For example, in client device 102 that can access multiple augmented reality content items, a user can use a single video clip with multiple augmented reality content items to see how different augmented reality content items would modify the stored clip. For example, by selecting different augmented reality content items for the same content, multiple augmented reality content items applying different pseudo-random motion models can be applied to the same content. Similarly, real-time video capture can be used with the shown modifications to display how the video images currently captured by the sensors of client device 102 would modify the captured data. Such data can be simply displayed on the screen without being stored in memory, or the content captured by the device sensors can be recorded and stored in memory with or without modifications (or both). In some systems, a preview function can display how different augmented reality content items would be displayed simultaneously in different windows of the display. For example, this can enable viewing multiple windows with different pseudo-random animations on the display at the same time.
[0082] Thus, systems that use data of augmented reality content items and various other such transformation systems that use such data to modify content can involve the detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in video frames, tracking these objects as they leave the field of view, enter the field of view, and move around in the field of view, and modifying or transforming these objects while tracking them. In various examples, different methods can be used to implement such transformations. Some examples can involve generating a 3D mesh model of one or more objects, and using the transformation and animated textures of the model within the video to implement the transformation. In other examples, tracking of points on the object can be used to place an image or texture (which can be two-dimensional or three-dimensional) at the tracked location. In further examples, neural network analysis of video frames can be used to place an image, model, or texture in the content (e.g., an image or video frame). Thus, augmented reality content items refer both to the images, models, and textures used to create transformations in content and to the additional modeling and analysis information required to implement such transformations through object detection, tracking, and placement.
[0083] Real-time video processing can be performed using any type of video data (e.g., video stream, video file, etc.) stored in the memory of any type of computerized system. For example, a user can load a video file and save it in the device's memory, or can use the device's sensors to generate a video stream. Additionally, any object, such as parts of a human face and body, an animal, or a non-biological object (e.g., a chair, a car, or other object) can be processed using a computer animation model.
[0084] In some examples, when a specific modification is selected along with the content to be transformed, the elements to be transformed are identified by a computing device and then detected and tracked if they are present in the frames of the video. The elements of the object are modified according to the modification request, thereby transforming the frames of the video stream. For different kinds of transformations, the transformation of the frames of the video stream can be performed by different methods. For example, for a frame transformation that mainly refers to the variation form of the elements of an object, the feature points of each element of the object are calculated (e.g., using an active shape model (ASM) or other known methods). Then, a grid based on the feature points is generated for each of at least one element of the object. This grid is used for the subsequent stage of tracking the elements of the object in the video stream. During the tracking process, the mentioned grid of each element is aligned with the position of each element. Then, additional points are generated on the grid. A first set of first points is generated for each element based on the modification request, and a set of second points is generated for each element based on this set of first points and the modification request. Then, the elements of the object can be modified based on this set of first points, this set of second points, and the grid to transform the frames of the video stream. In this method, the background of the modified object can also be changed or distorted by tracking and modifying the background.
[0085] In some examples, the transformation of changing some regions of an object using the elements of the object can be performed by calculating the feature points of each element of the object and generating a grid based on the calculated feature points. Points are generated on the grid, and then various regions are generated based on these points. Then, the elements of the object are tracked by aligning the regions of each element with the position of each of at least one element, and the attributes of the regions can be modified based on the modification request, thereby transforming the frames of the video stream. Depending on the specific modification request, the attributes of the mentioned regions can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least part of the region from the frames of the video stream; including one or more new objects in the region based on the modification request; and modifying or distorting the region or the elements of the object. In various examples, any combination of such modifications or other similar modifications can be used. For some models to be animated, some feature points can be selected as control points for determining the entire state space of the options for model animation.
[0086] In some examples of computer animation models that use face detection to transform image data, a specific face detection algorithm (e.g., Viola-Jones) is used to detect faces in an image. Then, an Active Shape Model (ASM) algorithm is applied to the face region of the image to detect facial feature reference points.
[0087] Other methods and algorithms suitable for face detection can be used. For example, in some examples, landmarks are used to locate features, which represent distinguishable points that exist in most of the images under consideration. For example, for face landmarks, the location of the left eye pupil can be used. If the initial landmark is not recognizable (e.g., if a person has an eye patch), secondary landmarks can be used. Such a landmark recognition process can be used for any such object. In some examples, a set of landmarks forms a shape. The shape can be represented as a vector using the coordinates of the points in the shape. One shape is aligned with another shape through a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the shape points. The average shape is the average of the aligned training shapes.
[0088] In some examples, the search for landmarks starts from an average shape that is aligned with the position and size of the face determined by the global face detector. Then, this search repeats the following steps: a tentative shape is proposed by adjusting the positioning of the shape points through template matching of the image texture around each point, and then the tentative shape is made to conform to the global shape model until convergence occurs. In some systems, individual template matches are unreliable, and the shape model pools the results of weak template matches to form a stronger overall classifier. The entire search is repeated at each level of the image pyramid from coarse resolution to fine resolution.
[0089] The transformation system can capture an image or video stream on a client device (e.g., client device 102) and perform complex image manipulations locally on the client device 102 while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion transformation (e.g., changing a face from a frown to a smile), state transformation (e.g., making an object age, reducing apparent age, changing gender), style transformation, application of graphical elements, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on the client device 102.
[0090] In some examples, a computer animation model for transforming image data can be used by a system in which a user can use a client device 102 having a neural network to capture an image or video stream of the user (e.g., a selfie), where the neural network operates as part of a messaging client 104 operating on the client device 102. A transformation system operating within the messaging client 104 determines the presence of a face in the image or video stream and provides a modification icon associated with the computer animation model to transform the image data, or the computer animation model can be presented as associated with the interfaces described herein. The modification icon includes a change that can be the basis for modifying the user's face in the image or video stream as part of a modification operation. Once the modification icon is selected, the transformation system initiates a process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiling face on the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video stream can be presented in a graphical user interface displayed on the client device 102. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, once the modification icon is selected, the user can capture the image or video stream and the modified result can be presented in real time or near real time. Additionally, when a video stream is being captured, the modification can be persistent and the selected modification icon remains toggled. A neural network of machine learning can be used to implement such modifications.
[0091] A graphical user interface presenting the modifications performed by the transformation system can provide additional interaction options for the user. Such options can be based on the interface used to initiate content capture and selection for a particular computer animation model (e.g., initiated from a content creator user interface). In various examples, the modification can persist after an initial selection of the modification icon. The user can open or close the modification by tapping or otherwise selecting the face modified by the transformation system and store it for later viewing or browse to other areas of the imaging application. In cases where the transformation system modifies multiple faces, the user can globally open or close the modification by tapping or selecting an individual face modified and displayed within the graphical user interface. In some examples, individual faces within a group of multiple faces can be modified separately, or such modification can be toggled individually by tapping or selecting each individual face or a series of individual faces displayed within the graphical user interface.
[0092] The story table 314 stores data related to a collection of images, videos, or audio data associated with messages, which are compiled into a collection (e.g., a story or a library). The creation of a particular collection can be initiated by a particular user (e.g., each user whose record is maintained in the entity table 306). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcast by that user. To this end, the user interface of the message client 104 can include user-selectable icons to enable the sending user to add specific content to his or her personal story.
[0093] The collection can also constitute a "live story", which is a collection of content from multiple users and is created manually, automatically, or using a combination of manual and automatic techniques. For example, a "live story" can constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and are at a co-location event at a particular time can be presented, for example, via the user interface of the message client 104, with the option to contribute content to a particular live story. The live story can be identified to him or her by the message client 104 based on the user's location. The end result is a "live story" told from a community perspective.
[0094] Another type of content collection is called a "location story", which enables users whose client devices 102 are located within a particular geolocation (e.g., on a college or university campus) to contribute to a particular collection. In some examples, contributing to a location story may require secondary authentication to verify that the end user belongs to a particular organization or other entity (e.g., is a student on a university campus).
[0095] As mentioned above, the video table 304 stores video data, which, in one example, is associated with messages whose records are stored in the message table 302. Similarly, the image table 312 stores image data associated with messages whose message data is stored in the entity table 306. The entity table 306 can associate various enhancements from the enhancement table 310 with the various images and videos stored in the image table 312 and the video table 304.
[0096] Data communication architecture
[0097] Figure 4FIG. 0 is a schematic diagram showing the structure of a message 400 according to some examples, which is generated by a message client 104 for transmission to another message client 104 or a message server 118. The content of a particular message 400 is used to populate a message table 302 stored in a database 126, which can be accessed by the message server 118. Similarly, the content of the message 400 is stored in memory as "in-transit" or "in-flight" data of the client device 102 or the application server 114. The message 400 is shown as including the following example components
[0098] ● Message identifier 402: A unique identifier that identifies the message 400.
[0099] ● Message text payload 404: Text to be generated by a user via a user interface of the client device 102 and included in the message 400.
[0100] ● Message image payload 406: Image data captured by a camera device component of the client device 102 or retrieved from the memory of the client device 102 and included in the message 400. The image data of the message 400 for sending or receiving can be stored in an image table 312.
[0101] ● Message video payload 408: Video data captured by a camera device component or retrieved from a memory component of the client device 102 and included in the message 400. The video data of the message 400 for sending or receiving can be stored in a video table 304.
[0102] ● Message audio payload 410: Audio data captured by a microphone or retrieved from a memory component of the client device 102 and included in the message 400.
[0103] ● Message enhancement data 412: Enhancement data (e.g., filters, stickers, or other enhancements) representing enhancements to be applied to the message image payload 406, the message video payload 408, or the message audio payload 410 of the message 400. The enhancement data of the message 400 for sending or receiving can be stored in an enhancement table 310.
[0104] ● Message duration parameter 414: A parameter value that indicates the amount of time in seconds for which the content of the message (e.g., the message image payload 406, the message video payload 408, the message audio payload 410) will be presented to a user or made accessible to the user via the message client 104.
[0105] ● Message Geolocation Parameter 416: Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of a message. Multiple Message Geolocation Parameter 416 values may be included in the payload, where each of these parameter values is associated with a content item included in the content (e.g., a specific image within a Message Image Payload 406, or a specific video within a Message Video Payload 408).
[0106] ● Message Story Identifier 418: An identifier value that identifies one or more content collections (e.g., a “story”) through which a specific content item within a Message Image Payload 406 of a Message 400 is associated with the one or more content collections. For example, an identifier value may be used to associate each of multiple images within a Message Image Payload 406 with multiple content collections.
[0107] ● Message Tag 420: Each Message 400 may be tagged with multiple tags, each of which indicates a theme of the content included in the message payload. For example, in a case where a specific image included in a Message Image Payload 406 depicts an animal (e.g., a lion), a tag value indicating the relevant animal may be included in the Message Tag 420. The tag values may be generated manually based on user input, or may be generated automatically using, for example, image recognition.
[0108] ● Message Sender Identifier 422: An identifier (e.g., a message system identifier, an email address, or a device identifier) that indicates the user of the client device 102 on which the Message 400 was generated and from which the Message 400 was sent.
[0109] ● Message Recipient Identifier 424: An identifier (e.g., a message system identifier, an email address, or a device identifier) that indicates the user of the client device 102 to which the Message 400 is addressed.
[0110] The content (e.g., values) of the various components of Message 400 may be pointers to the locations of stored content data values in a table. For example, an image value within a Message Image Payload 406 may be a pointer to a location (or address) within an Image Table 312. Similarly, values within a Message Video Payload 408 may point to data stored within a Video Table 304, values stored in a Message Annotation 412 may point to data stored in an Annotation Table 310, values stored in a Message Story Identifier 418 may point to data stored in a Story Table 314, and values stored in a Message Sender Identifier 422 and a Message Recipient Identifier 424 may point to user records stored in an Entity Table 306.
[0111] Figure 5FIG. 500 is a graphical representation of operations performed by a 3D body model generation system 230 according to some embodiments. Specifically, the top row shows bone synthesis: starting from a standard bind pose 510, first the bone lengths are scaled and then joint transformations are applied. The bottom row shows shape control: the standard mesh template is affected by the bone scale transformation through bone scale blend shapes and then further updated to capture identity-specific shape variations. The bones drive the deformation of the generated template through LBS to produce a posed shape 580.
[0112] As an example, the model starts with an average skeleton 510 in a rest or bind pose. This average skeleton 510 represents an average population or overall population size. A template 550 or mesh representing the appearance and feel (identity or skin) of the overall or average population is also provided. As a first control, bone scale parameters 512 (referred to as b) are calculated for a given person (e.g., the person depicted in a 2D image). These bone scale parameters 512 are used to adjust the length of each individual bone (e.g., increase the size of the arms and decrease the size of the legs) to output a scaled skeleton 520. In parallel with or after scaling the skeleton, the template of the overall population is corrected based on the bone scale parameters by applying a bone scale correction 514 (referred to as V b ). This outputs a template 560 having a size and scale corresponding to the scaled skeleton 520.
[0113] The parameter θ is applied to the scaled skeleton to kinematically adjust the bones and joints of the scaled skeleton to a specific pose. This outputs a scaled and posed skeleton 530. A shape update parameter 562 (referred to as V S ) is applied to the template that has already been scaled to account for person-specific variability, and a template 570 is generated after the scale correction and shape update. Linear blend skinning (LBS) 540 is applied to the scaled and posed skeleton 530 and the template 570 after the scale correction and shape update to output a posed shape template 580. That is, the template representing the identity of the person is posed based on the shape and pose of the skeleton 530 to output a posed shape. This posed shape can then be used to animate or avatarize the representation of the person.
[0114] The skeleton 510 is determined by a tree structure diagram that links human bones together through joints. Starting from a single bone, its "bind pose" is expressed by a template rotation matrix Rt and a translation vector Ot representing the displacement and rotation between the coordinate systems at the two bone joints. The transformation is modeled relative to the bind pose by a rotation matrix R and a scaling factor s, which are bundled together in a 4×4 matrix T:
[0115]
[0116] Common models for character modeling use s = 1 and only allow limb rotation. Any changes in object scaling or bone length are modeled by modifying the displacement O at the bind pose. This is done only implicitly by regressing the bind pose joints from the 3D synthetic shape. The disclosed embodiments provide a handle on limb scaling through the parameter s, making the synthesis of the human skeleton explicitly controllable.
[0117] The complete skeleton is recursively constructed, propagating from the root node to the leaf nodes along the kinematic chain. Each bone transformation encodes the displacement, rotation, and scaling between two adjacent bones i and j, where i is the parent node and j is the child node. To simplify the notation, the modeling is described along a single kinematic chain, i.e., j = i + 1, and the local transformation of the bone is denoted as T i . The global transformation Tj from the local coordinates of bone j to world coordinates is given by: T j = Π i≤j T i , where the transformation consists of each bone on the path from the root to the j-th node. This product accumulates the effects of successive transformations: for example, a change in bone scaling will have the same scaling on all its descendants. These descendants can in turn have their own scaling parameters, which are combined with those of their ancestors. The 3D position of each bone j can be read from the last column of Tj, while the upper left 3×3 part of Tj provides the scaling and orientation of its coordinate system.
[0118] The human proportions are explicitly modeled by scaling each bone. In the example, principal component analysis (PCA) is performed on the bone lengths and the individual bone scalings are represented using principal component analysis as: (Equation 3), where c b is the bone scaling coefficient, P b is the bone scaling matrix, and b- is the average bone scaling. Referring back Figure 5 , the bone scaling parameter 512 scales the average skeleton 510 according to equation 2 above by adjusting the c b for a specific person (e.g., the person depicted in a 2D image). The parameter P b is a learned parameter set for the population and will be discussed in more detail in conjunction with Figure 6 . The bone scaling s that appears in equation 1 is intended to be used recursively through the kinematic chain, meaning that the product of the parent scalings gives the actual bone scaling, bj = Π i≤j s i . This can be used to transform the prediction of equation 3 into a form that can be used in equation 1:
[0119]
[0120] To adjust the pose of the skeleton to create a scaled and posed skeleton 530, the joint angles are modeled to account for the kinematic constraints of the human body. For example, the knee has one degree of freedom, the wrist has two degrees of freedom, and the neck has three degrees of freedom. For each joint, the invalid degrees of freedom are set to be equivalent to zero, and the remaining angles are restricted within a reasonable range (e.g., ±45 degrees for the elbow). Figure 8 An example mesh synthesized by posing the template along one valid degree of freedom is shown. For each such degree of freedom, an unconstrained variable x ∈ ℝ is used and mapped to a valid Euler angle θ ∈ [θmin, θmax] by using a hyperbolic tangent unit:
[0121]
[0122] This enables unconstrained optimization when fitting the model to the data while obtaining kinematically feasible poses. The Euler angles of each resulting joint are converted into a rotation matrix, obtaining the matrix R in Equation 1.
[0123] The template 550 is modeled considering the bone lengths of a person and identity-specific variability. Bone lengths can be used to explain a large part of the body shape variability. For example, longer bones are associated with a male body type, while the proportions of the limbs may be related to the variability of the lanky, stocky, and medium body types. The bone length-related deformations on the template surface are represented by a linear update: V b = c b P bc , where P bc is the matrix of bone correction blend shapes and is a learned parameter (as discussed in conjunction with Figure 6 ).
[0124] Figure 7 The effect of bone length variation on the template is shown. Simple linear blend skinning results in artifacts 710 and 720. The linear bone correction blend shapes eliminate these artifacts and capture the correlation between bone lengths and gender and body type. For example, the artifacts 710 and 720 no longer exist in the templates 712 and 722. Specifically, if the size of a given template is reduced based on smaller bones, the usual template generation model introduces the artifacts 710 and 720, while the templates 712 and 722 generated according to the disclosed embodiments do not have the artifacts.
[0125] After considering the bone length-related part of the shape variability to generate an intermediate template 560, the person-specific variability is considered, and the template 570 is generated according to V s In one embodiment, V is linearly calculated according to V s = c s P s , where c s is the shape coefficient and P s is...s is a matrix of shape components.
[0126] Use LBS to synthesize template 570 based on the skeleton, where the deformation of the template mesh V is determined by the transformation of the skeleton. As considered by the matrix describing the binding pose of the skeleton, where the 3D mesh vertices take their standard values v i ∈ V, and the target pose is described by T j According to LBS, each vertex is affected by each bone j according to the weight w ij ; the vertex position of the target pose is given by:
[0127]
[0128] This equation can be understood as drawing each point v k (by multiplying it by ), and then transferring it to the target bone (by multiplying by T j ).
[0129] Figure 6 FIG. shows a 3D body model generation system 230 according to some embodiments and more specifically shows the training of the 3D body model generation system 230. The 3D body model generation system 230 includes a pose and bone scaling module 620, a bone correction and average shape module 630, and a vertex displacement module 640. The 3D body model generation system 230 receives skeleton data 610 (e.g., provided by landmark positions, which are generated based on 2D images depicting a person and processed by a neural network). In some cases, the 3D body model generation system 230 operates on avatar equipment data 614 to adjust the avatar or character to correspond to the physical attributes, shape, and size of the person depicted in the 2D image.
[0130] To train the 3D body model generation system 230, a set of 3D scans of various people depicting different poses and shapes is processed. The training is carried out in multiple stages. The result of the training is to learn the above parameters of the model from the data. In some cases, the 3D scans include high-resolution 3D scans 612 of 4400 subjects wearing tight clothes. The training process solves a continuously required optimization problem and uses automatic differentiation during optimization to efficiently calculate the derivatives.
[0131] Each 3D scan 612S n is associated with 73 anatomical landmarks L that have been located in 3D n The training first fits the template to these landmarks by gradient descent on bone scaling S n and joint angles to minimize the 3D distance between the landmark positions and the corresponding template vertices. More specifically, the pose and bone scaling module 620 solves the following optimization problem:
[0132]
[0133] Among them, A selects a subset of landmarks from the template. When processing 4400 3D scans 612, the result of solving the optimization problem of Equation 4 above is 4400 θ n and bone scaling S n parameter set.
[0134] Specifically, for each 3D scan 612, the pose and bone scaling module 620 applies LBS to the template according to the skeleton to find the joint angles θ and bone scaling s parameters that minimize the distance to the landmarks of the given 3D scan 612. This results in an initial fit, which is further refined by registering the prediction to each scan 612.
[0135] In some cases, the bone scaling basis is used as a regularizer to re-estimate the pose θ n and bone scaling coefficients to match the template V T with each registration by solving the optimization problem of Equation 5:[[]]
[0136]
[0137] The output of Equation 5 is provided to the bone correction and average shape module 630 to calculate and determine the bone correction parameter P bc and average shape parameters To determine these parameters, the following optimization problem defined by Equation 6 is solved:[[]]
[0138]
[0139] As an example, the bone correction and average shape module 630 processes each 3D scan 612 and minimizes the difference in the LBS function between the objective true template and the template calculated using the corresponding pose θ n and bone scaling coefficients calculated for a specific 3D scan 612.
[0140] Once the bone correction blend shapes are used to improve the fit of the model to the registered shapes the residuals in the reconstruction are only due to identity-specific shape variability. The vertex displacement module 640 adjusts the parameters by deriving the function V D to account for identity-specific shape variability. Specifically, the residuals are modeled as vertex displacements and are estimated for each registration by setting to ensure that the residuals are defined in the T-pose coordinate system.
[0141] Given a new 3D scan, the pose and bone scaling module 620 adjusts the parameters θ n and the bone scaling factors And the bone correction and mean shape module 630 uses these adapted parameters to adjust the template. Then, the template is refined to represent the person depicted in the 3D scan by applying the landmarks of the 3D scan to the decoder D.
[0142] In some cases, the blending weights of the LBS formula are manually initialized. To improve this process, registrations for various identities and poses are processed. For each registration in the dataset, first the parameters are estimated, i.e., as well as the residuals The residuals are the errors on the T-pose coordinate system after taking into account the shape blend shapes. Then, the blending weights are optimized to minimize the following error:
[0143]
[0144] The mapping where is used to freely optimize W', while ensuring that the output weights W satisfy the LBS blending weight constraints:
[0145] ∑ j W ij = 1, and \(\mathbf{W}\) ij ≥ 0.
[0146] In some cases, the vertices on the torso are first fitted by optimizing the shape coefficients and joint angles of the torso bones (e.g., by blending the weights of the torso bones to those defined on the template). Then, for the second and third stages, the upper and lower limbs are added respectively. In the final stage, all vertices are used to fine-tune the fitting parameters.
[0147] Figure 9Shows an example of matching an avatar or character to a person in an image according to some embodiments. In some embodiments, a rigged character 910 is received. The rigged character 910 is received by selecting a character from a plurality of rigged characters. The rigged character 910 is applied to the skeleton provided by the 3D body model generation system 230. Given an image of a person (e.g., as shown in 930), the model is fitted to the person to estimate the bone transformation (scaling and rotation) of the person in the image. The bone transformation is applied to the rigged character to achieve accurate image-driven character animation. In the example, as shown in 930, the rigged character adapted according to the bone transformation of the person is displayed with the person. Alternatively, as shown in 920, the rigged character is presented in place of the person. In some cases, if the model determines that the bone transformation of the person depicted in the 2D image is greater than the average skeleton, the arms of the rigged character are extended. In some cases, if the model determines that the bone transformation of the person depicted in the 2D image is less than the average skeleton, the arms of the rigged character are shortened. The identity attributes of the character are also adapted based on the shape attributes determined by the model. This allows the system to transform any person into a selected avatar or character while preserving the pose and body shape of the person in the image.
[0148] Figure 10 Is a flowchart showing an example operation of the message client 104 in the execution process 1000 according to an example embodiment. The process 1000 can be implemented by computer-readable instructions executed by one or more processors, such that the operations of the process 1000 can be partially or fully executed by the functional components of the message server system 108; thus, the process 1000 is described below by way of example. However, in other embodiments, at least some of the operations in the process 1000 can be deployed on various other hardware configurations. The operations in the process 1000 can be executed in any order, executed in parallel, or can be completely skipped and omitted.
[0149] At operation 1001, the image processing server 122 receives a plurality of bone scaling factors, each bone scaling factor corresponding to a respective bone of the bone model. For example, the image processing server 122 receives the bone scaling parameter 512. In some cases, the bone scaling parameter 512 is received as the bone scaling factor c b is received.
[0150] At operation 1002, the image processing server 122 receives a plurality of joint angle factors that jointly define the pose of the bone model. For example, the image processing server 122 receives the pose parameter θ n .
[0151] At operation 1003, the image processing server 122 generates a bone model based on the received bone scaling factors and the received joint angle factors. For example, the image processing server 122 generates the scaled and posed skeleton 530.
[0152] At operation 1004, the image processing server 122 generates a base surface based on a plurality of bone scaling factors. For example, the image processing server 122 generates a template 550.
[0153] At operation 1005, the image processing server 122 generates an identity surface by deforming the base surface. For example, the image processing server 122 adjusts the template 550 based on bone parameters and applies identity information using the shape update parameter 562 to generate a template 570.
[0154] At operation 1006, the image processing server 122 generates a 3D body model by mapping the identity surface to the posed bone model. For example, the image processing server 122 uses the LBS 540 to adjust the template 570 based on the scaled and posed bone 530 to output a posed template 580.
[0155] Machine architecture
[0156] Figure 11is an illustrative representation of a machine 1100, in which instructions 1108 (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine 1100 to perform any one or more of the methods discussed herein. For example, the instructions 1108 can cause the machine 1100 to perform any one or more of the methods described herein. The instructions 1108 transform the general, non-programmed machine 1100 into a particular machine 1100 programmed to perform the described and illustrated functions in the described manner. The machine 1100 can operate as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1100 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1100 can include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular telephones, smart phones, mobile devices, wearable devices (e.g., smart watches), smart home devices (e.g., smart appliances), other smart devices, web appliances, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing the instructions 1108 specifying the actions to be taken by the machine 1100. Further, while only a single machine 1100 is shown, the term "machine" shall also be taken to include a collection of machines that individually or jointly execute the instructions 1108 to perform any one or more of the methods discussed herein. For example, the machine 1100 can include a client device 102 or any one of a number of server devices forming part of a message server system 108. In some examples, the machine 1100 can also include both a client and a server system, where certain operations of a particular method or algorithm are executed on the server side and certain operations of the particular method or algorithm are executed on the client side.
[0157] Machine 1100 may include a processor 1102, a memory 1104, and input / output components 1038, and the processor 1102, the memory 1104, and the input / output components 1138 may be configured to communicate with each other via a bus 1040. In an example, the processor 1102 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 1006 and a processor 1110 that may execute instructions 1108. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Figure 11 a multi-processor 1102 is shown, machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0158] The memory 1104 includes a main memory 1112, a static memory 1114, and a storage unit 1116, all of which may be accessed by the processor 1102 via a bus 1140. The main memory 1104, the static memory 1114, and the storage unit 1116 store instructions 1108 that embody any one or more of the methods or functions described herein. The instructions 1108 may also reside, completely or partially, within the main memory 1112, within the static memory 1114, within the machine-readable medium within the storage unit 1116, within at least one of the processors 1102 (e.g., within the cache memory of the processor), or within any suitable combination thereof during execution by the machine 1100.
[0159] The I / O components 1138 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, etc. The specific I / O components 1138 included in a particular machine will depend on the type of the machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is likely not to include such a touch input device. It will be understood that the I / O components 1138 may include Figure 11Many other components not shown. In various examples, I / O component 1138 may include user output component 1124 and user input component 1126. User output component 1124 may include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), auditory components (e.g., speakers), tactile components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. User input component 1126 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or another pointing instrument), tactile input components (e.g., physical buttons, a touch screen that provides the location and force of a touch or a touch gesture, or other tactile input components), audio input components (e.g., a microphone), etc.
[0160] In additional examples, I / O component 1138 may include biometric component 1128, motion component 1130, environmental component 1132, or location component 1134, as well as a wide array of other components. For example, biometric component 1128 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, face recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. Motion component 1130 may include: an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope).
[0161] Environmental component 1132 may include, for example, one or more camera devices (with still image / photo and video capabilities), a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect the ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor that detects the concentration of a hazardous gas for safety or measures pollutants in the atmosphere), or other components that can provide an indication, measurement, or signal corresponding to the surrounding physical environment.
[0162] Regarding the imaging device, the client device 102 may have an imaging device system that includes, for example, a front imaging device on the front surface of the client device 102 and a rear imaging device on the rear surface of the client device 102. The front imaging device may be used, for example, to capture still images and videos of the user of the client device 102 (e.g., "selfies"), and then the still images and videos may be enhanced with the above-described enhancement data (e.g., filters). For example, the rear imaging device may be used to capture still images and videos in a more traditional imaging device mode, and these images are similarly enhanced using the enhancement data. In addition to the front and rear imaging devices, the client device 102 may further include a 360° imaging device for capturing 360° photos and videos.
[0163] In addition, the imaging device system of the client device 102 may include a dual rear imaging device (e.g., a main imaging device and a depth-sensing imaging device), or even a triple, quadruple, or quintuple rear imaging device configuration on the front and rear sides of the client device 102. For example, these multi-imaging device systems may include a wide-angle imaging device, an ultra-wide-angle imaging device, a telephoto imaging device, a macro imaging device, and a depth sensor.
[0164] The location component 1134 may include a positioning sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure, from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), and the like.
[0165] A variety of techniques may be used to implement communication. The I / O component 1138 also includes a communication component 1136 that is operable to couple the machine 1100 to the network 1120 or the device 1122 via corresponding couplings or connections, respectively. For example, the communication component 1136 may include a network interface component or another suitable device to interface with the network 1120. In another example, the communication component 1136 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power), components, and other communication components that provide communication via other modalities. The device 1122 may be any one of other machines or various peripheral devices (e.g., a peripheral device coupled via USB).
[0166] In addition, the communication component 1136 can detect an identifier or include components operable to detect an identifier. For example, the communication component 1136 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be obtained via the communication component 1136, such as a location via Internet Protocol (IP) geolocation, a location via signal triangulation, a location via detecting an NFC beacon signal that can indicate a specific location, and the like.
[0167] Various memories (e.g., main memory 1112, static memory 1114, and the memory of the processor 1102) and the storage unit 1116 can store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. When executed by the processor 1102, these instructions (e.g., instruction 1108) cause various operations to implement the disclosed examples.
[0168] Instructions 1108 can be sent or received via a network interface device (e.g., the network interface component included in the communication component 1136), using a transmission medium and using any one of a plurality of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)), over the network 1120. Similarly, instructions 1108 can be sent or received using a transmission medium via a coupling (e.g., a peer-to-peer coupling) to the device 1122.
[0169] Software Architecture
[0170] Figure 12 is a block diagram 1200 showing a software architecture 1204 that can be installed on any one or more of the devices described herein. The software architecture 1204 is supported by hardware such as a machine 1202 including a processor 1220, a memory 1226, and an I / O component 1238. In this example, the software architecture 1204 can be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 1204 includes layers such as an operating system 1212, libraries 1210, frameworks 1208, and applications 1206. In operation, the application 1206 activates API calls 1250 through the software stack and receives messages 1252 in response to the API calls 1250.
[0171] The operating system 1212 manages hardware resources and provides common services. The operating system 1212 includes, for example: a kernel 1214, services 1216, and drivers 1222. The kernel 1214 acts as an abstraction layer between the hardware layer and other software layers. For example, the kernel 1214 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. The services 1216 can provide other common services for other software layers. The drivers 1222 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 1222 can include a display driver, a camera device driver, or a low-power driver, a flash driver, a serial communication driver (e.g., a USB driver), drivers, an audio driver, a power management driver, etc.
[0172] The library 1210 provides common low-level infrastructure used by the application 1206. The library 1210 can include a system library 1218 (e.g., a C standard library), and the system library 1218 provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, the library 1210 can include an API library 1224, such as a media library (e.g., a library for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group 4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer 3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., the OpenGL framework for presenting graphical content in two-dimensional (2D) and three-dimensional (3D) on a display), a database library (e.g., SQLite that provides various relational database functions), a web library (e.g., WebKit that provides web browsing functions), etc. The library 1210 can also include various other libraries 1228 to provide many other APIs to the application 1206.
[0173] The framework 1208 provides common high-level infrastructure used by the application 1206. For example, the framework 1208 provides various Graphical User Interface (GUI) functions, advanced resource management, and advanced location services. The framework 1208 can provide a wide range of other APIs that can be used by the application 1206, and some of these APIs can be specific to a particular operating system or platform.
[0174] In an example, the application 1206 can include a home application 1236, a contacts application 1230, a browser application 1232, a book reader application 1234, a location application 1242, a media application 1244, a messaging application 1246, a gaming application 1248, and various other applications such as an external application 1240. The application 1206 is a program that executes functions defined in a program. One or more of the applications 1206 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C language or assembly language). In a specific example, the external application 1240 (e.g., an application developed using an ANDROID TM or IOS TM software development kit (SDK) by an entity other than the vendor of a specific platform) can be mobile software that runs on a mobile operating system such as IOS TM 、ANDROID TM 、 Phone or another mobile operating system. In this example, the external application 1240 can activate an API call 1250 provided by the operating system 1212 to facilitate the functions described herein.
[0175] Glossary
[0176] "Carrier signal" means any intangible medium that can store, encode, or carry instructions executed by a machine and includes digital or analog communication signals or other intangible media to facilitate the communication of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.
[0177] "Client device" means any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device can be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smart phone, a tablet computer, a superbook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a gaming console, a set-top box, or any other communication device that a user can use to access a network.
[0178] "Communication network" means one or more portions of a network, which can be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a network, other types of networks, or a combination of two or more such networks. For example, a network or a part of a network can include a wireless network or a cellular network, and the coupling can be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless couplings. In this example, the coupling can implement any one of various types of data transmission technologies, such as Single-Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data Rate for GSM Evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, the 4th Generation Wireless (4G) network, Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long-Term Evolution (LTE) standard, other standards defined by various standards-setting organizations, other long-range protocols, or other data transmission technologies.
[0179] "Component" refers to a device, physical entity, or logic having boundaries defined by a function or subroutine call, a branch point, an API, or other technologies that provide partitioning or modularization for a particular processing or control function. A component can be combined with other components via its interfaces to perform machine processing. A component can be an encapsulated functional hardware unit designed to be used with other components and can be part of a program that typically performs a specific function among related functions.
[0180] A component can constitute a software component (e.g., code implemented on a machine-readable medium) or a hardware component. A "hardware component" is a tangible unit capable of performing certain operations and can be configured or arranged in a physical manner. In various example embodiments, one or more computer systems (e.g., stand-alone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or a part of an application) to be a hardware component for performing certain operations described herein.
[0181] The hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, the hardware components can include dedicated circuits or logic that are permanently configured to perform certain operations. The hardware components can be, for example, dedicated processors such as field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs). The hardware components can also include programmable logic or circuits that are temporarily configured by software to perform certain operations. For example, the hardware components can include software executed by a general purpose processor or other programmable processor. Once configured by such software, the hardware components become a particular machine (or a particular component of a machine) that is uniquely customized to perform the configured functions and is no longer a general purpose processor. It will be appreciated that the decision to implement the hardware components mechanically, in dedicated and permanently configured circuits, or in temporarily configured (e.g., software-configured) circuits can be made for cost and time considerations. Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, i.e., an entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in some manner or to perform certain operations described herein.
[0182] Considering an implementation where the hardware components are temporarily configured (e.g., programmed), it is not necessary to configure or instantiate each of the hardware components at any given time. For example, in a case where the hardware components include a general purpose processor that is configured by software as a dedicated processor, the general purpose processor can be configured as different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures one or more specific processors to, for example, constitute a particular hardware component at one time and different hardware components at different times.
[0183] The hardware components can provide information to and receive information from other hardware components. Thus, the described hardware components can be considered to be communicatively coupled. In a case where there are multiple hardware components present, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuits and buses). In an implementation where multiple hardware components are configured or instantiated at different times, such communication between the hardware components can be achieved, for example, by storing information in a memory structure accessible by the multiple hardware components and retrieving the information from the memory structure. For example, one hardware component can perform an operation and store the output of the operation in a memory device to which it is communicatively coupled. Then, other hardware components can access the memory device at a subsequent time to retrieve the stored output and process it. The hardware components can also initiate communication with an input device or an output device and can operate on resources (e.g., collect information).
[0184] The various operations of the example methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented components that operate to perform one or more of the operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be at least in part processor-implemented, where a particular one or more processors are examples of hardware. For example, at least some of the operations of the method can be performed by one or more processors 1102 or processor-implemented components. Additionally, one or more processors can also operate to support the performance of relevant operations in a "cloud computing" environment or operate as "software as a service" (SaaS). For example, at least some of the operations can be performed by a group of computers (as an example of a machine including processors), where the operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The performance of certain operations can be distributed among the processors, not residing only within a single machine but deployed across multiple machines. In some example embodiments, the processor or processor-implemented components can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the processor or processor-implemented components can be distributed across several geographical locations.
[0185] "Computer-readable storage medium" refers to both machine storage media and transmission media. Thus, the term includes both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium", "computer-readable medium", and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure.
[0186] "Ephemeral message" refers to a message that is accessible for a limited duration. An email message can be text, an image, a video, etc. The access time of an ephemeral message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is temporary.
[0187] "Machine storage medium" means a single or multiple storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Thus, the term should be considered to include, but not be limited to, solid-state memories as well as optical and magnetic media, including memories internal or external to a processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memories, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks (such as internal hard disks and removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium".
[0188] "Non-transitory computer-readable storage medium" means a tangible medium that can store, encode, or carry instructions executable by a machine.
[0189] "Signal medium" means any intangible medium that can store, encode, or carry instructions executable by a machine, and "signal medium" includes digital or analog communication signals or other intangible media to facilitate the communication of software or data. The term "signal medium" should be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal in which one or more of the signal characteristics are set or changed to encode information therein. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.
[0190] Without departing from the scope of this disclosure, changes and modifications may be made to the disclosed embodiments. These and other changes or modifications are intended to be included within the scope of the disclosure expressed in the appended claims.
Claims
1. A computer-implemented method for generating a 3D body model, the method comprising: Receive a plurality of bone scaling factors, each bone scaling factor corresponding to a respective bone of a bone model; Receive a plurality of joint angle factors that jointly define a pose of the bone model; Generate the bone model based on the received bone scaling factors and the received joint angle factors; Generate a base surface based on the plurality of bone scaling factors; Generate an identity surface by deforming the base surface; Generate the 3D body model by mapping the identity surface to the posed bone model; And Train a 3D model generator to generate the 3D body model by performing a training operation, the training operation including: Obtain a collection of 3D scans depicting various people in different poses and shapes, each 3D scan in the collection of 3D scans being associated with a set of anatomical landmarks that have been located in 3D; Generate a set of bone scalings and parameters by fitting an individual template to the set of anatomical landmarks by gradient descent on joint angles and bone scalings to minimize the 3D distance between the landmark positions and the corresponding template vertices; Refine the fit of the individual template by predicting the registration of each of the parameters to each 3D scan; and Minimize a function based on the difference between the ground truth template and the refined fitted individual template.
2. The method according to claim 1, wherein, The bone model includes a tree structure diagram, and generating the bone model includes: recursively generating a root bone element and a plurality of leaf bone elements.
3. The method according to claim 2, wherein, The bone model includes a rest position of each bone element and a scaling factor and a rotation factor for each bone element, wherein the rotation factor and the scaling factor are recursively applied to each bone element in sequence.
4. The method according to claim 3, wherein, The rest position of each bone element is represented by a translation vector and a template rotation matrix of the corresponding bone element.
5. The method according to claim 1, wherein, Each of the plurality of joint angle factors is restricted to a kinematically valid angular range.
6. The method according to claim 5, wherein, Constrain each of the plurality of joint angle factors by mapping a corresponding unconstrained variable to the kinematically valid angular range.
7. The method according to claim 1, wherein, The joint angle factors include 47 joint angle factors.
8. The method according to claim 1, wherein, Generating the base surface includes: generating an average surface and applying a correction to the average surface based on a deformation related to bone length.
9. The method according to claim 1, wherein, Generating the identity surface includes: deforming the base surface based on a plurality of linear identity parameters.
10. The method according to claim 1, wherein, Mapping the identity surface to the posed bone model includes: performing the mapping using a linear blend skinning process, wherein the 3D body model corresponds to a human or an animal.
11. The method according to claim 10, wherein, The linear blend skinning process includes mapping surface points on the identity surface based on the bone elements of the bone model by: drawing the surface points relative to the rest position of the bone element; and transporting the drawn points based on the posed position of the corresponding bone element.
12. The method according to claim 11, wherein, The linear blend skinning process includes: mapping surface points on the identity surface based on each bone element of the bone model and calculating a weighting factor for each corresponding mapping.
13. The method according to claim 1, wherein, The plurality of bone scaling factors and the plurality of joint angle factors are determined based on image processing of an input 2D image that includes a representation of at least one body.
14. The method according to claim 13, further comprising: Process the input 2D image using a deep convolutional neural network to detect the presence of at least one body, and estimate the plurality of bone scaling coefficients and the plurality of joint angle coefficients for the detected at least one body.
15. The method according to claim 14, further comprising: Generate an output 2D image based on the input 2D image and the generated 3D body model, wherein a representation of the body in the input 2D image is replaced with a corresponding 2D representation of the 3D body model.
16. The method according to claim 15, wherein,Generating the output 2D image includes: Generating the 3D body model based on the plurality of bone scaling coefficients and the plurality of joint angle coefficients determined from the input 2D image; Generating a 2D projection of the 3D body model based on the input 2D image; and Superimposing the 2D projection onto the representation of the body in the input 2D image.
17. A system for generating a 3D body model, comprising: A processor, configured to perform operations, the operations including: Receiving a plurality of bone scaling coefficients, each bone scaling coefficient corresponding to a respective bone of a bone model; Receiving a plurality of joint angle coefficients that jointly define a pose of the bone model; Generating the bone model based on the received bone scaling coefficients and the received joint angle coefficients; Generating a base surface based on the plurality of bone scaling coefficients; Generating an identity surface by deforming the base surface; Generating the 3D body model by mapping the identity surface onto the posed bone model; and Training a 3D model generator to generate the 3D body model by performing training operations, the training operations including: Obtaining a set of 3D scans depicting various people in different poses and shapes, each 3D scan in the set of 3D scans being associated with a set of anatomical landmarks that have been located in 3D; Generating a set of bone scalings and parameters by fitting an individual template to the set of anatomical landmarks by gradient descent on joint angles and bone scalings to minimize the 3D distance between landmark positions and corresponding template vertices; Refining the fit of the individual template by predicting the registration of each of the parameters to each 3D scan; and Minimizing a function based on the difference between the ground truth template and the refined fit individual template.
18. The system according to claim 17, wherein, The bone model includes a tree-structured graph, and generating the bone model includes: recursively generating a root bone element and a plurality of leaf bone elements.
19. The system according to claim 18, wherein the bone model includes a rest position of each bone element and a scaling factor and a rotation factor of each bone element, wherein, The rotation factor and the scaling factor are recursively applied to each bone element in sequence.
20. A non-transitory machine-readable storage medium, the non-transitory machine-readable storage medium includes instructions that, when executed by one or more processors of a machine, cause the machine to perform operations for generating a 3D body model, the operations including: Receiving a plurality of bone scaling coefficients, each bone scaling coefficient corresponding to a respective bone of a bone model; Receiving a plurality of joint angle coefficients that jointly define a pose of the bone model; Generating the bone model based on the received bone scaling coefficients and the received joint angle coefficients; Generating a base surface based on the plurality of bone scaling coefficients; Generating an identity surface by deforming the base surface; Generating the 3D body model by mapping the identity surface onto the posed bone model; And Training a 3D model generator to generate the 3D body model by performing training operations, the training operations including: Obtain a collection of 3D scans of various people depicting different poses and shapes, where each 3D scan in the collection of 3D scans is associated with a set of anatomical landmarks that have been located in 3D; Generate a set of bone scalings and parameters by fitting an individual template to the set of anatomical landmarks via gradient descent on joint angles and bone scalings to minimize the 3D distance between the landmark positions and the corresponding template vertices, Refine the fit of the individual template by predicting the registration of each of the parameters to each 3D scan; and Minimize a function based on the difference between the ground truth template and the refined fitted individual template.
Citation Information
Patent Citations
Systems and methods for avatar creation
US20140078144A1
Methods and systems for interpolation of disparate inputs
US20200005138A1
Cited By
3D object model reconstruction from 2D images
US12524963B2