Object reconstruction using media data
Patent Information
- Application Number
- CN202280054693.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-16
- Filing Date
- 2022-08-11
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-08-11
Smart Images

Figure CN117795559B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is a continuation-in-part of U.S. Patent Application No. 17 / 158,909, filed January 26, 2021, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to systems and techniques for constructing three-dimensional (3D) models based on a single video. Background Technology
[0004] Many devices and systems allow a scene to be captured by generating frames (also called images) and / or video data (including multiple images or frames). For example, a camera or a computing device that includes a camera (e.g., a mobile device including one or more cameras, such as a mobile phone or smartphone) can capture a sequence of frames of a scene. Frame and / or video data can be captured and processed by such devices and systems (e.g., mobile devices, IP cameras, etc.) and can be output for consumption (e.g., displayed on the device and / or other devices). In some cases, frame and / or video data can be captured by such devices and systems and output for processing and / or consumption by other devices.
[0005] Frames can be processed (e.g., using object detection, recognition, segmentation, etc.) to determine the objects present in the frame, which can be useful for many applications. For example, a model can be determined to represent the objects in the frame, and this model can be used to facilitate the efficient operation of various systems. Examples of such applications and systems, among many others, include: augmented reality (AR), robotics, automotive and aerospace, 3D scene understanding, object grasping, and object tracking. Summary of the Invention
[0006] In some examples, this document describes systems and techniques for generating one or more models. According to at least one example, a process for generating one or more models includes: generating a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object; generating a mask for the one or more frames, the mask including indications of one or more regions of the object; generating a 3D base model based on the 3D model of the first portion of the object and the mask, the 3D base model representing the first portion of the object and a second portion of the object; and generating a 3D model of the second portion of the object based on the mask and the 3D base model.
[0007] In another example, an apparatus for generating one or more models is provided, the apparatus including a memory (e.g., configured to store data such as virtual content data, one or more images, etc.) and one or more processors (e.g., implemented in a circuit) coupled to the memory. The one or more processors are configured and capable of: generating a three-dimensional (3D) model of a first portion of the object based on one or more frames depicting the object; generating a mask for the one or more frames, the mask including indications of one or more regions of the object; generating a 3D base model based on the 3D model of the first portion of the object and the mask, the 3D base model representing the first portion of the object and a second portion of the object; and generating a 3D model of the second portion of the object based on the mask and the 3D base model.
[0008] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: generate a three-dimensional (3D) model of a first portion of the object based on one or more frames depicting the object; generate a mask for the one or more frames, the mask including indications of one or more regions of the object; generate a 3D base model based on the 3D model of the first portion of the object and the mask, the 3D base model representing the first portion of the object and a second portion of the object; and generate a 3D model of the second portion of the object based on the mask and the 3D base model.
[0009] In another example, an apparatus for generating one or more models is provided. The apparatus includes: a unit for generating a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object; a unit for generating a mask for the one or more frames, the mask including indications of one or more regions of the object; a unit for generating a 3D base model based on the 3D model of the first portion of the object and the mask, the 3D base model representing the first portion of the object and a second portion of the object; and a unit for generating a 3D model of the second portion of the object based on the mask and the 3D base model.
[0010] In some aspects, the 3D model of the second part corresponds to an item that is part of the object. For example, in some aspects, the object is a person, the first part of the object corresponds to the person's head, and the second part of the object corresponds to the hair on the person's head.
[0011] In some aspects, the 3D model of the second part corresponds to an item that is at least one of the following: separable from the object and movable relative to the object. For example, in some aspects, the object is a person, the first part of the object corresponds to a body area of the person, and the second part of the object corresponds to accessories or clothing worn by the person.
[0012] In some aspects, the 3D model of the second part of the object is adjacent to at least a portion of the 3D model of the first part of the object. In some aspects, the 3D model of the second part of the object does not visually collide with the 3D model of the first part of the object.
[0013] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: selecting one or more frames from a frame sequence as keyframes, wherein each keyframe depicts an object from a different angle.
[0014] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: determining that a first keyframe does not meet a quality threshold; outputting feedback to facilitate the positioning of the object corresponding to the first keyframe; capturing at least one frame based on the feedback; and inserting a frame from the at least one frame into the keyframe.
[0015] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: generating a first bitmap from a 3D model of a first portion of the object for a first angle selected along an axis; generating a first metric at least in part by comparing the first bitmap with a reference frame in the frame sequence; and selecting a first keyframe based on the result of the comparison.
[0016] In some aspects, comparing the first bitmap with the reference frame includes performing an intersection-over-union (IoU) comparison between the first bitmap and the bitmaps of the reference frame.
[0017] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: generating a second metric by at least partly comparing the reference frame with a bitmap of a second frame of the frame sequence; and selecting the second frame as the first keyframe based on the second metric.
[0018] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: segmenting each of the one or more frames into one or more regions; and generating a mask for each of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.
[0019] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: projecting each vertex of an initial 3D model onto a mask associated with the frame based on pose information associated with a frame in one or more of the frames; determining whether each vertex of the first portion of the 3D model is located within a first region of the mask associated with the frame; and extracting the 3D base model based on the vertices of the first portion of the 3D model being within the first region of the mask associated with the frame.
[0020] In some aspects, the object is a person, and the first region corresponds to the person's facial region and the person's hair region.
[0021] In some respects, the object is a person, and the first area corresponds to the area of the person's body and the area of the person's clothing.
[0022] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: removing one or more vertices from the 3D base model based on the probability of each vertex in a region of one or more regions from the one or more frames.
[0023] In some aspects, the mask for each of the one or more frames includes a first mask identifying a first region and a second mask identifying a second region.
[0024] In some respects, the first region is the facial region, and the second region is the hair region.
[0025] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: initializing the value of each vertex of the 3D base model to an initial value, wherein the initial value indicates that the corresponding vertex is located within the first region; projecting a first vertex of the 3D base model onto a keyframe in one or more frames; determining whether the vertex of the 3D base model is projected onto a first mask or a second mask of the first keyframe; and adjusting the value of each vertex based on whether the corresponding vertex is projected onto the first mask or the second mask. In some aspects, when the first vertex is projected onto the second region, the value of the first vertex is increased, and when the first vertex is projected onto the first region, the value of the first vertex is decreased. In some aspects, the process, apparatus, and non-transitory computer-readable medium include: determining an average probability based on the value of each vertex, wherein the probability that a vertex corresponds to a 3D model of a first portion of the object is based on a comparison of the vertex's value with the average probability.
[0026] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: generating animation in an application using a 3D model of the first portion and a 3D model of the second portion, wherein the object includes a person, the 3D model of the first portion corresponding to a person's head, and the 3D model of the second portion corresponding to a person's hair.
[0027] In some aspects, the application includes the ability to send and receive at least one of audio and text.
[0028] In some respects, the 3D model of the first part and the 3D model of the second part depict the user of the application.
[0029] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: receiving input corresponding to the selection of at least one graphical control for modifying a 3D model of the second part; and modifying the 3D model of the second part based on the received input.
[0030] In some aspects, the one or more frames are associated with rotation of the object along a first axis. In some aspects, the one or more frames are associated with rotation of the object along a second axis. In some aspects, the first axis corresponds to the yaw axis, and the second axis corresponds to the pitch axis.
[0031] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: performing pose refinement of pose information associated with frames in the one or more frames.
[0032] In some aspects, in order to perform pose refinement of pose information associated with the frame, the process, apparatus, and non-transitory computer-readable medium include minimizing the difference between one or more landmarks in a distorted reference frame model and one or more landmarks of the frame.
[0033] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: determining that the coordinate values of fewer than a threshold number of vertices of a 3D model of a second portion of the object are less than predetermined coordinate values; and, based on determining that the coordinate values of fewer than a threshold number of vertices of the 3D model are less than the predetermined coordinate values, removing one or more vertices of the 3D model that are less than the predetermined coordinate values.
[0034] According to at least one other example, a process for generating one or more models is provided. The process includes: generating a three-dimensional (3D) model of a person's head based on one or more frames depicting a person; generating a mask for the one or more frames, the mask including indications of one or more regions of the person; generating a 3D base model based on the 3D model of a first portion of the person and the mask, the 3D base model representing the person's head and the person's hair; and generating a 3D model of the person's hair based on the mask and the 3D base model.
[0035] In another example, an apparatus for generating one or more models is provided, the apparatus including a memory (e.g., configured to store data such as virtual content data, one or more images, etc.) and one or more processors (e.g., implemented in a circuit) coupled to the memory. The one or more processors are configured and capable of: generating a three-dimensional (3D) model of a person's head based on one or more frames depicting a person; generating a mask for the one or more frames, the mask including indications of one or more regions of the person; generating a 3D base model based on a 3D model of a first portion of the person and the mask, the 3D base model representing the person's head and the person's hair; and generating a 3D model of the person's hair based on the mask and the 3D base model.
[0036] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: generate a three-dimensional (3D) model of a person's head based on one or more frames depicting a person; generate a mask for the one or more frames, the mask including indications of one or more regions of the person; generate a 3D base model based on a 3D model of a first portion of the person and the mask, the 3D base model representing the person's head and the person's hair; and generate a 3D model of the person's hair based on the mask and the 3D base model.
[0037] In another example, an apparatus for generating one or more models is provided. The apparatus includes: a unit for generating a three-dimensional (3D) model of a person's head based on one or more frames depicting a person; a unit for generating a mask for the one or more frames, the mask including indications of one or more regions of the person; a unit for generating a 3D base model based on a 3D model of a first portion of the person and the mask, the 3D base model representing the person's head and the person's hair; and a unit for generating a 3D model of the person's hair based on the mask and the 3D base model.
[0038] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: selecting one or more frames from a frame sequence as keyframes, wherein each keyframe depicts the person from a different angle.
[0039] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: determining that a first keyframe does not meet a quality threshold; outputting feedback to facilitate the localization of the person corresponding to the first keyframe, capturing at least one frame based on the feedback; and inserting a frame from the at least one frame into the keyframe.
[0040] In some aspects, generating a mask for the one or more frames includes: segmenting each of the one or more frames into one or more regions; and generating a mask for each of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.
[0041] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: projecting each vertex of an initial 3D model onto a mask associated with the frame based on pose information associated with a frame in one or more of the frames; determining whether each vertex of the 3D model of the head is located within a head region of the mask associated with the frame; and extracting the 3D base model based on the vertices of the 3D model of the head being within the head region of the mask associated with the frame.
[0042] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: removing one or more vertices from the 3D base model based on the probability that each of the one or more vertices is outside the head region.
[0043] In some respects, the 3D model of the person's hair does not visually collide with the 3D model of the person's head.
[0044] In some aspects, generating a 3D model of the person's hair includes: initializing the value of each vertex of the 3D base model to an initial value, wherein the initial value indicates that the corresponding vertex is located within a hair region; projecting a first vertex of the 3D base model onto a keyframe in one or more frames; determining whether the vertex of the 3D base model is projected onto a first mask or a second mask of the first keyframe, wherein the first mask corresponds to a face region and the second mask corresponds to the hair region; and adjusting the value of each vertex based on whether the corresponding vertex is projected onto the first mask or the second mask.
[0045] In some aspects, when the first vertex is projected onto the second mask (corresponding to the hair region), the value of the first vertex is increased. In some aspects, when the first vertex is projected onto the first mask (corresponding to the face region), the value of the first vertex is decreased.
[0046] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: determining an average probability based on the value of each vertex, wherein the probability that a vertex corresponds to a 3D model of the hair is based on a comparison of the value of the vertex with the average probability.
[0047] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: performing pose refinement of pose information associated with frames in the one or more frames.
[0048] In some aspects, in order to perform pose refinement of pose information associated with the frame, the process, apparatus, and non-transitory computer-readable medium include minimizing the difference between one or more landmarks in a distorted reference frame model and one or more landmarks of the frame.
[0049] In some aspects, the process, apparatus, and non-transitory computer-readable medium include: determining that the coordinate values of fewer than a threshold number of vertices in a 3D model of the human hair are less than predetermined coordinate values; and, based on determining that the coordinate values of fewer than a threshold number of vertices in the 3D model are less than the predetermined coordinate values, removing one or more vertices of the 3D model that are less than the predetermined coordinate values.
[0050] In some aspects, one or more of the devices described above are one or more of the following: vehicles (e.g., computing devices of vehicles), mobile devices (e.g., mobile phones or so-called "smartphones" or other mobile devices), wearable devices, extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), personal computers, laptop computers, server computers, or other devices. In some aspects, a device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device may include one or more sensors that can be used to determine the position and / or orientation of the device, the state of the device, and / or for other purposes.
[0051] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to define the scope of the claimed subject matter. The subject matter should be understood by referring to the appropriate portions of the entire specification, any or all of the drawings, and each claim.
[0052] The foregoing, along with other features and embodiments, will become more apparent from the following description, claims, and drawings. Attached Figure Description
[0053] The illustrative embodiments of this application are described in detail below with reference to the following figures:
[0054] Figure 1 An example of a 3D model of the first part of an object colliding with the second part of the object is shown;
[0055] Figure 2 This is a schematic diagram illustrating examples of systems that can perform detailed three-dimensional (3D) face reconstruction from a single red-green-blue (RGB) image, based on some examples;
[0056] Figure 3 Three example keyframes are shown, selected from a frame sequence by a keyframe selector, based on some examples.
[0057] Figure 4A This is a schematic diagram illustrating an example of a process for performing object reconstruction based on 3D deformable model (3DMM) technology, according to some examples;
[0058] Figure 4B This is a schematic diagram illustrating an example of a 3DMM based on some sample objects;
[0059] Figure 5 This is a flowchart illustrating an example of a process for selecting keyframes from a sequence of frames of an object, based on some examples;
[0060] Figure 6A The following are examples. Figure 3 The results of object analysis of a portion of the keyframe;
[0061] Figure 6B It shows some examples that can be obtained from Figure 3 Different object parsing masks are generated from the results of object analysis of keyframes;
[0062] Figure 6C An example of a 3DMM generated from a reference frame in a front view is shown;
[0063] Figure 6D An example of the current keyframe is shown, based on some examples;
[0064] Figure 7 This is a flowchart illustrating an example of a process for generating a 3D base model corresponding to a first region and a second region of an object, based on some examples;
[0065] Figure 8A This is a schematic diagram illustrating the creation of an object parsing mask, based on some examples, that will be used to create a 3D base model;
[0066] Figure 8B This is a schematic diagram illustrating an example of a 3D scene, including a 3D model that can be projected into a two-dimensional (2D) coordinate system for comparison with an object resolution mask, based on some examples.
[0067] Figure 8C This is a schematic diagram illustrating an example of a 3D base model corresponding to a first region and a second region of an object, based on some examples;
[0068] Figure 8D This shows, based on some examples, the post-processing process. Figure 8C A schematic diagram of an example 3D basic model;
[0069] Figure 9 This is a flowchart illustrating an example of a process for generating a 3D model of a second object using a 3D base model, based on some examples;
[0070] Figure 10 This is a schematic diagram illustrating an example model extractor that can generate a second object from a 3D base model, based on some examples.
[0071] Figure 11 The 3D model of the second region, generated using a 3D base model, is shown based on some examples.
[0072] Figure 12 This is a flowchart illustrating an example of a process for generating one or more models based on some examples;
[0073] Figure 13 This is a flowchart illustrating another example of a process for generating one or more models, based on some examples; and
[0074] Figure 14 This is a schematic diagram illustrating an example of a system used to implement certain aspects of this technology. Detailed Implementation
[0075] The following provides certain aspects and embodiments of this disclosure. Some of these aspects and embodiments can be applied independently, and some can be applied in combination, as will be apparent to those skilled in the art. In the following description, specific details are set forth for ease of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The drawings and specification are not intended to be limiting.
[0076] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with enabling descriptions for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0077] The generation of three-dimensional (3D) models of physical objects can be useful for many systems and applications, such as extended reality (XR) (e.g., including augmented reality (AR), virtual reality (VR), mixed reality (MR), etc.), robotics, automotive, aviation, 3D scene understanding, object grasping, object tracking, and many other systems and applications. For example, in an AR environment, a user can view an image (also known as a frame) that integrates artificial or virtual graphics with the user's natural surroundings. AR applications allow processing of real-world images to add virtual objects to the image and align or register virtual objects to the image in multiple dimensions. For example, real-world objects that exist in the real world can be represented using models that are similar to or exactly match real-world objects. In one example, as a user continues to view his or her natural surroundings in an AR environment, a model representing a real aircraft located on a runway can be presented in the view of an AR device (e.g., AR glasses, AR head-mounted displays (HMDs), or other devices). The viewer can be able to manipulate the model while viewing the real-world scene. In another example, a model with different colors or different physical properties in an AR environment can be used to identify and render an actual object sitting on a table. In some cases, computer-generated copies of artificial virtual objects that do not exist in reality or actual objects or structures in the user's natural surroundings can also be added to the AR environment.
[0078] The increasing number of applications utilizing facial data (e.g., for XR systems, 3D graphics, security, etc.) has led to a significant demand for systems capable of generating detailed 3D facial models (and 3D models of other objects) efficiently and with high quality. There is also substantial demand for generating 3D models of other types of objects, such as 3D models of vehicles (e.g., for autonomous driving systems), 3D models of room layouts (e.g., for XR applications, for navigation by devices, robots, etc.), and so on.
[0079] Generating detailed 3D models of objects (e.g., 3D facial models) often requires expensive equipment and multiple cameras in environments with controlled lighting, hindering large-scale data collection and potentially necessitating multiple 3D models to reconstruct different aspects. A face is an example of an object for which 3D models can be generated, requiring different models for different aspects. For instance, a person's hair might need to be modeled separately because hair movement is non-rigid and can be independent of facial movement. In another example, a person's clothing (e.g., coat, jacket, etc.) might need to be modeled differently because the material and geometric properties of clothing may change in ways different from those of a living object. Other aspects of living objects can also be modeled separately because the movement of one or more features of an object can often be independent of other features (e.g., an elephant's trunk, a dog's tail, a lion's mane). In some examples, inanimate objects (e.g., trees, vehicles, buildings, environments, or scenes, etc.) might require a model that typically depicts static aspects (e.g., the trunk and main branches) as well as another model that depicts dynamic aspects (e.g., branches and leaves, different parts of a vehicle, different parts of a building, different parts of an environment or scene, etc.).
[0080] Performing 3D object reconstruction (e.g., to generate 3D models of objects such as facial and hair models) from one or more images (e.g., a single video) can be challenging. For example, 3D object reconstruction based on reconstructions involving geometry, albedo textures, and lighting estimations can be difficult. 3D object reconstruction may not reflect other aspects associated with the object. In an illustrative example, the shape of a head typically does not change during head movement (e.g., rigid movement). However, the shape of the hair associated with the head may change during head movement (e.g., non-rigid movement). Because hair and head move in different ways, a person's hair can be modeled separately from the person's head.
[0081] There are different types of 3D models for hair, such as partial derivative equation (PDE) models, data-driven and deformable models, 3D models generated using deep learning models (e.g., one or more neural networks, such as deep neural networks, generative adversarial networks, convolutional neural networks, etc.), and so on. However, problems can arise from using these types of 3D models.
[0082] For example, a PDE model uses at least one frame to estimate the depth (e.g., volume) of a 3D hair model. A frame including a half-body portrait of a person (e.g., the head region and hair region) is used as initial conditions, and the hair depth is estimated based on the half-body portrait model. Facial resolution algorithms can be implemented to segment different regions into depths corresponding to volumes. Depths are divided into different segments (e.g., regions), and smoothing of depths between different segments must be enforced in the PDE model. When a PDE model is created using a single frame, the hair and face regions may be over-segmented or under-segmented, resulting in inaccurate boundaries between the hair and face regions. In cases where a PDE model is determined based on multiple frames, lighting conditions and rapid motion can significantly affect and limit the visual fidelity of the PDE model. Specifically, inter-frame correspondences (e.g., references) need to be established based on photometric information (e.g., RGB). However, photometric consistency across multiple frames depends on ideal lighting conditions that cannot be guaranteed. Furthermore, hair will move non-rigidly when the object is moving, and identifying correspondences between non-rigid content in different frames is challenging.
[0083] In driven and deformable models, artists create different 3D representations of various hairstyles, and people (e.g., users of the application) select a suitable 3D hairstyle for the head within the application. In this type of hair model, the quality of the 3D hair model is primarily based on the similarity between the source shape (e.g., the created 3D hairstyle) and the target shape (e.g., the hair of a person in a frame). For example, if a person is building a 3D model of that person's head based on a 2D reference frame, that person must search for 3D hairstyle models similar to those in the 2D reference frame. Searching for hairstyles can be challenging because not every hairstyle can be represented, and identifying the ideal hairstyle based on description and classification is difficult. Furthermore, any identified 3D hairstyle will not precisely match the hairstyle in the 2D reference frame. A person can deform the 3D hairstyle within the application to match the 3D frame, but deformation can introduce unexpected fitting problems. As a result, the quality of a 3D hair model is acceptable when there is a high correlation between the 3D hairstyle and the hairstyle in the 2D reference frame.
[0084] A deep learning-based method extracts the 2D orientation of hair from frames and the confidence level of the orientation from the frames. A half-body model is then fitted to (multiple) frames to create a fitted depth map. The orientation, confidence level, and depth map are used as input to a machine learning network to create a 3D hair model, which typically has outputs for an occupancy field and an orientation field. The occupancy field determines whether a voxel in 3D space belongs to hair, and if it does, the orientation field determines whether to determine the 3D orientation (e.g., direction) of the current voxel. The 3D hair model is constructed by combining the occupancy field and the orientation field. The occupancy field determines where hair grows, and the orientation field determines the growth direction (e.g., grouping of hair voxels). However, deep learning 3D hair models will be inaccurate because the deep learning process is not based on a real hair dataset. Furthermore, there is no process for detecting collisions between the deep learning 3D hair model and the 3D head model, which can produce unwanted visual artifacts (e.g., collisions, such as skull overlap with hair).
[0085] Furthermore, in current techniques for modeling different parts of an object (such as a model of a human head and different models for human hair), the resulting 3D models (e.g., head model and hair model) are generated separately and are misaligned in the 3D coordinate system. In conventional methods, a user interface can be implemented to locate and orient the 3D head model and 3D hair model to create a complete 3D head model. The 3D hair model controls or influences the cranial region of the 3D head model. In some cases, the cranial region of the 3D head model may differ from that of a human head. When the 3D head model and 3D hair model are rasterized into 2D frames for display (e.g., converted from 3D vector and texture representations to 2D pixel bitmaps), the 3D head model and 3D hair model may collide. Collisions of 3D geometry occur when two solid objects occupy the same space, which is physically impossible in the real world. As a result, collisions in 3D space produce undesirable visual artifacts.
[0086] Figure 1 A 3D model 100 is shown, comprising a head model 105 and a hair model 110 colliding in a 3D coordinate system. In the illustrative example, existing techniques are used to position the hair model 110 and the head model 105. Therefore, the fitting (e.g., positioning and alignment) of the hair model 110 to the head model 105 does not control how the shape of the skull region of the head model 105 affects the hair model 110. In this case, when the head model 105 and the hair model 110 are converted from a 3D coordinate system to a 2D frame (e.g., rasterization, rendering, etc.), the head model 105 visually appears to overlap with and be outside the hair model 110.
[0087] To address this issue, application designers (e.g., applications displaying human avatars) can create functionality to align 3D hair models with 3D head models. Furthermore, the avatar may move within the application, causing collisions between the 3D hair model and the avatar, resulting in undesirable visual artifacts. In this case, designers will also need to create a collision detection process to detect collisions with the 3D models. Designers will also need to apply a collision resolution process to determine how to display the 3D models without producing undesirable visual artifacts.
[0088] This document describes systems, apparatuses, processes (or methods), and computer-readable media (collectively, the “Systems and Techniques”) for generating 3D models of specific portions of an object from video (e.g., RGB video comprising a sequence of frames) or from multiple images. In some examples, as described in more detail below, the Systems and Techniques can generate a 3D model of a first portion of an object (e.g., a 3D model of a human head) based on one or more frames depicting the object. The Systems and Techniques can generate a mask for one or more frames. In some examples, the mask may include indications of one or more regions of the object. The Systems and Techniques can also generate a 3D base model based on the 3D model of the first portion of the object and the mask. In some cases, the 3D base model depicts both the first portion and a second portion of the object. The Systems and Techniques can generate a 3D model of a second portion of the object (e.g., a 3D model of human hair) based on the mask and the 3D base model.
[0089] The system and techniques described can also automatically align the 3D model of the first part and the 3D model of the second part. In some examples, because the 3D model of the second part of the object is created based on removing content associated with the first part of the object, the 3D model of the second part is adjacent to (e.g., aligned with) at least a portion of the 3D model of the first part. In an illustrative example, the system and techniques described can be applied to generate 3D models of a human face and human hair. Because the 3D model of the hair is created independently of the 3D model of the head, the 3D model of the hair can be applied to different hairstyles (e.g., to recreate different hairstyles). Furthermore, because the 3D model of the hair can be extracted with fewer requirements or constraints than conventional processes, many different hairstyles can be modeled. In some examples, the different hairstyles can be used in deep learning processes to improve, for example, generative adversarial networks (GANs) used to create hair-related content using machine learning techniques.
[0090] Among other benefits, the systems and techniques described herein address current problems in 3D model reconstruction techniques (e.g., shortcomings associated with base items, complexity, etc.). For example, the systems and techniques can use a single video to create a 3D model of a first part of an object. The systems and techniques can also implement novel algorithms to select frames from a single video and, guided by an object parsing mask (e.g., an indicator bitmap corresponding to features of the object) (such as a facial parsing mask), create a 3D base model, and remove several parts of the 3D base model to create a second 3D model. Guided by the object parsing mask, unwanted content in the second 3D model can be identified and removed.
[0091] For illustrative purposes, the head and hair are used herein as illustrative examples of parts of an object. However, those skilled in the art will understand that the systems and techniques described herein can be performed on any type of object captured in one or more frames and on any part of such an object. In one illustrative example, similar systems and techniques can be applied to generate a first model of a human body and a second model of an item worn by the person (e.g., a coat, hat, etc.). In another illustrative example, similar systems and techniques can be applied to generate a first model of a vehicle or a scene observed by one or more sensors of a vehicle, and a second model of a second part of a scene observed by one or more sensors of a vehicle. In some examples, the systems and techniques can be applied to the deformation of non-rigid objects such as clothing, accessories, fabrics, etc.
[0092] Figure 2 This is a schematic diagram illustrating an example of a system 200 that can generate one or more models using at least one frame (e.g., a frame sequence). Figure 2 As shown, system 200 includes: a 3D object modeler 205, a keyframe selector 210, an object analyzer 215, a 3D base model generator 220, and a model extractor 225. In some aspects, system 200 may include a pose refinement engine 230 (in... Figure 2 (The dashed outline is used to indicate that the pose refinement engine is optional). In some aspects, system 200 may include an object cleaning engine. Figure 2 (Not shown in the image).
[0093] The 3D object modeler 205 can acquire at least one frame (not shown). In some cases, the at least one frame may include a frame sequence. The frame sequence may be video, a set of consecutively captured images, or other frame sequences. In an illustrative example, each frame in the frame sequence may include red (R), green (G), and blue (B) components per pixel (referred to as RGB video including RGB frames). Other examples of frames include frames having luminance, chroma-blue, chroma-red (YUV, YCbCr, or Y'Cb) components per pixel and / or any other suitable type of image. The frame sequence may be captured by one or more cameras of system 200 or another system, obtained from a storage device, received from another device (e.g., a camera or a device including a camera), or obtained from another source.
[0094] In some examples, the frame sequence depicts the movement of an object to be converted into one or more 3D models by System 200. In one illustrative example, the frame sequence begins with a front view of a person with a neutral head position, where the head is not rotated relative to the neck region (e.g., 0°). Figure 3 An example of frame 310 is shown as a frontal view of a person with a neutral head position. The frame sequence can then show a perspective view of the left side of the face of a person with their head rotated to the right by an angle of +60° on a first axis (e.g., the yaw axis). Figure 3 Frame 320 also shows a person with their head rotated +60° to their right. The person can then rotate their head -60° to the left on the first axis to show a perspective view of the right side of their face. For example, Figure 3 Frame 330 further shows a perspective view of a person's head rotated -60°. In some examples, the frame sequence may show the person also rotating their head on a second axis (e.g., the pitch axis).
[0095] The 3D object modeler 205 can process one or more frames from a frame sequence to generate a 3D model of the first part of an object. In an illustrative example, the 3D model of the first part of the object could be a human head. In some examples, the 3D object modeler 205 can generate a 3D deformable model (3DMM) of the human head for each frame in the frame sequence. The 3DMM model generated using 3DMM fitting is a statistical model representing the 3D geometry and texture of the object. For example, a 3DMM could be constructed from a model with a shape X... shape , Expression X expression and texture X albedo The coefficients are represented by a linear combination of the basic terms, for example as follows:
[0096] Vertices 3D_coordinate =X shape Basis shape +Xexpression Basis expression (1)
[0097] Vertices color =Color mean_albedo +X albedo Basis albedo (2)
[0098] Equation (1) is used to determine the position of each vertex of the 3DMM model, and Equation (2) is used to determine the color of each vertex of the 3DMM model.
[0099] Figure 4A This is a schematic diagram illustrating an example of a process 400 for performing object reconstruction based on 3DMM technology. At operation 402, process 400 includes: obtaining input, including an image (e.g., an RGB image) and landmarks (e.g., facial landmarks or other landmarks that uniquely identify the object). At operation 404, process 400 performs a 3DMM fitting technique to generate a 3DMM model. 3DMM fitting includes: solving for the shape (e.g., X-shape) of the 3DMM model of the object (e.g., face). shape ), facial expressions (e.g., X) expression ) and albedo (e.g., X albedo The fitting may also include solving for the camera matrix and spherical harmonic illumination coefficients.
[0100] At operation 406, process 400 includes performing a Laplacian deformation on the 3DMM model. For example, a Laplacian deformation can be applied to the vertices of the 3DMM model to improve landmark fitting. In some cases, another type of deformation can be performed to improve landmark fitting. At operation 408, process 400 includes solving for albedo. For example, process 400 can fine-tune the albedo coefficients to separate colors that do not belong to the spherical harmonic lighting model. At operation 410, process 400 solves for depth. For example, process 400 can determine the depth displacement per pixel based on the shape-from-shading formula or other similar functions. The shape-from-shading formula defines the color of each point in the 3DMM model as the product of the albedo color and the light coefficient. For example, the color seen in an image is represented as the albedo color multiplied by the light coefficient. The light coefficient of a given point is based on the surface normal of that point. At operation 412, process 400 includes outputting a depth map and / or a 3D model (e.g., outputting a 3DMM).
[0101] In some cases, 3DMM models can be used to describe object spaces with principal component analysis (PCA) (e.g., 3D face space). Below is an example equation (3) that can be used to describe the shape of a 3D object (e.g., 3D head shape):
[0102]
[0103] Using a head as an example of a 3D object, S is the 3D head shape. It is the average facial shape, A id It is a feature vector (or principal component) trained on a 3D facial scan with a neutral expression, α id It is the shape factor, A exp It is a feature vector trained on the offset between facial expression scanning and neutral scanning, and α exp This is the facial expression coefficient. The shape of the 3DMM head can be projected onto the image plane using projection techniques (such as weak perspective projection). The following example equations (4) and (5) can be used to calculate the aligned facial shape:
[0104] I=mPR(α,β,γ)S+t (4)
[0105]
[0106] Where I is the aligned facial shape, S is the 3D facial model, R(α,β,γ) is a 3×3 rotation matrix with rotation angles α,β,γ, m is the scaling parameter, t is the translation vector, and P is a weak perspective transformation.
[0107] Each 3DMM can be fitted to an object in each frame of a frame sequence. Because the 3DMM is fitted to each frame in the frame sequence, the accuracy of the 3DMM model may vary and it may be misaligned. Furthermore, the object may change between frames. For example, a person's head may change between frames due to factors such as jitter. Therefore, the 3DMM model may vary between frames.
[0108] Figure 4B This is a schematic diagram illustrating an example of a 3DMM 411 object. Figure 4B In the example shown, the object corresponds to a person and the head area is modeled. 3DMM 411 does not include non-rigid aspects of a person, such as the hair or clothing areas. 3DMM 411 can deform to depict various facial movements of the object (e.g., the nose area, eye area, mouth area, etc.). As an example, across multiple frames, 3DMM 411 can be modified to allow showing a person speaking or facial expressions, such as a person smiling. In other examples, the object can be any physical object, such as a person, attachment, vehicle, building, animal, plant, clothing, and / or other objects.
[0109] As described above, a 3DMM can be generated for each frame in a frame sequence. Different 3DMMs correspond to the movement of objects in the frame sequence. In some examples, the 3DMM model may include positional information related to the position of a first region of the object. In an example where the object is a human head, the positional information may include pose information related to the head's posture. For example, the pose information may indicate the angular rotation of the head relative to its neutral position. The rotation may be along a first axis (e.g., the yaw axis) and / or a second axis (e.g., the pitch axis). In some examples, the pose information may include any physical properties that can be used to determine motion, such as displacement (e.g., movement), velocity, acceleration, rotation, etc.
[0110] Keyframe selector 210 can select one or more frames (referred to herein as keyframes) from a frame sequence. In an illustrative example, keyframe selector 210 can select a reference frame from the frame sequence. Keyframe selector 210 can compare the similarity of the 3DMM model of the reference frame with that of other 3DMM models in other frames to determine whether the 3DMM models are aligned spatially and temporally. For example, keyframe selector 210 can analyze frames based on the alignment of the object when the object moves in one direction during a first time period (or when the camera moves relative to the object while the object is stationary), and can analyze frames when the object moves in the opposite direction during a second time period (or the camera is moving) (e.g., returning to a neutral position). In some cases, keyframe selector 210 may analyze fewer frames when the object moves in the opposite direction (or the camera is moving) during a second time period. A frame may be selected as a keyframe if a comparison of another 3DMM of a frame is appropriately associated with the reference frame. Each keyframe can be selected to depict an object rotating at different angles, and to represent a specific range of motion of the object (e.g., a rotation range such as 10° to 15°, 15° to 20°, etc.). See below for an example. Figure 5 and Figure 6A Describe further details about the process used to select keyframes.
[0111] As described above, in some aspects, system 200 may include a pose refinement engine 230. For example, head shape variations and / or pose errors can cause pose alignment problems, such as misalignment of head shapes between different views or images (e.g., the same area from different views corresponds to different areas in 3D). In some cases, pose errors may have a greater impact because they may affect the alignment of the entire shape (e.g., head shape). The pose refinement engine 230 can perform global pose refinement (e.g., using a global pose refinement algorithm) to mitigate misalignment. When using the pose refinement engine 230, it can process one or more keyframes output from keyframe selector 210 to refine the pose determined by 3D object modeler 205. The pose refinement engine 230 can output the refined pose to 3D base model generator 220. Without using the pose refinement engine 230, keyframe selector 210 can output keyframes to 3D base model generator 220. References below. Figure 6C and Figure 6D Further details about the pose refinement engine 230 are described.
[0112] Object analyzer 215 analyzes keyframes selected by keyframe selector 210 to identify different regions (e.g., segments) of an object. For example, object analyzer 215 can perform object segmentation to divide the object into different regions. In some examples, object analyzer 215 can be implemented by a face parser, which can resolve a person's head in a frame into different regions (e.g., face region, hair region, etc.). Object analyzer 215 can generate one or more object parsing masks for different regions. Examples of object parsing masks are provided in... Figure 6B As shown in the diagram (and described in more detail below), an object resolution mask can indicate whether a pixel in a frame corresponds to one or more regions in different areas. In some cases, the object resolution mask can include a 2D bitmap. In an illustrative example, object analyzer 215 can generate an object resolution mask corresponding to a first region of the head (e.g., a face region) and an object resolution mask corresponding to a second region of the head (e.g., a hair region). In some cases, object analyzer 215 can generate object resolution masks corresponding to both the first and second regions (e.g., the face region and the hair region). For example, the object resolution masks for the first and second regions can show a silhouette of a person. In other cases, different masks can be generated based on the properties of specific parts of the object (e.g., based on a person's hair).
[0113] In some aspects, object parsing masks can be used to generate 3D models of specific parts of an object (e.g., a 3D hair model of a person). In some examples, a 3D base model generator 220 can generate a 3D base model based on the object parsing mask and pose information provided from the 3D model of a first part of the object. In an illustrative example, the 3D base model generator 220 generates an additional object parsing mask based on the 3D model of the first part of the object (e.g., a 3DMM model) and the object parsing mask from the object analyzer 215. The additional object parsing mask is created by the union of the 3DMM model and different regions (e.g., hair region, face region), which results in the alignment (e.g., adjacency) of the resulting 3D model. After creating the object parsing mask, the 3D base model generator 220 then initializes the 3D model. Vertices of the initial 3D model can be selectively removed based on the additional object parsing mask to create the 3D base model. For example, the initial 3D model can be projected onto a 2D frame, and the object parsing mask can be compared to the projected 2D frame. See below for reference. Figure 8B Describe further details regarding projection from 3D to 2D.
[0114] In some examples, the 3D base model includes a 3D model of a first part of the object and a 3D model of a second part of the object. In the illustrative example above, the first part of the object is a person's head, and the second part is a person's hair. As described in more detail below, generating the 3D base model using the 3D models of the first and second parts ensures that the 3D model of the second part is automatically aligned (e.g., adjacent) with the 3D model of the first part and does not collide with it. As a result, unwanted visual artifacts can be reduced. See below for reference. Figure 7 , Figure 8A , Figure 8B , Figure 8C and Figure 8D Describe further details regarding the generation of the 3D base model.
[0115] Model extractor 225 can extract a 3D model of a second region from a 3D base model (e.g., extract a 3D model of hair from a 3D base model). For example, using a 3D base model created by 3D base model generator 220, an object parsing mask, and pose information, vertices in the 3D base model can be removed to obtain a 3D model of the second region. See below for reference. Figure 9 and Figure 10 Further details regarding the extraction of the 3D model of the second region are described.
[0116] In some respects, as described above, system 200 may include an object cleanup engine ( Figure 2(Not shown in the image). For example, an object cleanup engine can implement a priori-based object cleanup (e.g., hair cleanup) algorithm to remove certain vertices from a 3D model of a second region (e.g., a 3D model of hair). In one example, predefined or pre-selected landmarks can be defined. The y-values of the predefined landmarks can be used as priors or threshold y-values to determine whether certain vertices will be removed (or cleaned). For example, if more than K vertices belonging to the hair (as an example of the second region) have y-values less than the prior threshold y-value, the object cleanup engine can determine that the hairstyle corresponds to long hair and can leave the hair unchanged. If the object cleanup engine determines that fewer than K vertices of the hair have y-values less than the prior threshold y-value, the object cleanup engine can determine that the hairstyle is short. In this case, the object cleanup engine can remove any vertices with y-values less than the prior.
[0117] Figure 2 The system 200 shown illustrates a functional block diagram that can be implemented using hardware, software, or any combination of hardware and software. As an example, the illustrated functional blocks can be identified using the illustrated relationships, and can be converted into a Universal Modeling Language (UML) diagram to at least partially identify an example implementation of system 200 as an object-oriented arrangement in software. However, system 200 can be implemented without abstraction, and for example, as a static functional implementation.
[0118] Figure 5 The process 500 for selecting keyframes from a frame sequence of an object using a 3D model of the first part of the object is illustrated. As described above, the 3D object modeler 205 creates a 3D model of the first part of the object for each frame. In some examples, process 500 may be... Figure 2 The keyframe selector 210 shown is implemented to create a set of keyframes from a frame sequence. As described below, keyframes are selected based on a comparison of the 3D model of the first part of the object.
[0119] At box 510, process 500 can select a reference frame from the frame sequence. The reference frame can be identified based on a 3D model of a first portion having a neutral or non-biased position. In an illustrative example where the object is a human head, the reference frame could be the first frame in the frame sequence and depict the head in a neutral position (e.g., 0° rotation). Figure 3 An example of frame 310 that can be considered a reference frame is shown, since frame 310 is the first frame in a frame sequence. In some examples, the first frame of the frame sequence may have a corresponding 3D model of a first portion rotated at a small angle (e.g., 0.3°), and subsequent frames with a 3D model without rotation (e.g., 0.0°) may be selected as reference frames. In this case, subsequent frames can be considered reference frames.
[0120] At box 520, process 500 can divide the frame sequence into frame groups based on the positional information of the 3D model from the first part. In an illustrative example of a frame sequence each having a 3DMM model, each frame can be associated with pose information (such as the rotation angle of the 3DMM model). The frame sequence can be divided into groups based on the pose information. In an example where the object is the head of a person rotating in the frame sequence, the groups can be separated based on the object's position relative to its rotation toward its right (e.g., +60°) and / or left (e.g., -60°) (e.g., a first group for 0° to 5°, a second group for 5° to 10°, a third group for 10° to 15°, and so on).
[0121] For example, frames from a frame sequence are grouped into different groups based on the range of motion depicted in the frames. Keyframes are then determined from each group, as described below. In this example, the number of groups is based on the desired visual fidelity and time consumption required by the generative model. For example, more keyframes will result in better visual fidelity but will require more computation time. Furthermore, the number of groups should be large enough to cover the range of motion of the object. In some examples, each frame may be considered a keyframe.
[0122] Figure 3 An example frame 320 depicting a +60° rotation of a human head is shown. This example frame 320 can be selected as a keyframe for different frame groups based on the number of keyframes (e.g., +55° to +65°, +59.5° to +63°, +60° to +62°, etc.). Figure 3 An example keyframe 330 depicting a -60° rotation of a human head is shown. This example keyframe 330 is selected for another set of frames (e.g., -55° to -65°, -59.5° to -63°, -60° to -62°, etc.).
[0123] The frames selected at box 520 may not necessarily be sequential in time. In an illustrative example, the frame sequence could show a person rotating their head toward their right and then toward their left. For example, the frame sequence could include a first set of frames depicting the head rotating from 5° to 10° at the first moment. The frame sequence could also include a second set of frames depicting the head rotating from 10° to 5° as it moves toward the left at the second moment.
[0124] At box 530, process 500 can select keyframes in each frame group based on a comparison of frames in the frame group with a reference frame. In an illustrative example, process 500 identifies the best frame in the reference frame for which a 3DMM in the group is optimally assigned.
[0125] In an illustrative example, frame comparison can be performed using Intersection over Union (IOU) based on the 3D model of the first part. As described above, a 3DMM can be generated and fitted to each frame in a frame sequence. This can result in alignment differences in the 3DMM model. For example, a reference frame (e.g., with a neutral head position) can be generated (e.g., rasterized) as a reference 2D frame, and the first frame in the group can also be generated as a first 2D frame. For example, to rasterize the 3D model of the reference 2D frame, a camera (e.g., a camera model) and lights can be positioned in the 3D coordinate system of the reference 3D model. The rendering engine can calculate the effect of the lights projected onto the vertices of the reference 3D model based on the camera position to generate a 2D bitmap that visually depicts the reference 3D model. In some examples, 3D models can be compared in a 3D coordinate system without rasterizing the 3D model to a 2D bitmap.
[0126] An Intersection over Union (IOU) is performed to determine the similarity (e.g., correlation or value) between a reference 2D frame and a first 2D frame. For example, IOU can be used to determine whether an object detected in the current frame matches an object detected in a previous frame. The IOU comprises the intersection (I) and union (U) of two bounding boxes, which include the first bounding box of an object in the current frame and the second bounding box of an object in the previous frame. The intersection region includes the overlapping area between the first and second bounding boxes. The union region includes the union of the first and second bounding boxes. If the area of overlap between the first and second bounding boxes divided by the union of the bounding boxes is greater than an IOU threshold, then the first and second bounding boxes can be determined to match.
[0127] IOU comparison is just one example comparison, and other techniques can be used. For example, a differential mask can be generated using XOR based on values in a reference 2D frame and a first 2D frame. The resulting bitmap will indicate the area of the difference between the frames and can be measured to determine the similarity between the reference 2D frame and the first 2D frame.
[0128] In some examples, additional comparisons can be made with other frames in the frame group. If a second frame has high similarity, it can be selected as the keyframe for that frame group. In this case, the second frame will replace the first frame as the keyframe.
[0129] After selecting each keyframe for each frame group, at box 540, process 500 adds the reference frame and the keyframes selected from each frame group to the keyframe set. In an illustrative example, the keyframe set is provided to object analyzer 215. Object analyzer 215 can analyze the keyframes to identify different regions of an object (e.g., segments). For example, object analyzer 215 can perform object segmentation to divide the object into different regions.
[0130] Figure 6AThe results of object analysis of frame 310 and keyframe 330, which can be performed by object analyzer 215, are shown. In this case, segmentation result 610 of frame 310 shows the human hair region 612 and the human face region 614. Region 616 corresponds to the region other than the human head. Segmentation result 620 of frame 330 shows the human hair region 622 and the human face region 624. Region 626 corresponds to the region other than the human head.
[0131] Although Figure 6A The segmentation of object analyzer 215 is shown, but objects can be segmented in more detail based on the features being modeled. For example, if a scarf is to be modeled, the segmentation of features might need to detect physical objects close to a person's head and could identify finer details such as the nose region, eye region, mouth region, etc.
[0132] In some examples, object analyzer 215 can be implemented by a face parser, which can resolve a head in a frame into different regions (e.g., face region, hair region, etc.). Object analyzer 215 can generate one or more object parsing masks indicating whether a pixel in the frame corresponds to one or more regions. In an illustrative example, the object parsing mask may include a 2D bitmap. In some cases, as described in more detail herein, the object parsing mask can be used to generate a 3D model of a specific part of an object (e.g., a 3D hair model of a person). In some cases where a person's head and hair are used as parts of a person, object analyzer 215 can generate an object parsing mask corresponding to a person's face, an object parsing mask corresponding to a person's hair, and in some cases, an object parsing mask corresponding to a person's hair and face. In other cases, different masks can be generated based on the attributes of a specific part of the object (e.g., based on a person's hair).
[0133] Figure 6B Different object resolution masks generated from the results of object analysis of frames 310 and 330 are shown. Each object resolution mask indicates the presence of a region of an object identified by object analyzer 215. Specifically, white areas of the object resolution mask (e.g., white value 255) indicate that the coordinates of the object resolution mask correspond to that region of the object. Black areas (e.g., white value 0) indicate that the coordinates of the object resolution mask frame do not correspond to that region of the object.
[0134] exist Figure 6BIn the example shown, object resolution masks 630, 650, and 670 are derived from the object analysis of keyframe 310. As described above, keyframe 310 (which, as mentioned above, is also considered a reference frame) shows a frontal view of a person with their head not rotated on the yaw axis (i.e., 0° rotation). Similarly, object resolution masks 640, 660, and 680 are derived from the object analysis of keyframe 330. As described above, keyframe 330 shows a person with their head rotated -60° on the yaw axis.
[0135] Object parse mask 630 identifies the facial region of a person with the head not rotated relative to the neck region. Object parse mask 640 identifies the facial region when the head is rotated -60° relative to the neck region. Object parse mask 650 identifies a front view of the hair region of a person without head rotation. Object parse mask 660 identifies the hair region when the head is rotated -60° relative to the neck region. Object parse mask 670 identifies the face and head region of a person (e.g., silhouette) with the head not rotated relative to the neck region. Object parse mask 680 identifies the face and hair region when the head is rotated -60° relative to the neck region.
[0136] As described above, in some example implementations, system 200 may include a pose refinement engine 230. The pose refinement engine 230 may perform global pose refinement (e.g., using a global pose refinement algorithm) to mitigate misalignment that may be caused by head shape variance and / or pose errors (e.g., when the same region from different views corresponds to different regions in 3D). An illustrative example of a global pose refinement algorithm is provided in equation (6) below:
[0137]
[0138] For example, pose refinement engine 230 can refine (e.g., determined by keyframe selector 21) the pose of each keyframe by minimizing the difference between the landmarks of the distorted canonical 3DMM model (e.g., the reference frame 3DMM model) and the landmarks of the current keyframe, as shown in equation (6). In one example, the canonical 3DMM model can be defined by 3D object modeler 205 for a front view including a person (e.g., ... Figure 6C The model is generated from the image of the front view shown. In equation (6), term v represents a 3D vertex of a given landmark, and term x represents a 2D pixel of a given landmark. Each 3D vertex corresponds to a specific 2D pixel; for example, the 3D vertex of the landmark represented by v is projected onto the 2D pixel of the same landmark represented by x. Furthermore, in equation (6), term T represents rigid deformation, П is the projection matrix, and This represents the weights of the distances between landmarks in the viewpoint-based and / or distorted canonical 3DMM model and their corresponding landmarks in the current keyframe. Term i refers to a specific landmark (where i = 0, 1, 2, etc.), and term k refers to a specific keyframe (where k = 0, 1, 2, etc.).
[0139] Figure 6C This is an illustrative example of the reference frame 3DMM 682 in the front view. Figure 6D This is an illustrative example of the current keyframe 684. Landmark 683 is shown in reference frame 3DMM 382. Landmark 685, corresponding to landmark 683 (e.g., associated with the same location on the face), is shown in the current keyframe 684. In some cases, certain landmarks may be invisible when the face is rotated (e.g., turned to a side view). For example, when a person turns to the right... Figure 6D The landmark on the outer corner of the right eye of the person shown in keyframe 684 may not be visible. In some cases, certain expressions may make the distance between the landmark in the distorted canonical 3DMM model and the landmark in the current keyframe appear larger. As mentioned above, weights can be set or adjusted based on viewpoint and / or distance. For example, a lower weight (e.g., weight value 0) can be assigned to the landmark corresponding to the outer corner of the right eye when a person's face is turned to the right (and therefore the outer corner of the right eye is not visible in the keyframe) compared to a higher weight applied to a landmark that is visible in the keyframe. In another example, a lower weight (e.g., weight value 0) can be assigned to a landmark when the distance between a landmark in the distorted canonical model and a landmark in the keyframe is less than a distance threshold (e.g., 5 pixels, 10 pixels, or other threshold).
[0140] Figure 7 A flowchart of process 700 for generating a 3D base model corresponding to a first region and a second region of an object is shown. In one example, process 700 can be performed by... Figure 2 The 3D base model generator 220 shown is executed.
[0141] At box 710, process 700 initializes the 3D model and generates an object resolution mask for each keyframe. In some examples, the initial 3D model can be any shape that can be applied to process 700 to create the 3D base model. In the illustrative example, the 3D model is a default shape, such as a 3D cube. As will be described in further detail below, using the object resolution mask generated at box 710, the vertices of the initial 3D model can be removed to create a 3D base model that encapsulates the first part of the 3D model and the second part of the 3D model.
[0142] At box 720, process 700 projects the vertices of the initial 3D model onto an object resolution mask based on the pose information associated with the 3D model of the first part. In the illustrative example of the head, the pose information of the 3DMM can indicate the angular rotation of the head. Each vertex of the initial 3D model is projected from the 3D coordinate system to the 2D coordinate system based on the camera positioned according to the pose information. For example, if the object's pose information is rotated 15° on the yaw axis, the camera is positioned at 15° along the arc, and each vertex of the initial 3D model is projected into the camera to generate a 2D bitmap. Figure 8B The text further explains in detail the projection from the 3D coordinate system to the 2D coordinate system.
[0143] At frame 730, process 700 determines the region in the object resolution mask onto which each vertex of the initial 3D model is projected. In the example where the initial 3D model is projected into keyframe 310 and the camera is positioned based on the pose information of the 3DMM (e.g., 0°), the vertices of the key initial 3D model are projected onto the object resolution mask to determine whether the vertices are within the head and hair region (e.g., silhouette) or outside the head and hair region.
[0144] At frame 740, process 700 removes vertices from the initial 3D model based on the regions of the object resolution mask into which the vertices are projected. In the example where the initial 3D model is projected into keyframe 310, vertices projected outside the object resolution mask are removed. Vertices projected inside the object resolution mask are retained.
[0145] Figure 8A This is a schematic diagram illustrating the creation of an object resolution mask 802, which will be used to create a 3D base model. For example, Figure 2 The 3D base model generator can be based on the output from the object analyzer 215 (e.g., object parsing mask 806) and the first part of the object 804 from the 3D object modeler 205. Figure 8A In some examples, the object resolution mask 802 is created based on a 3D model of the first part of object 804 (e.g., a 3DMM model) and the rasterized union of object resolution mask 806. As will be described further below, object resolution mask 802 represents the second part of object 804 (in the example of a human head). Figure 8A In the example, the outer boundary of the 3D model of the first part of object 804 (hair on a human head) is ensured, and alignment relative to the first part of the object is guaranteed (e.g., alignment exists between the model representing the head and the model representing the hair). The rasterization of the 3D model of the first part of object 804 is aligned such that the head or skull region of the 3D model of the first part of object 804 is surrounded by object resolution mask 806 (e.g., within object resolution mask 806). Figure 8A In this example, the object resolution mask 806 identifies both a first region (e.g., the face region) and a second region (e.g., the hair region) of a person depicted in the image. In this example, the 3D model of the first part of the object ( Figure 8A The object parsing mask 802 is created by the union of the head of object 804 and object parsing mask 806, which will automatically align the 3D model of the first part of the object (e.g., the 3D model of the head) with the 3D model of the second part of the object (e.g., the 3D model of the hair).
[0146] In the example shown, the 3D model of the head portion (first part) of object 804 is almost entirely surrounded by the object resolution mask 806 because the hair region of the object resolution mask 806 is extensive. Hair regions can be much smaller (e.g., when the object is rotated, when a person has a smaller and less dense hair region, etc.), and segmenting the hair region using object resolution may be inaccurate. Therefore, aligning the 3D model of the head (first part) of object 804 with the object resolution mask 806 ensures alignment and accuracy when the object resolution mask 802 is used to create the 3D base model.
[0147] Figure 8B A top view of a 3D scene to be projected onto a 2D bitmap is shown. In an illustrative example, an initial 3D model 808 is statically positioned at the center point of arc 809. In the illustrative example, a camera 810 can be positioned along arc 809 to render the scene based on pose information of a 3DMM model representing a part of a person (e.g., the head or upper body and head). For example, camera 810 is positioned at... Figure 8B The camera is positioned at 0° and will correspond to the camera position in keyframe 310, since the person's head is rotated 0°. The camera can be an ideal pinhole camera model that describes the mathematical relationship between the coordinates of the camera position in 3D space and its projection onto the image plane 812. For example, a representation of the initial 3D model 808 is projected onto the image plane 812 and represented as a 2D bitmap.
[0148] A bitmap on image plane 812 can be compared with bitmap 814 to determine whether a vertex corresponds to a specific region. For example, because both the bitmap on image plane 812 and bitmap 814 are 2D, the bitmap on image plane 812 can be aligned with bitmap 814 and compared based on 2D positions. In the example where bitmap 814 is an object resolution mask identifying a region of a frame (e.g., a first region, such as a hair region identified by a specific pixel value, such as a value 1 or 0), each vertex in the initial 3D model can be mapped to a region within the object resolution mask. In the example where bitmap 814 is object resolution mask 670, a vertex can be determined to be inside or outside the silhouette of a person in keyframe 310. For example, when vertex 816 of the initial 3D model 808 is projected onto image plane 812, vertex 816 can be identified as inside or outside the silhouette based on bitmap 814. As described above, this occurs for each vertex of the initial 3D model 808.
[0149] In addition, the camera 820 can Figure 8B The camera 820 is positioned at +60° to project the initial 3D model 808 onto another keyframe. Camera 820 can be the same camera as camera 810 or a different camera. For example, camera 820 can correspond to the camera position used to capture keyframe 330. The different representations of 3D space are projected onto image plane 822 by camera 820 because camera 820 is located at a different position compared to camera 810. The projection onto image plane 822 can be compared with another image 824 to determine whether a vertex (e.g., vertex 816) is inside or outside the object resolving mask. In some examples, vertex 816 can be mapped to a different position in image plane 822 than in image plane 812.
[0150] In addition, the camera 830 can Figure 8B The camera 830 is positioned at -60° to project the initial 3D model 808 into other keyframes. Camera 830 can be the same as or different from camera 810 and / or camera 820. For example, camera 830 can correspond to the camera position used to capture keyframe 320. Different representations of 3D space are projected from camera 830 onto image plane 832. The projection into image plane 832 can be compared with another figure 834 to determine whether vertex 816 is inside or outside the object resolution mask. In some examples, vertex 816 can be mapped to a position in image plane 832 different from image planes 812 and 822.
[0151] Figure 8CAn example 3D base model 850 generated by the system and techniques described herein is shown. The 3D base model 850 includes a first region of the object, a second region of the object, and a rasterized union of the 3D models of the first portion (e.g., a 3DMM model). A portion of the 3D base model 800 includes an outer boundary 852 that at least partially corresponds to the outer boundary of the object. In the case where a human face corresponds to the object, the outer boundary may correspond to the human's hair region. For example, the 2D representation of the 3D base model has a high correlation with the object's resolving mask 670. However, in the case where hair is absent, the outer boundary may also correspond to the human's skin.
[0152] As described above, process 700 selectively removes several parts of the initial 3D model. In the illustrative example of a human head, a silhouette of a person is used to remove areas of the initial 3D model that are not associated with the head. As a result, the 3D base model 850 appears to be a cone encompassing both the human head and hair. Therefore, process 700 removes the outer boundary of the initial 3D model.
[0153] In some examples, the 3D base model 850 may include unintentionally retained vertices 854. In some examples, post-processing can be implemented to remove the errors. For example, a Laplacian smoothing operation can be implemented to remove vertex 854 to produce... Figure 8D The 3D basic model 850 described in the document.
[0154] In some examples, because the 3D base model 850 is constructed based on a combination of an object resolution mask and pose information corresponding to the object resolution mask, the 3D base model 850 is aligned with and scaled to the 3D model of the first part. Furthermore, at least a portion of the outer boundary of the 3D base model 850 is determined based on the boundary identified by the object resolution mask.
[0155] Figure 9 A flowchart of a process 900 for generating a 3D model of a second object is shown. In one example, process 900 can be performed by... Figure 2 The model extractor 225 shown is executed.
[0156] At box 910, procedure 900 includes: initializing the value of each vertex of the 3D base model to a default value. For example, Figure 2The model extractor 225 shown can initialize the value of each vertex of the 3D base model to a default value. In some examples, the initial value of the vertex is 1, but any value can be selected. As will be described in detail below, the vertex values of the 3D base model correspond to whether the vertex is retained in the 3D model of the second region (i.e., the vertex corresponds to the 3D model of the second object) or not present (e.g., removed) in the 3D model of the second region. Each vertex is set to a default value to indicate that the vertex corresponds to the second region of the 3D model, which allows unmodeled and not visible content in the 3D model to be associated with the second region. For example, in the case of a human head, the frame sequence may show ±60° rotation of the human head, and the occipital region (i.e., the back of the human head) cannot be directly modeled based on the range of motion to be modeled and the available frame sequence. In this case, setting the default value will force at least a portion of the occipital region to be covered by the model of the second region (i.e., the hair region).
[0157] At frame 920, procedure 900 projects the vertices of the 3D base model into the keyframe. For example, Figure 2 The model extractor 225 shown can project the vertices of the 3D base model into the object resolution mask of each keyframe. As described above, the vertices of the 3D base model can be projected from 3D to 2D based on the pose information corresponding to the 3DMM model of the keyframe.
[0158] At box 920, process 900 determines the region where each vertex of the 3D underlying model is projected onto a keyframe based on an object resolution mask. In some examples, the object resolution mask may identify the region where vertices are projected. In illustrative examples, the object resolution mask may identify the hair region of a person, such as object resolution masks 650 and 660. In some examples, the object resolution mask may identify another region (e.g., a facial region).
[0159] At box 930, process 900 adjusts the value of each vertex based on the region of the keyframe into which the vertex is projected. In the illustrative example, the vertex value can be incremented when the vertex is projected into the hair region. If the vertex falls into a region other than the hair region, the vertex value can be decremented. In other examples, the face region can be used to adjust the vertex values.
[0160] At box 940, process 900 includes: generating a predictive model based on vertex values. In some examples, the predictive model is generated by averaging all vertices of the 3D base model based on the projections described above in boxes 910 through 930.
[0161] At box 950, process 900 includes: removing vertices from the 3D base model based on the predictive model and vertex values. In the illustrative example, each vertex with a value less than the average is removed, while the others are retained. In effect, vertices from the 3D base model that are identified as not associated with the second region are removed to create the 3D model for the second part.
[0162] In the example of a human head, the second region corresponds to the hair area in the different object resolution mask, and the 3D model of the second part corresponds to the human hair. For example, content that does not correspond to the second region in the 3D base model is removed from the 3D base model to create the 3D model of the second part.
[0163] In an illustrative example, the 3D model of the second region is automatically aligned with the 3D model of the first part (e.g., without visible collisions) because the 3D base model is derived from the rasterized union of the object's first region, the object's second region, and the 3D model of the first part (e.g., a 3DMM model). In some examples, the 3D model of the second region is adjacent to at least a portion of the 3D model of the first part. For example, the outer boundary of the 3D model of the second region is determined based on an object resolution mask so that the 3D model of the second region overlaps with the 3D model of the first part. Therefore, removing content from the 3D base model as described above will result in the resulting 3D model of the second region being aligned with and proportional to the 3D model of the first part. The 3D model of the second part is also formed in a manner that will prevent any visible collisions with the 3D model of the first part.
[0164] Figure 10 This is a schematic diagram illustrating an example model extractor 1000 that can generate a 3D model of a second object from a 3D base model based on process 900. (See diagram 900.) Figure 10 As shown, the model extractor 1000 includes: a vertex initializer 1010, a vertex projector 1020, a region recognizer 1030, and a vertex remover 1040.
[0165] In some examples, vertex initializer 1010 can initialize the values of 3D base model 1012 to default values. Each vertex is set to a default value to indicate that the vertex corresponds to a second region of the 3D model, which allows unmodeled content in the 3D model that should not be visible to be associated with the second region. Vertex projector 1020 projects the vertices of 3D base model 1012 into the keyframe using, for example, pose information corresponding to the 3DMM model of the keyframe.
[0166] The region recognizer 1030 determines the region where each vertex of the 3D base model is projected onto a keyframe based on the object parsing mask. The region recognizer 1030 can also modify vertex values based on the region of the keyframe. For example, if a vertex is projected onto the hair region of the object parsing mask, the vertex value can be increased to retain vertices corresponding to the 3D model of the second part. Conversely, if a vertex is projected onto the face region of the object parsing mask, the vertex value can be decreased to remove vertices that do not correspond to the 3D model of the second part.
[0167] Vertex remover 1040 can perform the function of removing vertices from the 3D base model to create the 3D model of the second part 1042. In some examples, vertex remover 1040 can generate a predictive model by averaging all vertices of the 3D base model based on the projection and then removing vertices of the 3D base model based on the predictive model and the values of the vertices. In some examples, vertices with values less than the average are removed, while vertices with values greater than the average are retained. The resulting 3D model of the second part 1042 is aligned with the 3D model of the first part because the 3D base model is created based on the union of the 3D models of the first part (e.g., head region), the second part (e.g., hair region), and the first part.
[0168] Figure 11 A 3D model of a second part 1100 generated by the system and techniques described herein is shown. As described in detail above, the 3D base model is created by removing content from the outer surface. In an illustrative example, the 3D model of the second part 1100 includes a surface 1102, which is at least partially derived from an object resolution mask used to create the 3D base model. Furthermore, the 3D model is created by removing content from within the 3D base model based on an object resolution mask associated with (e.g., encapsulated therein) the 3D base model. In some examples, the inner surface of the 3D model of the second part is similar to the outer surface of the 3D model of the first part and does not collide with the outer surface of the 3D model of the first part. In the illustrative example of a human head and hair, the 3D model of the first part (e.g., the head) will not collide with the 3D model of the second part (e.g., the hair).
[0169] In some cases, as described above, the object clearing engine may implement a prior-based object clearing (e.g., hair clearing) algorithm. For example, when a person registers with the system 200, the person may be instructed to rotate their head left and right, which may allow the system to capture multiple frames of the person's face. Due to constraints on the mobility of a person's head (e.g., a person can generally only rotate their head approximately 55° in the yaw direction), the back of the person's hair region is not visible. This may lead to the problem that invisible vertices in the image cannot be removed (e.g., removed from a 3D base model). To solve these problems, a prior-based object clearing (e.g., hair clearing) algorithm is provided. In some cases, the object clearing algorithm may be implemented by the object clearing engine described above. For example, predefined or preselected landmarks may be defined, such as landmarks on the back of a person's head. In an illustrative example, the predefined landmarks may be selected based on analysis of a dataset including multiple frames. For example, based on analysis of the dataset, an optimal landmark (e.g., a landmark that generally corresponds to the hairline of a person with short hair) may be determined to be used as the predefined landmark. The preselected landmark may have an x-value (corresponding to the horizontal direction), a y-value (corresponding to the vertical direction), and a z-value (corresponding to the depth direction). The y-value of the predefined landmark may be used as a prior or a threshold y-value to determine whether certain vertices are to be removed (or cleared). If the y-values of more than K vertices belonging to hair (e.g., K may be equal to 20, 25, 30, or other suitable values) are less than the prior y-value, the object clearing engine may determine that the hairstyle corresponds to long hair. Whether a vertex belongs to a hair region may be determined by projecting the vertex onto a semantic label image of a key frame, as Figure 6A shown. If the projection belongs to the hair region, the object clearing engine may determine that the vertex also belongs to the hair region. If the object clearing engine determines that the hair corresponds to long hair (e.g., the y-values of more than K vertices of the hair are less than the prior), the object clearing engine may keep the hairstyle unchanged (the object clearing engine will not modify or remove any vertices corresponding to the hair). If the object clearing engine determines that the y-values of less than K hair vertices are less than the prior or the y-values of more than K vertices of the hair are greater than the prior, the object clearing engine may determine that the hair corresponds to short hair. In a case where the object clearing engine determines that the hair corresponds to short hair, the object clearing engine may remove any vertices whose y-values are less than the prior (e.g., remove any vertices with y < prior_y).
[0170] Figure 12 An example of a process 1200 for generating one or more models is shown. At block 1202, process 1200 includes generating a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object. For example, Figure 2The 3D object modeler 205 shown can generate a 3D model of a first part of an object based on one or more frames. In some cases, the one or more frames can depict the rotation of the object along a first axis. In some cases, the one or more frames can also depict the rotation of the object along a second axis. For example, the first axis can correspond to the yaw axis, and the second axis can correspond to the pitch axis.
[0171] According to some examples, process 1200 includes selecting one or more frames from a frame sequence as keyframes. For example, Figure 2 The keyframe selector 210 shown can select one or more frames as keyframes from a frame sequence. In some examples, as described above, each keyframe depicts an object at a different angle (e.g., based on keyframes captured at different angles relative to the object). In some cases, to select one or more frames as keyframes, process 1200 may include generating a first map from a 3D model of a first portion of the object for a first angle selected along an axis. Process 1200 may include generating a first metric, at least in part, by comparing the first map with a reference frame in the frame sequence. In some cases, the comparison is performed by determining the intersection-to-union (IoU) ratio of the bitmaps of the first map and the reference frame. For example, as described above, the first keyframe may be selected based on an IoU greater than an IoU threshold. After selecting the first keyframe, process 1200 may include generating a second metric, at least in part, by comparing the reference frame with a bitmap of a second frame in the frame sequence. In this case, process 1200 may identify a better frame (e.g., one with a larger IoU compared to the first keyframe) and use that better frame as a keyframe. In this case, process 1200 may include: selecting a second frame as the first keyframe based on a second metric.
[0172] In some cases, selecting one or more frames may include features that facilitate the capture of additional content. For example, process 1200 may also include: determining that a first keyframe does not meet a quality threshold. If the first keyframe does not meet the quality threshold, process 1200 may include: outputting feedback to facilitate the positioning of an object corresponding to the first keyframe, capturing at least one frame based on the feedback, and inserting frames from the at least one frame into the keyframe.
[0173] At frame 1204, process 1200 includes generating a mask for the one or more frames. For example, Figure 2The object analyzer 215 shown can generate an object resolution mask for the one or more frames, as described above. The mask includes indications of one or more regions of an object. In some cases, process 1200 may include generating a first mask identifying a first region and a second mask identifying a second region. In one illustrative example, the object is a person, the first region is the person's face (or facial) region, and the second region is the person's hair region. In another illustrative example, the first region may correspond to the person's body region and the region of clothing worn by the person.
[0174] At box 1206, process 1200 includes: generating a 3D base model based on a 3D model and a mask of a first part of the object. The 3D base model can represent the first part of the object and a second part of the object. In some examples, Figure 2 The 3D base model generator 220 shown can generate a 3D base model based on the 3D model and mask of the first part of the object.
[0175] In some aspects, to generate a 3D base model, process 1200 may include: projecting each vertex of an initial 3D model onto a mask associated with a frame based on pose information associated with a frame in the one or more frames. Process 1200 may include: determining whether each vertex of the first portion of the 3D model lies within a first region of the mask associated with the frame. As mentioned above, the first region of the mask may correspond to a human face region. In some examples, the 3D base model may be generated by combining the mask with rasterization of the first portion of the 3D model. In some cases, process 1200 may include: extracting the 3D base model based on the vertices of the first portion of the 3D model lying within the first region of the mask associated with the frame, such as by at least referencing... Figure 10 As stated above.
[0176] At box 1208, process 1200 includes: generating a 3D model of the second part of the object based on a mask and a 3D base model. For example, Figure 2 The model extractor 225 shown can generate a 3D model of the second part of an object based on a mask and a 3D base model, such as using a reference. Figure 10 The described technique involves generating a 3D model of the second part of an object such that it is aligned (e.g., adjacent) to the 3D model of the first part of the object. As described above, the 3D model of the second part of the object visually does not collide with the 3D model of the first part of the object. For example, as referenced... Figure 8A As mentioned above, because the mask is generated based on a combination of an object resolution mask (e.g., from a keyframe) and a rasterization of the 3D model of the first part, the alignment of the 3D model of the second part with respect to the 3D model of the first part can be ensured.
[0177] In some examples, the 3D model of the second part corresponds to an object that is part of the object. In one illustrative example, the object is a person, the first part of the object corresponds to the person's head, and the second part of the object corresponds to the hair on the person's head (e.g., as shown in the reference above). Figures 8A to 10 (As described above). In some examples, the second part of the 3D model corresponds to an item that can be separated from the object and / or moved relative to the object. In one illustrative example, the object is a person, the first part of the 3D model corresponds to the person's body area, and the second part of the 3D model corresponds to accessories or clothing worn by the person.
[0178] In some examples, to generate a 3D model of the second part of an object, process 1200 may include: initializing the value of each vertex of the 3D base model to an initial value. Each vertex of the 3D base model may be projected onto a keyframe in one or more frames. Process 1200 may include: determining whether a vertex of the 3D base model is projected onto a first mask or a second mask of the first keyframe. In some cases, process 1200 may include: adjusting the value of each vertex based on whether the corresponding vertex is projected onto the first mask or the second mask. For example, when the first vertex is projected onto the second region, the value of the first vertex may be increased, and when the first vertex is projected onto the first region, the value of the first vertex may be decreased.
[0179] In some aspects, process 1200 may include determining an average probability based on the value of each vertex. Process 1200 may further determine the probability that a vertex corresponds to a 3D model of a first or second part of an object. In some examples, process 1200 may determine the probability that a vertex corresponds to a 3D model of a first or second part of an object at least in part by comparing the vertex value to the average probability. Vertices are removed from the 3D base model based on the probability that each of one or more vertices corresponds to a first or second part of an object (e.g., based on whether one or more vertices are within a region of one or more regions from the one or more frames). For example, process 1200 may include removing vertices based on whether a corresponding vertex is identified as corresponding to a second region (e.g., a hair region) or a first region (e.g., a face region). In one example, vertices corresponding to a first region (e.g., a face region) may be removed from the 3D base model, while vertices corresponding to a second region (e.g., a hair region) may be retained and used for the 3D model of a second part of the object (e.g., for the hair).
[0180] In some cases, process 1200 may include: performing pose refinement on pose information associated with a frame among the one or more frames. In an illustrative example, pose refinement engine 230 may perform pose refinement on the pose information. For example, process 1200 (e.g., implemented by pose refinement engine 230) may include: minimizing the difference between one or more landmarks of a warped reference frame model and one or more landmarks of said frame, as described above.
[0181] In some examples, process 1200 may include: determining that coordinate values less than a predetermined coordinate value are held by fewer than a threshold number of vertices of the 3D model of the second part of the object. Process 1200 may include: removing one or more vertices of the 3D model with coordinate values less than the predetermined coordinate value based on the determination that fewer than the threshold number of vertices of the 3D model have coordinate values less than the predetermined coordinate value. For example, as described above, an object cleaning engine may determine that y-values of fewer than K vertices of an object (e.g., hair) are less than a prior y-value (e.g., a y-value coordinate based on a predetermined landmark), or that y-values of more than K vertices of the object are greater than the prior y-value. Based on the determination that y-values of fewer than K vertices of the object are less than the prior y-value or that y-values of more than K vertices of the object are greater than the prior y-value, the object cleaning engine may determine that the object corresponds to a specific type of object (e.g., the hair corresponds to short hair), and may remove any vertices with y-values less than the prior y-value (e.g., remove any vertices satisfying y < prior_y).
[0182] In some examples, process 1200 may include: generating an animation in an application using the 3D model of the first part and the 3D model of the second part. The application may include functionality for transmitting and receiving at least one of audio and text, and may display the 3D model of the first part and the 3D model of the second part. For example, when displaying the 3D model of the first part and the 3D model of the second part simultaneously, the application may depict a user of the application.
[0183] In some aspects, process 1200 includes: receiving an input corresponding to a selection of at least one graphical control for modifying the 3D model of the second part. Process 1200 may include: modifying the 3D model of the second part based on the received input. In one example, a user of the application may modify the 3D model of the second part, for example to increase or decrease the length of hair.
[0184] Figure 13Another example of a process 1300 for generating one or more models is shown. At box 1302, process 1300 includes generating a three-dimensional (3D) model of a human head based on one or more frames depicting a human. In some cases, similar to that described for process 1200, process 1300 may include selecting one or more frames from a frame sequence as keyframes. For example, process 1300 may include determining that a first keyframe does not meet a quality threshold. When the first keyframe does not meet the quality threshold, process 1300 may include functionality related to capturing additional images. In an illustrative example, process 1300 may include outputting feedback to facilitate positioning the human corresponding to the first keyframe, capturing at least one frame based on the feedback, and inserting frames from the at least one frame into the keyframe.
[0185] At box 1304, process 1300 includes: generating a mask for the one or more frames, the mask including indications of one or more regions of a person. For example, process 1300 may include: segmenting each of the one or more frames into one or more regions and generating one or more masks for each frame. The one or more masks include indications of the one or more regions. In some examples, a first mask may include indications of a first region (e.g., a person's facial region), and a second mask may include indications of a second region (e.g., a person's hair region).
[0186] At box 1306, process 1300 includes generating a 3D base model based on a 3D model of a first part of a person and a mask. The 3D base model may correspond to the outer boundaries of a person's head and hair. To generate the 3D base model, process 1300 may include: projecting each vertex of an initial 3D model onto a mask associated with the frame based on pose information associated with the frame or more frames; determining whether each vertex of the 3D model of the head is located within a head region of the mask associated with the frame; and extracting the 3D base model based on the vertices of the 3D model of the head being within the head region of the mask associated with the frame. In some examples, the one or more vertices may be removed from the 3D base model based on the probability that each vertex of the one or more vertices is outside the head region.
[0187] At box 1308, process 1300 includes generating a 3D model of human hair based on a mask and a 3D base model. In an illustrative example, the 3D model of human hair visually does not collide with the 3D model of a human head. To generate the 3D model of human hair, process 1300 may include: initializing the values of each vertex of the 3D base model to initial values, projecting a first vertex of the 3D base model into a keyframe in one or more frames, and determining whether the vertices of the 3D base model are projected into a first mask or a second mask of the first keyframe. In an illustrative example, as described above, the first mask may correspond to a facial region, and the second mask may correspond to a hair region.
[0188] In some examples, generating a 3D model of human hair may include adjusting the value of each vertex based on whether the corresponding vertex is projected onto a first mask or a second mask. For example, the value of the first vertex may be increased when it is projected onto the second mask (corresponding to the hair region), and decreased when it is projected onto the first mask (corresponding to the face region).
[0189] In some cases, process 1300 may include determining an average probability based on the value of each vertex. The probability that a vertex corresponds to a 3D model of hair is based on a comparison of the vertex's value with the average probability. For example, when the probability indicates that a vertex does not correspond to a 3D model of hair, the vertex is removed because the vertex may correspond to a 3D model of the head. When a vertex corresponds to a 3D model of hair, the vertex is retained for the 3D model of hair.
[0190] In some examples, process 1300 may include performing pose refinement of pose information associated with the frames in the one or more frames. In an illustrative example, pose refinement engine 230 may perform pose refinement of the pose information. For example, process 1300 (e.g., implemented by pose refinement engine 230) may include minimizing the difference between one or more landmarks in the distorted reference frame model and one or more landmarks of the frame, as described above.
[0191] In some cases, process 1300 may include: determining that coordinate values of less than a threshold number of vertices of a 3D model of a person's hair are less than a predetermined coordinate value. Process 1300 may include: removing one or more vertices of the 3D model whose coordinate values are less than the predetermined coordinate value based on the determination that coordinate values of less than the threshold number of vertices of the 3D model are less than the predetermined coordinate value. For example, as described above, an object cleaning engine may determine that y-values of less than K vertices of the hair of the 3D model are less than a prior y-value (e.g., based on a y-coordinate of a predetermined landmark) or that y-values of more than K vertices of the hair of the 3D model are greater than the prior y-value. Based on determining that y-values of less than K vertices of the hair of the 3D model are less than the prior y-value or that y-values of more than K vertices of the hair of the 3D model are greater than the prior y-value, the object cleaning engine may determine that the hair of the 3D model corresponds to short hair, and may remove any vertex whose y-value is less than the prior y-value (e.g., remove any vertex with y<prior_y).
[0192] In some examples, the 3D model of an object may be included in various applications executed by a device. The device may execute an application that displays a 3D model (e.g., an avatar) corresponding to a user of the application. The application may include a function of sending communications input by a person (e.g., voice, text, etc.). In an illustrative example, the application may be a text messaging application that displays a user corresponding to an avatar, and the user may include additional context by providing input to animate the user's avatar. The user may provide input to animate and / or otherwise move a first portion of the 3D model of the user's avatar, such as to provide non-verbal cues and context (e.g., smiling, laughing, shaking the head, nodding, etc.). In an illustrative example, the 3D hair model may be positioned with and move together with the head model without generating any visible collision.
[0193] The application may also include a user interface for modifying a 3D model of a second portion (e.g., a 3D hair model). The application may: display at least one graphical control to modify the 3D model of the second portion, receive a signal indicating that the user has selected the at least one graphical control, and visually modify the 3D model of the second portion based on the at least one graphical control. The graphical control may be any suitable function, such as size adjustment or shape adjustment of the 3D hair model. However, the graphical control may also control specific functions related to the second portion of the 3D model. By way of example, the graphical control may increase the hair volume of the 3D hair model or change the hair length (e.g., shorten or lengthen).
[0194] In other examples, the device may include an application or function for capturing a sequence of frames and performing some of the processes described herein (e.g., process 500, process 700, process 900, process 1200, process 1300, and / or other processes described herein). The application or function may be able to determine that a particular keyframe does not meet a quality threshold (e.g., the relevance or similarity of the keyframe is less than a minimum threshold). In this case, the application or function may output user interface elements (e.g., notifications, messages, etc.) or otherwise provide feedback (e.g., visual, auditory, tactile, and / or other feedback) to facilitate the positioning of an object corresponding to the particular keyframe. For example, the application may provide feedback information to a person performing the application or function audibly, visually, or physically (e.g., via tactile feedback) to help align the object. The application may cause the camera to capture at least one image or frame while providing feedback information. If the application determines that a particular frame in at least one frame meets a quality threshold or is an improvement on the original keyframe, the application may remove the original keyframe and insert the particular frame into the keyframe set.
[0195] In some examples, the processes described herein (e.g., process 500, process 700, process 900, process 1200, process 1300, and / or other processes described herein) may be executed by a computing device or apparatus. In some examples, processes 500, 700, 900, 1200, or 1300 may be executed by system 200. In another example, processes 500, 700, 900, 1200, or 1300 may be executed by a system having Figure 14 The computing device or system of the architecture shown in the computing system 1400 is used to perform the operation.
[0196] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, extended reality (XR) devices or systems (e.g., VR headsets, AR headsets, AR glasses, or other XR devices or systems), wearable devices (e.g., connected watches or smartwatches or other wearable devices), server computers or systems, vehicles or vehicles with computing capabilities (e.g., autonomous vehicles), robotic devices, televisions, and / or any other computing device having the resource capability to perform the processes described herein (including processes 500, 700, 900, 1200, or 1300). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP) based data or other types of data.
[0197] Components of a computing device can be implemented in a circuit system. For example, components may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or components may include computer software, firmware, or combinations thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or combinations thereof for performing the various operations described herein.
[0198] Processes 500, 700, 900, 1200, and 1300 are shown as logic flowcharts, whose operations represent a series of operations that can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media, which, when executed by one or more processors, performs the described operation. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or can be parallelized to implement these processes.
[0199] Furthermore, the processes 500, 700, 900, 1200, and 1300 and / or any other processes described herein can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors via hardware or a combination thereof. As mentioned above, the code can be stored, for example, on a computer-readable or machine-readable storage medium in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0200] Figure 14 This is a schematic diagram illustrating an example of a system used to implement certain aspects of this technology. Specifically, Figure 14 An example of computing system 1400 is shown, which can be any computing device, such as constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1405. Connection 1405 can be a physical connection using a bus, or a direct connection to processor 1410 (such as in a chipset architecture). Connection 1405 can also be a virtual connection, a networking connection, or a logical connection.
[0201] In some embodiments, the computing system 1400 is a distributed system, wherein the functions described herein may be distributed across data centers, multiple data centers, peer-to-peer networks, etc. In some embodiments, one or more of the described system components represent a plurality of such components, each performing some or all of the functions described for that component. In some embodiments, a component may be a physical device or a virtual device.
[0202] Example system 1400 includes: at least one processing unit (CPU or processor) 1410, and connections 1405 that couple various system components, including system memory 1415 (such as read-only memory (ROM) 1420 and random access memory (RAM) 1425), to processor 1410. Computing system 1400 may include a cache 1412 of high-speed memory that is directly connected to, adjacent to, or integrated into processor 1410.
[0203] Processor 1410 may include any general-purpose processor and hardware or software services configured to control processor 1410 (such as services 1432, 1434, and 1436 stored in storage device 1430), as well as dedicated processors in which software instructions are incorporated into the actual processor design. Processor 1410 may essentially be a fully self-contained computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0204] To enable user interaction, the computing system 1400 includes an input device 1445, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keypad, a mouse, motion input, voice, etc. The computing system 1400 may also include an output device 1435, which can be one or more of multiple output mechanisms. In some examples, a multimodal system allows the user to provide multiple types of input / output to communicate with the computing system 1400. The computing system 1400 may include a communication interface 1440, which typically manages and controls user input and system output. The communication interface may use wired and / or wireless transceivers to perform or facilitate the reception and / or transmission of wired or wireless communications, including using audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, etc. Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, proprietary wired ports / plugs Wireless signal transmission Low-power (BLE) wireless signal transmission Wired and / or wireless transceivers for wireless signal transmission, including radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), global microwave access interoperability (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 1440 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1400 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to: the US-based Global Positioning System (GPS), the Russian-based Global Navigation Satellite System (GLONASS), the Chinese-based BeiDou Navigation Satellite System (BDS), and the European-based Galileo GNSS. There are no restrictions on operation for any particular hardware configuration, therefore the basic features described here can be easily replaced with developed, improved hardware or firmware configurations.
[0205] Storage device 1430 may be a non-volatile and / or non-transitory and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as magnetic tape, flash memory cards, solid-state storage devices, digital versatile optical discs, magnetic tape, floppy disks, flexible disks, hard disks, magnetic tape, magnetic stripes / strips, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, CD-ROM, rewritable CD, digital video optical disc (DVD), Blu-ray disc (BDD), holographic disc, another optical medium, secure digital storage (SD) cards, microsecure digital storage (microSD) cards, etc. Cards, smart card chips, EMV chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase-change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or magnetic tape, and / or combinations thereof.
[0206] Storage device 1430 may include software services, servers, etc., which enable the system to perform functions when the code defining such software is executed by processor 1410. In some embodiments, hardware services that perform specific functions may include software components stored in a computer-readable medium that are connected to necessary hardware components (such as processor 1410, connection 1405, output device 1435, etc.) to perform said functions.
[0207] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media can include non-transitory media capable of storing data, but excludes carrier waves and / or transient electronic signals that are propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to: magnetic disks or magnetic tapes, optical disc storage media (such as compact discs (CDs) or digital multipurpose discs (DVDs)), flash memory, memory, or storage devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent any combination of procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameter data, etc., can be passed, forwarded, or transmitted using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.
[0208] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.
[0209] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that these embodiments can be practiced without these specific details. For clarity, in some instances, the techniques described herein may be presented as comprising individual functional blocks including devices, device components, steps or routines in a software-embodied method, or a combination of hardware and software. Other components may be used in addition to those shown in the accompanying drawings and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure these embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring these embodiments.
[0210] The various embodiments described above can be presented as processes or methods, depicted as flowcharts, schematic diagrams, data flow diagrams, structural diagrams, or block diagrams. While a flowchart can describe operations as a sequential process, many operations within an operation can be executed in parallel or concurrently. Additionally, the order of these operations can be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process can correspond to a method, function, process, subroutine, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0211] The processes and methods described in the examples above can be implemented using computer-executable instructions stored in or otherwise accessible from a computer-readable medium. For example, such instructions may include instructions and data that cause a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions, or otherwise configure such a computer, special-purpose computer, or processing device. Part of the computer resources used may be accessible via a network. The computer-executable instructions may be, for example, binary, intermediate-format instructions (such as assembly language, firmware, source code, etc.). Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include hard disks or optical disks, flash memory, USB devices provided with non-volatile memory, network storage devices, and so on.
[0212] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may employ any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored on a computer-readable or machine-readable medium. One (or more) processors may perform the necessary tasks. Typical examples of form factors include: laptop computers, mobile phones (e.g., smartphones or other types of mobile phones), tablet devices or other small form factor personal computers, personal digital assistants, rack-mount devices, standalone devices, and so on. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By further example, such functionality may also be implemented between different processes executed on different chips on a circuit board or in a single device.
[0213] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example units for providing the functionality described in this disclosure.
[0214] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art should recognize that this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it should be understood that these inventive concepts may be embodied and employed differently in other ways, and the appended claims are intended to be construed as including such variations beyond those limited by the prior art. Various features and aspects of the above applications may be used individually or in combination. Furthermore, embodiments may be used in any number of settings and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods have been described in a specific order. It should be recognized that, in alternative embodiments, the methods may be performed in a different order than that described.
[0215] Those skilled in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.
[0216] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0217] The phrase “coupled to” refers to any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0218] The use of declarative language or other languages that refer to "at least one" and / or "one or more" in a set indicates that one or more members of that set (with any combination) satisfy the claim. For example, the claim language stating "at least one of A and B" or "at least one of A or B" refers to A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" refers to A, B, C, or A and B, or A and C, or B and C, or A and B and C. The use of the language set "at least one" and / or the set "one or more" is not limited to the set of items listed in that set. For example, the claim language stating "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0219] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above in relation to their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.
[0220] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication handheld devices, or integrated circuit devices with multiple uses, including applications in wireless communication handheld devices and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as a discrete but interoperable logic device. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, etc. Alternatively or alternatively, these technologies can be implemented at least in part through computer-readable communication media that carry or transmit program code in the form of instructions or data structures (such as propagated signals or waves), and the program code can be accessed, read, and / or executed by a computer.
[0221] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein.
[0222] The illustrative aspects of this disclosure include:
[0223] Aspect 1: An apparatus for generating one or more models, comprising a memory and one or more processors (e.g., implemented in a circuit) coupled to the memory. The memory may be configured to store data, such as one or more frames, one or more three-dimensional models, and / or other data. The one or more processors are configured to: generate a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object; generate a mask for the one or more frames, the mask including indications of one or more regions of the object; generate a 3D base model based on the 3D model of the first portion of the object and the mask, the 3D base model representing the first portion of the object and a second portion of the object; and generate a 3D model of the second portion of the object based on the mask and the 3D base model.
[0224] Aspect 2: The apparatus according to aspect 1, wherein the 3D model of the second part corresponds to an article that is part of the object.
[0225] Aspect 3: The apparatus according to any one of Aspect 1 or 2, wherein the object is a person, a first portion of the object corresponds to the head of the person, and a second portion of the object corresponds to the hair on the head of the person.
[0226] Aspect 4: The apparatus according to aspect 1, wherein the 3D model of the second part corresponds to an article that is at least one of the following: separable from the object and movable relative to the object.
[0227] Aspect 5: The apparatus according to any one of Aspect 1 or 4, wherein the object is a person, a first portion of the 3D model corresponds to a body region of the person, and a second portion of the 3D model corresponds to an accessory or clothing worn by the person.
[0228] Aspect 6: The apparatus according to any one of aspects 1 to 5, wherein the 3D model of the second part of the object is adjacent to at least a portion of the 3D model of the first part of the object.
[0229] Aspect 7: The apparatus according to any one of aspects 1 to 6, wherein the 3D model of the second part of the object does not visually collide with the 3D model of the first part of the object.
[0230] Aspect 8: An apparatus according to any one of aspects 1 to 7, wherein the one or more processors are configured to: select the one or more frames from a frame sequence as keyframes, wherein each keyframe depicts the object from a different angle.
[0231] Aspect 9: An apparatus according to any one of Aspects 1 to 8, wherein the one or more processors are configured to: determine that a first keyframe does not meet a quality threshold; output feedback to facilitate the positioning of the object corresponding to the first keyframe; capture at least one frame based on the feedback; and insert a frame from the at least one frame into the keyframe.
[0232] Aspect 10: An apparatus according to any one of Aspects 1 to 9, wherein the one or more processors are configured to: generate a first bitmap from a 3D model of a first portion of the object for a first angle selected along an axis; generate a first metric at least in part by comparing the first bitmap with a reference frame in the frame sequence; and select a first keyframe based on the result of the comparison.
[0233] Aspect 11: The apparatus according to aspect 10, wherein, in order to compare the first bitmap with the reference frame, the one or more processors are configured to: perform an intersection-over-conclusion (IoC) ratio of the first bitmap with the bitmap of the reference frame.
[0234] Aspect 12: An apparatus according to any one of aspects 10 or 11, wherein the one or more processors are configured to: generate a second metric at least in part by comparing the reference frame with a bitmap of a second frame in the frame sequence; and select the second frame as the first keyframe based on the second metric.
[0235] Aspect 13: An apparatus according to any one of aspects 1 to 12, wherein the one or more processors are configured to: segment each of the one or more frames into one or more regions; and generate a mask for each of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.
[0236] Aspect 14: An apparatus according to any one of aspects 1 to 13, wherein the one or more processors are configured to: determine a rasterization of a 3D model of a first portion of the object of a frame, a union between a first mask for a first region of the object and a second mask for a second region of the object; and generate a mask for the frame based on the determined union.
[0237] Aspect 15: The apparatus according to any one of aspects 1 to 14, wherein the mask for each of the one or more frames comprises: a first mask identifying a first region of the object and a second mask identifying a second region of the object.
[0238] Aspect 16: An apparatus according to any one of Aspects 1 to 15, wherein the one or more processors are configured to: initialize the value of each vertex of the 3D base model to an initial value, wherein the initial value indicates that the corresponding vertex is located within the first region; project a first vertex of the 3D base model into a keyframe in the one or more frames; determine whether the vertex of the 3D base model is projected into a first mask or a second mask of the first keyframe; and adjust the value of each vertex based on whether the corresponding vertex is projected into the first mask or the second mask.
[0239] Aspect 17: The apparatus according to any one of aspects 14 to 16, wherein the first region is a facial region and the second region is a hair region.
[0240] Aspect 18: The apparatus according to any one of Aspects 16 or 17, wherein when the first vertex is projected onto the second region, the value of the first vertex is increased, and when the first vertex is projected onto the second region, the value of the first vertex is increased, and...
[0241] Aspect 19: An apparatus according to any one of Aspects 16 to 18, wherein the one or more processors are configured to: determine an average probability based on the value of each vertex, wherein the probability of a vertex corresponding to a 3D model of a first part of the object is based on a comparison of the value of the vertex with the average probability.
[0242] Aspect 20: An apparatus according to any one of aspects 1 to 19, wherein the one or more processors are configured to: project each vertex of an initial 3D model onto a mask associated with the frame based on pose information associated with a frame in the one or more frames; determine whether each vertex of the first portion of the 3D model is located within a first region of the mask associated with the frame; and extract the 3D base model based on the vertices of the first portion of the 3D model being within the first region of the mask associated with the frame.
[0243] Aspect 21: The apparatus according to aspect 20, wherein the object is a person, and the first region corresponds to the person's facial region and the person's hair region.
[0244] Aspect 22: The apparatus according to aspect 20, wherein the object is a person, and the first region corresponds to the body region of the person and the outer garment region worn by the person.
[0245] Aspect 23: An apparatus according to any one of aspects 13 to 22, wherein the one or more processors are configured to remove the one or more vertices from the 3D base model based on the probability of each of the one or more vertices being in a region of the one or more regions of a frame from the one or more frames.
[0246] Aspect 24: The apparatus according to any one of aspects 1 to 23, wherein the one or more processors are configured to generate animation in an application using a 3D model of the first portion and a 3D model of the second portion, wherein the object includes a person, the 3D model of the first portion corresponding to the head of the person, and the 3D model of the second portion corresponding to the hair of the person.
[0247] Aspect 25: The apparatus according to any one of aspects 1 to 24, wherein the application includes a function for transmitting and receiving at least one of audio and text.
[0248] Aspect 26: The apparatus according to any one of aspects 1 to 25, wherein the 3D model of the first part and the 3D model of the second part depict the user of the application.
[0249] Aspect 27: An apparatus according to any one of aspects 1 to 26, wherein the one or more processors are configured to: receive input corresponding to the selection of at least one graphical control for modifying the 3D model of the second part; and modify the 3D model of the second part based on the received input.
[0250] Aspect 28: The apparatus according to any one of aspects 1 to 27, wherein the one or more frames are associated with rotation of the object along a first axis.
[0251] Aspect 29: The apparatus according to aspect 28, wherein the one or more frames are associated with rotation of the object along a second axis.
[0252] Aspect 30: The apparatus according to aspect 29, wherein the first axis corresponds to the yaw axis and the second axis corresponds to the pitch axis.
[0253] Aspect 31: The apparatus according to any one of aspects 1 to 30, wherein the one or more processors are configured to: perform pose refinement of pose information associated with frames in the one or more frames.
[0254] Aspect 32: The apparatus according to aspect 31, wherein, in order to perform pose refinement of pose information associated with the frame, the one or more processors are configured to minimize the difference between one or more landmarks of the distorted reference frame model and one or more landmarks of the frame.
[0255] Aspect 33: An apparatus according to any one of aspects 1 to 31, wherein the one or more processors are configured to: determine that the coordinate values of fewer than a threshold number of vertices of a 3D model of a second portion of the object are less than a predetermined coordinate value; and based on determining that the coordinate values of fewer than a threshold number of vertices of the 3D model are less than the predetermined coordinate value, remove one or more vertices of the 3D model that are less than the predetermined coordinate value.
[0256] Aspect 34: A method for generating one or more models. The method includes: generating a three-dimensional (3D) model of a first portion of an object based on one or more frames depicting the object; generating a mask for the one or more frames, the mask including indications of one or more regions of the object; generating a 3D base model based on the 3D model of the first portion of the object and the mask, the 3D base model representing the first portion of the object and a second portion of the object; and generating a 3D model of the second portion of the object based on the mask and the 3D base model.
[0257] Aspect 35: According to the method of aspect 34, wherein the 3D model of the second part corresponds to an item that is part of the object.
[0258] Aspect 36: The method according to aspect 34 or any one of aspect 34, wherein the object is a person, a first portion of the object corresponds to the head of the person, and a second portion of the object corresponds to the hair on the head of the person.
[0259] Aspect 37: According to the method of aspect 36, wherein the 3D model of the second part corresponds to an item that is at least one of the following: separable from the object and movable relative to the object.
[0260] Aspect 38: The method according to any one of Aspects 34 to 37, wherein the object is a person, a first portion of the 3D model corresponds to a body region of the person, and a second portion of the 3D model corresponds to an accessory or clothing worn by the person.
[0261] Aspect 39: The method according to any one of aspects 34 to 38, wherein the 3D model of the second part of the object is adjacent to at least a portion of the 3D model of the first part of the object.
[0262] Aspect 40: The method according to any one of aspects 34 to 39, wherein the 3D model of the second part of the object does not visually collide with the 3D model of the first part of the object.
[0263] Aspect 41: The method according to any one of aspects 34 to 40 further includes: selecting one or more frames from a frame sequence as keyframes, wherein each keyframe depicts the object from a different angle.
[0264] Aspect 42: The method according to any one of aspects 34 to 41 further includes: determining that a first keyframe does not meet a quality threshold; outputting feedback to facilitate the positioning of the object corresponding to the first keyframe; capturing at least one frame based on the feedback; and inserting a frame from the at least one frame into the keyframe.
[0265] Aspect 43: The method according to any one of aspects 34 to 42 further includes: generating a first bitmap from a 3D model of a first portion of the object for a first angle selected along an axis; generating a first metric at least in part by comparing the first bitmap with a reference frame in the frame sequence; and selecting a first keyframe based on the result of the comparison.
[0266] Aspect 44: According to the method of aspect 43, comparing the first bitmap with the reference frame includes: performing an intersection-over-union ratio of the first bitmap and the bitmaps of the reference frame.
[0267] Aspect 45: The method according to any one of aspects 42 or 44 further includes: generating a second metric by at least partly comparing the reference frame with a bitmap of a second frame in the frame sequence; and selecting the second frame as the first keyframe based on the second metric.
[0268] Aspect 46: The method according to any one of aspects 34 to 45 further includes: segmenting each of the one or more frames into one or more regions; and generating a mask for each of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.
[0269] Aspect 47: The method according to any one of aspects 34 to 46 further includes: determining a rasterization of a 3D model of a first portion of the object of the frame, a union between a first mask for a first region of the object and a second mask for a second region of the object; and generating a mask for the frame based on the determined union.
[0270] Aspect 48: The method according to any one of aspects 34 to 47, wherein the mask for each of the one or more frames comprises: a first mask identifying a first region of the object and a second mask identifying a second region of the object.
[0271] Aspect 49: The method according to any one of Aspects 34 to 48 further includes: initializing the value of each vertex of the 3D base model to an initial value, wherein the initial value indicates that the corresponding vertex is located within the first region; projecting a first vertex of the 3D base model into a keyframe in the one or more frames; determining whether the vertex of the 3D base model is projected into a first mask or a second mask of the first keyframe; and adjusting the value of each vertex based on whether the corresponding vertex is projected into the first mask or the second mask.
[0272] Aspect 50: The method according to any one of aspects 46 to 49, wherein the first region is a facial region and the second region is a hair region.
[0273] Aspect 51: The method according to any one of Aspects 48 or 50, wherein when the first vertex is projected onto the second region, the value of the first vertex is increased, and when the first vertex is projected onto the second region, the value of the first vertex is increased, and...
[0274] Aspect 52: The method according to any one of aspects 34 to 51 further includes: determining an average probability based on the value of each vertex, wherein the probability that a vertex corresponds to a 3D model of a first part of the object is based on a comparison of the value of the vertex with the average probability.
[0275] Aspect 53: The method according to any one of aspects 34 to 52 further includes: projecting each vertex of an initial 3D model onto a mask associated with the frame based on pose information associated with a frame in the one or more frames; determining whether each vertex of the first portion of the 3D model is located within a first region of the mask associated with the frame; and extracting the 3D base model based on the vertices of the first portion of the 3D model being within the first region of the mask associated with the frame.
[0276] Aspect 54: The method according to aspect 53, wherein the object is a person, and the first region corresponds to the person's facial region and the person's hair region.
[0277] Aspect 55: The method according to aspect 54, wherein the object is a person, and the first region corresponds to the body region of the person and the outer garment region worn by the person.
[0278] Aspect 56: The method according to any one of aspects 45 to 55 further includes: removing the one or more vertices from the 3D base model based on the probability of each vertex in a region of the one or more regions of a frame from the one or more frames.
[0279] Aspect 57: The method according to any one of Aspects 34 to 56 further includes: generating an animation in an application using the 3D model of the first portion and the 3D model of the second portion, wherein the object includes a person, the 3D model of the first portion corresponds to the head of the person, and the 3D model of the second portion corresponds to the hair of the person.
[0280] Aspect 58: The method according to any one of Aspects 34 to 57, wherein the application includes functionality for sending and receiving at least one of audio and text.
[0281] Aspect 59: The method according to any one of Aspects 34 to 58, wherein the 3D model of the first part and the 3D model of the second part depict the user of the application.
[0282] Aspect 60: The method according to any one of aspects 34 to 59 further includes: receiving input corresponding to the selection of at least one graphical control for modifying the 3D model of the second part; and modifying the 3D model of the second part based on the received input.
[0283] Aspect 61: The method according to any one of aspects 34 to 60, wherein the one or more frames are associated with a rotation of the object along a first axis.
[0284] Aspect 62: According to the method of aspect 61, wherein the one or more frames are associated with a rotation of the object along a second axis.
[0285] Aspect 63: According to the method of aspect 62, wherein the first axis corresponds to the yaw axis and the second axis corresponds to the pitch axis.
[0286] Aspect 64: The method according to any one of aspects 34 to 63 further includes: performing pose refinement of pose information associated with the frames in the one or more frames.
[0287] Aspect 65: According to the method of aspect 64, wherein performing pose refinement of pose information associated with the frame includes: minimizing the difference between one or more landmarks in the distorted reference frame model and one or more landmarks of the frame.
[0288] Aspect 66: The method according to any one of Aspects 34 to 65 further includes: determining that the coordinate values of fewer than a threshold number of vertices of a 3D model of a second portion of the object are less than a predetermined coordinate value; and removing one or more vertices of the 3D model that are less than the predetermined coordinate value based on determining that the coordinate values of fewer than a threshold number of vertices of the 3D model are less than the predetermined coordinate value.
[0289] Aspect 67: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the operation according to any one of aspects 1 to 67.
[0290] Aspect 68: An apparatus for digital imaging, the apparatus comprising a unit for performing operations according to any one of aspects 1 to 67.
[0291] Aspect 69: An apparatus for generating one or more models, comprising a memory and one or more processors (e.g., implemented in a circuit) coupled to the memory. The memory may be configured to store data, such as one or more frames, one or more three-dimensional models, and / or other data. The one or more processors are configured to: generate a three-dimensional (3D) model of a person's head based on one or more frames depicting a person; generate a mask for the one or more frames, the mask including indications of one or more regions of the person; generate a 3D base model based on a 3D model of a first portion of the person and the mask, the 3D base model representing the person's head and the person's hair; and generate a 3D model of the person's hair based on the mask and the 3D base model.
[0292] Aspect 70: The apparatus according to aspect 69, wherein the one or more processors are configured to: select the one or more frames from a frame sequence as keyframes, wherein each keyframe depicts the person from a different angle.
[0293] Aspect 71: An apparatus according to any one of aspects 69 or 70, wherein the one or more processors are configured to: determine that a first keyframe does not meet a quality threshold; output feedback to facilitate the localization of the person corresponding to the first keyframe; capture at least one frame based on the feedback; and insert a frame from the at least one frame into the keyframe.
[0294] Aspect 72: The apparatus according to any one of aspects 69 to 71, wherein, in order to generate the mask for the one or more frames, the one or more processors are configured to: segment each of the one or more frames into one or more regions; and generate a mask for each of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.
[0295] Aspect 73: The apparatus according to any one of aspects 69 to 72, wherein, in order to generate the 3D base model, the one or more processors are configured to: project each vertex of an initial 3D model onto a mask associated with the frame based on pose information associated with a frame in the one or more frames; determine whether each vertex of the 3D model of the head is located within a head region of the mask associated with the frame; and extract the 3D base model based on the vertices of the 3D model of the head being within the head region of the mask associated with the frame.
[0296] Aspect 74: An apparatus according to any one of aspects 69 to 73, wherein the one or more processors are configured to remove the one or more vertices from the 3D base model based on the probability that each of the one or more vertices is outside the head region.
[0297] Aspect 75: The apparatus according to any one of aspects 69 to 74, wherein a 3D model of the second portion of the object is adjacent to at least a portion of a 3D model of the first portion of the object.
[0298] Aspect 76: The apparatus according to any one of aspects 69 to 75, wherein the 3D model of the human hair does not visually collide with the 3D model of the human head.
[0299] Aspect 77: The apparatus according to any one of aspects 69 to 76, wherein, in order to generate a 3D model of the human hair, the one or more processors are configured to: initialize the value of each vertex of the 3D base model to an initial value, wherein the initial value indicates that the corresponding vertex is located within a hair region; project a first vertex of the 3D base model into a keyframe in the one or more frames; determine whether the vertex of the 3D base model is projected into a first mask or a second mask of the first keyframe, wherein the first mask corresponds to a face region and the second mask corresponds to the hair region; and adjust the value of each vertex based on whether the corresponding vertex is projected into the first mask or the second mask.
[0300] Aspect 78: The apparatus according to aspect 77, wherein when the first vertex is projected onto the second mask (corresponding to the hair region), the value of the first vertex is increased, and wherein when the first vertex is projected onto the first mask (corresponding to the face region), the value of the first vertex is decreased.
[0301] Aspect 79: The apparatus according to any one of Aspects 67 to 78, wherein the one or more processors are configured to: determine an average probability based on the value of each vertex, wherein the probability that a vertex corresponds to a 3D model of the hair is based on a comparison of the value of the vertex with the average probability.
[0302] Aspect 80: An apparatus according to any one of aspects 67 to 79, wherein the one or more processors are configured to: perform pose refinement of pose information associated with frames in the one or more frames.
[0303] Aspect 81: The apparatus according to aspect 80, wherein, in order to perform pose refinement of pose information associated with the frame, the one or more processors are configured to: minimize the difference between one or more landmarks of the distorted reference frame model and one or more landmarks of the frame.
[0304] Aspect 82: An apparatus according to any one of aspects 67 to 81, wherein the one or more processors are configured to: determine that the coordinate values of fewer than a threshold number of vertices of the 3D model of the human hair are less than a predetermined coordinate value; and based on determining that the coordinate values of fewer than a threshold number of vertices of the 3D model are less than the predetermined coordinate value, remove one or more vertices of the 3D model that are less than the predetermined coordinate value.
[0305] Aspect 83: A method for generating one or more models, comprising: generating a three-dimensional (3D) model of a person's head based on one or more frames depicting a person; generating a mask for the one or more frames, the mask including indications of one or more regions of the person; generating a 3D base model based on a 3D model of a first portion of the person and the mask, the 3D base model representing the person's head and the person's hair; and generating a 3D model of the person's hair based on the mask and the 3D base model.
[0306] Aspect 84: The method according to aspect 83 further includes: selecting one or more frames from a frame sequence as keyframes, wherein each keyframe depicts the person from a different angle.
[0307] Aspect 85: The method according to any one of aspects 83 to 84 further includes: determining that a first keyframe does not meet a quality threshold; outputting feedback to facilitate the localization of the person corresponding to the first keyframe; capturing at least one frame based on the feedback; and inserting a frame from the at least one frame into the keyframe.
[0308] Aspect 86: The method according to any one of aspects 83 to 85, wherein generating the mask for the one or more frames comprises: segmenting each of the one or more frames into one or more regions; and generating a mask for each of the one or more frames, wherein the mask for each frame includes an indication of the one or more regions.
[0309] Aspect 87: The method according to any one of Aspects 83 to 86, wherein generating the 3D base model comprises: projecting each vertex of an initial 3D model onto a mask associated with the frame based on pose information associated with a frame in one or more frames; determining whether each vertex of the 3D model of the head is located within a head region of the mask associated with the frame; and extracting the 3D base model based on the vertices of the 3D model of the head being within the head region of the mask associated with the frame.
[0310] Aspect 88: The method according to any one of aspects 83 to 87 further includes: removing the one or more vertices from the 3D base model based on the probability that each of the one or more vertices is outside the head region.
[0311] Aspect 89: The method according to any one of aspects 83 to 88, wherein the 3D model of the second part of the object is adjacent to at least a portion of the 3D model of the first part of the object.
[0312] Aspect 90: The method according to any one of aspects 83 to 89, wherein the 3D model of the person's hair does not visually collide with the 3D model of the person's head.
[0313] Aspect 91: The method according to any one of Aspects 83 to 90, wherein generating a 3D model of the human hair comprises: initializing the value of each vertex of the 3D base model to an initial value, wherein the initial value indicates that the corresponding vertex is located within a hair region; projecting a first vertex of the 3D base model onto a keyframe in one or more frames; determining whether the vertex of the 3D base model is projected onto a first mask or a second mask of the first keyframe, wherein the first mask corresponds to a face region and the second mask corresponds to the hair region; and adjusting the value of each vertex based on whether the corresponding vertex is projected onto the first mask or the second mask.
[0314] Aspect 92: According to the method of aspect 91, wherein when the first vertex is projected onto the second mask (corresponding to the hair region), the value of the first vertex is increased, and wherein when the first vertex is projected onto the first mask (corresponding to the face region), the value of the first vertex is decreased.
[0315] Aspect 93: The method according to any one of aspects 83 to 92 further includes: determining an average probability based on the value of each vertex, wherein the probability that a vertex corresponds to a 3D model of the hair is based on a comparison of the value of the vertex with the average probability.
[0316] Aspect 94: The method according to any one of aspects 83 to 92 further includes: performing pose refinement of pose information associated with the frames in the one or more frames.
[0317] Aspect 95: According to the method of aspect 94, wherein performing pose refinement of pose information associated with the frame includes: minimizing the difference between one or more landmarks of the distorted reference frame model and one or more landmarks of the frame.
[0318] Aspect 96: The method according to any one of aspects 83 to 95 further includes: determining that the coordinate values of fewer than a threshold number of vertices of the 3D model of the human hair are less than a predetermined coordinate value; and removing one or more vertices of the 3D model that are less than the predetermined coordinate value based on determining that the coordinate values of fewer than a threshold number of vertices of the 3D model are less than the predetermined coordinate value.
[0319] Aspect 97: A computer-readable medium comprising at least one instruction for causing a computer or processor to perform operations according to any one of aspects 69 to 96.
[0320] Aspect 93: An apparatus for generating one or more images, the apparatus comprising a unit for performing operations according to any one of aspects 69 to 96.
[0321] Aspect 94: An apparatus for generating one or more models. The apparatus includes at least one memory and at least one processor coupled to said at least one memory. The at least one processor is configured to perform operations according to any one of aspects 1 to 66 and any one of aspects 69 to 96.
[0322] Aspect 95: A method for generating one or more models, the method comprising operations according to any one of aspects 1 to 66 and any one of aspects 69 to 96.
[0323] Aspect 96: A computer-readable medium comprising at least one instruction for causing a computer or processor to perform operations according to any one of aspects 1 to 66 and any one of aspects 69 to 96.
[0324] Aspect 97: An apparatus for generating one or more models, the apparatus comprising a unit for performing operations according to any one of aspects 1 to 66 and any one of aspects 69 to 96.
Claims
1. An apparatus for generating one or more models, comprising: Memory; as well as One or more processors coupled to the memory, the one or more processors being configured to: A three-dimensional 3D model of the first part of the object is generated based on one or more frames depicting the object. Perform pose refinement of the pose information associated with the frames in the one or more frames depicting the object; A mask is generated for the one or more frames, the mask including indications for one or more regions in the one or more frames corresponding to a first portion of the object and one or more regions in the one or more frames corresponding to a second portion of the object; A 3D base model is generated based on the 3D model of the first part of the object and the mask, the 3D base model representing the first part of the object and the second part of the object; as well as Based on the mask and the 3D base model, a 3D model of the second part of the object is generated, at least in part, by removing one or more vertices corresponding to the first part of the object from the 3D base model.
2. The apparatus according to claim 1, wherein, The 3D model in the second part corresponds to an item that is part of the object.
3. The apparatus according to claim 1, wherein, The object is a person, the first part of the object corresponds to the person's head, and the second part of the object corresponds to the hair on the person's head.
4. The apparatus according to claim 1, wherein, The 3D model of the second part corresponds to an item that is at least one of the following: separable from the object and movable relative to the object.
5. The apparatus according to claim 1, wherein, The object is a person, the first part of the object corresponds to a body area of the person, and the second part of the object corresponds to an accessory or clothing worn by the person.
6. The apparatus according to claim 1, wherein, The 3D model of the second part of the object is adjacent to at least a portion of the 3D model of the first part of the object.
7. The apparatus according to claim 1, wherein, The one or more processors are configured to: Each of the one or more frames is divided into one or more regions; as well as A mask is generated for each of the one or more frames, wherein the mask for each frame includes the indication.
8. The apparatus according to claim 1, wherein, The one or more processors are configured to: Determine the rasterization of the 3D model of the first portion of the object in the frame, the union of the first mask for the first region of the object and the second mask for the second region of the object; as well as A mask for the frame is generated based on the determined union.
9. The apparatus according to claim 8, wherein, The first region is the facial region of the object, and the second region is the hair region of the object.
10. The apparatus according to claim 1, wherein, The one or more processors are configured to: Based on the pose information associated with a frame in one or more of the frames, each vertex of the initial 3D model is projected onto a mask associated with the frame; Determine whether each vertex of the 3D model in the first part is located within a first region of the mask associated with the frame; as well as Based on the vertices of the 3D model in the first part, the 3D base model is extracted within the first region of the mask associated with the frame.
11. The apparatus according to claim 10, wherein, The object is a person, and the first region corresponds to the person's facial region and the person's hair region.
12. The apparatus according to claim 10, wherein, The object is a person, and the first area corresponds to the person's body area and the area of the clothing worn by the person.
13. The apparatus according to claim 10, wherein, The one or more processors are configured to: The one or more vertices are removed from the 3D base model based on the probability of each vertex in a region of the one or more regions of the frame from the one or more frames.
14. The apparatus according to claim 1, wherein, The one or more processors are configured to: Animations are generated in the application using the 3D model of the first part and the 3D model of the second part, wherein the object includes a person, the 3D model of the first part corresponds to the person's head, and the 3D model of the second part corresponds to the person's hair.
15. The apparatus according to claim 14, wherein, The application includes functionality for sending and receiving at least one of audio and text.
16. The apparatus according to claim 14, wherein, The 3D model in the first part and the 3D model in the second part depict the user of the application.
17. The apparatus according to claim 1, wherein, The one or more processors are configured to: Receive input corresponding to the selection of at least one graphical control for modifying the 3D model of the second part; as well as The 3D model in the second part is modified based on the received input.
18. The apparatus according to claim 1, wherein, In order to perform pose refinement of the pose information associated with the frame, the one or more processors are configured to: Minimize the difference between one or more landmarks in the distorted reference frame model and one or more landmarks in the frame.
19. The apparatus according to claim 1, wherein, The one or more processors are configured to: The coordinate values of fewer than a threshold number of vertices in the 3D model of the second part of the object are determined to be less than a predetermined coordinate value. as well as Based on the determination that the coordinate values of fewer than a threshold number of vertices in the 3D model are less than the predetermined coordinate values, one or more vertices in the 3D model whose coordinate values are less than the predetermined coordinate values are removed.
20. A method for generating one or more models, comprising: A three-dimensional 3D model of the first part of the object is generated based on one or more frames depicting the object. Perform pose refinement of the pose information associated with the frames in the one or more frames depicting the object; A mask is generated for the one or more frames, the mask including indications for one or more regions in the one or more frames corresponding to a first portion of the object and one or more regions in the one or more frames corresponding to a second portion of the object; A 3D base model is generated based on the 3D model of the first part of the object and the mask, the 3D base model representing the first part of the object and the second part of the object; as well as Based on the mask and the 3D base model, a 3D model of the second part of the object is generated, at least in part, by removing one or more vertices corresponding to the first part of the object from the 3D base model.
21. The method according to claim 20, wherein, Generating the mask for the one or more frames includes: Each of the one or more frames is divided into one or more regions; and A mask is generated for each of the one or more frames, wherein the mask for each frame includes the indication.
22. The method according to claim 20, wherein, Generating the mask for the one or more frames includes: Determine the rasterization of the 3D model of the first portion of the object in the frame, the union between the first mask for the first region of the object and the second mask for the second region of the object; and A mask for the frame is generated based on the determined union.
23. The method of claim 20, wherein, Generating the 3D base model includes: Based on the pose information associated with a frame in one or more of the frames, each vertex of the initial 3D model is projected onto a mask associated with the frame; Determine whether each vertex of the 3D model in the first part is located within a first region of the mask associated with the frame; and Based on the vertices of the 3D model in the first part, the 3D base model is extracted within the first region of the mask associated with the frame.
24. The method of claim 23, further comprising: The one or more vertices are removed from the 3D base model based on the probability of each vertex in a region of the one or more regions of the frame from the one or more frames.
25. The method of claim 20, further comprising: Animations are generated in the application using the 3D model of the first part and the 3D model of the second part, wherein the object includes a person, the 3D model of the first part corresponds to the person's head, and the 3D model of the second part corresponds to the person's hair.
26. The method of claim 20, further comprising: Receive input corresponding to the selection of at least one graphical control for modifying the 3D model of the second part; as well as The 3D model in the second part is modified based on the received input.
27. The method of claim 20, wherein, Performing pose refinement of the pose information associated with the frame includes: Minimize the difference between one or more landmarks in the distorted reference frame model and one or more landmarks in the frame.
28. The method of claim 20, further comprising: The coordinate values of fewer than a threshold number of vertices in the 3D model of the second part of the object are determined to be less than a predetermined coordinate value. as well as Based on the determination that the coordinate values of fewer than a threshold number of vertices in the 3D model are less than the predetermined coordinate values, one or more vertices in the 3D model whose coordinate values are less than the predetermined coordinate values are removed.
29. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform the steps of the method according to any one of claims 20-28.
30. An apparatus for generating one or more models, comprising: Units for performing the steps of the method according to any one of claims 20-28.
Citation Information
Patent Citations
Object reconstruction using media data
US11354860B1
Method for single-view hair modeling and portrait editing
US20140233849A1