Interoperable avatars

The implementation of a common latent representation and mapping systems addresses interoperability issues between blend shape-based and neural code-based 3D avatar technologies, enabling high-quality, seamless animation across diverse devices and systems.

WO2025159799A1PCT designated stage expired Publication Date: 2025-07-31QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/047106
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-23
Filing Date
2024-09-17
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current technologies face challenges in interoperability between blend shape-based and neural code-based approaches for generating 3D avatars, as well as inconsistencies across different neural code representations, leading to difficulties in seamless communication and animation across various devices and systems.

Method used

Implementing a common latent representation and forward/backward mapping systems to enable communication between blend shape-based and neural code-based approaches, allowing for interoperable avatars by projecting neural codes to a common latent space and back-projecting them to neural codes, using machine learning models for mapping.

Benefits of technology

Facilitates seamless interoperability of 3D avatars across different devices and systems, ensuring high-quality animation and fidelity without loss, by establishing a common format for various avatar representations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024047106_31072025_PF_FP_ABST
    Figure US2024047106_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are systems, apparatuses, processes, and computer-readable media for generating three-dimensional (3D) models, such as avatars. For example, a first device can receive, from a second device, a common latent representation of a 3D model of a user. The first device can back-project, using a backward mapper, the common latent representation to neural codes representing the 3D model of the user. The first device can generate the 3D model of the user using the neural codes.
Need to check novelty before this filing date? Find Prior Art

Description

PATENT Qualcomm Ref. No.2402226WO INTEROPERABLE AVATARS FIELD

[0001] The present disclosure generally relates to generating three-dimensional (3D) models, such as avatars. For example, aspects of the present disclosure relate to a technique for interoperable avatars. BACKGROUND

[0002] Many devices and systems allow a scene to be captured by generating frames (also referred to as images) and / or video data (including multiple images or frames) of the scene. For example, a camera or a computing device including a camera (e.g., a mobile device such as a mobile telephone or smartphone including one or more cameras) can capture a sequence of frames of a scene. The frames and / or video data can be captured and processed by such devices and systems (e.g., mobile devices, IP cameras, etc.) and can be output for consumption (e.g., displayed on the device and / or other device). In some cases, the frame and / or video data can be captured by such devices and systems and output for processing and / or consumption by other devices.

[0003] A frame can be processed (e.g., using object detection, recognition, segmentation, etc.) to determine objects that are present in the frame, which can be useful for many applications. For instance, a model can be determined for representing an object in a frame and can be used to facilitate effective operation of various systems. Examples of such applications and systems include augmented reality (AR), robotics, automotive and aviation, three- dimensional scene understanding, object grasping, object tracking, in addition to many other applications and systems. SUMMARY

[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose 1 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO to present certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.

[0005] Disclosed are systems, apparatuses, methods and computer-readable media for interoperable avatars. According to at least one example, A first device for generating one or more three-dimensional (3D) models, the first device comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive, from a second device, a common latent representation of a 3D model of a user; back-project, using a backward mapper, the common latent representation to neural codes representing the 3D model of the user; and generate the 3D model of the user using the neural codes.

[0006] In another illustrative example, a method is provided for generating one or more 3D models. The method includes: receiving, from a second device, a common latent representation of a 3D model of a user; back-projecting, by a backward mapper, the common latent representation to neural codes representing the 3D model of the user; and generating the 3D model of the user using the neural codes.

[0007] In another illustrative example, a non-transitory computer-readable storage medium is provided having instructions stored thereon which, when executed by at least one processor, causes the at least one processor to: receive, from a second device, a common latent representation of a 3D model of a user; back-project, using a backward mapper, the common latent representation to neural codes representing the 3D model of the user; and generate the 3D model of the user using the neural codes.

[0008] In another illustrative example, an apparatus is provided for generating one or more 3D models. The apparatus includes: means for receiving, from a second device, a common latent representation of a 3D model of a user; means for back-projecting the common latent representation to neural codes representing the 3D model of the user; and means for generating the 3D model of the user using the neural codes.

[0009] Aspects generally include a method, apparatus, system, computer program product, non-transitory computer-readable medium, user equipment, base station, wireless 2 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO communication device, and / or processing system as substantially described herein with reference to and as illustrated by the drawings and specification.

[0010] In some aspects, each of the apparatuses described herein is, can be part of, or can include an audio device, a mobile device, a smart or connected device, a camera system, and / or an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device). In some examples, each of the apparatuses can include or be part of a vehicle, a mobile device (e.g., a mobile telephone or so-called “smart phone” or other mobile device), a wearable device, a personal computer, a laptop computer, a tablet computer, a server computer, a robotics device or system, an aviation system, or other device. In some aspects, the apparatus includes an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each of the apparatuses can include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each of the apparatuses can include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each of the apparatuses can include one or more sensors. In some cases, the one or more sensors can be used for determining a location of the apparatuses, a state of the apparatuses (e.g., a tracking state, an operating state, a temperature, a humidity level, and / or other state), and / or for other purposes.

[0011] Some aspects include a device having a processor configured to perform one or more operations of any of the methods summarized above. Further aspects include processing devices for use in a device configured with processor-executable instructions to perform operations of any of the methods summarized above. Further aspects include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processor of a device to perform operations of any of the methods summarized above. Further aspects include a device having means for performing functions of any of the methods summarized above.

[0012] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The 3 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO conception and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages, will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims.

[0013] While aspects are described in the present disclosure by illustration to some examples, those skilled in the art will understand that such aspects may be implemented in many different arrangements and scenarios. Techniques described herein may be implemented using different platform types, devices, systems, shapes, sizes, and / or packaging arrangements. For example, some aspects may be implemented via integrated chip implementations or other non-module-component based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchasing devices, medical devices, and / or artificial intelligence devices). Aspects may be implemented in chip-level components, modular components, non-modular components, non-chip-level components, device-level components, and / or system-level components. Devices incorporating described aspects and features may include additional components and features for implementation and practice of claimed and described aspects. For example, transmission and reception of wireless signals may include one or more components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders, and / or summers). It is intended that aspects described herein may be practiced in a wide variety of devices, components, systems, distributed arrangements, and / or end-user devices of varying size, shape, and constitution.

[0014] Other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art based on the accompanying drawings and detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim. 4 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0015] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Illustrative aspects of the present application are described in detail below with reference to the following figures:

[0017] FIG. 1 is a diagram illustrating an example of an extended reality (XR) system, in accordance with some aspects of the present disclosure.

[0018] FIG.2 is a diagram illustrating an example of a three-dimensional (3D) collaborative virtual environment, in accordance with some aspects of the present disclosure.

[0019] FIG.3 is an image with a virtual representation (an avatar) of a user, in accordance with some aspects of the present disclosure.

[0020] FIG.4 is a diagram illustrating another example of an XR system, in accordance with some aspects of the present disclosure.

[0021] FIG. 5 is a diagram illustrating an example configuration of a client device, in accordance with some aspects of the present disclosure.

[0022] FIG.6 is a diagram illustrating an example of a normal map, an albedo map, and a specular reflection map, in accordance with some aspects of the present disclosure.

[0023] FIG. 7 is a diagram illustrating an example of one technique for performing avatar animation, in accordance with some aspects of the present disclosure.

[0024] FIG.8 is a diagram illustrating an example of performing facial animation with blend shapes, in accordance with some aspects of the present disclosure.

[0025] FIG. 9 is a diagram illustrating an example of a system that can generate a 3D Morphable Model (3DMM) face mesh, in accordance with some aspects of the present disclosure. 5 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0026] FIG. 10 is a diagram illustrating an example of animating an avatar, in accordance with some aspects of the present disclosure.

[0027] FIG.11 illustrates an example 3D facial model and corresponding two-dimensional (2D) facial images overlaid with landmarks projected from the 3D facial model, in accordance with some aspects of the present disclosure.

[0028] FIG.12 illustrates an example head mounted XR system with user facing cameras for generating a 3D facial model, in accordance with some aspects of the present disclosure.

[0029] FIG.13 is a diagram illustrating an example of a 3D modeling system, in accordance with some aspects of the present disclosure.

[0030] FIG.14 is a diagram illustrating an example of a system for generating a 3D model of a facial avatar of a face of a user, in accordance with some aspects of the present disclosure.

[0031] FIG. 15 is a diagram illustrating an example of a system for generating a blend shaped-based facial avatar of a face of a user, in accordance with some aspects of the present disclosure.

[0032] FIG. 16 is a diagram illustrating an example of a system for generating a photorealistic facial avatar of a face of a user, in accordance with some aspects of the present disclosure.

[0033] FIG.17 is a diagram illustrating examples of problems with representations, such as inconsistencies across rendering engines, in accordance with some aspects of the present disclosure.

[0034] FIG.18 is a diagram illustrating examples of problems with representations, such as inconsistencies across different modalities, in accordance with some aspects of the present disclosure.

[0035] FIG. 19 is a diagram illustrating an example of a system that provides a common latent representation and forward and backwards mappings between the representations, in accordance with some aspects of the present disclosure. 6 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0036] FIG.20 is a diagram illustrating an example of a solution for inconsistencies across different rendering engines or modalities, in accordance with some aspects of the present disclosure.

[0037] FIG.21 is a diagram illustrating examples of base meshes, in accordance with some aspects of the present disclosure.

[0038] FIG. 22 is a diagram illustrating examples of UV maps, in accordance with some aspects of the present disclosure.

[0039] FIG.23 is a diagram illustrating examples of blend shapes, in accordance with some aspects of the present disclosure.

[0040] FIG. 24 is a flow chart illustrating an example of a process for wireless communications, in accordance with some aspects of the present disclosure.

[0041] FIG.25 is a block diagram illustrating an example computing system, in accordance with some aspects of the present disclosure. DETAILED DESCRIPTION

[0042] Certain aspects of this disclosure are provided below for illustration purposes. Alternate aspects may be devised without departing from the scope of the disclosure. Additionally, well-known elements of the disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of the disclosure. Some of the aspects described herein can be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.

[0043] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the example aspects will provide those skilled in the art with an enabling description for implementing an example aspect. It should be understood that various changes may be made 7 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.

[0044] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage or mode of operation.

[0045] The generation of three-dimensional (3D) models for physical objects can be useful for many systems and applications, such as for extended reality (XR) (e.g., including augmented reality (AR), virtual reality (VR), mixed reality (MR), etc.), robotics, automotive, aviation, 3D scene understanding, object grasping, object tracking, in addition to many other systems and applications. In AR environments, for example, a user may view images (also referred to as frames) that include an integration of artificial or virtual graphics with the user’s natural surroundings. AR applications allow real images to be processed to add virtual objects to the images or to display virtual objects on a see-through display (so that the virtual objects appear to be overlaid over the real-world environment). AR applications can align or register the virtual objects to real-world objects (e.g., as observed in the images) in multiple dimensions. For instance, a real-world object that exists in reality can be represented using a model that resembles or is an exact match of the real-world object. In one example, a model of a virtual airplane representing a real airplane sitting on a runway may be presented by the display of an AR device (e.g., AR glasses, AR head-mounted display (HMD), or other device) while the user continues to view his or her natural surroundings through the display. The viewer may be able to manipulate the model while viewing the real-world scene. In another example, an actual object sitting on a table may be identified and rendered with a model that has a different color or different physical attributes in the AR environment. In some cases, artificial virtual objects that do not exist in reality or computer-generated copies of actual objects or structures of the user’s natural surroundings can also be added to the AR environment.

[0046] There is an increasing number of applications that use face data (e.g., for XR systems, for 3D graphics, for security, among others), leading to a large demand for systems 8 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO with the ability to generate detailed 3D face models (as well as 3D models of other objects) in an efficient and high-quality manner. There also exists a large demand for generating 3D models of other types of objects, such as 3D models of vehicles (e.g., for autonomous driving systems), 3D models of room layouts (e.g., for XR applications, for navigation by devices, robots, etc.), among others.

[0047] For generating a three-dimensional (3D) model, such as for an avatar, a linear 3D morphable model (3DMM) can be used to represent the geometry of a user’s head, body, and / or hands for the avatar. A linear 3DMM estimates blend shapes weights. These blend shape weights are sufficient to be able to retarget an avatar in any game engine (e.g., animation of the avatar, such as facial animation, can be performed using blend shapes). Recent approaches extend this model to employ neural-based expression codes (e.g., which can simply be referred to as neural codes or expression codes) to enhance photorealism of the avatar. Switching between these two approaches (e.g., a blend shapes approach and a neural codes approach) along with the variants of the neural representations (e.g., different types of neural codes) can be a problem as they are not compatible with each other. Currently, a blend shape-based approach cannot directly communicate with a neural code-based approach, and vice versa. Also, as mentioned, neural codes (e.g., from different modalities) can have various incompatible representations, such as efficient geometry-aware 3D (EG3D), Stable Diffusion, and Style generative adversarial network (StyleGAN) representations.

[0048] For example, avatars are used in various types of immersive or virtual environments, such as in immersive / virtual communications to represent a user in a call. The user may use a particular application or service to generate an animatable avatar for the user. During a setup of the call, the user may offer their preferred avatar format (e.g., the format in which the user has captured and stored the avatar representation) to other users. Such a situation can raise inter-operability issues for service providers, given that there is a large number of proprietary avatar formats.

[0049] An advantage that can be achieved by offering the avatar representation in which the user has captured / stored their avatar representation can be in the optimal usage of the animation streams and base model characteristics during the avatar animation. For instance, an 9 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO animation process that matches the avatar format will ensure no fidelity loss and high quality. As an example, an avatar format may support a blend-shape based animation with the ability to animate several blend shapes (e.g., 87 blend shapes or other number of blend shapes) for an accurate facial expression. If a receiver supports the corresponding base avatar and the animation stream formats, then the receiver can generate and output a high-fidelity avatar animation.

[0050] However, to ensure interoperability of the service, a format that can serve as a common denominator for all avatar representation formats is needed. As such, improved systems and techniques that allow for a blend shape-based approach to directly communicate with a neural code-based approach, and vice versa, as well as for different neural code representations (e.g., from different modalities) to be able to directly communicate with each other can be beneficial.

[0051] In some aspects of the present disclosure, systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively referred to herein as “systems and techniques”) are described herein for interoperable avatars.

[0052] Various aspects relate generally to generating 3D models, such as avatars. Some aspects more specifically relate to systems and techniques that provide solutions for interoperable avatars. In one or more aspects, the systems and techniques allow for communication between a blend shape-based approach and a neural code-based approach (and vice versa) and between different neural representations (e.g., different types of neural codes) by defining a forward and backward mapping between the approaches and / or representations, and provide a common latent space representation between different systems (e.g., systems with different OEMs and, as such, with different types of neural codes). In one or more examples, by employing a forward mapper, the systems and techniques can project neural codes to a common latent representation (e.g., blend shapes). By adding a backward mapper, the systems and techniques can back-project the common latent representation (e.g., blend shapes) to the neural codes. In some examples, the systems and techniques can implement a common latent representation or format (e.g., to have a common blend shape basis and base mesh). 10 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0053] The common latent representation or format can serve as a common denominator for all avatar representation formats. In some cases, the common latent representation format can be offered as a fallback by all participants that wish to participate in an immersive session of a collaborative environment (e.g., a collaborative call). For instance, in the case of multiple participants in the session, a media gateway may be responsible for injecting the fallback representation format into offers and answers by the different participants. The media gateway can then be responsible for converting between the different formats whenever needed.

[0054] The forward mapper can convert any proprietary avatar representation format into the common latent representation. This can apply to a base model format and to animation streams (e.g., represented by neural codes, blend shapes, or other representation) that animate the base model. For example, a common format can be a 3DMM representation with blend shape weights as animation streams and a proprietary format can use neural codes as animation streams.

[0055] In one or more examples, during operation for generating 3D models at a first device (e.g., such as an HMD or mobile phone), the first device can receive, from a second device (e.g., an HMD or a mobile phone), a common latent representation of a 3D model of a user. A backward mapper, of the first device, can back-project the common latent representation to neural codes representing the 3D model of the user. The first device can generate the 3D model of the user using the neural codes.

[0056] In one or more examples, a forward mapper of the first device can project the neural codes to the common latent representation. The first device can transmit, to the second device or a third device, the common latent representation. In some examples, the forward mapper is a machine learning model trained to map the neural codes to the common latent representation. In one or more examples, the backward mapper is a machine learning model trained to map the common latent representation to the neural codes.

[0057] In one or more examples, the common latent representation is a space designed to represent avatars. In some examples, the space comprises at least one of efficient geometry- aware 3D (EG3D) latent variables, Stable Diffusion latent variables, Style generative adversarial network (StyleGAN) latent variables, or common blend shapes with a base mesh. 11 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0058] In some examples, a machine learning encoder of the first device can generate the neural codes by encoding one or more two-dimensional (2D) images of the user. In one or more examples, the first device is a head mounted device (HMD) or a mobile phone. In some examples, the 3D model is an avatar of the user. In one or more examples, the 3D model is a facial avatar of a face of the user.

[0059] Additional aspects of the present disclosure are described in more detail below. Illustrative aspects are also described in Appendix A accompanying this disclosure.

[0060] FIG. 1 illustrates an example of an extended reality system 100. As shown, the extended reality system 100 includes a device 105, a network 120, and a communication link 125. In some cases, the device 105 may be an extended reality (XR) device, which may generally implement aspects of extended reality, including virtual reality (VR), augmented reality (AR), mixed reality (MR), etc. Systems including a device 105, a network 120, or other elements in extended reality system 100 may be referred to as extended reality systems.

[0061] The device 105 may overlay virtual objects with real-world objects in a view 130. For example, the view 130 may generally refer to visual input to a user 110 via the device 105, a display generated by the device 105, a configuration of virtual objects generated by the device 105, etc. For example, view 130-A may refer to visible real-world objects (also referred to as physical objects) and visible virtual objects, overlaid on or coexisting with the real-world objects, at some initial time. View 130-B may refer to visible real-world objects and visible virtual objects, overlaid on or coexisting with the real-world objects, at some later time. Positional differences in real-world objects (e.g., and thus overlaid virtual objects) may arise from view 130-A shifting to view 130-B at 135 due to head motion 115. In another example, view 130-A may refer to a completely virtual environment or scene at the initial time and view 130-B may refer to the virtual environment or scene at the later time.

[0062] Generally, device 105 may generate, display, project, etc. virtual objects and / or a virtual environment to be viewed by a user 110 (e.g., where virtual objects and / or a portion of the virtual environment may be displayed based on user 110 head pose prediction in accordance with the techniques described herein). In some examples, the device 105 may include a transparent surface (e.g., optical glass) such that virtual objects may be displayed on the 12 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO transparent surface to overlay virtual objects on real word objects viewed through the transparent surface. Additionally or alternatively, the device 105 may project virtual objects onto the real-world environment. In some cases, the device 105 may include a camera and may display both real-world objects (e.g., as frames or images captured by the camera) and virtual objects overlaid on displayed real-world objects. In various examples, device 105 may include aspects of a virtual reality headset, smart glasses, a live feed video camera, a GPU, one or more sensors (e.g., such as one or more IMUs, image sensors, microphones, etc.), one or more output devices (e.g., such as speakers, display, smart glass, etc.), etc.

[0063] In some cases, head motion 115 may include user 110 head rotations, translational head movement, etc. The device 105 may update the view 130 of the user 110 according to the head motion 115. For example, the device 105 may display view 130-A for the user 110 before the head motion 115. In some cases, after the head motion 115, the device 105 may display view 130-B to the user 110. The extended reality system (e.g., device 105) may render or update the virtual objects and / or other portions of the virtual environment for display as the view 130- A shifts to view 130-B.

[0064] In some cases, the extended reality system 100 may provide various types of virtual experiences, such as a three-dimensional (3D) gaming experiences, social media experiences, collaborative virtual environment for a group of users (e.g., including the user 110), among others. While some examples provided herein apply to 3D collaborative virtual environments, the systems and techniques described herein apply to any type of virtual environment or experience in which a virtual representation (or avatar) can be used to represent a user or participant of the virtual environment / experience.

[0065] FIG. 2 is a diagram illustrating an example of a 3D collaborative virtual environment 200 in which various users interact with one another in a virtual session via virtual representations (or avatars) of the users in the virtual environment 200. The virtual representations include including a virtual representation 202 of a first user, a virtual representation 204 of a second user, a virtual representation 206 of a third user, a virtual representation 208 of a fourth user, and a virtual representation 210 of a fifth user. Other background information of the virtual environment 200 is also shown, including a virtual 13 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO calendar 212, a virtual web page 214, and a virtual video conference interface 216. The users may visually, audibly, haptically, or otherwise experience the virtual environment from each user’s perspective while interacting with the virtual representations of the other users. For example, the virtual environment 200 is shown from the perspective of the first user (represented by the virtual representation 202).

[0066] FIG.3 is an image 300 illustrating an example of virtual representations of various users, including a virtual representation 302 of one of the users. For instance, the virtual representation 302 may be used in the 3D collaborative virtual environment 200 of FIG.2.

[0067] FIG. 4 is a diagram illustrating an example of a system 400 that can be used to perform the systems and techniques described herein. As shown, the system 400 includes client devices 405, an animation and scene rendering system 410, and storage 415. Although the system 400 illustrates two devices 405, a single animation and scene rendering system 410, a single storage 415, and a single network 420, the present disclosure applies to any system architecture having one or more devices 405, animation and scene rendering systems 410, storage 415, and networks 420. In some cases, the storage 415 may be part of the animation and scene rendering system 410. The devices 405, the animation and scene rendering system 410, and the storage 415 may communicate with each other and exchange information that supports generation of virtual content for XR, such as multimedia packets, multimedia data, multimedia control information, pose prediction parameters, via network 420 using communications links 425. In some cases, a portion of the techniques described herein for providing distributed generation of virtual content may be performed by one or more of the devices 405 and a portion of the techniques may be performed by the animation and scene rendering system 410, or both.

[0068] A device 405 may be an XR device (e.g., a head-mounted display (HMD), XR glasses such as virtual reality (VR) glasses, augmented reality (AR) glasses, etc.), a mobile device (e.g., a cellular phone, a smartphone, a personal digital assistant (PDA), etc.), a wireless communication device, a tablet computer, a laptop computer, and / or other device that supports various types of communication and functional features related to multimedia (e.g., transmitting, receiving, broadcasting, streaming, sinking, capturing, storing, and recording 14 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO multimedia data). A device 405 may, additionally or alternatively, be referred to by those skilled in the art as a user equipment (UE), a user device, a smartphone, a Bluetooth device, a Wi-Fi device, a mobile station, a subscriber station, a mobile unit, a subscriber unit, a wireless unit, a remote unit, a mobile device, a wireless device, a wireless communications device, a remote device, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a user agent, a mobile client, a client, and / or some other suitable terminology. In some cases, the devices 405 may also be able to communicate directly with another device (e.g., using a peer-to-peer (P2P) or device-to-device (D2D) protocol, such as using sidelink communications). For example, a device 405 may be able to receive from or transmit to another device 405 variety of information, such as instructions or commands (e.g., multimedia-related information).

[0069] The devices 405 may include an application 430 and a multimedia manager 435. While the system 400 illustrates the devices 405 including both the application 430 and the multimedia manager 435, the application 430 and the multimedia manager 435 may be an optional feature for the devices 405. In some cases, the application 430 may be a multimedia- based application that can receive (e.g., download, stream, broadcast) from the animation and scene rendering systems 410, storage 415 or another device 405, or transmit (e.g., upload) multimedia data to the animation and scene rendering systems 410, the storage 415, or to another device 405 via using communications links 425.

[0070] The multimedia manager 435 may be part of a general-purpose processor, a digital signal processor (DSP), an image signal processor (ISP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a discrete gate or transistor logic component, a discrete hardware component, or any combination thereof, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described in the present disclosure, and / or the like. For example, the multimedia manager 435 may process multimedia (e.g., image data, video data, audio data) from and / or write multimedia data to a local memory of the device 405 or to the storage 415. 15 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0071] The multimedia manager 435 may also be configured to provide multimedia enhancements, multimedia restoration, multimedia analysis, multimedia compression, multimedia streaming, and multimedia synthesis, among other functionality. For example, the multimedia manager 435 may perform white balancing, cropping, scaling (e.g., multimedia compression), adjusting a resolution, multimedia stitching, color processing, multimedia filtering, spatial multimedia filtering, artifact removal, frame rate adjustments, multimedia encoding, multimedia decoding, and multimedia filtering. By further example, the multimedia manager 435 may process multimedia data to support server-based pose prediction for XR, according to the techniques described herein.

[0072] The animation and scene rendering system 410 may be a server device, such as a data server, a cloud server, a server associated with a multimedia subscription provider, proxy server, web server, application server, communications server, home server, mobile server, edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a network router, any combination thereof, or other server device. The animation and scene rendering system 410 may in some cases include a multimedia distribution platform 440. In some cases, the multimedia distribution platform 440 may be a separate device or system from the animation and scene rendering system 410. The multimedia distribution platform 440 may allow the devices 405 to discover, browse, share, and download multimedia via network 420 using communications links 425, and therefore provide a digital distribution of the multimedia from the multimedia distribution platform 440. As such, a digital distribution may be a form of delivering media content such as audio, video, images, without the use of physical media but over online delivery mediums, such as the Internet. For example, the devices 405 may upload or download multimedia-related applications for streaming, downloading, uploading, processing, enhancing, etc. multimedia (e.g., images, audio, video). The animation and scene rendering system 410 or the multimedia distribution platform 440 may also transmit to the devices 405 a variety of information, such as instructions or commands (e.g., multimedia- related information) to download multimedia-related applications on the device 405.

[0073] The storage 415 may store a variety of information, such as instructions or commands (e.g., multimedia-related information). For example, the storage 415 may store 16 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO multimedia 445, information from devices 405 (e.g., pose information, representation information for virtual representations or avatars of users, such as codes or features related to facial representations, body representations, hand representations, etc., and / or other information). A device 405 and / or the animation and scene rendering system 410 may retrieve the stored data from the storage 415 and / or more send data to the storage 415 via the network 420 using communication links 425. In some examples, the storage 415 may be a memory device (e.g., read only memory (ROM), random access memory (RAM), cache memory, buffer memory, etc.), a relational database (e.g., a relational database management system (RDBMS) or a Structured Query Language (SQL) database), a non-relational database, a network database, an object-oriented database, or other type of database, that stores the variety of information, such as instructions or commands (e.g., multimedia-related information).

[0074] The network 420 may provide encryption, access authorization, tracking, Internet Protocol (IP) connectivity, and other access, computation, modification, and / or functions. Examples of network 420 may include any combination of cloud networks, local area networks (LAN), wide area networks (WAN), virtual private networks (VPN), wireless networks (using 802.11, for example), cellular networks (using third generation (3G), fourth generation (4G), long-term evolved (LTE), or new radio (NR) systems (e.g., fifth generation (5G)), etc. Network 420 may include the Internet.

[0075] The communications links 425 shown in the system 400 may include uplink transmissions from the device 405 to the animation and scene rendering systems 410 and the storage 415, and / or downlink transmissions, from the animation and scene rendering systems 410 and the storage 415 to the device 405. The communications links 425 may transmit bidirectional communications and / or unidirectional communications. In some examples, the communication links 425 may be a wired connection or a wireless connection, or both. For example, the communications links 425 may include one or more connections, including but not limited to, Wi-Fi, Bluetooth, Bluetooth low-energy (BLE), cellular, Z-WAVE, 802.11, peer-to-peer, LAN, wireless local area network (WLAN), Ethernet, FireWire, fiber optic, and / or other connection types related to wireless communication systems. 17 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0076] In some aspects, a user of the device 405 (referred to as a first user) may be participating in a virtual session with one or more other users (including a second user of an additional device). In such examples, the animation and scene rendering systems 410 may process information received from the device 405 (e.g., received directly from the device 405, received from storage 415, etc.) to generate and / or animate a virtual representation (or avatar) for the first user. The animation and scene rendering systems 410 may compose a virtual scene that includes the virtual representation of the user and in some cases background virtual information from a perspective of the second user of the additional device. The animation and scene rendering systems 410 may transmit (e.g., via network 120) a frame of the virtual scene to the additional device. Further details regarding such aspects are provided below.

[0077] FIG.5 is a diagram illustrating an example of a device 500. The device 500 can be implemented as a client device (e.g., device 405 of FIG. 4) or as an animation and scene rendering system (e.g., the animation and scene rendering system 410). As shown, the device 500 includes a central processing unit (CPU) 510 having CPU memory 515, a GPU 525 having GPU memory 530, a display 545, a display buffer 535 storing data associated with rendering, a user interface unit 505, and a system memory 540. For example, system memory 540 may store a GPU driver 520 (illustrated as being contained within CPU 510 as described below) having a compiler, a GPU program, a locally-compiled GPU program, and the like. User interface unit 505, CPU 510, GPU 525, system memory 540, display 545, and extended reality manager 550 may communicate with each other (e.g., using a system bus).

[0078] Examples of CPU 510 include, but are not limited to, a digital signal processor (DSP), general purpose microprocessor, application specific integrated circuit (ASIC), field programmable logic array (FPGA), or other equivalent integrated or discrete logic circuitry. Although CPU 510 and GPU 525 are illustrated as separate units in the example of FIG.5, in some examples, CPU 510 and GPU 525 may be integrated into a single unit. CPU 510 may execute one or more software applications. Examples of the applications may include operating systems, word processors, web browsers, e-mail applications, spreadsheets, video games, audio and / or video capture, playback or editing applications, or other such applications that initiate the generation of image data to be presented via display 545. As illustrated, CPU 510 may include CPU memory 515. For example, CPU memory 515 may represent on-chip storage or 18 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO memory used in executing machine or object code. CPU memory 515 may include one or more volatile or non-volatile memories or storage devices, such as flash memory, a magnetic data media, an optical storage media, etc. CPU 510 may be able to read values from or write values to CPU memory 515 more quickly than reading values from or writing values to system memory 540, which may be accessed, e.g., over a system bus.

[0079] GPU 525 may represent one or more dedicated processors for performing graphical operations. For example, GPU 525 may be a dedicated hardware unit having fixed function and programmable components for rendering graphics and executing GPU applications. GPU 525 may also include a DSP, a general purpose microprocessor, an ASIC, an FPGA, or other equivalent integrated or discrete logic circuitry. GPU 525 may be built with a highly-parallel structure that provides more efficient processing of complex graphic-related operations than CPU 510. For example, GPU 525 may include a plurality of processing elements that are configured to operate on multiple vertices or pixels in a parallel manner. The highly parallel nature of GPU 525 may allow GPU 525 to generate graphic images (e.g., graphical user interfaces and two-dimensional or three-dimensional graphics scenes) for display 545 more quickly than CPU 510.

[0080] GPU 525 may, in some instances, be integrated into a motherboard of device 500. In other instances, GPU 525 may be present on a graphics card or other device or component that is installed in a port in the motherboard of device 500 or may be otherwise incorporated within a peripheral device configured to interoperate with device 500. As illustrated, GPU 525 may include GPU memory 530. For example, GPU memory 530 may represent on-chip storage or memory used in executing machine or object code. GPU memory 530 may include one or more volatile or non-volatile memories or storage devices, such as flash memory, a magnetic data media, an optical storage media, etc. GPU 525 may be able to read values from or write values to GPU memory 530 more quickly than reading values from or writing values to system memory 540, which may be accessed, e.g., over a system bus. That is, GPU 525 may read data from and write data to GPU memory 530 without using the system bus to access off-chip memory. This operation may allow GPU 525 to operate in a more efficient manner by reducing the need for GPU 525 to read and write data via the system bus, which may experience heavy bus traffic. 19 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0081] Display 545 represents a unit capable of displaying video, images, text or any other type of data for consumption by a viewer. In some cases, such as when the device 500 is implemented as an animation and scene rendering system, the device 500 may not include the display 545. The display 545 may include a liquid-crystal display (LCD), a light emitting diode (LED) display, an organic LED (OLED), an active-matrix OLED (AMOLED), or the like. Display buffer 535 represents a memory or storage device dedicated to storing data for presentation of imagery, such as computer-generated graphics, still images, video frames, or the like for display 545. Display buffer 535 may represent a two-dimensional buffer that includes a plurality of storage locations. The number of storage locations within display buffer 535 may, in some cases, generally correspond to the number of pixels to be displayed on display 545. For example, if display 545 is configured to include 640x480 pixels, display buffer 535 may include 640x480 storage locations storing pixel color and intensity information, such as red, green, and blue pixel values, or other color values. Display buffer 535 may store the final pixel values for each of the pixels processed by GPU 525. Display 545 may retrieve the final pixel values from display buffer 535 and display the final image based on the pixel values stored in display buffer 535.

[0082] User interface unit 505 represents a unit with which a user may interact with or otherwise interface to communicate with other units of device 500, such as CPU 510. Examples of user interface unit 505 include, but are not limited to, a trackball, a mouse, a keyboard, and other types of input devices. User interface unit 505 may also be, or include, a touch screen and the touch screen may be incorporated as part of display 545.

[0083] System memory 540 may include one or more computer-readable storage media. Examples of system memory 540 include, but are not limited to, a random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disc storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer or a processor. System memory 540 may store program modules and / or instructions that are accessible for execution by CPU 510. Additionally, system memory 540 may store user 20 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO applications and application surface data associated with the applications. System memory 540 may in some cases store information for use by and / or information generated by other components of device 500. For example, system memory 540 may act as a device memory for GPU 525 and may store data to be operated on by GPU 555 as well as data resulting from operations performed by GPU 525

[0084] In some examples, system memory 540 may include instructions that cause CPU 510 or GPU 525 to perform the functions ascribed to CPU 510 or GPU 525 in aspects of the present disclosure. System memory 540 may, in some examples, be considered as a non- transitory storage medium. The term “non-transitory” should not be interpreted to mean that system memory 540 is non-movable. As one example, system memory 540 may be removed from device 500 and moved to another device. As another example, a system memory substantially similar to system memory 540 may be inserted into device 500. In certain examples, a non-transitory storage medium may store data that can, over time, change (e.g., in RAM).

[0085] System memory 540 may store a GPU driver 520 and compiler, a GPU program, and a locally-compiled GPU program. The GPU driver 520 may represent a computer program or executable code that provides an interface to access GPU 525. CPU 510 may execute the GPU driver 520 or portions thereof to interface with GPU 525 and, for this reason, GPU driver 520 is shown in the example of FIG.5 within CPU 510. GPU driver 520 may be accessible to programs or other executables executed by CPU 510, including the GPU program stored in system memory 540. Thus, when one of the software applications executing on CPU 510 requires graphics processing, CPU 510 may provide graphics commands and graphics data to GPU 525 for rendering to display 545 (e.g., via GPU driver 520).

[0086] In some cases, the GPU program may include code written in a high level (HL) programming language, e.g., using an application programming interface (API). Examples of APIs include Open Graphics Library (“OpenGL”), DirectX, Render-Man, WebGL, or any other public or proprietary standard graphics API. The instructions may also conform to so- called heterogeneous computing libraries, such as Open-Computing Language (“OpenCL”), DirectCompute, etc. In general, an API includes a predetermined, standardized set of 21 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO commands that are executed by associated hardware. API commands allow a user to instruct hardware components of a GPU 525 to execute commands without user knowledge as to the specifics of the hardware components. In order to process the graphics rendering instructions, CPU 510 may issue one or more rendering commands to GPU 525 (e.g., through GPU driver 520) to cause GPU 525 to perform some or all of the rendering of the graphics data. In some examples, the graphics data to be rendered may include a list of graphics primitives (e.g., points, lines, triangles, quadrilaterals, etc.).

[0087] The GPU program stored in system memory 540 may invoke or otherwise include one or more functions provided by GPU driver 520. CPU 510 generally executes the program in which the GPU program is embedded and, upon encountering the GPU program, passes the GPU program to GPU driver 520. CPU 510 executes GPU driver 520 in this context to process the GPU program. That is, for example, GPU driver 520 may process the GPU program by compiling the GPU program into object or machine code executable by GPU 525. This object code may be referred to as a locally-compiled GPU program. In some examples, a compiler associated with GPU driver 520 may operate in real-time or near-real-time to compile the GPU program during the execution of the program in which the GPU program is embedded. For example, the compiler generally represents a unit that reduces HL instructions defined in accordance with a HL programming language to low-level (LL) instructions of a LL programming language. After compilation, these LL instructions are capable of being executed by specific types of processors or other types of hardware, such as FPGAs, ASICs, and the like (including, but not limited to, CPU 510 and GPU 525).

[0088] In the example of FIG. 5, the compiler may receive the GPU program from CPU 510 when executing HL code that includes the GPU program. That is, a software application being executed by CPU 510 may invoke GPU driver 520 (e.g., via a graphics API) to issue one or more commands to GPU 525 for rendering one or more graphics primitives into displayable graphics images. The compiler may compile the GPU program to generate the locally-compiled GPU program that conforms to a LL programming language. The compiler may then output the locally-compiled GPU program that includes the LL instructions. In some examples, the LL instructions may be provided to GPU 525 in the form of a list of drawing primitives (e.g., triangles, rectangles, etc.). 22 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0089] The LL instructions (e.g., which may alternatively be referred to as primitive definitions) may include vertex specifications that specify one or more vertices associated with the primitives to be rendered. The vertex specifications may include positional coordinates for each vertex and, in some instances, other attributes associated with the vertex, such as color coordinates, normal vectors, and texture coordinates. The primitive definitions may include primitive type information, scaling information, rotation information, and the like. Based on the instructions issued by the software application (e.g., the program in which the GPU program is embedded), GPU driver 520 may formulate one or more commands that specify one or more operations for GPU 525 to perform in order to render the primitive. When GPU 525 receives a command from CPU 510, it may decode the command and configure one or more processing elements to perform the specified operation and may output the rendered data to display buffer 535.

[0090] GPU 525 may receive the locally-compiled GPU program, and then, in some instances, GPU 525 renders one or more images and outputs the rendered images to display buffer 535. For example, GPU 525 may generate a number of primitives to be displayed at display 545. Primitives may include one or more of a line (including curves, splines, etc.), a point, a circle, an ellipse, a polygon (e.g., a triangle), or any other two-dimensional primitive. The term “primitive” may also refer to three-dimensional primitives, such as cubes, cylinders, sphere, cone, pyramid, torus, or the like. Generally, the term “primitive” refers to any basic geometric shape or element capable of being rendered by GPU 525 for display as an image (or frame in the context of video data) via display 545. GPU 525 may transform primitives and other attributes (e.g., that define a color, texture, lighting, camera configuration, or other aspect) of the primitives into a so-called “world space” by applying one or more model transforms (which may also be specified in the state data). Once transformed, GPU 525 may apply a view transform for the active camera (which again may also be specified in the state data defining the camera) to transform the coordinates of the primitives and lights into the camera or eye space. GPU 525 may also perform vertex shading to render the appearance of the primitives in view of any active lights. GPU 525 may perform vertex shading in one or more of the above model, world, or view space. 23 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0091] Once the primitives are shaded, GPU 525 may perform projections to project the image into a canonical view volume. After transforming the model from the eye space to the canonical view volume, GPU 525 may perform clipping to remove any primitives that do not at least partially reside within the canonical view volume. For example, GPU 525 may remove any primitives that are not within the frame of the camera. GPU 525 may then map the coordinates of the primitives from the view volume to the screen space, effectively reducing the three-dimensional coordinates of the primitives to the two-dimensional coordinates of the screen. Given the transformed and projected vertices defining the primitives with their associated shading data, GPU 525 may then rasterize the primitives. Generally, rasterization may refer to the task of taking an image described in a vector graphics format and converting it to a raster image (e.g., a pixelated image) for output on a video display or for storage in a bitmap file format.

[0092] A GPU 525 may include a dedicated fast bin buffer (e.g., a fast memory buffer, such as GMEM, which may be referred to by GPU memory 530). As discussed herein, a rendering surface may be divided into bins. In some cases, the bin size is determined by format (e.g., pixel color and depth information) and render target resolution divided by the total amount of GMEM. The number of bins may vary based on device 500 hardware, target resolution size, and target display format. A rendering pass may draw (e.g., render, write, etc.) pixels into GMEM (e.g., with a high bandwidth that matches the capabilities of the GPU). The GPU 525 may then resolve the GMEM (e.g., burst write blended pixel values from the GMEM, as a single layer, to a display buffer 535 or a frame buffer in system memory 540). Such may be referred to as bin-based or tile-based rendering. When all bins are complete, the driver may swap buffers and start the binning process again for a next frame.

[0093] For example, GPU 525 may implement a tile-based architecture that renders an image or rendering target by breaking the image into multiple portions, referred to as tiles or bins. The bins may be sized based on the size of GPU memory 530 (e.g., which may alternatively be referred to herein as GMEM or a cache), the resolution of display 545, the color or Z precision of the render target, etc. When implementing tile-based rendering, GPU 525 may perform a binning pass and one or more rendering passes. For example, with respect 24 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO to the binning pass, GPU 525 may process an entire image and sort rasterized primitives into bins.

[0094] The device 500 may use sensor data, sensor statistics, or other data from one or more sensors. Some examples of the monitored sensors may include IMUs, eye trackers, tremor sensors, heart rate sensors, etc. In some cases, an IMU may be included in the device 500, and may measure and report a body's specific force, angular rate, and sometimes the orientation of the body, using some combination of accelerometers, gyroscopes, or magnetometers.

[0095] As shown, device 500 may include an extended reality manager 550. The extended reality manager 550 may implement aspects of extended reality, augmented reality, virtual reality, etc. In some cases, such as when the device 500 is implemented as a client device (e.g., device 405 of FIG.4), the extended reality manager 550 may determine information associated with a user of the device and / or a physical environment in which the device 500 is located, such as facial information, body information, hand information, device pose information, audio information, etc. The device 500 may transmit the information to an animation and scene rendering system (e.g., animation and scene rendering system 410). In some cases, such as when the device 500 is implemented as an animation and scene rendering system (e.g., the animation and scene rendering system 410 of FIG.4), the extended reality manager 550 may process the information provided by a client device as input information to generate and / or animate a virtual representation for a user of the client device.

[0096] Virtual representations (e.g., avatars) are an important component of virtual environments. A virtual representation (or avatar) is a 3D representation of a user and allows the user to interact with the virtual scene As noted previously, there are different ways to represent a virtual representation of a user (e.g., an avatar) and corresponding animation data. For example, avatars may be purely synthetic or may be an accurate representation of the user (e.g., as shown by the virtual representation 302 shown in the image of FIG. 3). A virtual representation (or avatar) may need to be real-time captured or retargeted to reflect the user’s actual motion, body pose, facial expression, etc. Because of the many ways to represent an avatar and corresponding animation data, it can be difficult to integrate every single variant of these representations into a scene description. 25 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0097] Various animation assets may be needed to model an avatar, including a mesh (e.g., a 3D mesh, such as a triangle mesh, including a plurality of vertices and line segments connected the vertices), a diffuse or albedo texture, normals specular reflection texture, and in some cases other types of textures. These various assets may be available from enrollment or offline reconstruction. FIG. 6 is a diagram illustrating an example of a normal map 602, an albedo map 604, and a specular reflection map 606.

[0098] Animation of a virtual representation (e.g., avatar) can be performed using various techniques. FIG. 7 is a diagram 700 illustrating an example of one technique for performing avatar animation. As shown, camera sensors of a head-mounted display (HMD) 702 are used to capture images of a user’s face, including eye cameras used to capture images 704 of the user’s eyes, face cameras used to capture images 706 of the visible part of the face (e.g., mouth, chin, cheeks, part of the nose, etc.), and other sensors for capturing other sensor data 708 (e.g., audio, etc.). Facial animation can then be performed by a facial animation engine 710 to generate a 3D mesh 712 and texture 714 for a 3D facial avatar 716 of the user. The mesh 712 and texture 714 can then be rendered by a rendering engine 718 to generate a rendered image 720.

[0099] In some cases, facial animation can be performed with or using blend shapes. FIG. 8 is a diagram 800 illustrating an example of performing facial animation with blend shapes. As shown, a system can estimate a rough or course 3D mesh 806 and blend shapes from images 802 (e.g., captured using sensors of an HMD or other XR device) using 3D Morphable Model (3DMM) encoding of a 3DMM encoder 804. The system can generate texture using one or more techniques, such as using a machine learning system 808 (e.g., one or more neural networks) or computer graphics techniques (e.g., Metahumans).

[0100] A 3DMM is a 3D face mesh representation of known topology. A 3DMM can be linear or non-linear. FIG. 9 is a diagram illustrating an example of a system 900 that can generate a 3DMM face model or mesh 904. The system 900 can obtain a dataset of 3D and / or color images for various persons (and in some cases grayscale images) from a database 902. The system 900 can also obtain known mesh topologies of face mesh models 906 corresponding to the faces of the images in the database 902. In some cases, Principal 26 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO Component Analysis (PCA) can be used to find a representation of identifiers (IDs) / expressions in case of linear representations. Expressions can also be modeled via blend shapes (e.g., defining meshes at various states or expressions). Using these parameters, the system can manipulate or steer the mesh. A 3D mesh can be represented using linear 3DMM, which may be represented as follows: ^^ ൌ ^^ ^ ∑ ே ^^ ெ^ ^ୀ^ ^ ∙ ^^^ ^ ∑^ୀ^ ^^^ ∙ ^^^

[0101] The 3D vertices can include the mean Shape ^^^(e.g., a mean face shape), a shape parameter ^^^, a shape basis ^^^, an expression parameter ^^^, and an expression basis or blend shape ^^^. In the illustrative example Equation (1), there are N facial shape coefficients ^^^and N facial shape basis vectors ^^^where N is an integer greater than or equal to 1. In some implementations, each of the mean shape ^^^, facial shape basis vectors ^^^, and facial expression vectors ^^^can include position information for 3D vertices (e.g., x, y, and z coordinates) that can be combined to form the 3D model ^^. In some implementations, facial shape basis vectors ^^^, and facial expression vectors ^^^can be expressed as positional offsets from the mean face ^^^, where the coefficients for facial shape ^^^and facial expressions ^^^provide a scaling factor for corresponding offset vectors. In one illustrative example, the 3D model S includes three thousand 3D vertices. In one illustrative example, M is equal to 219, which corresponds to 219 facial shape basis vectors ^^^and facial shape coefficients ^^^. In some implementations, the shape basis vectors ^^^can include principal component analysis eigenvectors. In some cases, there are M facial expression coefficients ^^^and M facial expression basis vectors where M is an integer greater than or equal to 1. In some cases, the facial expression vectors ^^^can include blend shape vectors. In one illustrative example, M is equal to 39, which corresponds to 39 facial expression basis vectors ^^^and 39 facial expression coefficients ^^^(e.g., 39 blend shapes and 39 blend shape coefficients). In some cases, the result of the linear combination shown in Equation (1) can be a 3D model (e.g., a 3DMM) of a face in a neutral pose. In some examples, the 3D model can be rotated with pose information such 27 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO as yaw, pitch, and roll values to match the pose of the face in an image frame (e.g., image frame 1302 of FIG.13 described below).

[0102] In some cases, blend shape weights can be determined using 3DMM encoding. The blend shapes weights can then be used to reconstruct a deformed mesh, such as to animate an avatar. For instance, animating an avatar can be summarized as determining the weight of each blend shape given an input image.

[0103] FIG.10 illustrates a two-dimensional (2D) facial image 1002 and FIG.11 illustrates a corresponding 3D facial model 1104 generated from the 2D facial image 1002 of FIG. 10 using a 3D morphable model (3DMM). In some aspects, the 3D facial model 1104 can include a representation of a facial expression in the 2D facial image 1002. In one illustrative example, the facial expression representation can be formed from blend shapes. Blend shapes can semantically represent movement of muscles or portions of facial features (e.g., opening / closing of the jaw, raising / lowering of an eyebrow, opening / closing eyes, etc.). In some cases, each blend shape can be represented by a blend shape coefficient paired with a corresponding blend shape vector.

[0104] In some examples, the 3D facial model 1104 can include a representation of the facial shape in the 2D facial image 1002. In some cases, the facial shape can be represented by a facial shape coefficient paired with a corresponding facial shape vector. In some implementations a 3D model engine (e.g., a machine learning model) can be trained (e.g., during a training process) to enforce a consistent facial shape (e.g., consistent facial shape coefficients) for a 3D facial model regardless of a pose (e.g., pitch, yaw, and roll) associated with the 3D facial model. For example, when the 3D facial model is rendered into a 2D image for display, the 3D facial model can be projected onto a 2D image using a projection technique. While a 3D model engine that enforces a consistent facial shape independent of pose, the projected 2D image may have varying degrees of accuracy based on the pose of the 3D facial model captured in the projected 2D image.

[0105] As shown in FIG. 12, the 3D model generator can utilize input frames such as oblique frames 1204A, 1204B, 1204C, and / or frame 1208 to generate the 3D facial model 1210. As shown in FIG.12, a 3D model fitting engine 1206 can generate and / or apply a texture 28 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO to the underlying 3D model (e.g., the 3D facial model 1104 of FIG. 11) to provide a digital representation of the user wearing the head mounted XR system 1202. In one illustrative example, a 3D morphable model (3DMM) can be used to represent the expression and geometry of the user. In some cases, a 3DMM may lack capability to accurately reproduce the inner mouth and eyeballs of the user. In some cases, the resulting 3D facial model 1210 can produce unrealistic results in the eye and mouth regions.

[0106] FIG.13 is a diagram illustrating an example of a 3D modeling system 1300 that can generate a 3D model (e.g., a 3D morphable model (3DMM)) using at least one image frame 1302. The 3D modeling system 1300 also obtains local frames (e.g., frames from a user facing camera of the head mounted XR system 1202 of FIG. 12). As shown in FIG. 13, the 3D modeling system 1300 includes an image frame engine 1304, a 3D model fitting engine 1306, and a face reconstruction engine 1310. While the 3D modeling system 1300 is shown to include certain components, one of ordinary skill will appreciate that the 3D modeling system 1300 can include more components than those shown in FIG. 13. The components of the 3D modeling system 1300 can include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of the 3D modeling system 1300 can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The software and / or firmware can include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the 3D modeling system 1300.

[0107] The image frame engine 1304 can obtain or receive an image frame 1302 and / or local frames 1303 captured by an image sensor, from storage, from memory, from an external source (e.g., a server, an external memory accessed via a network, or other external source), or the like. In some cases, the image frame can be included in a sequence of frames (e.g., a video, a sequence of standalone or still images, etc.). In one illustrative example, each frame of the sequence of frames can include a grayscale component per pixel. Other examples of frames 29 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO include frames having red (R), green (G), and blue (B) components per pixel (referred to as an RGB video including RGB frames), luma, chroma-blue, chroma-red (YUV, YCbCr, or Y’CbCr) components per pixel and / or any other suitable type of image. The sequence of frames can be captured by one or more cameras, obtained from storage, received from another device (e.g., a camera or device including a camera), or obtained from another source. In some implementations, the image frame engine 1304 can convert the image frame 1302 to grayscale. The image frame engine 1304 can, in some cases, crop a portion of the image frame 1302 that corresponds to a face. In some examples, the image frame engine 1304 can perform a face detection process and / or face recognition process to detect and / recognize a face within the image frame 1302. The image frame engine 1304 can generate or apply a bounding box (e.g., bounding box 1030 shown in FIG.10) around the face and can crop out the image data within the bounding box to generate an input image for the 3D model fitting engine 1306.

[0108] The 3D model fitting engine 1306 can receive an input image (e.g., the image frame 1302, the cropped bounding box around the face in the image frame 1302, etc.) from the image frame engine 1304. Using the input image, the 3D model fitting engine 1306 can perform a 3D model fitting technique to generate a 3D model (e.g., a 3DMM model) of the face (which can include the head of the person in the image frame 1302). The 3D model fitting technique can include solving for shape coefficients ^^^and expression coefficients ^^^. In some examples, the 3D model fitting can include solving for positional information related to the object. In the example of the object being a head of a person, the positional information may include pose information related to a pose of the head. For example, the pose information may indicate an angular rotation of the head with respect to a neutral position of the head. The rotation may be along a first axis (e.g., a yaw axis), a second axis (e.g., a pitch axis), and / or a third axis (e.g., a roll axis). In some cases, the 3D model fitting can also include a focal length for projection of the 3D model onto a 2D image using any suitable projection technique. In some examples, a weak perspective model can use the focal length produced by the 3D model fitting engine 1306 to project the 3D vertices of the 3D model (e.g., the 3DMM) onto a 2D image. In some examples, a full perspective model can use the focal length produced by the 3D model fitting engine 1306 to project the 3D vertices of the 3D model (e.g., the 3DMM) onto a 2D image. 30 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0109] The local feature engine 1308 can receive one or more input frames (e.g., local frames 1303) from the image frame engine 1304. In some implementations, the local feature engine 1308, can be implemented as a machine learning model (e.g., a deep learning neural network). In some examples, the machine learning model can be trained to generate local textures for portions of the face such as the eyes and the mouth that can be combined with a full facial texture for the 3DMM generated by the 3D model fitting engine 1306. In one example, the local frames 1303 can include oblique frames 1204A, 1204B, 1204C captured by user facing cameras on a head mounted XR system as illustrated in FIG.12.

[0110] In FIG. 13, the face reconstruction engine 1310 can receive the coefficients generated by the 3D model fitting engine 1306 to generate the 3D model (e.g., the 3DMM). The 3D model can be generated or constructed as a linear combination of a mean face (sometimes referred to as a neutral face), facial shape basis vectors, and facial expression basis vectors. The mean face can represent an average face that can be transformed (e.g., by the shape basis vectors and expression basis vectors) to achieve the desired final 3D face shape of the 3D model. The facial shape basis vectors can be used to scale proportions of the mean face. In some cases, the facial shape basis vectors may be used to represent a fat or thin face, a small or large nose, and any adjustment to the basic facial shape. In some implementations, the facial shape basis vectors are determined based on principal component analysis (PCA). In some cases, the facial expression basis vectors can represent facial expressions, such as smiling, lifting an eyebrow, blinking, winking, frowning, etc.

[0111] Blend shapes are an illustrative example of facial expression basis vectors. As used herein, a blend shape can correspond to an approximate semantic parametrization of all or a portion of a facial expression. For example, a blend shape can correspond to a complete facial expression, or correspond to a “partial” (e.g., “delta”) facial expression. Examples of partial expressions include raising one eyebrow, closing one eye, moving one side of the face, etc. In one example, an individual blend shape can approximate a linearized effect of the movement of an individual facial muscle. In some cases, the semantic representation can be modeled to correspond with movements of one or more facial muscles. 31 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0112] As noted previously, a 3D model ^^ generated using a 3D model fitting technique (e.g., a 3DMM generated using a 3DMM fitting technique) can be a statistical model representing 3D geometry of an object (e.g., a face), as described above with respect to Equation (1).

[0113] In some implementations, the face reconstruction engine 1310 can generate a UV face position map (also referred to as a UV position map or UV texture) for an input image frame 1302 that includes a first portion of a face. For example, the first portion of the face may be a view of substantially all of the face of a person. In some examples, the local UV textures output by the local feature engine 1400 can be combined with full face features represented in a texture map generated by the face reconstruction engine 1310 to improve the quality of the appearance of the portions of the face represented in the local UV textures (e.g., the mouth and eyes). In some cases, the local UV textures may be obtained from images of at least a second portion of the face, such as oblique images, where the second portion of the face at least part overlaps the first portion of the face. In some cases, the UV face position map can be applied as a texture to the 3DMM generated by the 3D model fitting engine 1306. In some examples, the 3DMM with the texture provided from the UV face position map can be used to render a 3D digital representation of the input image. In some implementations, a machine learning model (e.g., a neural network) can be used to generate the UV face position map.

[0114] FIG. 14 is a diagram illustrating an example of a system 1400 for generating a 3D model of a facial avatar of a face of a user based on images 1402 (e.g., captured using cameras of an HMD or other XR device), which can be 2D images. In FIG.14, the system 1400 includes an encoder 1404 (e.g., a neural network encoder, also referred to as a neural encoder), a decoder 1406 (e.g., a neural network decoder, also referred to as a neural decoder), and a rendering engine 1408 (e.g., a game engine, such as Metahuman). During operation of the system, the images 1402 of the face of the user may be obtained or captured (e.g., from one or more cameras of the HMD or other XR device). The images 1402 can be input into the encoder 1404. The encoder 1404 can encode the images to output neural codes 1405 (e.g., expression codes), which can be a latent representation (e.g., a feature representation) of the input images 1402. A server (or other sender) 1410 can send user enroll data 1412 to the system 1400 via a network. The enroll data 1412 can include data related to a specific avatar selected by the user. 32 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0115] The decoder 1406 can decode the neural codes 1405. In one or more examples, the decoder 1406 can use the neural codes along with the user enroll data 1412 to construct the facial avatar 1416 of the face of the user. In some examples, the rendering engine 1408 can use the user enroll data 1412 to construct the facial avatar 1416 of the face of the user.

[0116] As previously mentioned, for generating a 3D model, such as for an avatar, a linear 3DMM can be used to represent the geometry of a user’s head, body, and / or hands for the avatar. A linear 3DMM estimates blend shapes weights. The blend shape weights are sufficient to be able to retarget an avatar in any game engine (e.g., ARKit, Movement SDK, Metahuman, ReadyPlayerMe, etc.) As such, animation of the avatar, such as facial animation, can be performed using blend shapes.

[0117] FIG.15 is a diagram illustrating an example of a system 1500 for generating a blend shaped-based facial avatar of a face of a user. In FIG. 15, the system is shown to include a 3DMM encoder 1504, a mapping table 1506, and a rendering engine 1508 (e.g., a game engine, such as Metahuman). In one or more examples, during operation of the system 1500, the system 1500 can obtain images 1502 (e.g., captured using cameras of an HMD or other XR device) of a face of the user. The images 1502 can be provided as input to the 3DMM encoder 1504. The 3DMM encoder 1504 can encode the images 1502 to output blend shape weights 1505 (and in some cases neural codes). In some cases, the blend shape weights 1505 can be generated per image frame.

[0118] The system 1500 can process the blend shape weights 1505 from the 3DMM encoder 1504 using the mapping table 1506. The mapping table 1506 can map the blend shape weights 1505 to rendering engine weights (e.g., corresponding to animation curves to animate the avatar). The rendering engine 1508 can use the mappings along with user enroll data 1512 to construct a blend shape based-facial avatar of the face of the user. In one or more examples, the user enroll data 1512 (e.g., including mesh and attribute maps of an avatar selected by the user) may be sent once to the system (e.g., to the rendering engine 1508).

[0119] FIG. 16 is a diagram illustrating an example of a system 1600 for generating a photorealistic facial avatar of a face of a user. The system 1600 of FIG.16 uses a mesh-based approach to improve photorealism of an avatar. In FIG.16, the system 1600 is shown to include 33 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO an encoder 1604 (e.g., a neural network encoder) and a decoder 1606 (e.g., a neural network decoder).

[0120] During operation of the system 1600, the system 1600 can obtain images 1602 of the face of the user (e.g., from a camera of an HMD or other XR device). In some cases, the images 1602 can be scaled by a scaling engine 1603. For example, the scaling engine 1603 can downscale the images to a size for which the neural encoder 1604 is trained to process. The scaling engine 1603 is optional in the system 1600. The images 1602 (or the scaled images) can be input to the encoder 1604 for processing. The encoder 1604 can encode the images to output neural codes 1605 (e.g., a neural code representation).

[0121] The neural decoder 1606 can obtain (e.g., retrieve from storage, receive from the encoder 1604, etc.) the neural codes 1605. In one or more examples, the decoder 1606 can also obtain user enroll data 1612 (e.g., including mesh and attribute maps of an avatar selected by the user). The decoder 1606 can decode the neural codes. In some cases, the user enroll data 1612 may be sent once to the system 1600 (e.g., to the decoder 1606), such as from a server or other sender. In one or more examples, the decoder 1606 can use the neural codes 1605, the user enroll data 1612, and a view 1614 of the user to generate a view dependent texture 1616 (e.g., a UV map) and a personalized mesh 1618 for a photorealistic facial avatar of the face of the user.

[0122] FIG.17 is a diagram illustrating examples of problems with representations, such as inconsistencies across rendering engines. As shown in FIG. 17, switching between a neural renderer (e.g., a neural codec approach, which is shown on the bottom half 1704 of FIG.17) and a rendering engine such as a game engine (e.g., a blend shapes approach, which is shown on the top half 1702 of FIG. 17) can be problematic because the two approaches are not compatible with each other. A neural render requires a neural code. Conversely, a rendering engine (e.g., a game engine) requires blend shapes (e.g., also including other body parameters, such as linear blend skinning (LBS)). Currently, a blend shape-based approach cannot directly communicate with a neural code-based approach, and vice versa. Base assets (e.g., neural data) could be the same between the two different approaches. However, currently, it is not possible to inter-change (e.g., communicate) between the approaches. 34 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0123] FIG.18 is a diagram illustrating examples of problems with representations, such as inconsistencies across different modalities. As shown in FIG. 18, moving between different modalities (e.g., such as from an HMD 1802 to a mobile phone 1804, or from a mobile phone 1806 of one OEM to another mobile phone 1808 of another OEM) can be problematic because they are not compatible with each other. Currently, different modalities cannot directly communicate with each other because the different modalities use different neural code representations.

[0124] As such, improved systems and techniques that allow for a blend shape-based approach to directly communicate with a neural code-based approach, and vice versa, as well as for different neural code representations (e.g., from different modalities) to be able to directly communicate with each other can be useful.

[0125] In one or more aspects, the systems and techniques provide solutions for interoperable avatars. In one or more examples, the systems and techniques allow for communication between a blend shape-based approach and a neural code-based approach (and vice versa) and between different neural representations (e.g., different types of neural codes output by different encoders) by defining a forward and backward mapping between the approaches and / or representations, and providing a common latent space representation between different systems (e.g., systems with different OEMs and, as such, with different types of neural codes). In one or more examples, by employing a forward mapper, the systems and techniques may project neural codes to a common latent representation (e.g., blend shapes). By adding a backward mapper, the systems and techniques may back-project the common latent representation (e.g., blend shapes) to the neural codes. In some examples, the systems and techniques may use a common latent space representation to have a common blend shape basis and base mesh.

[0126] FIG.19 is a diagram illustrating an example of a system 1900 that provides a common latent representation 1925 and forward and backwards mappings between different representations (e.g., between a blend shape representation and a neural code-based representation). In FIG. 19, the system is shown to include an encoder 1904 (e.g., a neural network encoder) and a decoder 1906 (e.g., a neural network decoder). The system also 35 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO includes a forward mapper 1922 (to determine forward mappings to the common latent representation 1925) and a backward mapper 1926 (to determine backward mappings to the common latent representation 1925).

[0127] In order for a blend shape-based approach to be able to directly communicate with a neural code-based approach, and vice versa, as well as for different neural code representations (e.g., from different modalities) to be able to directly communicate with each other, forward and backward mappings between the different representations need to be learned by the forward mapper 1922 and the backward mapper 1926, respectively. For example, the system 1900 can obtain images 1902 of a face of a user. In some cases, the images 1902 can be captured by an HMD or other XR device. The images 1902 can include images of eyes of the user (e.g., captured by eye facing cameras, such as cameras of the HMD or XR device), images a visible part of a face of the user such as mouth, chin, cheeks, part of the nose, etc. (e.g., captured by face cameras, such as cameras of the HMD or XR device), and / or other images.

[0128] The images 1902 can be provided as input to the encoder 1904. The encoder 1904 can encode the images 1902 of the face of a user to generate neural codes 1905. The neural codes 1905 can be expression codes, which are a coded representation (or latent representation, such as a feature representation) of expressions of the face in the input images 1902. The decoder 1906 can use the neural codes 1905, the user enroll data 1912, and a view 1914 of the user to generate a view dependent texture 1916 (e.g., a UV map) and a personalized mesh 1918 for a photorealistic facial avatar of the face of the user.

[0129] The forward mapper 1922 and the backward mapper 1926 can be trained to map to and from the common latent representation 1925 (e.g., a common latent variable representation). For instance, forward mapper 1922 and the backward mapper 1926 can include an encoder-decoder architecture (e.g., an auto encoder) that can be trained, where the common latent representation 1925 may be defined. In one or more examples, the forward mapper 1922 can be learned or trained (e.g., using machine learning training, such as backpropagation or forward training) to map the neural codes 1905 to the common latent representation 1925. The backward mapper 1926 can be learned / trained (e.g., using machine learning) to map the 36 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO common latent representation 1925 back to the neural codes 1905. In some examples, the forward mapper 1922 and the backward mapper 1926 can be jointly or separately learned.

[0130] FIG.20 is a diagram illustrating an example of a system 2000 that provides forward and backwards mappings between different representations between a common latent representation 2025 and a neural code-based representation 2005. The system 2000 can resolve inconsistencies across different rendering engines or modalities. The system 2000 can obtain images 2002 of a face of a user. In some cases, the images 2002 can be captured by an HMD or other XR device. The images 2002 can include images of eyes of the user (e.g., captured by eye facing cameras, such as cameras of the HMD or XR device), images a visible part of a face of the user such as mouth, chin, cheeks, part of the nose, etc. (e.g., captured by face cameras, such as cameras of the HMD or XR device), and / or other images.

[0131] The system images 2002 can be provided as input toa 3DMM encoder 2024. The 3DMM encoder 2024 can process the input images 2002 to generate a common latent representation 2025. The common latent representation 2025 can be output to a rendering engine 2008 (e.g., a game engine, such as Metahuman) are shown. The rendering engine 2008 can use the common latent representation 2025 to generate a blend shape-based avatar 2014.

[0132] The system 2000 also includes an encoder 2004 (e.g., a neural network encoder) and a decoder 2006 (e.g., a neural network decoder). The images 2002 can also be provided as input to the encoder 2004. The encoder 2004 can encode the images 2002 of the face of the user to generate neural codes 2005. The neural codes 2005 can be expression codes, which are a coded representation (or latent representation, such as a feature representation) of expressions of the face in the input images 2002. The decoder 2006 can use the neural codes 2005, the user enroll data 2012 (and in some cases a view of the user to generate a texture 2016 (e.g., a UV map), which may be a view dependent texture, and a personalized mesh 2018 for a photorealistic facial avatar of the face of the user.

[0133] The system 2000 includes a forward mapper 2022 and a backward mapper 2026. The forward mapper 2022 and the backward mapper 2026 can be trained to map to and from the common latent representation 2024 (e.g., a common latent variable representation). For instance, forward mapper 2022 and the backward mapper 2026 can include an encoder-decoder 37 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO architecture (e.g., an auto encoder) that can be trained, where the common latent representation 2024 may be defined. In one or more examples, the forward mapper 2022 can be trained (e.g., using machine learning training, such as backpropagation or forward training) to map the neural codes 2005 to the common latent representation 2025. The backward mapper 2026 can be learned / trained (e.g., using machine learning) to map the common latent representation 2025 back to the neural codes 2005. In some examples, the forward mapper 2022 and the backward mapper 2026 can be jointly or separately learned.

[0134] With the addition of the forward mapper 2022 and the backward mapper 2026, it is possible for interoperability between different rendering engines or modalities. For example, as noted previously, avatars can be used in various types of immersive or virtual environments, such as in immersive / virtual communications to represent a user in a call. The user may use a particular application or service to generate an animatable avatar for the user. During setup of an application or service (e.g., a setup of a call), the user may offer their preferred avatar format (e.g., the format in which the user has captured and stored the avatar representation) to other users. Such a situation can raise inter-operability issues for service providers, given that there is a large number of proprietary avatar formats. By adding the forward mapper 2022, neural codes can be projected to the common latent representation 2025. By adding the backward mapper 2026, the common latent representation 2025 can be back-projected to the neural codes 2005.

[0135] In one or more examples, a common latent representation (e.g., the common latent representation 1925 of FIG.19 or the common latent representation 2025 of FIG. 20) can be defined by any space designed to represent 3D models (e.g., avatars). For example, efficient geometry-aware 3D (EG3D) latent variables, Stable Diffusion latent variables, and / or Style generative adversarial network (StyleGAN) latent variables may be considered as a common latent space representation. The mappings may have some difficulties (e.g., inconsistencies) depending upon the space being used. One example implementation for the common latent space is to have a common blend shape basis and base mesh (e.g., examples and details of different blend shapes and base meshes are shown in FIGS.21 to 23). 38 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0136] In one illustrative implementation, a 3DMM model can have a total of 87 different blend shapes, three UV maps (e.g., a UV map for the head, a UV map for the eyes, and a UV map for the mouth), and three resolutions of meshes (e.g., as a base mesh). In some cases, landmarks can be added. These definitions can also be extended. FIGS.21 – 23 show examples of QC base meshes, UV maps, and blend shapes.

[0137] FIG.21 is a diagram illustrating examples of a base mesh. A base mesh can be used to map any user wanted avatar. A base mesh is one aspect that can also simplify recognition of the various facial parts of the avatar. For example, a person’s avatar can be mapped to the base mesh by the process of retopology (e.g., using a tool such as Wrap4D) as shown in FIG. 21. Currently, there are three different resolutions for the base meshes, which include 3210, 12840, or 51200 vertices.

[0138] FIG. 22 is a diagram illustrating examples of UV maps. The UV maps can include three distinct parts. The three parts can include one part for the head, one part for the eye balls, and one part for the mouth, teeth, and tongue.

[0139] FIG.23 is a diagram illustrating examples of blend shapes. As described previously, a possible implementation of the common latent space is to have a common blend shape basis and base mesh. The example blend shapes shown in FIG. 23 are examples of blend shapes (applied to a base mesh) that can be included in a common blend shape basis used as the common latent space. In some examples, several blend shapes can be used for the blend shape basis.

[0140] Various message protocols and formats can be used to exchange animation steams (e.g., represented using neural codes and / or blend shapes). The formats can be a payload format for the carrying the timed animation streams. Examples of transport protocols can include a Real-time Transport Protocol (RTP) transport mechanism or using a data channel or a separate transport channel (e.g., over the QUIC protocol). An SDP declaration for an animation stream can includes a description of the supported avatar formats in the dcmap or dcsa attributes. The attribute can contain a unique reference number (URN) that uniquely identifies the used schema. A receiver that does support the offered representation can remove it from an offer and only leave the identifier for the common avatar representation format (e.g., the common 39 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO latent representation). The following is an example of SDP signaling that indicates both a proprietary avatar representation and the common avatar representation format (e.g., the common latent representation): m=application 52718 UDP / DTLS / SCTP webrtc-datachannel c=IN IP6 fe80::6676:baff:fe9c:ee4a b=AS:500 a=candidate:11 UDP 2130706431 fe80::6676:baff:fe9c:ee4a 52718 typ host a=ice-ufrag:8hhY a=ice-pwd:asd88fgpdd777uzjYhagZg a=max-message-size:1024 a=sctp-port:5000 a=setup:actpass a=fingerprint:SHA-1 4A:AD:B9:B1:3F:82:18:3B:54:02:12:DF:3E:5D:49:6B:19:E5:7C:AB a=tls-id: abc3de65cddef001be82 a=dcmap:1000 subprotocol="avatar-representation" a=dcsa:1000 format="urn:3gpp:avatar:base-format" a=dcsa:1001 format="urn:qualcomm:avatar:avatar-format"

[0141] FIG. 24 is a flow chart illustrating an example of a process 2400 for interoperable avatars. The process 2400 can be performed by a first device (e.g., computing system 2500 of FIG.25) or by a component or system (e.g., a chipset, one or more processors such as one or more central processing units (CPUs), digital signal processors (DSPs), graphics processing units (GPUs), any combination thereof, and / or other type of processor(s), or other component or system) of the first device. For instance, the first device can be or may be part of an extended reality (XR) device (e.g., an AR, VR, and / or MR head mounted device (HMD)), a mobile device (e.g., a mobile phone), or other device. The operations of the process 2400 may be implemented as software components that are executed and run on one or more processors (e.g., processor 2510 of FIG.25 or other processor(s)). Further, the transmission and reception of signals by the computing device in the process 2400 may be enabled, for example, by one or more antennas and / or one or more transceivers (e.g., wireless transceiver(s)). 40 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0142] At block 2410, the first device (or component thereof) can receive, from a second device (e.g., an XR device such as an HMD, a mobile device such as mobile phone, or other device), a common latent representation of a 3D model of a user. In some cases, the common latent representation is a space designed to represent avatars. For instance, the space can include efficient geometry-aware 3D (EG3D) latent variables, Stable Diffusion latent variables, Style generative adversarial network (StyleGAN) latent variables, common blend shapes with a base mesh, any combination thereof, and / or other space.

[0143] At block 2420, the first device (or component thereof) can back-project, using a backward mapper, the common latent representation to neural codes representing the 3D model of the user. In some aspects, the backward mapper can be a machine learning model trained to map the common latent representation to the neural codes. In some examples, the first device (or component thereof) can generate, using a machine learning encoder, the neural codes by encoding one or more two-dimensional (2D) images of the user. For instance, the machine learning encoder can include a neural network encoder that is trained to extract features from the 2D images.

[0144] At block 2430, the first device (or component thereof) can generate the 3D model of the user using the neural codes. In some cases, the 3D model is an avatar of the user, such as a facial avatar of a face of the user.

[0145] In some aspects, the first device (or component thereof) can project, using a forward mapper, the neural codes to the common latent representation. In some cases, the forward mapper can be a machine learning model trained to map the neural codes to the common latent representation. The first device (or component thereof) can transmit, to the second device or to a third device, the common latent representation.

[0146] In some cases, the first device may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) that are configured to carry out the steps of processes described herein. In some examples, the computing device may include a display, one or more network interfaces configured to communicate and / or receive the data, any combination thereof, and / or 41 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO other component(s). The one or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data according to the 3G, 4G, 5G, and / or other cellular standard, data according to the Wi-Fi (802.11x) standards, data according to the BluetoothTMstandard, data according to the Internet Protocol (IP) standard, and / or other types of data.

[0147] The components of the computing device of process 2400can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The computing device may further include a display (as an example of the output device or in addition to the output device), a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component(s). The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.

[0148] The process 2400 is illustrated as a logical flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.

[0149] Additionally, process 2400 may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) 42 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0150] FIG. 25 is a block diagram illustrating an example of a computing system 2500, which may be employed for interoperable avatars. In particular, FIG.25 illustrates an example of computing system 2500, which can be for example any computing device making up internal computing system, a remote computing system, a camera, or any component thereof in which the components of the system are in communication with each other using connection 2505. Connection 2505 can be a physical connection using a bus, or a direct connection into processor 2510, such as in a chipset architecture. Connection 2505 can also be a virtual connection, networked connection, or logical connection.

[0151] In some aspects, computing system 2500 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple data centers, a peer network, etc. In some aspects, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some aspects, the components can be physical or virtual devices.

[0152] Example system 2500 includes at least one processing unit (CPU or processor) 2510 and connection 2505 that communicatively couples various system components including system memory 2515, such as read-only memory (ROM) 2520 and random access memory (RAM) 2525 to processor 2510. Computing system 2500 can include a cache 2512 of high- speed memory connected directly with, in close proximity to, or integrated as part of processor 2510.

[0153] Processor 2510 can include any general purpose processor and a hardware service or software service, such as services 2532, 2534, and 2536 stored in storage device 2530, configured to control processor 2510 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 2510 may essentially 43 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.

[0154] To enable user interaction, computing system 2500 includes an input device 2545, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 2500 can also include output device 2535, which can be one or more of a number of output mechanisms. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 2500.

[0155] Computing system 2500 can include communications interface 2540, which can generally govern and manage the user input and system output. The communication interface may perform or facilitate receipt and / or transmission wired or wireless communications using wired and / or wireless transceivers, including those making use of an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, an AppleTMLightningTMport / plug, an Ethernet port / plug, a fiber optic port / plug, a proprietary wired port / plug, 3G, 4G, 5G and / or other cellular data network wireless signal transfer, a BluetoothTMwireless signal transfer, a BluetoothTMlow energy (BLE) wireless signal transfer, an IBEACONTMwireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 802.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, ad-hoc network signal transfer, radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof.

[0156] The communications interface 2540 may also include one or more range sensors (e.g., LIDAR sensors, laser range finders, RF radars, ultrasonic sensors, and infrared (IR) sensors) configured to collect data and provide measurements to processor 2510, whereby processor 2510 can be configured to perform determinations and calculations needed to obtain 44 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO various measurements for the one or more range sensors. In some examples, the measurements can include time of flight, wavelengths, azimuth angle, elevation angle, range, linear velocity and / or angular velocity, or any combination thereof. The communications interface 2540 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to determine a location of the computing system 2500 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based GPS, the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based Galileo GNSS. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.

[0157] Storage device 2530 can be a non-volatile and / or non-transitory and / or computer- readable memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu-ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Stick® card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (e.g., Level 1 (L1) cache, Level 2 (L2) cache, Level 3 (L3) cache, Level 4 (L4) cache, Level 5 (L5) cache, or other (L#) cache), resistive random-access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT- RAM), another memory chip or cartridge, and / or a combination thereof. 45 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0158] The storage device 2530 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 2510, it causes the system to perform a function. In some aspects, a hardware service that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 2510, connection 2505, output device 2535, etc., to carry out the function. The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.

[0159] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. 46 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.

[0160] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.

[0161] Further, those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0162] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function. 47 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0163] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.

[0164] In some aspects the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bitstream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.

[0165] Those of skill in the art will appreciate that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof, in some cases depending in part on the particular application, in part on the desired design, in part on the corresponding technology, etc.

[0166] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks. Examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form 48 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.

[0167] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.

[0168] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.

[0169] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may 49 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.

[0170] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein can be replaced with less than or equal to (“^”) and greater than or equal to (“^”) symbols, respectively, without departing from the scope of this description.

[0171] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.

[0172] The phrase “coupled to” or “communicatively coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.

[0173] Claim language or other language reciting “at least one of” a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and 50 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO B, C and C, A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C. The language “at least one of” a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.

[0174] Claim language or other language reciting “at least one processor configured to,” “at least one processor being configured to,” “one or more processors configured to,” “one or more processors being configured to,” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation(s). For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.

[0175] Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.

[0176] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method), the entity may be configured to cause one or more elements (individually or collectively) to 51 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function).

[0177] The various illustrative logical blocks, modules, engines, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, engines, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0178] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as engines, modules, or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form 52 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO part of a computer program product, which may include packaging materials. The computer- readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.

[0179] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated software modules or hardware modules configured for encoding and decoding, or incorporated in a combined video encoder-decoder (CODEC).

[0180] Illustrative aspects of the disclosure include:

[0181] Aspect 1. A first device for generating one or more three-dimensional (3D) models, the first device comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive, from a second device, a common latent representation of a 3D model of a user; back-project, using a backward mapper, the common 53 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO latent representation to neural codes representing the 3D model of the user; and generate the 3D model of the user using the neural codes.

[0182] Aspect 2. The first device of Aspect 1, wherein the at least one processor is configured to: project, using a forward mapper, the neural codes to the common latent representation; and transmit, to the second device or a third device, the common latent representation.

[0183] Aspect 3. The first device of Aspect 2, wherein the forward mapper is a machine learning model trained to map the neural codes to the common latent representation.

[0184] Aspect 4. The first device of any one of Aspects 1 to 3, wherein the backward mapper is a machine learning model trained to map the common latent representation to the neural codes.

[0185] Aspect 5. The first device of any one of Aspects 1 to 4, wherein the common latent representation is a space designed to represent avatars.

[0186] Aspect 6. The first device of Aspect 5, wherein the space comprises at least one of efficient geometry-aware 3D (EG3D) latent variables, Stable Diffusion latent variables, Style generative adversarial network (StyleGAN) latent variables, or common blend shapes with a base mesh.

[0187] Aspect 7. The first device of any one of Aspects 1 to 6, wherein the at least one processor is configured to generate, using a machine learning encoder, the neural codes by encoding one or more two-dimensional (2D) images of the user.

[0188] Aspect 8. The first device of any one of Aspects 1 to 7, wherein the first device is a head mounted device (HMD) or a mobile phone.

[0189] Aspect 9. The first device of any one of Aspects 1 to 8, wherein the 3D model is an avatar of the user.

[0190] Aspect 10. The first device of any one of Aspects 1 to 9, wherein the 3D model is a facial avatar of a face of the user. 54 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0191] Aspect 11. A method for generating one or more three-dimensional (3D) models at a first device, the method comprising: receiving, from a second device, a common latent representation of a 3D model of a user; back-projecting, by a backward mapper, the common latent representation to neural codes representing the 3D model of the user; and generating the 3D model of the user using the neural codes.

[0192] Aspect 12. The method of Aspect 11, further comprising: projecting, by a forward mapper, the neural codes to the common latent representation; and transmitting, to the second device or a third device, the common latent representation.

[0193] Aspect 13. The method of Aspect 12, wherein the forward mapper is a machine learning model trained to map the neural codes to the common latent representation.

[0194] Aspect 14. The method of any one of Aspects 11 to 13, wherein the backward mapper is a machine learning model trained to map the common latent representation to the neural codes.

[0195] Aspect 15. The method of one of Aspects 11 to 14, wherein the common latent representation is a space designed to represent avatars.

[0196] Aspect 16. The method of Aspect 15, wherein the space comprises at least one of efficient geometry-aware 3D (EG3D) latent variables, Stable Diffusion latent variables, Style generative adversarial network (StyleGAN) latent variables, or common blend shapes with a base mesh.

[0197] Aspect 17. The method of one of Aspects 11 to 16, further comprising generating, by a machine learning encoder, the neural codes by encoding one or more two-dimensional (2D) images of the user.

[0198] Aspect 18. The method of one of Aspects 11 to 17, wherein the first device is a head mounted device (HMD) or a mobile phone.

[0199] Aspect 19. The method of one of Aspects 11 to 18, wherein the 3D model is an avatar of the user. 55 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO

[0200] Aspect 20. The method of one of Aspects 11 to 19, wherein the 3D model is a facial avatar of a face of the user.

[0201] Aspect 21. A non-transitory computer-readable storage medium comprising instructions stored thereon which, when executed by at least one processor, causes the at least one processor to perform operations according to any one of Aspects 11 to 20.

[0202] Aspect 22. An apparatus comprising one or more means for performing operations according to any one of Aspects 11 to 20.

[0203] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language claims, wherein reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” 56 Polsinelli Ref. No.094922-809002

Claims

PATENT Qualcomm Ref. No.2402226WO CLAIMS What is claimed is:

1. A first device for generating one or more three-dimensional (3D) models, the first device comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive, from a second device, a common latent representation of a 3D model of a user; back-project, using a backward mapper, the common latent representation to neural codes representing the 3D model of the user; and generate the 3D model of the user using the neural codes.

2. The first device of claim 1, wherein the at least one processor is configured to: project, using a forward mapper, the neural codes to the common latent representation; and transmit, to the second device or a third device, the common latent representation.

3. The first device of claim 2, wherein the forward mapper is a machine learning model trained to map the neural codes to the common latent representation.

4. The first device of claim 1, wherein the backward mapper is a machine learning model trained to map the common latent representation to the neural codes.

5. The first device of claim 1, wherein the common latent representation is a space designed to represent avatars.

6. The first device of claim 5, wherein the space comprises at least one of efficient geometry-aware 3D (EG3D) latent variables, Stable Diffusion latent variables, Style generative adversarial network (StyleGAN) latent variables, or common blend shapes with a base mesh. 57 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO 7. The first device of claim 1, wherein the at least one processor is configured to generate, using a machine learning encoder, the neural codes by encoding one or more two-dimensional (2D) images of the user.

8. The first device of claim 1, wherein the first device is a head mounted device (HMD) or a mobile phone.

9. The first device of claim 1, wherein the 3D model is an avatar of the user.

10. The first device of claim 1, wherein the 3D model is a facial avatar of a face of the user.

11. A method for generating one or more three-dimensional (3D) models at a first device, the method comprising: receiving, from a second device, a common latent representation of a 3D model of a user; back-projecting, by a backward mapper, the common latent representation to neural codes representing the 3D model of the user; and generating the 3D model of the user using the neural codes.

12. The method of claim 11, further comprising: projecting, by a forward mapper, the neural codes to the common latent representation; and transmitting, to the second device or a third device, the common latent representation.

13. The method of claim 12, wherein the forward mapper is a machine learning model trained to map the neural codes to the common latent representation.

14. The method of claim 11, wherein the backward mapper is a machine learning model trained to map the common latent representation to the neural codes. 58 Polsinelli Ref. No.094922-809002PATENT Qualcomm Ref. No.2402226WO 15. The method of claim 11, wherein the common latent representation is a space designed to represent avatars.

16. The method of claim 15, wherein the space comprises at least one of efficient geometry- aware 3D (EG3D) latent variables, Stable Diffusion latent variables, Style generative adversarial network (StyleGAN) latent variables, or common blend shapes with a base mesh.

17. The method of claim 11, further comprising generating, by a machine learning encoder, the neural codes by encoding one or more two-dimensional (2D) images of the user.

18. The method of claim 11, wherein the first device is a head mounted device (HMD) or a mobile phone.

19. The method of claim 11, wherein the 3D model is an avatar of the user.

20. The method of claim 11, wherein the 3D model is a facial avatar of a face of the user. 59 Polsinelli Ref. No.094922-809002