Virtual representation coding in scene description
By processing child node data in the hierarchical node set of virtual representations, generating and animation virtual representations are solved, and a high-quality and low-latency virtual environment experience is achieved.
Patent Information
- Application Number
- CN202380072048.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-13
- Filing Date
- 2023-10-16
- Publication Date
- 2025-05-23
AI Technical Summary
It is difficult for the prior art to effectively generate and animate high-quality virtual representations (incarnations) to achieve low latency and efficient virtual environment experience, especially in extended reality (XR) systems.
By obtaining information describing the hierarchical node set of virtual representations, identifying data associated with child nodes, and processing these data to generate segments of virtual representations, the user's virtual representation generation and animation are realized.
It realizes efficiently and low-latency generation and animation of high-quality virtual representations in virtual environments, improving the quality and real-timeness of user experience.
Smart Images

Figure CN120035846A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to virtual content for virtual environments or portions of virtual environments. For example, aspects of the present disclosure include systems and techniques for providing virtual representations (eg, avatars) encoded in scene descriptions. Background Art
[0002] Extended reality (XR) (e.g., virtual reality, augmented reality, mixed reality) systems can provide users with virtual experiences by immersing them in a fully virtual environment (composed of virtual content), and / or can provide users with augmented or mixed reality experiences by combining the real world or physical environment with a virtual environment.
[0003] One example use case for providing virtual, augmented, or mixed reality XR content to a user is to present a "metaverse" experience to the user. A metaverse is essentially a virtual world that includes one or more three-dimensional (3D) virtual worlds. For example, a metaverse virtual environment may allow a user to virtually interact with other users (e.g., in a social setting, in a virtual meeting, etc.), virtually purchase goods, services, property, or other items, play computer games, and / or experience other services.
[0004] In some cases, a user may be represented in a virtual environment (e.g., a metaverse virtual environment) as a virtual representation of the user, sometimes referred to as an avatar. In any virtual environment, it is important that the system generates high-quality avatars for representing people in an efficient and low-latency manner. Summary of the invention
[0005] The following is a simplified summary of the invention related to one or more aspects disclosed herein. Therefore, the following summary of the invention should not be regarded as an extensive review related to all contemplated aspects, nor should the following summary of the invention be regarded as identifying key or important elements related to all contemplated aspects or outlining the scope associated with any particular aspect. Therefore, the following summary of the invention presents certain concepts related to one or more aspects related to the mechanisms disclosed herein in a brief form before the detailed embodiments presented below.
[0006] In one illustrative example, a method for generating a virtual representation of a user is provided. The method includes: obtaining information describing a virtual representation of a user, the information including a hierarchical node set, wherein a first node in the hierarchical node set includes a root node for the hierarchical node set, wherein the first node includes a mapping configuration for mapping a child node in the hierarchical node set to a segment of the virtual representation of the user, and wherein the child node in the hierarchical node set includes data associated with the segment of the virtual representation of the user; identifying a portion of the data associated with the child node; and processing the portion of the data associated with the child node to generate the segment of the virtual representation of the user.
[0007] As another example, a device for generating a virtual representation of a user is provided. The device includes at least one memory and at least one processor coupled to the at least one memory. The at least one processor is configured to: obtain information describing the virtual representation of the user, the information including a hierarchical node set, wherein a first node in the hierarchical node set includes a root node for the hierarchical node set, wherein the first node includes a mapping configuration for mapping a child node in the hierarchical node set to a segment of the user's virtual representation, and wherein the child node in the hierarchical node set includes data associated with the segment of the user's virtual representation; identify a portion of the data associated with the child node; and process a portion of the data associated with the child node to generate a segment of the user's virtual representation.
[0008] In another example, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium has instructions stored thereon, which, when executed by at least one processor, cause the at least one processor to: obtain information describing a virtual representation of a user, the information comprising a hierarchical set of nodes, wherein a first node in the hierarchical set of nodes comprises a root node for the hierarchical set of nodes, wherein the first node comprises a mapping configuration for mapping a child node in the hierarchical set of nodes to a segment of the virtual representation of the user, and wherein the child node in the hierarchical set of nodes comprises data associated with the segment of the virtual representation of the user; identify a portion of the data associated with the child node; and process the portion of the data associated with the child node to generate the segment of the virtual representation of the user.
[0009] As another example, an apparatus for generating a virtual representation is provided. The apparatus includes: a unit for obtaining information describing a virtual representation of a user, the information including a hierarchical node set, wherein a first node in the hierarchical node set includes a root node for the hierarchical node set, wherein the first node includes a mapping configuration for mapping a child node in the hierarchical node set to a segment of the virtual representation of the user, and wherein the child node in the hierarchical node set includes data associated with the segment of the virtual representation of the user; a unit for identifying a portion of the data associated with the child node; and a unit for processing the portion of the data associated with the child node to generate the segment of the virtual representation of the user.
[0010] In some aspects, one or more of the apparatuses described herein are, are part of, and / or include an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device (e.g., a mobile phone or other mobile device), a wearable device, a wireless communication device, a camera, a personal computer, a laptop computer, a vehicle or a computing device or component of a vehicle, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device (such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the apparatus includes one or more cameras for capturing one or more images. In some aspects, the apparatus also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus described above may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors.
[0011] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood with reference to appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0012] The foregoing and other features and aspects will become more fully apparent after reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Illustrative examples of the present application are described in detail below with reference to the following drawings:
[0014] Figure 1is a diagram illustrating an example of an extended reality (XR) system according to aspects of the present disclosure;
[0015] Figure 2 is a diagram illustrating an example of a three-dimensional (3D) collaborative virtual environment according to aspects of the present disclosure;
[0016] Figure 3 is an image with a virtual representation (avatar) of a user according to aspects of the present disclosure;
[0017] Figure 4 is a diagram illustrating another example of an XR system according to aspects of the present disclosure;
[0018] Figure 5 is a diagram illustrating an example configuration of a client device according to aspects of the present disclosure;
[0019] Figure 6 is a diagram showing examples of a normal map, an albedo map, and a specular map according to aspects of the present disclosure;
[0020] Figure 7 is a diagram illustrating an example of a technique for performing avatar animation according to aspects of the present disclosure;
[0021] Figure 8 is a diagram illustrating an example of performing facial animation using blend shapes according to aspects of the present disclosure;
[0022] Fig. 9 is a diagram illustrating an example of a system that can generate a 3D deformable model (3DMM) facial mesh according to aspects of the present disclosure;
[0023] Fig.10 is a diagram showing an example of animating an avatar according to aspects of the present disclosure;
[0024] Fig.11 is a diagram showing an example of using a 3DMM fit curve to drive a virtual representation (or avatar) using a metahuman according to aspects of the present disclosure;
[0025] Fig.12 is a diagram illustrating an example of an end-to-end flow of a system according to aspects of the present disclosure;
[0026] Fig.13 is a diagram showing an example of performing avatar animation according to aspects of the present disclosure;
[0027] Fig.14 is a diagram illustrating an example of an example of an XR system configured with avatar call flow directly between client devices in accordance with aspects of the present disclosure;
[0028] Fig.15 is a diagram illustrating an example of an example of an XR system configured with avatar call flow directly between client devices in accordance with aspects of the present disclosure;
[0029] Fig.16 is a block diagram illustrating an example of a virtual representation (or avatar) reconstruction system or pipeline according to aspects of the present disclosure;
[0030] Fig.17 is a diagram illustrating the structure of a virtual representation (or avatar) in Graphics Language Transfer Format (glTF) according to aspects of the present disclosure;
[0031] Fig.18 is an example of a JavaScript Object Notation (JSON) schema for use with the systems and techniques described herein, according to some examples;
[0032] Fig.19 is a flow chart illustrating a process for generating virtual content in a distributed system according to aspects of the present disclosure; and
[0033] Fig. 20 is a diagram illustrating an example of a computing system according to aspects of the present disclosure. DETAILED DESCRIPTION
[0034] Certain aspects of the present disclosure are provided below. Some of these aspects can be applied independently, and some of them can be applied in combination, which will be apparent to those skilled in the art. In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of various aspects of the application. However, it will be apparent that various aspects can be implemented without these specific details. The accompanying drawings and description are not intended to be limiting.
[0035] The following description provides only example aspects and is not intended to limit the scope, applicability or configuration of the present disclosure. Rather, the subsequent description of the example aspects will provide those skilled in the art with an enabling description for implementing the example aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0036] As previously described, an extended reality (XR) system or device can provide an XR experience to a user by presenting virtual content to the user (e.g., for a fully immersive experience), and / or can combine a view of a real-world or physical environment with a display of a virtual environment (composed of virtual content). The real-world environment may include real-world objects (also called physical objects), such as people, vehicles, buildings, tables, chairs, and / or other real-world or physical objects. As used herein, the terms XR system and XR device are used interchangeably. Examples of XR systems or devices include head-mounted displays (HMDs), smart glasses (e.g., AR glasses, MR glasses, etc.), and the like.
[0037] XR systems may include VR systems that facilitate interaction with virtual reality (VR) environments, AR systems that facilitate interaction with augmented reality (AR) environments, MR systems that facilitate interaction with mixed reality (MR) environments, and / or other XR systems. For example, VR provides a complete immersive experience in a three-dimensional (3D) computer-generated VR environment or a video used to depict a virtual version of a real-world environment. In some cases, VR content may include VR videos that can be captured and rendered at very high quality, potentially providing a truly immersive virtual reality experience. Virtual reality applications may include games, training, education, sports videos, online shopping, and the like. VR content may be rendered and displayed using a VR system or device (such as a VR HMD or other VR head-mounted device) that completely covers the user's eyes during the VR experience.
[0038] AR is a technology that provides virtual or computer-generated content (referred to as AR content) on a user's view of a physical, real-world scene or environment. AR content can include any virtual content, such as video, images, graphical content, location data (e.g., global positioning system (GPS) data or other location data), sound, any combination thereof, and / or other enhanced content. AR systems are designed to enhance (or enhance) a person's current perception of reality, rather than replace a person's current perception of reality. For example, a user can see a real static or moving physical object through an AR device display, but the user's visual perception of the physical object can be enhanced or enhanced by a virtual image of the object (e.g., a real-world car replaced by a virtual image of a DeLorean), by AR content added to a physical object (e.g., virtual wings added to a living animal), by AR content displayed relative to a physical object (e.g., informative virtual content displayed near a sign on a building, a virtual coffee cup virtually anchored to a real-world table (e.g., placed on top of it) in one or more images, etc.), and / or by displaying other types of AR content. Various types of AR systems can be used for games, entertainment, and / or other applications.
[0039] MR technology can combine aspects of VR and AR to provide an immersive experience for users. For example, in an MR environment, real-world and computer-generated objects can interact (e.g., real people can interact with virtual people as if the virtual people were real people).
[0040] The XR environment can be interactive in a way that appears real or physical. As a user experiencing an XR environment (e.g., an immersive VR environment) moves in the real world, the rendered virtual content (e.g., images rendered in the virtual environment in the VR experience) also changes, giving the user a perception of the user's movement within the XR environment. For example, the user can turn left or right, look up or down, and / or move forward or backward, thereby changing the user's viewpoint of the XR environment. The XR content presented to the user can change accordingly, making the user's experience in the XR environment as seamless as in the real world.
[0041] In some cases, the XR system can match the relative pose and movement of objects and devices in the physical world. For example, the XR system can use tracking information to calculate the relative pose of features of devices, objects, and / or the real-world environment in order to match the relative position and movement of devices, objects, and / or the real-world environment. In some examples, the XR system can use the pose and movement of one or more devices, objects, and / or the real-world environment to render content relative to the real-world environment in a consistent manner. Relative pose information can be used to match virtual content to the user's perceived motion and spatiotemporal state of devices, objects, and the real-world environment. In some cases, the XR system can track parts of the user (e.g., the user's hands and / or fingertips) to allow the user to interact with virtual content items.
[0042] An XR system or device may facilitate interaction with different types of XR environments (e.g., a user may use an XR system or device to interact with an XR environment). An example of an XR environment is a virtual reality environment. Users may virtually interact with other users (e.g., in a social setting, in a virtual meeting, etc.), virtually purchase items (e.g., goods, services, property, etc.), play computer games, and / or experience other services in a virtual reality environment. In an illustrative example, an XR system may provide a 3D collaborative virtual environment for a group of users. Users may interact with each other via virtual representations of users in the virtual environment. Users may experience the virtual environment visually, auditorily, tactilely, or otherwise while interacting with virtual representations of other users.
[0043] A virtual representation of a user can be used to represent the user in a virtual environment. The virtual representation of a user is also referred to as an avatar in this article. The avatar representing the user can mimic the user's appearance, movement, adventures, and / or other characteristics. A virtual representation (or avatar) can be generated / animated in real time based on input captured from the user's device. The range of avatars can range from basic synthetic 3D representations to more realistic representations of the user. In some examples, the user may expect the avatar representing the person in the virtual environment to appear as a digital twin of the user. In any virtual environment, it is important that the XR system efficiently generates high-quality avatars (e.g., realistically representing the person's appearance, movement, etc.) in a low-latency manner. It is also important for the XR system to present audio in an effective manner to enhance the XR experience.
[0044] For example, in the example of the 3D collaborative virtual environment above, the XR system of a user from a user group can display virtual representations (or avatars) of other users sitting at a virtual table or at a specific location in a virtual room. The virtual representations of the users and the background of the virtual environment should be displayed in a realistic manner (e.g., as if the users were sitting together in the real world). The heads, bodies, arms, and hands of the users can be animated as they move in the real world. Audio may need to be rendered spatially or may be rendered in mono. In order to maintain a high-quality user experience, latency in rendering and animating virtual representations should be minimal.
[0045] The computational complexity of generating virtual environments by XR systems may impose significant power and resource requirements, which may be a limiting factor in achieving XR experiences (e.g., reducing the ability of XR devices to efficiently generate and animate virtual content in a low-latency manner). For example, the computational complexity of rendering and animating virtual representations of users and composing virtual scenes may impose significant power and resource requirements on devices when implementing XR applications. Such power and resource requirements are exacerbated by the recent trend of implementing such technologies in mobile and wearable devices (e.g., HMDs, XR glasses, etc.) and making such devices smaller, lighter, and more comfortable (e.g., by reducing the heat emitted by the devices) to be worn by users for longer periods of time. In view of such factors, a user's XR device (e.g., HMD) may have difficulty rendering and animating virtual representations of other users and having difficulty synthesizing scenes and generating target views of the virtual environment for display to the user of the XR device.
[0046] Furthermore, there are different ways of representing avatars and corresponding animation data, in which case it may be difficult to integrate every single variant of these representations into the scene description.A scene description is a file or document comprising information describing or defining a 3D scene.
[0047] Described herein are systems, apparatus, electronic devices, methods (also referred to as processes), and computer-readable media (collectively referred to herein as “systems and techniques”) for providing encoding of a virtual representation (e.g., an avatar) in a scene description. The systems and techniques may decouple the representation of a virtual representation (or avatar) and its animation data from the integration of the avatar in a scene description. For example, a virtual representation (or avatar) reconstruction step may be customized for a user's virtual representation (or avatar), which may be identified by a field (e.g., a type Uniform Resource Name (URN) field such as defined in RFC8141) and may be used to create a dynamic mesh for representing the user's virtual representation (or avatar). Such a solution may allow the systems and techniques to decompose the virtual representation (or avatar) into multiple mesh nodes, where each mesh node corresponds to a body part of the virtual representation (or avatar). Multiple mesh nodes enable the XR system to support interactivity with various parts of the virtual representation (e.g., with the avatar's hands).
[0048] Various aspects of the application will be described with reference to the accompanying drawings.
[0049] Figure 1 An example of an extended reality system 100 is shown. As shown, the extended reality system 100 includes a device 105, a network 120, and a communication link 125. In some cases, the device 105 may be an extended reality (XR) device, which may generally implement aspects of extended reality, including virtual reality (VR), augmented reality (AR), mixed reality (MR), etc. A system including the device 105, the network 120, or other elements in the extended reality system 100 may be referred to as an extended reality system.
[0050] Device 105 may overlay virtual objects with real-world objects in view 130. For example, view 130 may generally refer to visual input to user 110 via device 105, a display generated by device 105, a configuration of virtual objects generated by device 105, etc. For example, view 130-A may refer to visible real-world objects (also referred to as physical objects) at some initial time and visible virtual objects overlaid on or coexisting with the real-world objects. View 130-B may refer to visible real-world objects and visible virtual objects, which are overlaid on or coexisting with the real-world objects at some later time. The position difference in the real-world objects (e.g., and thus the overlaid virtual objects) may be caused by the head movement 115 to shift from view 130-A at 135 to view 130-B. In another example, view 130-A may refer to a completely virtual environment or scene at an initial time, and view 130-B may refer to a virtual environment or scene at a later time.
[0051] Typically, the device 105 can generate, display, project, etc. a virtual object and / or a virtual environment to be viewed by the user 110 (e.g., where a portion of the virtual object and / or virtual environment can be displayed based on the user 110 head posture prediction according to the techniques described herein). In some examples, the device 105 may include a transparent surface (e.g., optical glass) so that the virtual object can be displayed on the transparent surface to overlay the virtual object on the real word object viewed through the transparent surface. Additionally or alternatively, the device 105 can project the virtual object onto the real world environment. In some cases, the device 105 may include a camera and can display both the real world object (e.g., as a frame or image captured by the camera) and the virtual object overlaid on the displayed real world object. In various examples, the device 105 may include aspects of a virtual reality head-mounted device, smart glasses, a live feed video camera, a GPU, one or more sensors (e.g., such as one or more IMUs, image sensors, microphones, etc.), one or more output devices (e.g., such as speakers, displays, smart glasses, etc.), etc.
[0052] In some cases, head movement 115 may include rotation of user 110's head, translational head movement, etc. Device 105 may update view 130 of user 110 based on head movement 115. For example, device 105 may display view 130-A for user 110 before head movement 115. In some cases, after head movement 115, device 105 may display view 130-B to user 110. As view 130-A shifts to view 130-B, the extended reality system (e.g., device 105) may render or update virtual objects and / or other portions of the virtual environment for display.
[0053] In some cases, extended reality system 100 can provide various types of virtual experiences for a group of users (e.g., including user 110), such as three-dimensional (3D) gaming experiences, social media experiences, collaborative virtual environments, etc. Although some examples provided herein apply to 3D collaborative virtual environments, the systems and techniques described herein are applicable to any type of virtual environment or experience in which virtual representations (or avatars) can be used to represent users or participants of the virtual environment / experience.
[0054] Figure 22 is a diagram illustrating an example of a 3D collaborative virtual environment 200 in which various users interact with each other in a virtual session via virtual representations (or avatars) of the users in the virtual environment 200. The virtual representations include a virtual representation 202 of a first user, a virtual representation 204 of a second user, a virtual representation 206 of a third user, a virtual representation 208 of a fourth user, and a virtual representation 210 of a fifth user. Other contextual information of the virtual environment 200 is also shown, including a virtual calendar 212, a virtual web page 214, and a virtual video conferencing interface 216. Users can experience the virtual environment visually, audibly, tactilely, or otherwise from each user's perspective while interacting with the virtual representations of other users. For example, the virtual environment 200 (represented by virtual representation 202) is shown from the perspective of the first user.
[0055] Figure 3 3 is an image 300 showing examples of virtual representations of various users, including a virtual representation 302 of one of the users. For example, the virtual representation 302 may be Figure 2 The 3D collaborative virtual environment 200 is used.
[0056] Figure 4 4 is a diagram showing an example of a system 400 that can be used to perform the systems and techniques described herein according to various aspects of the present disclosure. As shown, the system 400 includes a client device 405, an animation and scene rendering system 410, and a storage 415. Although the system 400 shows two devices 405, a single animation and scene rendering system 410, a single storage 415, and a single network 420, the present disclosure is applicable to any system architecture with one or more devices 405, animation and scene rendering systems 410, storage 415, and networks 420. In some cases, the storage 415 can be part of the animation and scene rendering system 410. The devices 405, the animation and scene rendering system 410, and the storage 415 can communicate with each other and exchange information supporting the generation of virtual content for XR using a communication link 425 via the network 420, such as multimedia packets, multimedia data, multimedia control information, and gesture prediction parameters. In some cases, a portion of the technology described herein for providing distributed generation of virtual content can be performed by one or more of the devices 405, and a portion of the technology can be performed by the animation and scene rendering system 410, or both.
[0057] Device 405 may be an XR device (e.g., a head mounted display (HMD), XR glasses such as virtual reality (VR) glasses, augmented reality (AR) glasses, etc.), a mobile device (e.g., a cellular phone, a smart phone, a personal digital assistant (PDA), etc.), a wireless communication device, a tablet computer, a laptop computer, and / or other devices that support various types of communication and functional features related to multimedia (e.g., sending, receiving, broadcasting, streaming, sinking, capturing, storing, and recording multimedia data). Additionally or alternatively, device 405 may be referred to by those skilled in the art as user equipment (UE), user equipment, smart phone, Bluetooth device, Wi-Fi device, mobile station, subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, and / or some other suitable term. In some cases, device 405 is also capable of communicating directly with another device (e.g., using a peer-to-peer (P2P) or device-to-device (D2D) protocol, such as using sidelink communication). For example, device 405 can receive or send various information, such as instructions or commands (eg, multimedia-related information), from another device 405 .
[0058] Device 405 may include application 430 and multimedia manager 435. Although system 400 shows device 405 including both application 430 and multimedia manager 435, application 430 and multimedia manager 435 may be optional features for device 405. In some cases, application 430 may be a multimedia-based application that may receive (e.g., download, stream, broadcast) multimedia data from animation and scene rendering system 410, storage 415, or another device 405, or send (e.g., upload) multimedia data to animation and scene rendering system 410, storage 415, or another device 405 via use of communication link 425.
[0059] The multimedia manager 435 may be part of a general purpose processor, a digital signal processor (DSP), an image signal processor (ISP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete gate or transistor logic components, discrete hardware components, or any combination thereof, or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described in the present disclosure, etc. For example, the multimedia manager 435 may process multimedia (e.g., image data, video data, audio data) from the local memory or storage 415 of the device 405 and / or write multimedia data to the local memory or storage 415 of the device 405.
[0060] The multimedia manager 435 may also be configured to provide multimedia enhancement, multimedia restoration, multimedia analysis, multimedia compression, multimedia streaming, and multimedia synthesis, among other functions. For example, the multimedia manager 435 may perform white balancing, cropping, scaling (e.g., multimedia compression), adjusting resolution, multimedia stitching, color processing, multimedia filtering, spatial multimedia filtering, artifact removal, frame rate adjustment, multimedia encoding, multimedia decoding, and multimedia filtering. By further example, in accordance with the techniques described herein, the multimedia manager 435 may process multimedia data to support server-based gesture prediction for XR.
[0061] The animation and scene rendering system 410 may be a server device, such as a data server, a cloud server, a server associated with a multimedia subscription provider, a proxy server, a web server, an application server, a communication server, a home server, a mobile server, an edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a network router, any combination thereof, or other server devices. In some cases, the animation and scene rendering system 410 may include a multimedia distribution platform 440. In some cases, the multimedia distribution platform 440 may be a device or system separate from the animation and scene rendering system 410. The multimedia distribution platform 440 may allow the device 405 to discover, browse, share, and download multimedia via the network 420 using the communication link 425, and thus provide digital distribution of multimedia from the multimedia distribution platform 440. Therefore, digital distribution may be a form of delivering media content such as audio, video, and images without using physical media but through an online delivery medium such as the Internet. For example, the device 405 may upload or download multimedia-related applications for streaming, downloading, uploading, processing, enhancing, etc. multimedia (e.g., images, audio, video). The animation and scene rendering system 410 or the multimedia distribution platform 440 may also send various information to the device 405, such as instructions or commands (e.g., multimedia-related information) for downloading multimedia-related applications on the device 405.
[0062] Storage 415 can store various information, such as instructions or commands (e.g., multimedia-related information). For example, storage 415 can store multimedia 445, information from device 405 (e.g., gesture information, representation information for a virtual representation or avatar of a user, such as codes or features related to facial representation, body representation, hand representation, etc., and / or other information). Device 405 and / or animation and scene rendering system 410 can use communication link 425 to retrieve stored data from storage 415 via network 420 and / or send more data to storage 415. In some examples, storage 415 can be a memory device (e.g., read-only memory (ROM), random access memory (RAM), cache memory, buffer memory, etc.) that stores various information, such as instructions or commands (e.g., multimedia-related information), a relational database (e.g., a relational database management system (RDBMS) or a structured query language (SQL) database), a non-relational database, a network database, an object-oriented database, or other types of databases.
[0063] The network 420 may provide encryption, access authorization, tracking, Internet Protocol (IP) connections, and other access, calculation, modification, and / or functionality. Examples of the network 420 may include any combination of cloud networks, local area networks (LANs), wide area networks (WANs), virtual private networks (VPNs), wireless networks (e.g., using 802.11), cellular networks (using third generation (3G), fourth generation (4G), long term evolution (LTE), or new radio (NR) systems (e.g., fifth generation (5G)), etc.). The network 420 may include the Internet.
[0064] The communication link 425 shown in the system 400 may include an uplink transmission from the device 405 to the animation and scene rendering system 410 and storage 415, and / or a downlink transmission from the animation and scene rendering system 410 and storage 415 to the device 405. The communication link 425 can send two-way communication and / or one-way communication. In some examples, the communication link 425 can be a wired connection or a wireless connection or both. For example, the communication link 425 may include one or more connections, including but not limited to Wi-Fi, Bluetooth, Bluetooth Low Energy (BLE), cellular, Z-WAVE, 802.11, peer-to-peer, LAN, wireless local area network (WLAN), Ethernet, FireWire, fiber optics and / or other connection types associated with wireless communication systems.
[0065] In some aspects, a user of device 405 (referred to as a first user) can participate in a virtual session with one or more other users (including a second user of an additional device). In such an example, animation and scene rendering system 410 can process information received from device 405 (e.g., directly received from device 405, received from storage 415, etc.) to generate and / or animate a virtual representation (or avatar) for the first user. Animation and scene rendering system 410 can constitute a virtual scene, which includes a virtual representation of the user, and in some cases includes background virtual information from the perspective of a second user of an additional device. Animation and scene rendering system 410 can send a frame of a virtual scene (e.g., via network 120) to an additional device. Further details about such aspects are provided below.
[0066] Figure 5 5 is a diagram showing an example of a device 500. The device 500 may be implemented as a client device (eg, Figure 4405) or an animation and scene rendering system (e.g., animation and scene rendering system 410). As shown, device 500 includes a CPU 510 having a central processing unit (CPU) memory 515, a GPU 525 having a GPU memory 530, a display 545, a display buffer 535 storing data associated with rendering, a user interface unit 505, and a system memory 540. For example, system memory 540 can store a GPU driver 520 (shown as being included in CPU 510, as described below) having a compiler, a GPU program, a natively compiled GPU program, etc. The user interface unit 505, CPU 510, GPU 525, system memory 540, display 545, and extended reality manager 550 can communicate with each other (e.g., using a system bus).
[0067] Examples of CPU 510 include, but are not limited to, a digital signal processor (DSP), a general-purpose microprocessor, an application-specific integrated circuit (ASIC), a field-programmable logic array (FPGA), or other equivalent integrated or discrete logic circuits. Figure 5 510 and GPU 525 may be integrated into a single unit. The CPU 510 may execute one or more software applications. Examples of applications may include operating systems, word processors, web browsers, email applications, spreadsheets, video games, audio and / or video capture, playback or editing applications, or other such applications that initiate the generation of image data to be presented via display 545. As shown, the CPU 510 may include a CPU memory 515. For example, the CPU memory 515 may represent an on-chip storage or memory for executing machine or object code. The CPU memory 515 may include one or more volatile or non-volatile memories or storage devices, such as flash memory, magnetic data media, optical storage media, etc. The CPU 510 is capable of reading values from or writing values to the CPU memory 515 faster than reading values from or writing values to the system memory 540, which may be accessed, for example, via a system bus.
[0068] GPU 525 can represent one or more special processors for performing graphics operations. For example, GPU 525 can be a dedicated hardware unit with fixed functions and programmable components for rendering graphics and executing GPU applications. GPU 525 can also include DSP, general-purpose microprocessor, ASIC, FPGA or other equivalent integrated or discrete logic circuits. GPU 525 can be constructed with a highly parallel structure, which provides more efficient processing of operations related to complex graphics than CPU 510. For example, GPU 525 can include multiple processing elements configured to operate multiple vertices or pixels in a parallel manner. The highly parallel nature of GPU 525 can allow GPU 525 to generate graphics images (e.g., graphical user interfaces and two-dimensional or three-dimensional graphics scenes) for display 545 faster than CPU 510.
[0069] In some cases, GPU 525 may be integrated into the motherboard of device 500. In other cases, GPU 525 may be present on a graphics card or other device or component installed in a port in the motherboard of device 500, or may be otherwise incorporated in a peripheral device configured to interoperate with device 500. As shown, GPU 525 may include GPU memory 530. For example, GPU memory 530 may represent on-chip storage or memory for executing machine or object code. GPU memory 530 may include one or more volatile or non-volatile memory or storage devices, such as flash memory, magnetic data media, optical storage media, etc. GPU 525 is able to read values from GPU memory 530 or write values to GPU memory 530 faster than reading values from system memory 540 or writing values to system memory 540, which may be accessed, for example, via a system bus. That is, GPU 525 can read data from GPU memory 530 and write data to GPU memory 530 without using a system bus to access off-chip memory. This operation may allow GPU 525 to operate in a more efficient manner by reducing the need for GPU 525 to read and write data via the system bus, which may experience heavy bus traffic.
[0070] Display 545 represents a unit capable of displaying video, images, text, or any other type of data for consumption by a viewer. In some cases, such as when device 500 is implemented as an animation and scene rendering system, device 500 may not include display 545. Display 545 may include a liquid crystal display (LCD), a light emitting diode (LED) display, an organic LED (OLED), an active matrix OLED (AMOLED), etc. Display buffer 535 represents a memory or storage device dedicated to storing data for presenting images (such as computer-generated graphics, still images, video frames, etc. for display 545). Display buffer 535 may represent a two-dimensional buffer including multiple storage locations. In some cases, the number of storage locations within display buffer 535 may generally correspond to the number of pixels to be displayed on display 545. For example, if display 545 is configured to include 640x480 pixels, display buffer 535 may include 640x480 storage locations storing pixel color and intensity information (such as red, green, and blue pixel values or other color values). The display buffer 535 may store the final pixel value processed for each of the pixels by the GPU 525. The display 545 may retrieve the final pixel value from the display buffer 535 and display the final image based on the pixel value stored in the display buffer 535.
[0071] The user interface unit 505 represents a unit with which a user can interact or otherwise interface to communicate with other units (e.g., CPU 510) of the device 500. Examples of the user interface unit 505 include, but are not limited to, trackballs, mice, keyboards, and other types of input devices. The user interface unit 505 may also be or include a touch screen, and the touch screen may be incorporated as a part of the display 545.
[0072] System memory 540 may include one or more computer-readable storage media. Examples of system memory 540 include, but are not limited to, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disk read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer or processor. System memory 540 can store program modules and / or instructions that can be accessed for execution by CPU 510. In addition, system memory 540 can store user applications and application surface data associated with the application. In some cases, system memory 540 can store information used by other components of device 500 and / or information generated by other components of device 500. For example, system memory 540 can act as device memory for GPU 525, and can store data to be operated by GPU 555 and data generated by operations performed by GPU 525.
[0073] In some examples, system memory 540 may include instructions that cause CPU 510 or GPU 525 to perform functions attributed to CPU 510 or GPU 525 in aspects of the present invention. In some examples, system memory 540 may be considered a non-transitory storage medium. The term "non-transitory" should not be interpreted to mean that system memory 540 is non-removable. For one example, system memory 540 may be removed from device 500 and moved to another device. For another example, a memory substantially similar to system memory 540 may be inserted into device 500. In some examples, a non-transitory storage medium may store data that changes over time (e.g., in RAM).
[0074] The system memory 540 may store the GPU driver 520 and the compiler, GPU program, and natively compiled GPU program. The GPU driver 520 may represent a computer program or executable code for providing an interface to access the GPU 525. The CPU 510 may execute the GPU driver 520 or a portion thereof to interface with the GPU 525, and for this reason, in Figure 5In the example of FIG5 , a GPU driver 520 is shown within the CPU 510. The GPU driver 520 may be accessed by programs or other executable files executed by the CPU 510, including a GPU program stored in the system memory 540. Thus, when one of the software applications executing on the CPU 510 requires graphics processing, the CPU 510 may provide graphics commands and graphics data to the GPU 525 for rendering to the display 545 (e.g., via the GPU driver 520).
[0075] In some cases, the GPU program may include code written in a high-level (HL) programming language, for example, using an application programming interface (API). Examples of APIs include Open Graphics Library ("OpenGL"), DirectX, Render-Man, WebGL, or any other public or proprietary standard graphics API. Instructions may also conform to so-called heterogeneous computing libraries, such as Open Computing Language ("OpenCL"), DirectCompute, etc. Typically, an API includes a predetermined set of standardized commands executed by associated hardware. API commands allow a user to instruct the hardware components of GPU 525 to execute commands without the user knowing the details of the hardware components. In order to process graphics rendering instructions, CPU 510 may issue one or more rendering commands to GPU 525 (e.g., via GPU driver 520) to cause GPU 525 to perform some or all rendering of graphics data. In some examples, the graphics data to be rendered may include a list of graphics primitives (e.g., points, lines, triangles, quadrilaterals, etc.).
[0076] The GPU program stored in the system memory 540 may call or otherwise include one or more functions provided by the GPU driver 520. The CPU 510 typically executes a program in which the GPU program is embedded, and when encountering the GPU program, the GPU program is passed to the GPU driver 520. The CPU 510 executes the GPU driver 520 in this context to process the GPU program. That is, for example, the GPU driver 520 may process the GPU program by compiling the GPU program into an object or machine code that can be executed by the GPU 525. The object code may be referred to as a natively compiled GPU program. In some examples, a compiler associated with the GPU driver 520 may operate in real time or near real time to compile the GPU program during the execution of the program in which the GPU program is embedded. For example, a compiler typically represents a unit for reducing the HL instructions defined according to the HL programming language to the LL instructions of the low-level (LL) programming language. After compilation, these LL instructions can be executed by a specific type of processor or other type of hardware, such as FPGA, ASIC, etc. (including but not limited to CPU 510 and GPU 525).
[0077] exist Figure 5 In an example of , the compiler may receive the GPU program from the CPU 510 when executing the HL code including the GPU program. That is, the software application executed by the CPU 510 may call the GPU driver 520 (e.g., via a graphics API) to issue one or more commands to the GPU 525 for rendering one or more graphics primitives into a displayable graphics image. The compiler may compile the GPU program to generate a natively compiled GPU program that conforms to the LL programming language. The compiler may then output the natively compiled GPU program including LL instructions. In some examples, the LL instructions may be provided to the GPU 525 in the form of a list of drawing primitives (e.g., triangles, rectangles, etc.).
[0078] LL instructions (e.g., which may alternatively be referred to as primitive definitions) may include vertex specifications for specifying one or more vertices associated with the primitive to be rendered. The vertex specifications may include position coordinates for each vertex, and in some cases, other attributes associated with the vertex, such as color coordinates, normal vectors, and texture coordinates. Primitive definitions may include primitive type information, scaling information, rotation information, and the like. Based on instructions issued by a software application (e.g., a program in which a GPU program is embedded), the GPU driver 520 may formulate one or more commands that specify one or more operations for the GPU 525 to perform in order to render the primitive. When the GPU 525 receives a command from the CPU 510, it may decode the command and configure one or more processing elements to perform the specified operation, and may output the rendered data to the display buffer 535.
[0079] The GPU 525 may receive a locally compiled GPU program, and then in some cases, the GPU 525 renders one or more images and outputs the rendered images to the display buffer 535. For example, the GPU 525 may generate several primitives to be displayed at the display 545. Primitives may include one or more of lines (including curves, splines, etc.), points, circles, ellipses, polygons (e.g., triangles), or any other two-dimensional primitives. The term "primitive" may also refer to three-dimensional primitives, such as cubes, cylinders, spheres, cones, pyramids, tori, etc. In general, the term "primitive" refers to any basic geometric shape or element that can be rendered by the GPU 525 to be displayed as an image (or a frame in the context of video data) via the display 545. The GPU 525 may transform primitives and other attributes of primitives (e.g., for defining color, texture, lighting, camera configuration, or other aspects) into a so-called "world space" by applying one or more model transformations (which may also be specified in the state data). Once transformed, the GPU 525 can apply a view transform to the active camera (which can also be specified in the state data defining the camera) to transform the coordinates of the primitives and lights into camera or eye space. The GPU 525 can also perform vertex shading to render the appearance of the primitives according to any active lights. The GPU 525 can perform vertex shading in one or more of the above-mentioned model, world, or view space.
[0080] Once the primitives are shaded, the GPU 525 can perform projection to project the image into the canonical view volume. After transforming the model from the eye space to the canonical view volume, the GPU 525 can perform clipping to remove any primitives that do not reside at least partially within the canonical view volume. For example, the GPU 525 can remove any primitives that are not within the camera's frame. The GPU 525 can then map the coordinates of the primitives from the view volume to the screen space, effectively reducing the three-dimensional coordinates of the primitives to the two-dimensional coordinates of the screen. Given the transformed and projected vertices for defining the primitives with their associated shading data, the GPU 525 can then rasterize the primitives. In general, rasterization can refer to the task of taking an image described in a vector graphics format and converting it into a raster image (e.g., a pixelated image) to be output on a video display or stored in a bitmap file format.
[0081] The GPU 525 may include a dedicated fast bin buffer (e.g., a fast memory buffer such as GMEM, which may be referred to by the GPU memory 530). As discussed herein, the rendering surface may be divided into bins. In some cases, the bin size is determined by the format (e.g., pixel color and depth information) and the rendering target resolution divided by the total amount of GMEM. The number of bins may vary based on the device 500 hardware, the target resolution size, and the target display format. The renderer may draw (e.g., render, write, etc.) pixels into GMEM (e.g., with a high bandwidth that matches the capabilities of the GPU). The GPU 525 may then parse the GMEM (e.g., write the mixed pixel values as a single layer from the GMEM burst to the display buffer 535 or the frame buffer in the system memory 540). This may be referred to as bin-based or tile-based rendering. When all bins are completed, the driver may swap buffers and start the binning process again for the next frame.
[0082] For example, the GPU 525 may implement a tile-based architecture that renders an image or render target by dividing the image into multiple parts, which are referred to as tiles or bins. The size of the bins may be determined based on the size of the GPU memory 530 (e.g., which may alternatively be referred to herein as GMEM or cache), the resolution of the display 545, the color or Z precision of the render target, etc. When implementing tile-based rendering, the GPU 525 may perform a binning pass and one or more rendering passes. For example, with respect to the binning pass, the GPU 525 may process the entire image and sort the rasterized primitives into bins.
[0083] Device 500 may use sensor data, sensor statistics, or other data from one or more sensors. Some examples of monitored sensors may include an IMU, an eye tracker, a tremor sensor, a heart rate sensor, etc. In some cases, an IMU may be included in device 500 and may use some combination of an accelerometer, a gyroscope, or a magnetometer to measure and report specific forces, angular rates, and sometimes the orientation of the body.
[0084] As shown, device 500 may include an extended reality manager 550. Extended reality manager 550 may implement aspects of extended reality, augmented reality, virtual reality, etc. In some cases, such as when device 500 is implemented as a client device (e.g., Figure 4When the device 500 is located in a physical environment such as a device 405, the extended reality manager 550 can determine information associated with the user of the device and / or the physical environment in which the device 500 is located, such as facial information, body information, hand information, device posture information, audio information, etc. The device 500 can send the information to the animation and scene rendering system (e.g., the animation and scene rendering system 410). In some cases, such as when the device 500 is implemented as an animation and scene rendering system (e.g., Figure 4 When using the animation and scene rendering system 410), the extended reality manager 550 can process information provided by the client device as input information to generate and / or animate a virtual representation of a user of the client device.
[0085] A virtual representation (e.g., an avatar) is an important component of a virtual environment. A virtual representation (or avatar) is a 3D representation of a user and allows the user to interact with a virtual scene. As previously mentioned, there are different ways to represent a virtual representation (e.g., an avatar) of a user and the corresponding animation data. For example, an avatar can be purely synthetic, or it can be an accurate representation of the user (e.g., Figure 3 302 shown in the image of FIG. 1 ). The virtual representation (or avatar) may need to be captured or reoriented in real time to reflect the user's actual movements, body postures, facial expressions, etc. Due to the many ways of representing avatars and corresponding animation data, it may be difficult to integrate every single variation of these representations into the scene description.
[0086] As previously described, systems and techniques are described herein for providing encoding of a virtual representation (e.g., an avatar) in a scene description. As described herein, systems and techniques may decouple the representation of a virtual representation (or avatar) and its animation data from the integration of the avatar in the scene description. For example, systems and techniques may perform virtual representation (or avatar) reconstruction to generate a dynamic mesh for representing a user's virtual representation (or avatar), which may allow systems and techniques to deconstruct the virtual representation (or avatar) into multiple mesh nodes. Each mesh node may correspond to a body part of the virtual representation (or avatar). Multiple mesh nodes enable the XR system to support interactivity with various parts of the virtual representation (e.g., with the hands of an avatar).
[0087] Various animation assets may be needed to model an avatar, including a mesh (e.g., a 3D mesh, such as a triangle mesh, which includes multiple vertices and line segments connecting the vertices), a diffuse or albedo texture, a normal specular texture, and in some cases other types of textures. These various assets may be obtained from registration or offline reconstruction. Figure 6 is a diagram illustrating examples of a normal map 602 , an albedo map 604 , and a specular map 606 .
[0088] Animation of a virtual representation (eg, avatar) may be performed using various techniques. Figure 7 700 is a diagram illustrating an example of a technique for performing avatar animation. As shown, camera sensors of a head mounted display (HMD) are used to capture images of a user's face, including an eye camera for capturing images of the user's eyes, a face camera for capturing visible portions of the face (e.g., mouth, chin, cheeks, a portion of the nose, etc.), and other sensors for capturing other sensor data (e.g., audio, etc.). Facial animation can then be performed to generate a 3D mesh and texture for a 3D facial avatar. The mesh and texture can then be rendered by a rendering engine to generate a rendered image.
[0089] In some cases, blend shapes may be utilized or used to perform facial animation. Figure 8 800 is a diagram illustrating an example of performing facial animation utilizing blend shapes. As shown, the system may estimate a coarse or approximate 3D mesh 806 and blend shapes from an image 802 (e.g., captured using a sensor of an HMD or other XR device) using a 3D deformable model (3DMM) encoder 804. The system may generate textures using one or more techniques, such as using a machine learning system 808 (e.g., one or more neural networks) or computer graphics techniques (e.g., Metahumans). In some cases, the system may need to compensate for misalignment due to coarse geometry, for example, as described in U.S. non-provisional application Ser. No. 17 / 845,884, filed on June 21, 2022, entitled “VIEW DEPENDENT THREE-DIMENSIONAL MORPHABLE MODELS,” the entire contents of which are incorporated herein by reference and for all purposes.
[0090] A 3DMM is a 3D facial mesh representation of known topology. A 3DMM can be linear or non-linear. Fig. 9 is a diagram showing an example of a system 900 that can generate a 3DMM facial model or mesh 904. The system 900 can obtain a dataset of 3D and / or color images (and in some cases grayscale images) for various people from a database 902. The system 900 can also obtain a known mesh topology of a facial mesh model 906 corresponding to the face of the image in the database 902. In some cases, principal component analysis (PCA) can be used to find a representation of an identifier (ID) / expression in the case of a linear representation. Expressions can also be modeled via mixed shapes (e.g., meshes) at various states or expressions. Using these parameters, the system can manipulate or guide the mesh. The 3DMM can be generated as follows:
[0091]
[0092] The output may include the average shape S0 , shape parameter a i 、Shape BasicsU i , expression parameter b j and Expression Basic or Blend Shape V j .
[0093] In some cases, 3DMM encoding can be used to determine blend shapes. The blend shapes can then be used to reconstruct the deformed mesh, such as to animate an avatar. For example, Fig.10 As shown, animating an avatar can be summarized as given an input image, determining the weights of each blend shape. This technique is described in U.S. non-provisional application Ser. No. 17 / 384,522, filed on July 23, 2021, entitled “ADAPTIVE BOUNDING FOR THREE-DIMENSIONAL MORPHABLE MODELS,” the entire contents of which are incorporated herein by reference and for all purposes. The 3DMM equation S from above is Fig.10 and are provided again below:
[0094]
[0095] And it can also be expressed as:
[0096] s=π(S·R+t)·F / z
[0097] Where S 0 is the average 3D shape, π is the chosen matrix used to obtain x,y coordinates, z is a constant, R is the rotation matrix according to pitch, yaw, roll, and t is the translation vector.
[0098] Fig.11 is a diagram showing an example of using a 3DMM fit curve to drive a virtual representation (or avatar) utilizing a metahuman using the techniques described above.
[0099] Fig.12 is a diagram showing an example of the end-to-end flow of the system. Fig.12 The process of may represent a general process capable of running end-to-end for 1-to-1 communication between user device A and user device B, such as described in U.S. Provisional Application No. 63 / 371,714, entitled “DISTRIBUTED GENERATION OF VIRTUAL CONTENT”, filed on August 17, 2022, which is incorporated herein by reference in its entirety and for all purposes. In some cases, Fig.12 The system may include additional server nodes that can be used in multi-user scenarios.
[0100] Fig.13is a diagram showing an example of performing avatar animation. For example, in order to improve the realism of the avatar animation, various techniques can be performed. For example, the system can estimate a coarse (or approximate) mesh and blend shapes from an image (e.g., an HMD image) using a 3DMM encoder. The system can use an additional facial part neural network to improve the realism of the texture, such as using the techniques described in U.S. non-provisional application 17 / 813,556, entitled "FACIAL TEXTURE SYNTHESIS FOR THREE-DIMENSINALMORPHABLE MODELS," filed on July 19, 2022, which is incorporated herein by reference in its entirety and for all purposes. As shown, the facial part network can be a neural network that includes an encoder and a decoder. In some cases, the facial part network can be a neural network that includes only an encoder (rather than an encoder-decoder), in which case the encoder will output features (e.g., a feature vector or an embedding vector) for representing the portion of the input image corresponding to the face. The system can improve the realism of the mesh by deforming the mesh vertices for a more personalized mesh, such as using the techniques described in U.S. non-provisional application Ser. No. 17 / 714,743, filed Apr. 6, 2022, which is incorporated herein by reference in its entirety and for all purposes. In some aspects, audio can also be fused to the Fig.13 to improve the authenticity of the input to the system, such as using the techniques described in U.S. non-provisional application Ser. No. 17 / 930,244, filed on September 7, 2022, and U.S. non-provisional application Ser. No. 17 / 930,257, filed on September 7, 2022, the entire contents of both applications are incorporated herein by reference and for all purposes.
[0101] Fig.14 is a diagram illustrating an example of an XR system 1400 configured with avatar call flows directly between client devices in accordance with aspects of the present disclosure. As shown, Fig.14 The first client device 1402 of the first user and the second client device 1404 of the second user are included. In one illustrative example, the first client device 1402 is a first XR device (e.g., an HMD configured to display VR, AR, and / or other XR content), and the second client device 1404 is a second XR device (e.g., an HMD configured to display VR, AR, and / or other XR content). Fig.14 Two client devices are shown, but Fig.14 The call flow can be between a first client device 1402 and a plurality of other client devices. The XR system 1400 shows a first client device 1402 and a second client device 1404 participating (in some cases, with Fig.14 ) virtual sessions of other client devices not shown in FIG. Fig.14 In the example, client device 1402 can be considered a source (e.g., a source of information for generating at least one frame for a virtual scene), and second client device 1404 can be considered a target (e.g., a target for receiving at least one frame generated for the virtual scene).
[0102] As described above, the information sent by the client device may include: information representing the face of the user of the first client device 1402 (e.g., a code or feature representing the appearance of the face or other information), information representing the body of the user of the first client device 1402 (e.g., a code or feature representing the appearance of the body, the posture of the body or other information), information representing one or more hands of the user of the first client device 1402 (e.g., a code or feature representing the appearance of the hand, the posture of the hand or other information), posture information of the first client device 1402 (e.g., a posture in six degrees of freedom (6-DOF), referred to as a 6-DOF posture), audio associated with the environment in which the first client device 1402 is located, any combination thereof, and / or other information.
[0103] For example, the first computing device 1402 may include a face encoder 1409, a geometry encoder engine 1411, a gesture engine 1412, a body encoder 1414, a hand engine 1416, and an audio decoder 1418. In some aspects, the first client device 1402 may include a body encoder 1414 configured to generate a virtual representation of the user's body. The first client device 1402 may include, in addition to Fig.14 Other components or engines other than those shown in Figure 5 One or more components on the device 500, Fig. 20 One or more components of the computing system 2000, etc.). The second client device 1404 is shown as including a user virtual representation system 1420, an audio decoder 1425, a spatial audio engine 1426, a lip synchronization engine 1428, a reprojection engine 1434, a display 1436, and a future pose prediction engine 1438. In some cases, each client device (e.g., the first client device 1402, the second client device 1402, other client devices) may include a face engine, a pose engine, a body engine, a hand engine, an audio decoder (and in some cases an audio decoder or a combined audio encoder-decoder), a video decoder (and in some cases a video encoder or a combined video encoder-decoder), a reprojection engine, a display, and a future pose prediction engine. Fig.14 These engines or components are not shown in relation to the first client device 1402 and the second client device 1404 because in Fig.14In the example of FIG. 14 , the first client device 1402 is the source device and the second client device 1404 is the target device.
[0104] The facial encoder 1409 of the first client device 1402 may receive one or more input frames 1415 from one or more cameras of the first client device 1402. For example, the input frames 1415 received by the facial encoder 1409 may include frames (or images) captured by one or more cameras having a field of view of the mouth of the user of the first client device 1402, the user's left eye, and the user's right eye. Other images may also be processed by the facial encoder 1409. In some cases, the input frames 1415 may be included in a sequence of frames (e.g., a video, a sequence of independent or still images, etc.). The facial encoder 1409 may generate and output a code (e.g., one or more feature vectors) for representing the face of the user of the first client device 1402. The facial encoder 1409 may send the code representing the user's face to the user virtual representation system 1420 of the second client device 1404. In an illustrative example, the facial encoder 1409 may include one or more facial encoders including one or more machine learning systems (e.g., deep learning networks, such as deep neural networks), which are trained to represent the user's face using codes or feature vectors. In some cases, the facial encoder 1409 may include a separate encoder for each type of image processed by the facial encoder 1409, such as a first encoder for frames or images of the mouth, a second encoder for frames or images of the right eye, and a third encoder for frames or images of the left eye. Training may include supervised learning (e.g., using labeled images and one or more loss functions, such as mean square error MSE), semi-supervised learning, unsupervised learning, etc.). For example, a deep neural network may generate and output a code (e.g., one or more feature vectors) for representing the face of a user of the first client device 1402. The code may be a latent code (or bitstream) that can be decoded by a facial decoder (not shown) of the user virtual representation system 1420, which is trained to decode the code (or feature vector) for representing the face of the user in order to generate a virtual representation of the face (e.g., a facial mesh). For example, the facial decoder of the user virtual representation system 1420 may decode the code received from the first client device 1402 to generate a virtual representation of the user's face.
[0105] The geometry encoder engine 1411 of the first client device 1402 can generate a 3D model (e.g., a 3D deformable model or 3DMM) of a user's head or face based on one or more frames 1415. In some aspects, the 3D model can include representations of facial expressions in frames from the one or more frames 1415. In an illustrative example, the facial expression representation can be formed by a blend shape. The blend shape can semantically represent the movement of a muscle or part of a facial feature (e.g., opening / closing of the jaw, raising / lowering of eyebrows, opening / closing of eyes, etc.). In some cases, each blend shape can be represented by a blend shape coefficient paired with a corresponding blend shape vector. In some examples, the facial model can include a representation of the facial shape of the user in the frame. In some cases, the facial shape can be represented by a facial shape coefficient paired with a corresponding facial shape vector. In some embodiments, the geometry encoder engine 1411 (e.g., a machine learning model) can be trained (e.g., during a training process) to implement a consistent facial shape (e.g., consistent facial shape coefficient) for the 3D facial model, regardless of the pose (e.g., pitch, yaw, and roll) associated with the 3D facial model. For example, when a 3D facial model is rendered into a 2D image or frame for display, a projection technique may be used to project the 3D facial model onto the 2D image or frame.
[0106] In some aspects, the blend shape coefficients can be further refined by additionally introducing a relative pose between the first client device 1402 and the second client device 1404. For example, a neural network can be used to process the relative pose information to generate a view-dependent 3DMM geometry that approximates the ground truth geometry. In another example, the facial geometry can be determined directly from the input frame 1415, such as by estimating additive vertex residuals from the input frame 1415 and producing more accurate facial expression details for texture synthesis.
[0107] The posture engine 1412 can determine the posture (e.g., 6-DOF posture) of the first client device 1402 in the 3D environment (and therefore determine the head posture of the user of the first client device 1402). The 6-DOF posture can be an absolute posture. In some cases, the posture engine 1412 may include a 6-DOF tracker, which can track three degrees of rotation data (e.g., including pitch, roll, and yaw) and three degrees of translation data (e.g., horizontal displacement, vertical displacement, and depth displacement relative to a reference point). The posture engine 1412 (e.g., 6-DOF tracker) can receive sensor data as input from one or more sensors. In some cases, the posture engine 1412 includes one or more sensors. In some examples, the one or more sensors may include one or more inertial measurement units (IMUs) (e.g., accelerometers, gyroscopes, etc.), and the sensor data may include IMU samples from one or more IMUs. The posture engine 1412 can determine raw posture data based on the sensor data. The raw gesture data may include 6DOF data representing the gesture of the first client device 1402, such as three-dimensional rotation data (e.g., including pitch, roll, and yaw) and three-dimensional translation data (e.g., horizontal displacement, vertical displacement, and depth displacement relative to a reference point).
[0108] The body encoder 1414 may receive one or more input frames 1415 from one or more cameras of the first client device 1402. The frames received by the body encoder 1414 and the camera used to capture the frames may be the same or different from the frames received by the face encoder 1409 and the camera used to capture those frames. For example, the input frames received by the body encoder 1414 may include frames (or images) of a field of view of the body of the user of the first client device 1402 (e.g., parts other than the face, such as the user's neck, shoulders, torso, lower body, feet, etc.) captured by one or more cameras. The body encoder 1414 may perform one or more techniques to output a representation of the body of the user of the first client device 1402. In one example, the body encoder 1414 may generate and output a 3D mesh representing the shape of the body (e.g., including multiple vertices, edges, and / or faces in 3D space). In another example, the body encoder 1414 may generate and output a code (e.g., a feature vector or multiple feature vectors) for representing the body of the user of the first client device 1402. In one illustrative example, the body encoder 1414 may include one or more body encoders including one or more machine learning systems (e.g., deep learning networks, such as deep neural networks) that are trained to represent the user's body using codes or feature vectors. Training may include supervised learning (e.g., using labeled images and one or more loss functions, such as MSE), semi-supervised learning, unsupervised learning, etc.). For example, the deep neural network may generate and output a code (e.g., one or more feature vectors) for representing the body of the user of the first client device 1402. The code may be a latent code (or bitstream) that can be decoded by a body decoder (not shown) of the user virtual representation system 1420, which is trained to decode the code (or feature vector) for representing the user's virtual body in order to generate a virtual representation of the body (e.g., a body mesh). For example, the body decoder of the user virtual representation system 1420 may decode the code received from the first client device 1402 to generate a virtual representation of the user's body.
[0109] The hand engine 1416 may receive one or more input frames 1415 from one or more cameras of the first client device 1402. The frames received by the hand engine 1416 and the camera used to capture the frames may be the same or different than the frames received by the face encoder 1409 and / or the body encoder 1414 and the camera used to capture those frames. For example, the input frames received by the hand engine 1416 may include frames (or images) of a field of view of the hands of the user of the first client device 1402 captured by one or more cameras. The hand engine 1416 may perform one or more techniques to output a representation of one or more hands of the user of the first client device 1402. In one example, the hand engine 1416 may generate and output a 3D mesh representing the shape of the hand (e.g., including multiple vertices, edges, and / or faces in 3D space). In another example, the hand engine 1416 may output a code (e.g., one or more feature vectors) representing one or more hands. For example, the hand engine 1416 may include one or more body encoders including one or more machine learning systems (e.g., deep learning networks, such as deep neural networks) that are trained (e.g., using supervised learning, semi-supervised learning, unsupervised learning, etc.) to represent the user's hands with codes or feature vectors. The code may be a latent code (or bitstream) that can be decoded by a hand decoder (not shown) of the animation and scene rendering system 1410, which is trained to decode the code (or feature vector) used to represent the user's virtual hand in order to generate a virtual representation of the body (e.g., a body mesh). For example, the hand decoder of the user virtual representation system 1420 may decode the code received from the first client device 1402 to generate a virtual representation of the user's hand (or in some cases, a hand).
[0110] An audio decoder 1418 (e.g., an audio encoder or a combined audio encoder-decoder) can receive audio data, such as audio obtained using one or more microphones 1417 of the first client device 1402. The audio decoder 1418 can encode or compress the audio data and send the encoded audio to an audio decoder 1425 of the second client device 1404. The audio decoder 1418 can perform any type of audio decoding to compress the audio data, such as a coding technique based on a modified discrete cosine transform (MDCT). The encoded audio can be decoded by an audio decoder 1425 of the second client device 1404, which is configured to perform an inverse process of the audio encoding process performed by the audio encoder 1418 to obtain decoded (or decompressed) audio.
[0111] The user virtual representation system 1420 of the second client device 1404 may receive input information received from the first client device 1402 and use the input information to generate and / or animate a virtual representation (or avatar) of the user of the first client device 1402. In some aspects, the user virtual representation system 1420 may also use the future predicted pose of the second client device 1404 to generate and / or animate the virtual representation of the user of the first client device 1402. In such aspects, the future pose prediction engine 1438 of the second client device 1404 may predict a pose of the second client device 1404 (e.g., corresponding to a predicted head pose, body pose, hand pose, etc. of the user) based on the target pose and send the predicted pose to the user virtual representation system 1420. For example, the future pose prediction engine 1438 may predict a future pose of the second client device 1404 at a future time (e.g., time T which may be the prediction time) based on the model (e.g., corresponding to a head position, head orientation, line of sight (such as a head position, ... Figure 1 The future time T may correspond to the time at which the second client device 1404 will output or display the target view frame. As used herein, references to a pose of a client device (e.g., second client device 1404) and references to a head pose, body pose, etc., of a user of the client device may be used interchangeably.
[0112] Predicting poses can be useful when generating virtual representations because, in some cases, virtual objects can appear delayed to a user when compared to the object's expected view to the user or compared to the real-world object that the user is viewing (e.g., in an AR, MR, or VR see-through scene). Figure 1 As an illustrative example, in the absence of head motion or pose prediction, updating of a virtual object in view 130-B from a previous view 130-A may be delayed until a head pose measurement is made so that the position, orientation, size, etc. of the virtual object can be updated accordingly. In some cases, the delay may be due to system latency (e.g., end-to-end system latency between the second client device 1404 and a system or device used to render virtual content (such as the user virtual representation system 1420)), which may be caused by rendering, time warping, or both. In some cases, such delays may be referred to as round-trip latency or dynamic registration error. In some cases, the error may be large enough that the user of the second client device 1404 may perform a head movement (e.g., a head gesture measurement may be made) before the temporal pose measurement can be ready for display. Figure 1130 -B). Therefore, it may be beneficial to predict the head movement 115 so that virtual objects associated with the view 130 -B can be determined and updated in real time based on the prediction (e.g., pattern) in the head movement 115.
[0113] As described above, the user virtual representation system 1420 can use the input information received from the first client device 1402 and the future predicted pose information from the future pose prediction engine 1438 to generate and / or animate a virtual representation of the user of the first client device 1402. As described herein, animating the virtual representation of the user can include modifying the position, movement, mannerisms, or other features of the virtual representation to match the corresponding position, movement, mannerisms, etc. of the user in the real world or physical space. In some aspects, the user virtual representation system 1420 can be implemented as a deep learning network (such as a neural network (e.g., a convolutional neural network (CNN), an autoencoder, or other type of neural network) trained based on input training data (e.g., a code representing the user's face, a code representing the user's body, a code representing the user's hands, pose information (such as 6-DOF pose information of the user's client device, inverse kinematics information, etc.)) using supervised learning, semi-supervised learning, unsupervised learning, etc.) to generate and / or animate the virtual representation of the user of the client device.
[0114] In some cases, the user virtual representation system 1420 may receive user registration data 1450 of a first user (user A) of the first client device 1402 to generate a virtual representation (e.g., an avatar) for the first user. The registration data 1450 for the first user may include mesh information. The mesh information may include information for defining a mesh of an avatar of the first user of the first client device 1402 and other assets associated with the avatar (e.g., a normal map, an albedo map, a specular reflection map, etc.). In some cases, the mesh animation parameters may include facial part codes from the facial encoder 1409, facial blend shapes from the geometry encoder engine 1411 (which may include a 3DMM head encoder in some cases), hand joint codes from the hand engine 1416, head poses from the pose engine 1412, and in some cases an audio stream from the audio encoder 1418.
[0115] In some cases, spatial audio engine 1426 may receive as input decoded audio generated by audio decoder 1425, and in some cases future predicted gesture information from future gesture prediction engine 1438. The input is used to generate audio that is spatially directed according to the gesture of the user of second client device 1404. Lip sync engine 1428 may synchronize the animation of the lips of the virtual representation of the user of first client device 1402 depicted in display 1436 with the spatial audio output by spatial audio engine 1426.
[0116] The reprojection engine 1434 may perform reprojection to reproject the virtual content of the decoded target view frame according to the predicted pose determined by the future pose prediction engine 1438. The reprojected target view frame may then be displayed on the display 1436 so that a user of the second client device 1404 may view the virtual scene from the user's perspective.
[0117] Fig.15 is a diagram illustrating an example of an example of an XR system 1500 configured with avatar call flows directly between client devices in accordance with aspects of the present disclosure. Fig.15 As shown, the XR system 1500 includes Fig.14 Some components of the XR system 1400 of the present invention are the same (with the same reference numerals). For example, the facial encoder 1409, the geometry encoder engine 1411, the posture engine 1412, the body encoder 1414, the hand engine 1416, the audio decoder 1418, the user virtual representation system 1420 that can receive the user registration data 1450, the audio decoder 1425, the spatial audio engine 1426, the lip synchronization engine 1428, the reprojection engine 1434, the display 1436 and the future posture projection engine 1438 are configured to perform the same Fig.14 Same or similar operation of XR system 1400 .
[0118] The XR system 1500 also includes a scene composition system 1522, a video encoder 1530, and a video decoder 1532 that can receive background scene information 1519. In some cases, the user virtual representation system 1420 can output a virtual representation of a user of the first client device 1502 to the scene composition system 1522, which can be part of or implemented by the server device 1505. The background scene information 1519 and virtual representations of other users participating in the virtual session (if any) can also be provided to the scene composition system 1522. The background scene information 1519 can include information about the scene, such as lighting of the virtual scene, virtual objects in the scene (e.g., virtual buildings, virtual streets, virtual animals, etc.), and / or other details related to the virtual scene (e.g., sky, clouds, etc.). In some cases, such as in an AR, MR, or VR see-through setting (wherein a video frame of a real-world environment is displayed to a user), the lighting information may include the lighting of the real-world environment in which the first client device 1502 and / or the second client device 1504 (and / or other client devices participating in the virtual session) are located. In some cases, the future predicted poses from the future pose prediction engine 1438 of the second client device 1504 may also be input to the scene composition system 1522.
[0119] Using the virtual representation of the user of first client device 1502, background scene information 1519, and virtual representations of other users (if any) (and in some cases, future predicted gestures), scene synthesis system 1522 can synthesize a target view frame for the virtual scene using a view of the virtual scene from the perspective of the user of second client device 1504 based on the relative difference between the gesture at second client device 1504 and each corresponding gesture of first client device 1502 and any other client devices of the user participating in the virtual session. For example, the synthesized target view frame can include a mixture of the virtual representation of the user of first client device 1502, background scene information 1519 (e.g., lighting, background objects, sky, etc.), and virtual representations of other users that may be involved in the virtual session. The gestures of the virtual representations of any other users are also based on the gesture of second client device 1504 (corresponding to the gesture of the user of second client device 1504).
[0120] The video encoder 1530 may encode (or compress) the target view frame from the scene composition system 1522 using a video coding technique (e.g., according to any suitable video codec, such as Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), Moving Picture Experts Group (MPEG), etc.). The video encoder 1530 may then send the encoded target view frame to the second client device 1504 via the network.
[0121] The video decoder 1532 of the second client device 1504 can obtain the encoded target view frame and can decode the encoded target view frame using the inverse of the video decoding technique performed by the video encoder 1530 (e.g., according to a video codec such as AVC, HEVC, VVC, etc.). The reprojection engine 1434 can perform reprojection to reproject the virtual content of the decoded target view frame according to the predicted pose determined by the future pose prediction engine 1438. The reprojected target view frame can then be displayed on the display 1436 so that a user of the second client device 1504 can view the virtual scene from the user's perspective.
[0122] In some cases, for full body pose estimation and avatar animation, the system can predict body shape / pose parameters. A parameterizable SMPL or ADAM body model like 3DMM can be used. The prediction can occur based on images captured using a camera on an XR device (e.g., HMD) and / or attached body sensors. The system can apply shape and pose deformations to the base model.
[0123] Various aspects of 3D reconstruction can be challenging, including mouth opening, hair and facial hair, eyes, emotions, interactivity with virtual representations or avatars, etc.
[0124] There are various standards related to animation. For example, W3D has created a standard on humanoid animation, https: / / www.web3d.org / documents / specifications / 19774-1 / V2.0 / index.html. This standard defines the joints of a humanoid model and their layering, and defines a layered model of a humanoid. The humanoid parts include body segments, joints, a skeleton, a skin with normals and coordinates, and transformations. It also defines where interactivity can be attached. Animation is defined in section 2 at https: / / www.web3d.org / documents / specifications / 19774-2 / V2.0 / index.html, which supports interpolation and moving objects.
[0125] There are various problems associated with virtual representations (e.g., avatars) of users in virtual environments. For example, different applications / platforms may use different representations of virtual representations (or avatars). Furthermore, in a shared space, the virtual representations (or avatars) from all participants must be synthesized into a single scene. It may be difficult to support a wide range of virtual representation (or avatar) representations in a scene description. Solutions to such problems should support a wide range of virtual representations (e.g., avatar representations), captured and synthesized avatars, animated and frame-by-frame avatars, and interactivity involving different parts of the avatar.
[0126] The systems and techniques described herein provide a way to integrate virtual representations (e.g., avatars) into scene descriptions. In some cases, a virtual scene can be described by a schema such as a graphics language transmission format (glTF). glTF can describe a virtual scene using multiple hierarchical tree structures that describe the environment of the scene, objects in the scene, and the like. In some cases, glTF can also be used to describe a virtual representation. For example, a virtual representation can include a mesh (e.g., a head mesh, a body mesh, and the like) onto which a texture can be overlaid. In some cases, it may be useful to map segments (e.g., portions) of a mesh to humanoid parts (such as body segments) that can be defined in glTF. For example, animations and interactions can be defined based on a hierarchical model of a humanoid, where certain animations and / or interactions are defined based on human components such as body segments, joints, and the like. For example, a glTF node can be associated with a hand that can be mapped to a portion of a mesh and associated with the ability to touch other objects in the environment (e.g., interact with it). These interactions and / or animations can define how portions of the body mesh can be deformed, moved, distorted, and the like.
[0127] Fig.16 1 is a block diagram illustrating an example of a virtual representation (or avatar) reconstruction system or pipeline 1600 according to aspects of the present disclosure. In some cases, the virtual representation reconstruction system 1600 may be included as Fig.14 and 15 The virtual representation reconstruction system 1600 may include a mesh generation engine 1602, a set of buffers 1604, and a rendering engine 1606.
[0128] The mesh generation engine 1602 may include an avatar reconstruction engine 1608 that receives input information. In some cases, the input information may be received as one or more data streams or channels. For example, the input information may be received via a set of data streams including a stream for a base model, which may be a generic mesh model of a virtual representation, texture information, a deformation map, parameterized data, and static metadata. The input information may be provided as input to the avatar reconstruction engine 1608.
[0129] The avatar reconstruction engine 1608 may generate components of the virtual representation, such as vertices for the geometry of the mesh, texture information, skinning information for the texture, positions of joints, interactivity information, etc., as a 3D mesh. In some cases, the avatar reconstruction engine 1608 may include a set of machine learning (ML) models and / or algorithms. The components of the virtual representation may be stored in the set of buffers 1604. The set of buffers 1604 may include buffers for various types of information, such as vertex information for the mesh, texture information for the mesh, skinning / joint information for the mesh, attribute information for the mesh, interactivity and / or metadata for the mesh, etc. In some cases, the components of the virtual representation may be stored in a single combined buffer.
[0130] The output of the avatar reconstruction engine 1608 can be output from the set of buffers 1604 to the rendering engine 1606. For example, the rendering engine 1606 can receive the 3D mesh from the set of buffers 1604 and render the virtual representation. As shown, the rendering engine 1606 is not dependent on specific input information for the avatar reconstruction engine 1608. For example, the rendering engine receives and processes the relevant mesh to be rendered from the avatar reconstruction engine 1608 without considering the input information provided to the avatar reconstruction engine 1608. The rendering engine 1606 can therefore render the data in the buffers, thereby allowing the specific input information for the avatar reconstruction engine 1608 to change without affecting the rendering engine 1606.
[0131] In some aspects, the input information to the avatar reconstruction engine 1608 may vary, for example, based on the format (or type) of the virtual representation. In some cases, the format of the virtual representation may vary based on the specific vendor responsible for the virtual representation, the complexity of the virtual representation, the specific computing system used, any combination thereof, and / or other information. For example, a virtual representation generated by a first vendor may include different input information streams that may be processed differently by the avatar reconstruction engine 1608 than a virtual representation generated by a second vendor. In another example, the input information provided for a virtual representation from one vendor may not include a specular texture, may use a different base model, may interact at a different location, etc., compared to a virtual representation from another vendor. In some cases, the format of the virtual representation may vary, for example, based on account level (e.g., premium account, regular account, etc.), the device being used, available bandwidth, etc.
[0132] By decoupling the rendering engine 1606 from specific input information, different formats of virtual representations can be accommodated by adapting the avatar reconstruction engine 1608 to virtual representations of different formats. For example, the avatar reconstruction engine 1608 can have multiple ML models, each ML model is trained to generate meshes from one or more virtual representations of different formats, or multiple avatar reconstruction engines 1608 can be used to generate meshes from one or more virtual representations of different formats. In some cases, while different input information can be received for different virtual representation formats, the input information itself can be arranged into and / or described by a common schema or format (such as gLTF).
[0133] In some cases, it may be useful to provide enhancement techniques to integrate virtual representations into such a schema by extending a mesh element (e.g., a gLTF mesh element) to represent a virtual representation or a portion of a virtual representation (e.g., an avatar or a portion of an avatar). For example, a common schema may be defined to accommodate virtual representations of different formats. Each part of a virtual representation (or avatar) may be associated with a part of a humanoid and may be associated with certain interactive behaviors. In some cases, a root node of a virtual representation (or avatar) may describe how the virtual representation is represented. The root node may have one or more child nodes, each associated with a humanoid part. For example, a child mesh node may indicate which humanoid part applies to the child mesh node by a path scheme (e.g., " / humanoid / arm / left / hand").
[0134] Fig.17 is a diagram illustrating the structure of a virtual representation (or avatar) in glTF 1700 according to aspects of the present disclosure. Fig.17 A root node 1702 (e.g., a parent node) is included, under which child nodes representing virtual representations are hierarchically arranged. In some cases, this hierarchical arrangement can be based on humanoid parts, such as body segments (such as body (e.g., torso), arms, hands, fingers, legs, etc.), with smaller sub-segments (such as fingers) represented by child nodes (e.g., sub-nodes) of larger parts (such as hands). In some cases, body segments (and sub-segments) and hierarchies for nodes can be defined based on a hierarchical model of humanoids such as provided in the W3D standard.
[0135] In some cases, a virtual representation framework for generating a virtual representation (e.g., for generating a virtual representation in a particular virtual representation format) specifies how information is laid out in input information and how segments (or processed segments) of the input information are mapped into the node structure of the glTF schema. For example, in order to allow different virtual representation formats to be used with a common rendering engine, a structure for interpreting different virtual representation formats can be provided. As an example of such a structure, a root node 1702 can include one or more fields 1704 that help allow for adaptation to different virtual representation formats. Field 1704 can include a type field 1706, a mapping field 1708, and a source field 1710. Although the root node 1702 is in Fig.17 1702 is shown as the first node of the virtual representation (e.g., a trunk node, a root node, a master node, etc.), but the root node 1702 can be a child node for the glTF representing the scene. In addition, while one or more fields 1704 are described as part of the root node 1702, it should be understood that one or more fields 1704 can be included in any defined segment of the node structure for the virtual representation. For example, one or more fields 1704 can be included in a specific child node of the node structure.
[0136] In some cases, the type field 1706 may be used to indicate a format for the virtual representation (e.g., a virtual representation framework for generating the virtual representation). As an example, the type field may include a uniform resource name (UN) or uniform resource locator (URL) for indicating a format for the virtual representation. The URN / URL may provide an indication of a format for the virtual representation, which may be used to determine how (e.g., which algorithm) to use to reconstruct the virtual representation.
[0137] Mapping field 1708 may indicate one or more child nodes (e.g., subnodes) corresponding to various body segments under root node 1702. For example, rendering engine 1606 may use mapping field 1708 to determine where to find information about certain body segments with interactivity in the glTF. Mapping field 1708 may indicate, for example, that information about the right hand is in a particular node in the node hierarchy, and that the node is associated with interactivity information that may be in another node in the node hierarchy. In some cases, the interactivity information for a node / subnode may include information about whether and / or how the body segment associated with the node / subnode may interact with other objects and / or the environment.
[0138] The source field 1710 may indicate where certain information may be located in the input information. For example, the input information may be received in one or more data streams, and the source field 1710 may include information indicating where certain information is located in the data streams. As a more specific example, the source field 1710 may indicate that body pose information may be available in a certain segment of the deformation graph data stream.
[0139] When reconstructing a virtual representation (or avatar) from a representation of the virtual representation (or avatar), the mesh data may be mapped to the signaled mesh structure. The system may map interactivity metadata into the reconstructed virtual representation (or avatar) mesh.
[0140] The scene description may include describing the 3D reconstructed virtual representation (or avatar) as a dynamic mesh and / or a skinned mesh. Fig.17 The system or pipeline can take different representations of a virtual representation (or avatar) and perform 3D reconstruction. The reconstructed avatar components can then be fed into Fig.16 The rendering engine for rendering. Fig.16 Examples of inputs to the system may include a base mesh, one or more different types of textures, a color image (e.g., a red-green-blue (RGB) image) and / or a depth image, deformation parameters, skin weights, any combination thereof, and / or other inputs.
[0141] Fig.18 is an example of a JavaScript Object Notation (JSON) schema 1800 for use with the systems and techniques described herein. In the example schema 1800, the type field includes a URN for indicating a format for a virtual representation. The mapping field includes an array of paths in a node hierarchy to nodes for representing the virtual representation, and the paths correspond to child nodes (e.g., subnodes) in the node hierarchy. The source field includes an array of pointers (e.g., links) that point to locations in an input information data stream where information for the virtual representation can be.
[0142] Fig.19 1 is a flow chart illustrating a process 1900 for generating virtual content in a distributed system according to aspects of the present disclosure. The process 1900 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., a chipset, a codec, etc.), such as Figure 5 CPU 510 and / or GPU 525 and / or Fig. 20The computing device may be an animation and scene rendering system (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device (such as a mobile phone acting as a server device), an XR device acting as a server device, a network router, or other device acting as a server or other device). The operations of process 1900 may be implemented as a processor (e.g., Figure 5 CPU 510 and / or GPU 525 and / or Fig. 20 A software component that is executed and runs on the processor 2012).
[0143] At block 1902, a computing device (or a component thereof) may obtain information describing a virtual representation of a user (e.g., Fig.17 The information includes a hierarchical node set. A first node of the hierarchical node set includes a root node for the hierarchical node set. The first node includes a mapping configuration (e.g., Fig.17 The child nodes in the hierarchical set of nodes include data associated with a segment of the virtual representation of the user. In some cases, the segment includes at least one body part of the virtual representation of the user. In some cases, the segment of the virtual representation of the user includes a body part of the virtual representation of the user. In some cases, the body part of the virtual representation of the user includes a humanoid component. In some cases, the first node includes a root node for the hierarchical set of nodes (e.g., Fig.17 In some examples, the first node includes type information, and the data associated with the child node is processed based on the type information (e.g., Fig.17In some cases, the type information includes information indicating how to represent the virtual representation of the user. In some examples, the type information includes a generic resource name for indicating the format of the virtual representation for the user. In some cases, the type information can be used to indicate the format of the virtual representation for the user (e.g., a virtual representation framework for generating the virtual representation of the user). In some examples, the computing device (or its components) can process the portion of the data associated with the child node by processing the portion of the data associated with the child node based on the indicated format of the virtual representation for the user. In some cases, the first node includes source information, and wherein the portion of the data associated with the child node is identified based on the source information. In some examples, the data includes one or more data streams, and wherein the one or more data streams are based on the format of the virtual representation for the user. For example, data is received via a data stream set including a stream for a base model, which can be a generic mesh model of the virtual representation of the user, texture information, deformation map, parameterized data, and static metadata. In some cases, the sub-node of the child node includes data associated with the sub-segment of the segment of the virtual representation of the user. In some examples, a map may indicate one or more child nodes (eg, sub-nodes) under a root node corresponding to various body segments.
[0144] At block 1904, the computing device (or a component thereof) may identify a portion of data associated with the child node. In some cases, the source information may indicate where certain information may be located in the input information. In some cases, the portion of the data associated with the child node includes interactivity information indicating whether a segment of the user's virtual representation may interact with other objects.
[0145] At block 1906, the computing device (or a component thereof) may process portions of the data associated with the child nodes to generate segments of the virtual representation of the user. In some cases, the generated segments of the virtual representation of the user include grid information. In some cases, the computing device (or a component thereof) may process the grid information to render the generated segments of the virtual representation of the user.
[0146] The systems and techniques described herein provide interactivity with a virtual representation (or avatar). For example, the systems and techniques provide a mapping scheme between input avatar components and output avatars. The systems and techniques also provide a standardized naming scheme for scene nodes that are used to map to avatar segments (e.g., thumb, right / left hand, etc.). Interactivity can be assigned to an avatar using a standardized naming scheme, such as / humanoid / arms / left / fingers / index triggers the light to turn on when it is near a light switch. The mapping can use one of the input components, such as a basic humanoid or a poseless geometry portion, a texture coordinate in a texture map (e.g., patch (x, y, w, h) maps to the left index finger), a combination thereof, or other inputs.
[0147] As described above, systems and techniques decouple virtual representations (e.g., avatar representations) from integration of virtual representations in scene descriptions. The systems and techniques allow reference to different types of avatar representations and mapping of reconstructed 3D avatars onto humanoid parts that can be associated with interactive behaviors.
[0148] In some cases, a computing device or device configured to perform the operations of one or more of the processes described herein may include a processor, a microprocessor, a microcomputer, or other components of a device configured to perform the steps of a process. In some instances, such a device or device may include one or more sensors configured to capture image data and / or other sensor measurements. In some examples, such a computing device or device may include one or more sensors and / or cameras configured to capture one or more images or videos. In some cases, such a device or device may include a display for displaying an image. In some examples, one or more sensors and / or cameras are separated from the device or device, in which case the device or device receives the sensed data. Such a device or device may further include a network interface configured to transmit data.
[0149] The components of the device or apparatus configured to perform one or more operations of the processes described herein may be implemented in a circuit. For example, the components may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or the components may include computer software, firmware, or a combination thereof for performing the various operations described herein and / or may be implemented using computer software, firmware, or a combination thereof for performing the various operations described herein. The computing device may also include a display (as an example of an output device or in addition to an output device), a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive data based on an Internet Protocol (IP) or other types of data.
[0150] The operations of the various processes may be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation refers to a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations may be combined in any order and / or in parallel to implement the described process.
[0151] In addition, the processes described herein may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented by hardware as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed together on one or more processors, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0152] Fig. 20 is a schematic diagram showing an example of a system for implementing certain aspects of the present technology. Specifically, Fig. 20An example of a computing system 2000 is shown, which may be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any part thereof, wherein the components of the system communicate with each other using connection 2005. Connection 2005 may be a physical connection using a bus, or a direct connection into processor 2010, such as in a chipset architecture. Connection 2005 may also be a virtual connection, a networked connection, or a logical connection.
[0153] In some aspects, computing system 2000 is a distributed system, where the various functions described in the present disclosure can be distributed in a data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent a number of such components, each of which performs some or all of the functions of the component. In some aspects, a component can be a physical or virtual device.
[0154] The example system 2000 includes at least one processing unit (CPU or processor) 2010 and connections 2005 coupling various system components including system memory 2015 to the processor 2010, such as read-only memory (ROM) 2020 and random access memory (RAM) 2025. The computing system 2000 may include a cache 2011 of high-speed memory directly connected to the processor 2010, in close proximity to the processor 1510, or integrated as part of the processor 1510.
[0155] Processor 2010 may include any general purpose processor and hardware or software services, such as services 2032, 2034, and 2036 stored in storage device 2030, configured to control processor 2010 as well as special purpose processors where software instructions are incorporated into the actual processor design. Processor 2010 may be essentially a fully self-contained computing system, including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0156] To enable user interaction, the computing system 2000 includes input devices 2045, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keypad, a mouse, motion input, voice, etc. The computing system 2000 can also include output devices 2035, which can be one or more of a plurality of output mechanisms. In some cases, a multimodal system can enable a user to provide multiple types of input / output to communicate with the computing system 2000. The computing system 2000 can include a communication interface 2040, which can generally govern and manage user input and system output.
[0157] The communication interface may use a wired and / or wireless transceiver to perform or facilitate receiving and / or transmitting wired or wireless communications, including using an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, Ports / plugs, Ethernet ports / plugs, Fiber optic ports / plugs, Proprietary wired ports / plugs, Wireless signal transmission, Low power (BLE) wireless signal transmission, Wired and / or wireless transceivers for wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, WLAN signal transmission, visible light communication (VLC), world interoperability for microwave access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / long term evolution (LTE) cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof.
[0158] The communication interface 2040 may also include one or more GNSS receivers or transceivers for determining the location of the computing system 2000 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the Global Positioning System (GPS) based on the United States, the Global Navigation Satellite System (GLONASS) based on Russia, the BeiDou Navigation Satellite System (BDS) based on China, and the Galileo GNSS based on Europe. There is no limitation to operation on any particular hardware arrangement, and thus the basic features herein may be easily replaced with improved hardware or firmware arrangements when developed.
[0159] The storage device 2030 may be a non-volatile and / or non-transitory and / or computer-readable memory device, and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as a magnetic tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disk, a magnetic cassette, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) disc, a rewritable compact disc (CD) disc, a digital video disc (DVD) disc, a Blu-ray disc (BDD) disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, card, a smart card chip, a Europay, Mastercard and Visa (EMV) chip, a Subscriber Identity Module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a RAM, a static RAM (SRAM), a dynamic RAM (DRAM), a ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin-transfer torque RAM (STT-RAM), another memory chip or box and / or a combination thereof.
[0160] Storage device 2030 may include software services, servers, services, etc., which, when the code defining such software is executed by processor 2010, causes the system to perform functions. In some aspects, hardware services that perform a particular function may include software components stored in a computer-readable medium that are connected to the necessary hardware components (such as processor 2010, connection 2005, output device 2035, etc.) to perform the function. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transient media in which data may be stored and does not include carrier waves and / or transient electronic signals that propagate wirelessly or over a wired connection.
[0161] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing or carrying instructions and / or data. Computer-readable media may include non-transient media in which data may be stored and do not include carrier waves and / or transient electronic signals that are propagated wirelessly or through a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memory, or memory devices. Computer-readable media may store thereon codes and / or machine-executable instructions that may represent a combination of a process, function, subroutine, program, routine, subroutine, module, software package, class, or any instruction, data structure, or program statement. A code segment may be coupled to another code segment or hardware circuit by transmitting and / or receiving information, data, variables, parameters, or memory contents. Information, variables, parameters, data, etc. may be transmitted, forwarded, or sent via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0162] In some aspects, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer readable storage media expressly excludes media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0163] Specific details are provided in the above description to provide a thorough understanding of the aspects and examples provided herein. However, it will be understood by those of ordinary skill in the art that aspects can also be practiced without these specific details. For clarity of explanation, in some instances, the present technology may be presented as including separate functional blocks, which include devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components other than those shown in the accompanying drawings and / or described herein may be used. For example, circuits, systems, networks, processes and other components may be shown as components in the form of block diagrams, so as not to obscure aspects with unnecessary details. In other cases, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary details, so as to avoid obscuring aspects.
[0164] Various aspects may be described above as a process or method depicted as a flow chart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flow chart may describe an operation as a sequential process, many operations may be performed in parallel or simultaneously. Additionally, the order of these operations may be rearranged. A process terminates when its operations are completed, but there may be other steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination may correspond to the function returning to a calling function or a main function.
[0165] The process and method according to the above-mentioned example can be implemented using computer executable instructions stored in a computer readable medium or otherwise obtainable from a computer readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a specific function or functional group. The part of the computer resources used may be accessible through a network. Computer executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer readable media that can be used to store instructions, information used, and / or information created during the method according to the described examples include disks or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0166] The equipment implementing the process and method according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language or any combination thereof, and may adopt any form factor in a variety of form factors. When implemented in software, firmware, middleware or microcode, the program code or code segment (e.g., computer program product) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. One (or more) processors may perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices or other small personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functions described herein may also be embodied in peripheral devices or plug-in cards. By further example, such functions may also be implemented on circuit boards between different processes performed in different chips or in a single device.
[0167] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0168] In the foregoing description, various aspects of the present application are described with reference to specific aspects of the present invention, but those skilled in the art will recognize that the present invention is not limited thereto. Therefore, although the illustrative aspects of the present application have been described in detail herein, it should be understood that these inventive concepts can be embodied and adopted differently in other ways, and the appended claims are intended to be interpreted as including such variations, except as limited by the prior art. The various features and aspects of the above-mentioned applications can be used individually or in combination. In addition, aspects can be used for any number of environments and applications beyond the environments and applications described herein, without departing from the broader spirit and scope of this specification. Therefore, the description and the accompanying drawings should be considered illustrative rather than restrictive. For the purpose of illustration, the method is described in a specific order. It should be understood that in alternative aspects, these methods can be performed in an order different from the described order.
[0169] A person of ordinary skill in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by the less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of the present specification.
[0170] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0171] The phrase "coupled to" refers to any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection, and / or other suitable communication interface).
[0172] Claim language or other language that recites "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that recites "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repeating information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, duplication, or combination of A, B, and C. The language "at least one" of a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language that recites "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.
[0173] Claim language or other language that recites "at least one processor is configured to," "at least one processor is configured to," "one or more processors are configured to," "one or more processors are configured to," etc. indicates that one processor or multiple processors (in any combination) can perform the associated operations. For example, claim language that recites "at least one processor is configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or multiple processors are each responsible for a subset of operations X, Y, and Z, so that multiple processors perform X, Y, and Z together; or a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language that recites "at least one processor is configured to: X, Y, and Z" can mean that any single processor can only perform at least a subset of operations X, Y, and Z.
[0174] In the case of reference to one or more elements that perform a function (e.g., steps of a method), one element may perform all functions, or more than one element may perform a function together. When more than one element performs a function together, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, in the case of reference to one or more elements that are configured to cause another element (e.g., a device) to perform a function, one element may be configured to cause another element to perform all functions, or more than one element may be configured together to cause another element to perform a function.
[0175] In the case of an entity (e.g., any entity or device described herein) that performs a function or is configured to perform a function (e.g., a step of a method), the entity may be configured to cause one or more elements to perform a function (individually or collectively). One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) functions of the function, and / or any combination thereof. In the case of referring to the entity performing a function, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform a function together. When an entity is configured to cause more than one component to perform a function together, each function does not need to be performed by each component in those components (e.g., different functions can be performed by different components) and / or each function does not need to be performed by only one component as a whole (e.g., different components can perform different sub-functions of a function).
[0176] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, the above has been generally described based on the functionality of various illustrative components, blocks, modules, circuits, and steps. Whether such functionality is implemented as hardware or software depends on specific applications and the design constraints imposed on the overall system. Technicians can implement the described functions in different ways for each specific application, but such implementation decisions should not be interpreted as causing deviations from the scope of the present application.
[0177] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication device handheld devices, or integrated circuit devices with multiple uses, including applications in wireless communication device handheld devices and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or implemented separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium including a program code, the program code including instructions for executing one or more of the methods, algorithms, and / or operations described above when executed. A computer-readable data storage medium may form part of a computer program product, which may include packaging materials. A computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a FLASH memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the techniques may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0178] The program code can be executed by a processor, which may include one or more processors, for example, one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any technology described in this disclosure. A general-purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the technology described herein.
[0179] Illustrative aspects of the disclosure include:
[0180] Aspect 1. A method for generating a virtual representation of a user, comprising: receiving data for describing the virtual representation of the user, the data comprising a hierarchical node set, wherein a first node in the node set comprises type information, source information, a mapping or a combination thereof, and wherein a child node in the hierarchical node set comprises data associated with a segment of the virtual representation of the user; identifying a format associated with the virtual representation of the user based on the type information; identifying the child nodes in the hierarchical node set based on the mapping; identifying a segment of the data associated with the child nodes based on the source information; and processing the data associated with the segment of the data associated with the child nodes based on a corresponding format for the virtual representation of the user to generate a segment of the virtual representation of the user.
[0181] Aspect 2. A method according to aspect 1, wherein the sub-node of the sub-node includes data associated with a sub-segment of the segment of the virtual representation of the user.
[0182] Aspect 3. A method according to any one of Aspects 1 or 2, wherein the segment of the virtual representation includes a body part of the virtual representation of the user.
[0183] Aspect 4. A method according to aspect 3, wherein the body part of the virtual representation of the user includes a humanoid component.
[0184] Aspect 5. A method according to any one of aspects 1 to 4, wherein the segment of the data associated with the child node includes interactivity information.
[0185] Aspect 6. The method according to any one of aspects 1 to 5, wherein the first node comprises a root node for the hierarchical set of nodes.
[0186] Aspect 7. The method according to any one of aspects 1 to 6, wherein the type information includes a generic resource name for indicating the format of the virtual representation for the user.
[0187] Aspect 8. The method according to any one of aspects 1 to 7, wherein the generated segments of the virtual representation of the user include grid information.
[0188] Aspect 9. The method according to aspect 8 further includes: processing the grid information to render segments of the virtual representation of the user.
[0189] Aspect 10. A method according to any one of aspects 1 to 9, wherein the data comprises one or more data streams, and wherein the one or more data streams can be varied based on the format of the virtual representation for the user.
[0190] Aspect 11. A device for generating a virtual representation of a user, comprising: at least one memory; and at least one processor, which is coupled to the at least one memory and is configured to: receive data for describing the virtual representation of the user, the data comprising a hierarchical node set, wherein the first node in the node set comprises type information, source information, a mapping, or a combination thereof, and wherein the child nodes in the hierarchical node set comprise data associated with a segment of the virtual representation of the user; identify a format associated with the virtual representation of the user based on the type information; identify the child nodes in the hierarchical node set based on the mapping; identify the segment of the data associated with the child nodes based on the source information; and process the data associated with the segment of the data associated with the child nodes based on the corresponding format for the virtual representation of the user to generate a segment of the virtual representation of the user.
[0191] Aspect 12. An apparatus according to aspect 11, wherein a sub-node of the sub-node includes data associated with a sub-segment of the segment of the virtual representation of the user.
[0192] Aspect 13. An apparatus according to any one of Aspects 11 or 12, wherein the segment of the virtual representation of the user includes a body part of the virtual representation of the user.
[0193] Aspect 14. An apparatus according to aspect 13, wherein the body portion of the virtual representation of the user includes a humanoid component.
[0194] Aspect 15. An apparatus according to any one of aspects 11 to 14, wherein the segment of the data associated with the child node includes interactivity information.
[0195] Aspect 16. An apparatus according to any one of aspects 11 to 15, wherein the first node comprises a root node for the hierarchical set of nodes.
[0196] Aspect 17. An apparatus according to any one of aspects 11 to 16, wherein the type information includes a universal resource name for indicating the format of the virtual representation for the user.
[0197] Aspect 18. An apparatus according to any one of aspects 11 to 17, wherein the generated segments of the virtual representation of the user include grid information.
[0198] Aspect 19. The apparatus of aspect 18, wherein the at least one processor is further configured to process the mesh information to render segments of the virtual representation of the user.
[0199] Aspect 20. An apparatus according to any one of aspects 11 to 19, wherein the data comprises one or more data streams, and wherein the one or more data streams can be varied based on the format of the virtual representation for the user.
[0200] Aspect 21. A non-temporary computer-readable medium having instructions stored thereon, the instructions, when executed by at least one processor, causes the at least one processor to: receive data describing a virtual representation of the user, the data comprising a hierarchical set of nodes, wherein a first node in the node set comprises type information, source information, a mapping, or a combination thereof, and wherein a child node in the hierarchical set of nodes comprises data associated with a segment of the virtual representation of the user; identify a format associated with the virtual representation of the user based on the type information; identify the child nodes in the hierarchical set of nodes based on the mapping; identify a segment of the data associated with the child nodes based on the source information; and process the data associated with the segment of the data associated with the child nodes based on a corresponding format for the virtual representation of the user to generate a segment of the virtual representation of the user.
[0201] Aspect 22. The non-transitory computer-readable medium of aspect 21, wherein the sub-node of the sub-node comprises data associated with a sub-segment of the segment of the virtual representation of the user.
[0202] Aspect 23. The non-transitory computer-readable medium of any one of Aspects 21 or 22, wherein the segment of the virtual representation of the user comprises a body part of the virtual representation of the user.
[0203] Aspect 24. The non-transitory computer-readable medium of aspect 23, wherein the body portion of the virtual representation of the user includes a humanoid component.
[0204] Aspect 25. The non-transitory computer-readable medium of any one of aspects 21 to 24, wherein the segment of the data associated with the child node includes interactivity information.
[0205] Aspect 26. The non-transitory computer-readable medium according to any one of aspects 21 to 25, wherein the first node comprises a root node for the hierarchical set of nodes.
[0206] Aspect 27. The non-transitory computer-readable medium of any one of aspects 21 to 26, wherein the type information comprises a generic resource name indicating the format of the virtual representation for the user.
[0207] Aspect 28. The non-transitory computer-readable medium of any one of aspects 21 to 27, wherein the generated segments of the virtual representation of the user include grid information.
[0208] Aspect 29. The non-transitory computer-readable medium of aspect 28, wherein the instructions cause the at least one processor to process the mesh information to render segments of the virtual representation of the user.
[0209] Aspect 30. A non-transitory computer-readable medium according to any one of Aspects 21 to 29, wherein the data comprises one or more data streams, and wherein the one or more data streams can vary based on the format of the virtual representation for the user.
[0210] Aspect 31. An apparatus for generating a virtual representation of a user, the apparatus comprising one or more units for performing the operations according to any one of aspects 1 to 10.
[0211] Aspect 41. A method for generating a virtual representation of a user, comprising: obtaining information for describing the virtual representation of the user, the information comprising a hierarchical node set, wherein a first node in the hierarchical node set comprises a root node for the hierarchical node set, wherein the first node comprises a mapping configuration for mapping child nodes in the hierarchical node set to segments of the virtual representation of the user, and wherein the child nodes in the hierarchical node set comprise data associated with the segments of the virtual representation of the user; identifying portions of the data associated with the child nodes; and processing the portions of the data associated with the child nodes to generate the segments of the virtual representation of the user.
[0212] Aspect 42. A method according to Aspect 41, wherein the segment includes at least one body part of the virtual representation of the user.
[0213] Aspect 43. A method according to aspect 42, wherein the at least one body part of the virtual representation of the user includes a humanoid component.
[0214] Aspect 44. A method according to any one of Aspects 41-43, wherein the first node includes type information, and further wherein the portion of the data associated with the child node is processed based on the type information.
[0215] Aspect 45. A method according to aspect 44, wherein the type information includes information for indicating how the virtual representation of the user is represented.
[0216] Aspect 46. A method according to any one of Aspects 44-45, wherein the type information includes a universal resource name for indicating the format of the virtual representation for the user.
[0217] Aspect 47. A method according to Aspect 46, wherein processing the portion of the data associated with the subnode includes processing the portion of the data associated with the subnode based on the indicated format for the virtual representation of the user.
[0218] Aspect 48. A method according to any one of Aspects 41-47, wherein the first node includes source information, and wherein the portion of the data associated with the child node is identified based on the source information.
[0219] Aspect 49. A method according to any one of aspects 41-48, wherein the sub-node of the sub-node includes data associated with a sub-segment of the segment of the virtual representation of the user.
[0220] Aspect 50. A method according to any one of Aspects 41-49, wherein the portion of the data associated with the child node includes interactivity information indicating whether the segment of the virtual representation of the user can interact with other objects.
[0221] Aspect 51. A method according to any one of aspects 41-50, wherein the generated segments of the virtual representation of the user include grid information.
[0222] Aspect 52. The method according to aspect 51 further comprises: processing the grid information to render the generated segments of the virtual representation of the user.
[0223] Aspect 53. A method according to any one of Aspects 41-52, wherein the data includes one or more data streams, and wherein the one or more data streams are based on the format of the virtual representation for the user.
[0224] Aspect 54. A device for generating a virtual representation of a user, comprising: at least one memory; and at least one processor, coupled to the at least one memory and configured to: obtain information for describing the virtual representation of the user, the information comprising a hierarchical set of nodes, wherein a first node in the hierarchical set of nodes comprises a root node for the hierarchical set of nodes, wherein the first node comprises a mapping configuration for mapping child nodes in the hierarchical set of nodes to segments of the virtual representation of the user, and wherein the child nodes in the hierarchical set of nodes comprise data associated with the segments of the virtual representation of the user; identify portions of the data associated with the child nodes; and process the portions of the data associated with the child nodes to generate the segments of the virtual representation of the user.
[0225] Aspect 55. An apparatus according to Aspect 54, wherein the segment includes at least one body part of the virtual representation of the user.
[0226] Aspect 56. An apparatus according to Aspect 55, wherein the at least one body part of the virtual representation of the user includes a humanoid component.
[0227] Aspect 57. An apparatus according to any one of aspects 54-56, wherein the first node includes type information, and further wherein the portion of the data associated with the child node is processed based on the type information.
[0228] Aspect 58. An apparatus according to Aspect 57, wherein the type information includes information for indicating how the virtual representation of the user is represented.
[0229] Aspect 59. An apparatus according to any one of Aspects 57-58, wherein the type information includes a universal resource name for indicating the format of the virtual representation for the user.
[0230] Aspect 60. An apparatus according to Aspect 59, wherein, in order to process the portion of the data associated with the subnode, the at least one processor is configured to: process the portion of the data associated with the subnode based on the indicated format of the virtual representation of the user.
[0231] Aspect 61. An apparatus according to any one of aspects 54-60, wherein the first node includes source information, and wherein the portion of the data associated with the child node is identified based on the source information.
[0232] Aspect 62. An apparatus according to any one of aspects 54-61, wherein the sub-node of the sub-node includes data associated with a sub-segment of the segment of the virtual representation of the user.
[0233] Aspect 63. An apparatus according to any one of Aspects 54-62, wherein the portion of the data associated with the child node includes interactivity information indicating whether the segment of the virtual representation of the user is able to interact with other objects.
[0234] Aspect 64. An apparatus according to any one of aspects 54-63, wherein the generated segment of the virtual representation of the user includes grid information.
[0235] Aspect 65. An apparatus according to Aspect 64, wherein the at least one processor is further configured to process the mesh information to render the generated segments of the virtual representation of the user.
[0236] Aspect 66. An apparatus according to aspects 64-68, wherein the data includes one or more data streams, and wherein the one or more data streams are based on the format of the virtual representation for the user.
[0237] Aspect 67. A non-temporary computer-readable medium having instructions stored thereon, which instructions, when executed by at least one processor, cause the at least one processor to: obtain information describing a virtual representation of a user, the information comprising a hierarchical set of nodes, wherein a first node in the hierarchical set of nodes comprises a root node for the hierarchical set of nodes, wherein the first node comprises a mapping configuration for mapping child nodes in the hierarchical set of nodes to segments of the virtual representation of the user, and wherein the child nodes in the hierarchical set of nodes comprise data associated with the segments of the virtual representation of the user; identify portions of the data associated with the child nodes; and process the portions of the data associated with the child nodes to generate the segments of the virtual representation of the user.
[0238] Aspect 68. The non-transitory computer-readable medium of aspect 67, wherein the segment comprises at least one body part of the virtual representation of the user.
[0239] Aspect 69. The non-transitory computer-readable medium of aspect 68, wherein the body portion of the virtual representation of the user comprises a humanoid component.
[0240] Aspect 70. The non-transitory computer-readable medium of aspects 67-69, wherein the first node comprises type information, and further wherein the portion of the data associated with the child node is processed based on the type information.
[0241] Aspect 71. The non-transitory computer-readable medium of aspect 70, wherein the type information comprises information indicating how the virtual representation of the user is represented.
[0242] Aspect 72. The non-transitory computer-readable medium of any one of aspects 70-71, wherein the type information comprises a generic resource name indicating a format of the virtual representation for the user.
[0243] Aspect 73. A non-temporary computer-readable medium according to Aspect 72, wherein, in order to process the portion of the data associated with the subnode, the instructions cause the at least one processor to process the portion of the data associated with the subnode based on an indicated format for the virtual representation of the user.
[0244] Aspect 74. A non-transitory computer-readable medium according to any one of Aspects 67-73, wherein the first node includes source information, and wherein the portion of the data associated with the child node is identified based on the source information.
[0245] Aspect 75. The non-transitory computer-readable medium of any one of Aspects 67-74, wherein the sub-node of the sub-node comprises data associated with a sub-segment of the segment of the virtual representation of the user.
[0246] Aspect 76. A non-transitory computer-readable medium according to any one of Aspects 67-75, wherein the portion of the data associated with the subnode includes interactivity information for indicating whether the segment of the virtual representation of the user is able to interact with other objects.
[0247] Aspect 77. The non-transitory computer-readable medium of any one of aspects 67-76, wherein the generated segments of the virtual representation of the user include grid information.
[0248] Aspect 78. The non-transitory computer-readable medium of any one of aspects 67-77, wherein the instructions cause the at least one processor to process the mesh information to render segments of the virtual representation of the user.
[0249] Aspect 79. A non-transitory computer-readable medium according to any one of Aspects 67-78, wherein the data includes one or more data streams, and wherein the one or more data streams vary based on the format of the virtual representation for the user.
[0250] Aspect 80. An apparatus for generating a virtual representation of a user, the apparatus comprising one or more means for performing the operations according to any one of aspects 41-53.
Claims
1. A method for generating a virtual representation of a user, include: obtaining information describing a virtual representation of the user, the information comprising a hierarchical set of nodes, wherein a first node in the hierarchical set of nodes comprises a root node for the hierarchical set of nodes, wherein the first node comprises a mapping configuration for mapping child nodes in the hierarchical set of nodes to segments of the virtual representation of the user, and wherein the child nodes in the hierarchical set of nodes comprise data associated with the segments of the virtual representation of the user; identifying a portion of the data associated with the child node; and The portion of the data associated with the child node is processed to generate the segment of the virtual representation of the user.
2. The method according to claim 1, in, The segment includes at least one body part of the virtual representation of the user.
3. The method according to claim 2, in, The at least one body portion of the virtual representation of the user includes a humanoid component.
4. The method according to claim 1, in, The first node includes type information, and further wherein the portion of the data associated with the child node is processed based on the type information.
5. The method according to claim 4, in, The type information includes information indicating how the virtual representation of the user is represented.
6. The method according to claim 4, in, The type information includes a generic resource name indicating a format of the virtual representation for the user.
7. The method according to claim 6, in, Processing the portion of the data associated with the child node includes processing the portion of the data associated with the child node based on the indicated format for the virtual representation of the user.
8. The method according to claim 1, in, The first node includes source information, and wherein the portion of the data associated with the child node is identified based on the source information.
9. The method according to claim 1, in, The sub-node of the child node includes data associated with a sub-segment of the segment of the virtual representation of the user.
10. The method according to claim 1, in, The portion of the data associated with the child node includes interactivity information indicating whether the segment of the virtual representation of the user is able to interact with other objects.
11. The method according to claim 1, in, The generated segments of the virtual representation of the user include grid information.
12. The method according to claim 11, further comprising: include: The mesh information is processed to render generated segments of the virtual representation of the user.
13. The method according to claim 1, in, The data comprises one or more data streams, and wherein the one or more data streams are based on a format of the virtual representation for the user.
14. A device for generating a virtual representation of a user, include: at least one memory; as well as at least one processor coupled to the at least one memory and configured to: obtaining information describing a virtual representation of the user, the information comprising a hierarchical set of nodes, wherein a first node in the hierarchical set of nodes comprises a root node for the hierarchical set of nodes, wherein the first node comprises a mapping configuration for mapping child nodes in the hierarchical set of nodes to segments of the virtual representation of the user, and wherein the child nodes in the hierarchical set of nodes comprise data associated with the segments of the virtual representation of the user; identifying a portion of the data associated with the child node; and The portion of the data associated with the child node is processed to generate the segment of the virtual representation of the user.
15. The device according to claim 14, in, The segment includes at least one body part of the virtual representation of the user.
16. The device according to claim 15, in, The at least one body portion of the virtual representation of the user includes a humanoid component.
17. The device according to claim 14, in, The first node includes type information, and further wherein the portion of the data associated with the child node is processed based on the type information.
18. The device according to claim 17, in, The type information includes information indicating how the virtual representation of the user is represented.
19. The device according to claim 17, in, The type information includes a generic resource name indicating a format of the virtual representation for the user.
20. The device according to claim 19, in, To process the portion of the data associated with the child node, the at least one processor is configured to process the portion of the data associated with the child node based on the indicated format for the virtual representation of the user.
21. The device according to claim 14, in, The first node includes source information, and wherein the portion of the data associated with the child node is identified based on the source information.
22. The device according to claim 14, in, The sub-node of the child node includes data associated with a sub-segment of the segment of the virtual representation of the user.
23. The device according to claim 14, in, The portion of the data associated with the child node includes interactivity information indicating whether the segment of the virtual representation of the user is able to interact with other objects.
24. The device according to claim 14, in, The generated segments of the virtual representation of the user include grid information.
25. The device according to claim 24, in, The at least one processor is further configured to process the mesh information to render the generated segments of the virtual representation of the user.
26. The device according to claim 14, in, The data comprises one or more data streams, and wherein the one or more data streams are based on a format of the virtual representation for the user.
27. A non-transitory computer readable medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to: obtaining information describing a virtual representation of a user, the information comprising a hierarchical set of nodes, in, A first node in the hierarchical set of nodes comprises a root node for the hierarchical set of nodes, wherein the first node comprises a mapping configuration for mapping child nodes in the hierarchical set of nodes to segments of the virtual representation of the user, and wherein the child nodes in the hierarchical set of nodes comprise data associated with the segments of the virtual representation of the user; identifying a portion of the data associated with the child node; and The portion of the data associated with the child node is processed to generate the segment of the virtual representation of the user.
28. The non-transitory computer readable medium of claim 27, in, The segment includes at least one body part of the virtual representation of the user.
29. The non-transitory computer readable medium of claim 28, in, The at least one body portion of the virtual representation of the user includes a humanoid component.
30. The non-transitory computer readable medium of claim 27, in, The first node includes type information, and further wherein the portion of the data associated with the child node is processed based on the type information.
31. The non-transitory computer readable medium of claim 30, in, The type information includes information indicating how the virtual representation of the user is represented.
32. The non-transitory computer readable medium of claim 30, in, The type information includes a generic resource name indicating a format of the virtual representation for the user.
33. The non-transitory computer readable medium of claim 32, in, To process the portion of the data associated with the child node, the instructions cause the at least one processor to process the portion of the data associated with the child node based on the indicated format for the virtual representation of the user.
34. The non-transitory computer readable medium of claim 27, in, The first node includes source information, and wherein the portion of the data associated with the child node is identified based on the source information.
35. The non-transitory computer readable medium of claim 27, in, The sub-node of the child node includes data associated with a sub-segment of the segment of the virtual representation of the user.
36. The non-transitory computer readable medium of claim 27, in, The portion of the data associated with the child node includes interactivity information indicating whether the segment of the virtual representation of the user is able to interact with other objects.
37. The non-transitory computer readable medium of claim 27, in, The generated segments of the virtual representation of the user include grid information.
38. The non-transitory computer readable medium of claim 37, in, The instructions cause the at least one processor to process the mesh information to render segments of the virtual representation of the user.
39. The non-transitory computer readable medium of claim 27, in, The data comprises one or more data streams, and wherein the one or more data streams are based on a format of the virtual representation for the user.
Citation Information
Patent Citations
Apparatus and methods for image reconstruction using machine learning processes
US12112433B2
Adaptive bounding for three-dimensional morphable models
US20230035282A1
View dependent three-dimensional morphable models
US20230410447A1
Facial texture synthesis for three-dimensional morphable models
US20240029354A1
Avatar representation and audio generation
US20240078731A1